Duplication Detection Method, Apparatus and Electronic Device
By obtaining multiple clips of the novel and their digital fingerprints and matching them with the pre-established fingerprint library, the problem of low repetition detection efficiency in the prior art is solved, and efficient repetition detection is achieved.
Patent Information
- Application Number
- CN202110166867.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-04
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-02-04
AI Technical Summary
When performing novel repetition detection, the calculation amount is large, resulting in low detection efficiency and inability to effectively process massive novel data.
By obtaining multiple fragments in the text to be detected and their corresponding digital fingerprints, and matching these digital fingerprints with a pre-established digital fingerprint library, the duplication of the text to be detected is detected.
The calculation amount during matching is reduced, the efficiency of novel repetition detection is improved, and massive novel data can be effectively processed.
Smart Images

Figure CN112861505B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method, apparatus, and electronic device for detecting duplication, which can be specifically used in artificial intelligence technology fields such as intelligent search technology and knowledge graphs. Background Art
[0002] In recent years, with the continuous increase in the number of texts on the Internet, the possibility of text duplication has become greater and greater. Identical texts are often scattered in the same database. The existence of duplicate texts may increase the risk of index updates and retrieval anomalies, resulting in unsolvable errors and maintenance-related problems. Therefore, it becomes particularly important to merge or delete duplicate texts. Especially for novels in online literature, the number of creations has increased exponentially. Detecting the duplication of novels is an important means to achieve effective management of novels.
[0003] In the prior art, when detecting the duplication of novels, the similarity between the novel to be detected and each novel in the novel database is calculated respectively, and it is determined whether there is duplication between the novel to be detected and the novels in the database according to the calculation results of the novel similarity.
[0004] However, using the method of calculating novel similarity has a large amount of calculation, which will lead to a low detection efficiency of novel duplication. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for detecting duplication, which improves the detection efficiency of text duplication when detecting text duplication.
[0006] According to the first aspect of this application, a method for detecting duplication is provided. The method for detecting duplication may include:
[0007] Obtain a plurality of segments in the text to be detected, and digital fingerprints corresponding to each segment among the plurality of segments.
[0008] Match the digital fingerprints corresponding to each segment with the digital fingerprints in a pre-established digital fingerprint library respectively, where the digital fingerprint library includes digital fingerprints corresponding to each of the plurality of segments included in each of the plurality of texts.
[0009] Detect the duplication degree of the text to be detected according to the matching result.
[0010] According to the second aspect of this application, a device for detecting duplication is provided. The device for detecting duplication may include:
[0011] An obtaining unit, configured to obtain a plurality of segments in the text to be detected, and digital fingerprints corresponding to each segment among the plurality of segments.
[0012] A processing unit, configured to respectively match the digital fingerprints corresponding to the respective segments with the digital fingerprints in a pre-established digital fingerprint library, where the digital fingerprint library includes the digital fingerprints corresponding to the respective segments included in each of multiple texts.
[0013] A detection unit, configured to detect the duplication degree of the text to be detected according to the matching result.
[0014] According to a third aspect of the present application, there is provided an electronic device, which may include:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the duplication degree detection method described in the first aspect above.
[0018] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the duplication degree detection method described in the first aspect above.
[0019] According to a fifth aspect of the present application, there is provided a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the duplication degree detection method described in the first aspect above.
[0020] According to the technical solution of the present application, when detecting the duplication degree of the text to be detected, by determining the digital fingerprints corresponding to the respective segments in multiple segments, respectively matching the digital fingerprints corresponding to the respective segments with the digital fingerprints in a pre-established digital fingerprint library, and then detecting the duplication degree of the text to be detected according to the matching result. In this way, the digital fingerprints corresponding to the segments of the text to be detected are determined in units of the segments, and the duplication degree of the text to be detected is detected based on the digital fingerprints corresponding to the respective segments, reducing the calculation amount during matching, thereby improving the detection efficiency of the duplication degree of the novel.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. Description of the Drawings
[0022] The accompanying drawings are used to better understand the present solution and do not limit the present application. Among them:
[0023] Figure 1 is a schematic flowchart of a repetition degree detection method provided according to the first embodiment of the present application;
[0024] Figure 2 is a schematic flowchart of a method for determining digital fingerprints corresponding to each segment among multiple segments provided according to the second embodiment of the present application;
[0025] Figure 3 is a schematic structural diagram of a repetition degree detection device provided according to the third embodiment of the present application;
[0026] Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0028] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In the written description of the present application, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0029] The technical solutions provided by the embodiments of the present application can be applied to scenarios of text repetition degree detection. For example, in scenarios of novel repetition degree detection. In the prior art, when performing novel repetition degree detection, the novel similarity between the novel to be detected and each novel in the novel database is calculated respectively, and it is determined whether there is a repetition between the novel to be detected and the novels in the database according to the calculation result of the novel similarity. However, due to the large number of chapters and the huge chapter content of novels, the calculation amount is large when using the novel similarity calculation method, which not only leads to low detection efficiency of novel repetition degree, but also cannot perform repetition degree detection on a large amount of novel data.
[0030] To improve the detection efficiency of the repetition degree of novels and enable the detection of the repetition degree for a large amount of novel data, it can be considered to perform matching based on the digital fingerprints corresponding to the novels to detect the repetition degree of the novels. Although the algorithm based on digital fingerprints is simpler than the algorithm for novel similarity, when performing matching based on the digital fingerprints corresponding to the novels, a digital fingerprint corresponding to the entire novel is obtained through a series of calculations on the feature words in the entire novel, and matching is performed based on this digital fingerprint to detect the repetition degree of the novel. Similarly, due to the large number of chapters and the huge chapter content in novels, the complexity of calculating the digital fingerprint corresponding to the novel is relatively high, resulting in a low detection efficiency for the repetition degree of novels.
[0031] Therefore, when performing matching based on the digital fingerprints corresponding to the novels, multiple chapter contents can be selected from the novels. These multiple chapter contents can be partial chapter contents among all the chapters included in the novel, or the chapter contents of all the chapters, and can be specifically set according to actual needs. After selecting multiple chapter contents, for each chapter content, calculate its corresponding digital fingerprint respectively, and then match the multiple digital fingerprints corresponding to the multiple chapter contents calculated respectively to detect the repetition degree of the novel. In this way, determining the corresponding digital fingerprint in units of chapter contents effectively reduces the amount of calculation during matching compared to determining the corresponding digital fingerprint in units of the entire novel, thereby improving the detection efficiency of the repetition degree of novels.
[0032] Based on the above concept, an embodiment of the present application provides a repetition degree detection method, which can be applied to artificial intelligence technical fields such as intelligent search technology and knowledge graph. The specific solution includes: obtaining multiple segments in the text to be detected, and the digital fingerprints corresponding to each segment among the multiple segments; respectively matching the digital fingerprints corresponding to each segment with the digital fingerprints in a pre-established digital fingerprint library, where the digital fingerprint library includes the digital fingerprints corresponding to each of the multiple segments included in each of the multiple texts; and detecting the repetition degree of the text to be detected according to the matching results.
[0033] Exemplarily, the text to be detected can be a novel, a paper, or other texts, and can be specifically set according to actual needs. When obtaining multiple segments in the text to be detected, the multiple segments can be all the segments included in the text to be detected, or partial segments among all the segments, and can be specifically set according to actual needs, as long as the multiple segments selected are sufficient to detect the repetition degree of the text to be detected.
[0034] Exemplarily, if the text to be detected is a novel, the above-mentioned "fragment" can be understood as a "chapter" in the novel, that is, taking "chapters" as units to obtain the digital fingerprints corresponding to each chapter; if the text to be detected is an article, the above-mentioned "fragment" can be understood as a "paragraph" in the article, that is, taking "paragraphs" as units to obtain the digital fingerprints corresponding to each paragraph. It can be seen that the definition of "fragment" is related to the type of the text to be detected and can be specifically set according to actual needs. It should be noted that the chapters described in this application can be chapter contents. For example, obtaining the first five chapters of the novel to be detected means obtaining the contents of the first five chapters of the novel to be detected.
[0035] It can be seen that in the embodiments of this application, when detecting the duplication degree of the text to be detected, by determining the digital fingerprints corresponding to each fragment among multiple fragments, and respectively matching the digital fingerprints corresponding to each fragment with the digital fingerprints in the pre-established digital fingerprint library, and then according to the matching results, detecting the duplication degree of the text to be detected. In this way, taking the fragments of the text to be detected as units to determine their corresponding digital fingerprints, and detecting the duplication degree of the text to be detected based on the digital fingerprints corresponding to each fragment, reduces the calculation amount during matching, thereby improving the detection efficiency of the duplication degree of the novel.
[0036] Next, the duplication degree detection method provided by this application will be described in detail through specific embodiments. It can be understood that the following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0037] Embodiment 1
[0038] Figure 1 is a schematic flowchart of the duplication degree detection method according to the first embodiment of this application. This duplication degree detection method can be executed by software and / or hardware devices. For example, this hardware device can be a terminal or a server. Exemplarily, please refer to Figure 1 as shown, this duplication degree detection method can include:
[0039] S101. Obtain multiple fragments in the text to be detected, and the digital fingerprints corresponding to each fragment among the multiple fragments.
[0040] Taking the text to be detected as a novel as an example, correspondingly, the "fragment" can be understood as a "chapter" in the novel. Taking the novel to be detected including 10 chapters as an example, exemplarily, the multiple chapters obtained of the novel to be detected can be 5 chapters, or 6 chapters, or 10 chapters, and can be specifically set according to the digital fingerprints stored in the pre-established digital fingerprint library, as long as the multiple chapters selected are sufficient for detecting the duplication degree of the novel to be detected. It can be understood that the more chapters are selected, correspondingly, the higher the accuracy of the duplication degree detection result.
[0041] When detecting the repetition degree of a novel to be detected, the novel name and chapter information of the novel to be detected input by the user can be received in advance. Taking the acquisition of 5 chapters in the novel to be detected as an example, the first five chapters of the novel to be detected can be obtained by using the chapter identifier of the novel. These 5 chapters can be the first five chapters of the novel to be detected, or the middle five chapters, or the last five chapters, or any 5 chapters randomly selected from the novel to be detected, which can be specifically set according to actual needs. In the embodiment of the present application, the first five chapters of the novel to be detected can be selected, and the digital fingerprint corresponding to each of the first five chapters can be further determined.
[0042] Exemplarily, when obtaining the first five chapters of the novel to be detected, the first five chapters of the novel to be detected can be obtained by using web crawlers, or other extraction methods can be used to obtain the first five chapters of the novel to be detected. It can be understood that in the embodiment of the present application, the reason for selecting the first five chapters of the novel to be detected is as follows: First, when using web crawlers to obtain the first five chapters of the novel to be detected, compared with obtaining other five chapters, the acquisition efficiency is higher; Second, generally, novels are written in sequence according to chapters. If the last five chapters of the novel to be detected are selected, the repetition degree of the novel to be detected can only be detected after all chapters of the novel to be detected are completed; while if the first five chapters of the novel to be detected are selected, the repetition degree of the novel to be detected can be detected as long as the content of the first five chapters is completed, without waiting until all chapters of the novel to be detected are completed, making the detection of the repetition degree of the novel to be detected more timely.
[0043] After obtaining multiple segments in the text to be detected and the digital fingerprints corresponding to each segment among the multiple segments, the digital fingerprints corresponding to each segment can be respectively matched with the digital fingerprints in the pre-established digital fingerprint library, and according to the matching results, the repetition degree of the text to be detected can be detected, that is, the following S102 and S103 are executed:
[0044] S102: Respectively match the digital fingerprints corresponding to each segment with the digital fingerprints in the pre-established digital fingerprint library.
[0045] Among them, the digital fingerprint library includes the digital fingerprints corresponding to each of the multiple segments included in each of the multiple texts.
[0046] Continuing with the example of the text to be detected being a novel, for example, before matching the digital fingerprints corresponding to each chapter with the digital fingerprints in the pre-established digital fingerprint library respectively, it is necessary to pre-establish a digital fingerprint library. For example, when pre-establishing the digital fingerprint library, novels that meet the screening conditions can be obtained from the novel material library. The screening conditions may include: genuine, in the online state, content cooperation method, and novels with the number of chapters greater than or equal to the number of segments of multiple segments, and each novel corresponds to a unique identifier; the unique chapter identifiers of the first five chapters of each novel can be obtained through the novel identifier; and using the novel identifier and the unique chapter identifier, multiple chapters of the novel that meet the screening conditions can be obtained from the novel material library, and then the digital fingerprints corresponding to each chapter in the multiple chapters are obtained respectively, and the digital fingerprints corresponding to each chapter are stored in the digital fingerprint library. It can be understood that in addition to storing the digital fingerprints corresponding to each chapter in the digital fingerprint library, the novel identifier to which each digital fingerprint belongs, the chapter corresponding to each digital fingerprint, and the chapter identifier, etc. can also be stored.
[0047] It should be noted that if the digital fingerprints corresponding to the first five chapters of the existing novel are stored in the digital fingerprint library, then when detecting the duplication degree of the novel to be detected, correspondingly, the first five chapters in the novel to be detected can also be obtained; if the digital fingerprints corresponding to the last five chapters of the existing novel are stored in the digital fingerprint library, then when detecting the duplication degree of the novel to be detected, correspondingly, the last five chapters in the novel to be detected can also be obtained; if the digital fingerprint library is powerful enough to store the digital fingerprints corresponding to each chapter of the existing novel, then any five chapters can be selected in the novel to be detected.
[0048] When matching the digital fingerprints corresponding to each segment with the digital fingerprints in the pre-established digital fingerprint library respectively, the Hamming distance between the digital fingerprints corresponding to each chapter and the digital fingerprints in the digital fingerprint library can be calculated, and it can be judged whether the calculated Hamming distance is greater than the set distance threshold. For example, in the embodiment of the present application, strong equality can be used as the standard for successful digital fingerprint matching, and the distance threshold is set to 0 in advance. The digital sequence can be a 64-bit binary sequence. That is, when each bit value in the 64-bit binary sequences of the two digital fingerprints performing the matching operation is exactly the same, it is considered a successful match; otherwise, even if there is one bit value that is different, it is considered a failed match; in this way, the matching results between the digital fingerprints corresponding to each chapter in the multiple chapters of the novel to be detected and the digital fingerprints in the digital fingerprint library can be obtained.
[0049] After obtaining the matching results between the digital fingerprints corresponding to each chapter and the digital fingerprints in the digital fingerprint library respectively, the duplication degree of the novel to be detected can be detected according to the matching results, that is, the following S103 is executed:
[0050] S103. Detect the duplication degree of the text to be detected according to the matching result.
[0051] For example, when detecting the duplication degree of the text to be detected according to the matching result, the number of segments in multiple segments of the text to be detected that match successfully with multiple segments of the same target text in the digital fingerprint database can be determined; and the duplication degree of the text to be detected and the target text can be determined according to the number of segments.
[0052] Combined with the description in S102 above, after obtaining the Hamming distances between the digital fingerprints corresponding to each chapter and the digital fingerprints in the digital fingerprint database respectively, and obtaining the matching results between the digital fingerprints corresponding to each chapter and the digital fingerprints in the digital fingerprint database, the number of chapters in multiple chapters of the novel to be detected that match successfully with multiple chapters of the same target novel in the digital fingerprint database can be determined according to the matching result; for example, assuming that among the first five chapters of the novel to be detected, if there are three or more chapters whose corresponding digital fingerprints respectively match the corresponding digital fingerprints of multiple chapters of the same target novel, it can be determined that the novel to be detected is duplicated with the target novel, and on the contrary, it can be determined that the novel to be detected is not duplicated with the target novel.
[0053] In addition, when it is determined that the novel to be detected is duplicated with the target novel, relevant information of the target novel can also be output, such as the identifier of the target novel, each chapter that matches multiple chapters of the novel to be detected, the chapter identifier, and the digital fingerprints corresponding to each chapter, etc.
[0054] It can be understood that in the embodiment of the present application, when determining whether the novel to be detected is duplicated with the target novel, although the digital fingerprints corresponding to the first five chapters of the novel to be detected are respectively matched with the digital fingerprints in the digital fingerprint database, as long as there are three or four chapters whose corresponding digital fingerprints respectively match the corresponding digital fingerprints of multiple chapters of the same target novel, it is determined that the novel to be detected is duplicated with the target novel, rather than requiring that the digital fingerprint corresponding to each of the first five chapters respectively match the corresponding digital fingerprints of multiple chapters of the same target novel. The reason is that the chapters to which the two or one unmatched digital fingerprints belong may have minor modifications, and such minor modifications are tolerable. Therefore, it is achieved that when there are minor modifications in the chapters of the novel to be detected, the detection result will not be affected, thus ensuring the accuracy of the duplication detection.
[0055] It can be seen that in the embodiments of the present application, when detecting the duplication degree of the text to be detected, by determining the digital fingerprints corresponding to each segment among multiple segments, and respectively matching the digital fingerprints corresponding to each segment with each digital fingerprint in the pre-established digital fingerprint library, and then detecting the duplication degree of the text to be detected according to the matching results. In this way, the digital fingerprints corresponding to the segments of the text to be detected are determined in units of segments, and the duplication degree of the text to be detected is detected based on the digital fingerprints corresponding to each segment, reducing the computational amount during matching, thereby improving the detection efficiency of the duplication degree of the novel.
[0056] Based on the above Figure 1 shown embodiments, for the convenience of understanding how to determine the digital fingerprints corresponding to each segment in the above S102, below, the following Figure 2 shown Embodiment 2 will be used to describe in detail how to determine the digital fingerprints corresponding to each segment among multiple segments.
[0057] Embodiment 2
[0058] Figure 2 is a schematic flowchart of a method for determining the digital fingerprints corresponding to each segment among multiple segments according to the second embodiment of the present application. This method for determining the digital fingerprints corresponding to each segment among multiple segments can also be executed by software and / or hardware devices. For example, please refer to Figure 2 shown. This method for determining the digital fingerprints corresponding to each segment among multiple segments may include:
[0059] S201. Determine the target statement corresponding to each segment, where the target statement is the statement with the longest sentence length in the segment.
[0060] For example, when determining the target statement corresponding to each segment, each segment can be segmented to obtain multiple segments of content; each segment of content can be clause-segmented to obtain multiple sentences corresponding to each segment of content; and according to the sentence lengths of the multiple sentences corresponding to each segment of content, the statement with the longest sentence length in the segment is determined as the target statement. In this way, by using the target statement to replace the chapter to which it belongs for matching, the data computational amount can be greatly reduced, thereby improving the matching efficiency of the duplication degree.
[0061] Continuing with the example of the first five chapters of the novel to be detected, when determining the target sentence corresponding to each of the first five chapters in the novel to be detected, for each chapter, the content of the chapter can first be split into paragraphs by the line break character '\n' in the chapter; after splitting into multiple paragraphs, for each paragraph, the paragraph can be split by the full stop, question mark, and exclamation mark in the paragraph; after splitting into multiple sentences, the sentence length of each sentence can be calculated respectively, and based on the text lengths of all the sentences corresponding to the chapter, the sentence with the longest length can be selected in the chapter. After selecting the sentence with the longest length, the sentence with the longest length can be directly determined as the target sentence corresponding to the chapter; alternatively, after selecting the sentence with the longest length, the sentence with the longest length can first be processed to remove impurities, and / or the letter case can be unified, and then the processed sentence with the longest length can be determined as the target sentence corresponding to the chapter.
[0062] For example, when processing the sentence with the longest length to remove impurities, the filtering function of regular expressions can be used to filter out punctuation marks other than Chinese characters, English, and numbers in the sentence with the longest length, such as commas, full stops, semicolons, quotation marks, spaces, tab characters, carriage return characters, line break characters, etc.
[0063] For example, before processing the sentence with the longest length to unify the letter case, it can first be determined whether there is English in the sentence with the longest length after processing to remove impurities. If there is no English in the sentence with the longest length after processing to remove impurities, the sentence with the longest length after processing to remove impurities can be directly determined as the target sentence; on the contrary, if there are English letters in the sentence with the longest length after processing to remove impurities, the function of converting capital letters to lowercase letters can be used to uniformly convert the capital letter format of the English letters in the sentence with the longest length after processing to remove impurities into the lowercase format of English letters, or the lowercase letter format of the English letters in the sentence with the longest length after processing to remove impurities can be uniformly converted into the capital letter format of English letters, and the processed sentence with the longest length can be determined as the target sentence.
[0064] After respectively determining the target sentences corresponding to each segment, for the target sentences corresponding to each segment, the target sentences can be respectively subjected to word segmentation processing to obtain multiple feature words corresponding to the target sentences, that is, the following S202 is executed:
[0065] S202: For the target sentences corresponding to each segment, the target sentences are respectively subjected to word segmentation processing to obtain multiple feature words corresponding to the target sentences.
[0066] For example, when performing word segmentation on the target statement, the jieba Chinese word segmentation component's accurate mode can be used to segment the target statement to obtain the word segmentation result set after segmenting the target statement. Since the word segmentation result set may include noise words, such as "yo", "oh", etc., the words in the word segmentation result set after segmentation can be filtered to remove the noise words such as "yo", "oh" that have been deactivated, thereby obtaining multiple feature words corresponding to the target statement.
[0067] After obtaining multiple feature words corresponding to the target statement respectively, the hash algorithm can be used to calculate the hash value of each feature word among the multiple feature words. The hash value can be a 64-bit binary number. In this way, a digital sequence corresponding to each feature word can be generated according to the hash value corresponding to each feature word, that is, the following S203 is executed:
[0068] S203: Generate digital sequences corresponding to each feature word according to the hash values corresponding to each feature word among the multiple feature words.
[0069] Among them, the hash value is a binary number.
[0070] For example, when generating digital sequences corresponding to each feature word according to the hash values corresponding to each feature word among the multiple feature words, for the binary number corresponding to each feature word, if the value of a bit in the binary number is 1, the value of the bit is set to the first value; if the value of a bit in the binary number is 0, the value of the bit is set to the second value. In this way, a digital sequence corresponding to each feature word can be obtained; among them, the second value and the first value are opposite to each other. For example, in the embodiments of the present application, the first value can be 1 and the second value can be -1.
[0071] For example, for the hash value corresponding to each feature word, in the order from left to right, first judge the value of the first bit in the 64-bit binary number. If the value of the first bit is 1, the value of the first bit is set to 1; if the value of the first bit is 0, the value of the first bit is set to -1; then judge the value of the second bit in the 64-bit binary number. If the value of the second bit is 1, the value of the second bit is set to 1; if the value of the second bit is 0, the value of the second bit is set to -1; and so on. The value of each bit in the 64-bit binary number can be set, thereby obtaining a digital sequence corresponding to each feature word. The digital sequence is a 64-bit digital sequence composed of 1 and -1. In addition, it can also be judged in the order from right to left, or it can also be judged in the order from the middle to both sides, thereby obtaining a digital sequence corresponding to each feature word. The implementation manner is similar to the implementation manner corresponding to the order from left to right. For the relevant description of the implementation manner corresponding to the order from left to right, refer to the above. Here, the embodiments of the present application will not be described in detail.
[0072] After obtaining the digital sequences corresponding to each feature word among multiple feature words, the digital fingerprint corresponding to the segment can be generated according to the digital sequences corresponding to each feature word, that is, execute the following S204:
[0073] S204. Generate the digital fingerprint corresponding to the segment according to the digital sequences corresponding to each feature word.
[0074] Exemplarily, when generating the digital fingerprint corresponding to the segment according to the digital sequences corresponding to each feature word, the digital sequences corresponding to each feature word can be first subjected to an accumulation process to obtain the target digital sequence corresponding to the target statement; for each bit in the target digital sequence, if the value of the bit is greater than 0, the value of the bit is dimension-reduced to 1; if the value of the bit is less than or equal to 0, the value of the bit is dimension-reduced to 0 to obtain the digital fingerprint corresponding to the segment.
[0075] Combined with the description in the above S202, for the target statement corresponding to each chapter in the novel to be detected, after obtaining the 64-bit digital sequences corresponding to each of the multiple feature words corresponding to the target statement, the 64-bit digital sequences corresponding to each of the multiple feature words can be vertically accumulated and merged to obtain the merged result corresponding to the target statement, and the merged result is the target digital sequence corresponding to the target statement, and the target digital sequence is still a 64-bit digital sequence; for each bit in the 64-bit target digital sequence, if the value of the bit is greater than 0, the value of the bit is dimension-reduced to 1; if the value of the bit is less than or equal to 0, the value of the bit is dimension-reduced to 0 to obtain the digital fingerprint corresponding to the chapter. It can be seen that the digital fingerprint corresponding to the chapter is a 64-bit binary number composed of 0 and 1.
[0076] It can be seen that when determining the digital fingerprints corresponding to each segment among multiple segments in the text to be detected, by determining the target statement corresponding to each segment and performing word segmentation processing on the target statement respectively, multiple feature words corresponding to the target statement are obtained; then, according to the hash values corresponding to each feature word among the multiple feature words, digital sequences corresponding to each feature word are generated, so as to generate the digital fingerprint corresponding to the segment according to the digital sequences corresponding to each feature word. In this way, the digital fingerprint corresponding to each segment of the text to be detected is determined as a unit, so that the repetition degree of the text to be detected can be detected based on the digital fingerprints corresponding to each segment, reducing the calculation amount during matching, thereby improving the detection efficiency of the repetition degree of the novel.
[0077] To verify the repeatability detection method provided in the embodiments of the present application, 1000 novels were randomly selected for recall rate evaluation. The repeatability detection method provided in the embodiments of the present application was used for repeatability detection. The recall rate of repeatability detection for these 1000 novels was 60%. While for the traditional repeatability detection using MD5 values, the recall rate of repeatability detection for novels was 50%. Thus, it can be seen that the repeatability detection method provided in the embodiments of the application can effectively improve the recall rate of novel repeatability detection when performing novel repeatability detection. As for the accuracy of repeatability detection, 1000 queries were randomly selected to evaluate its accuracy rate, and the current accuracy rate is 99%. While for the traditional repeatability detection method using MD5 values, the accuracy rate of repeatability detection for novels is 81%. Therefore, through the repeatability detection method provided in the embodiments of the present application, the problem in the prior art that it is impossible to meet the requirement of repeatability detection for a large amount of novel data can be solved, not only improving the recall rate and accuracy rate of novel repeat or near-repeat detection, but also improving the detection efficiency of novel repeatability.
[0078] Embodiment III
[0079] Figure 3 It is a schematic block diagram of a repeatability detection device 300 provided according to the third embodiment of the present application. For example, please refer to Figure 3 As shown, the repeatability detection device 300 may include:
[0080] An acquisition unit 301, configured to acquire multiple segments in the text to be detected, and digital fingerprints corresponding to each segment among the multiple segments.
[0081] A processing unit 302, configured to respectively match the digital fingerprints corresponding to each segment with the digital fingerprints in a pre-established digital fingerprint library, where the digital fingerprint library includes digital fingerprints corresponding to each of the multiple segments included in each text among multiple texts.
[0082] A detection unit 303, configured to detect the repeatability of the text to be detected according to the matching result.
[0083] Optionally, the acquisition unit 301 includes a first acquisition module and a second acquisition module.
[0084] The first acquisition module is configured to determine a target statement corresponding to each segment, and the target statement is the statement with the longest sentence length in the segment.
[0085] The second acquisition module is configured to generate digital fingerprints corresponding to each segment according to the target statement corresponding to each segment.
[0086] Optionally, the second acquisition module includes a first acquisition sub-module, a second acquisition sub-module, and a third acquisition sub-module.
[0087] The first acquisition sub-module is used to perform word segmentation on the target statements corresponding to each segment respectively, and obtain multiple feature words corresponding to the target statements.
[0088] The second acquisition sub-module is used to generate a digital sequence corresponding to each feature word according to the hash value corresponding to each feature word among the multiple feature words.
[0089] The third acquisition sub-module is used to generate a digital fingerprint corresponding to the segment according to the digital sequence corresponding to each feature word.
[0090] Optionally, the hash value is a binary number; the second acquisition sub-module is specifically used for the binary number corresponding to each feature word. If the value of a bit in the binary number is 1, the value of the bit is set to the first numerical value; if the value of a bit in the binary number is 0, the value of the bit is set to the second numerical value, so as to obtain the digital sequence corresponding to each feature word; wherein, the second numerical value and the first numerical value are opposite to each other.
[0091] Optionally, the third acquisition sub-module is specifically used to perform an accumulation process on the digital sequences corresponding to each feature word to obtain a target digital sequence corresponding to the target statement; and perform dimensionality reduction processing on the target digital sequence according to the value of each bit in the target digital sequence to obtain the digital fingerprint corresponding to the segment.
[0092] Optionally, the third acquisition sub-module is specifically used for each bit. If the value of the bit is greater than 0, the value of the bit is reduced in dimension to 1; if the value of the bit is less than or equal to 0, the value of the bit is reduced in dimension to 0, so as to obtain the digital fingerprint corresponding to the segment.
[0093] Optionally, the detection unit 303 includes a first detection module and a second detection module.
[0094] The first detection module is used to determine the number of segments that match successfully with multiple segments of the same target text in the multiple segments of the text to be detected.
[0095] The second detection module is used to determine the duplication degree of the text to be detected and the target text according to the number of segments.
[0096] Optionally, the first acquisition module includes a fourth acquisition sub-module, a fifth acquisition sub-module, and a sixth acquisition sub-module.
[0097] The fourth acquisition sub-module is used to perform segmentation processing on each segment to obtain multiple segments of content.
[0098] The fifth acquisition sub-module is used to perform clause segmentation on each segment of content among the multiple segments of content to obtain multiple sentences corresponding to each segment of content.
[0099] The sixth acquisition sub-module is used to determine the target statement in the segment according to the sentence length of the multiple sentences corresponding to each segment of content.
[0100] The repeatability detection device 300 provided by the embodiments of the present application can execute the technical solutions of the repeatability detection method shown in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the repeatability detection method. For details, refer to the implementation principle and beneficial effects of the repeatability detection method, which will not be elaborated here.
[0101] According to the embodiments of the present application, the present application also provides a computer program product, which includes: a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the execution of the computer program by at least one processor enables the electronic device to execute the solutions provided in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the repeatability detection method. For details, refer to the implementation principle and beneficial effects of the repeatability detection method, which will not be elaborated here.
[0102] According to the embodiments of the present application, the present application also provides an electronic device and a readable storage medium.
[0103] According to the embodiments of the present application, the present application also provides a computer program product, which includes: a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the execution of the computer program by at least one processor enables the electronic device to execute the solutions provided in any of the above embodiments.
[0104] Figure 4 It is a schematic block diagram of an electronic device 400 provided by the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0105] As Figure 4As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 402 or computer programs loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0106] Multiple components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0107] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the duplication degree detection method. For example, in some embodiments, the duplication degree detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the duplication degree detection method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the duplication degree detection method by any other appropriate means (such as by means of firmware).
[0108] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0109] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.
[0110] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0112] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0113] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0114] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is made herein.
[0115] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A method for detecting duplication, comprising: Obtaining multiple segments in the text to be detected; Determining the target sentence corresponding to each segment, where the target sentence is the sentence with the longest length in the segment; Generating a digital fingerprint corresponding to each segment according to the target sentence corresponding to each segment; Respectively matching the digital fingerprints corresponding to each segment with the digital fingerprints in a pre-established digital fingerprint library, where the digital fingerprint library includes digital fingerprints corresponding to multiple segments included in each text among multiple texts; wherein, the digital fingerprint library is a database constructed by obtaining multiple chapters of novels that meet the screening conditions from a novel material library, then respectively obtaining the digital fingerprints corresponding to each chapter among the multiple chapters, and storing the digital fingerprints corresponding to each chapter; determining the number of segments in the multiple segments of the text to be detected that match successfully with the multiple segments of the same target text in the digital fingerprint library; Determining the duplication degree of the text to be detected and the target text according to the number of segments; If the number of segments is greater than or equal to a preset threshold, determining that the text to be detected and the target text are duplicated; Wherein, the determining the target sentence corresponding to each segment includes: Performing segmentation processing on each segment to obtain multiple segments of content; Performing sentence splitting processing on each segment of content in the multiple segments of content to obtain multiple sentences corresponding to each segment of content; Determining the target sentence in the segment according to the sentence lengths of the multiple sentences corresponding to each segment of content.
2. The method according to claim 1, wherein The generating a digital fingerprint corresponding to each segment according to the target sentence corresponding to each segment includes: For the target sentence corresponding to each segment, performing word segmentation processing on the target sentence respectively to obtain multiple feature words corresponding to the target sentence; Generating a digital sequence corresponding to each feature word according to the hash value corresponding to each feature word in the multiple feature words; Generating a digital fingerprint corresponding to the segment according to the digital sequences corresponding to each feature word.
3. The method according to claim 2, wherein, The hash value is a binary number, and the generating a digital sequence corresponding to each feature word according to the hash value corresponding to each feature word in the multiple feature words includes: For the binary number corresponding to each feature word, if the value of a bit in the binary number is 1, setting the value of the bit to a first numerical value; if the value of a bit in the binary number is 0, setting the value of the bit to a second numerical value, to obtain the digital sequence corresponding to each feature word; Wherein, the second numerical value and the first numerical value are opposite to each other.
4. The method according to claim 2, wherein, The generating a digital fingerprint corresponding to the segment according to the digital sequences corresponding to each feature word includes: Performing an accumulation process on the digital sequences corresponding to each feature word to obtain a target digital sequence corresponding to the target sentence; Performing dimensionality reduction processing on the target digital sequence according to the values of each bit in the target digital sequence to obtain the digital fingerprint corresponding to the segment.
5. The method according to claim 4, wherein The performing dimensionality reduction processing on the target digital sequence according to the values of each bit in the target digital sequence to obtain the digital fingerprint corresponding to the segment includes: For each bit, if the value of the bit is greater than 0, reducing the value of the bit to 1; If the value of the bit is less than or equal to 0, the value of the bit is reduced to 0 through dimensionality reduction processing to obtain the digital fingerprint corresponding to the segment.
6. A duplication detection device, comprising: An acquisition unit, configured to acquire multiple segments in a text to be detected, and digital fingerprints corresponding to each of the multiple segments; A processing unit, configured to respectively match the digital fingerprints corresponding to each segment with the digital fingerprints in a pre-established digital fingerprint library, where the digital fingerprint library includes digital fingerprints corresponding to multiple segments included in each of multiple texts; wherein, the digital fingerprint library is a database constructed by acquiring multiple chapters of novels that meet screening conditions from a novel material library, then respectively acquiring digital fingerprints corresponding to each of the multiple chapters, and storing the digital fingerprints corresponding to each chapter; A detection unit, configured to detect the duplication degree of the text to be detected according to the matching result; The acquisition unit includes a first acquisition module and a second acquisition module; The first acquisition module is configured to determine a target sentence corresponding to each segment, where the target sentence is the sentence with the longest length in the segment; The second acquisition module is configured to generate digital fingerprints corresponding to each segment according to the target sentences corresponding to each segment; The detection unit includes a first detection module and a second detection module; The first detection module is configured to determine the number of segments that match successfully with multiple segments of the same target text in multiple segments of the text to be detected; The second detection module is configured to determine the duplication degree of the text to be detected and the target text according to the number of segments; if the number of segments is greater than or equal to a preset threshold, it is determined that the text to be detected and the target text are duplicated; Wherein, the first acquisition module includes a fourth acquisition sub-module, a fifth acquisition sub-module, and a sixth acquisition sub-module; The fourth acquisition sub-module is configured to perform segmentation processing on each segment to obtain multiple segments of content; The fifth acquisition sub-module is configured to perform clause separation processing on each segment of content in the multiple segments of content to obtain multiple sentences corresponding to each segment of content; The sixth acquisition sub-module is configured to determine the target sentence in the segment according to the sentence lengths of the multiple sentences corresponding to each segment of content.
7. The device according to claim 6, wherein The second acquisition module includes a first acquisition sub-module, a second acquisition sub-module, and a third acquisition sub-module; The first acquisition sub-module is configured to perform word segmentation processing on the target sentence corresponding to each segment respectively to obtain multiple feature words corresponding to the target sentence; The second acquisition sub-module is configured to generate digital sequences corresponding to each of the multiple feature words according to the hash values corresponding to each of the multiple feature words; The third acquisition sub-module is configured to generate digital fingerprints corresponding to the segment according to the digital sequences corresponding to each of the multiple feature words.
8. The device according to claim 7, wherein The hash value is a binary number; The second obtaining sub-module is specifically configured to, for the binary numbers corresponding to each feature word, if the value of a bit in the binary number is 1, set the value of the bit to a first numerical value; if the value of the bit in the binary number is 0, set the value of the bit to a second numerical value, so as to obtain the digital sequences corresponding to the respective feature words; wherein, the second numerical value and the first numerical value are opposite to each other.
9. The apparatus according to claim 7, wherein, The third obtaining sub-module is specifically configured to perform an accumulation process on the digital sequences corresponding to the respective feature words to obtain a target digital sequence corresponding to the target statement; and perform a dimensionality reduction process on the target digital sequence according to the values of the bits in the target digital sequence to obtain a digital fingerprint corresponding to the segment.
10. The apparatus according to claim 9, wherein, The third obtaining sub-module is specifically configured to, for each bit, if the value of the bit is greater than 0, perform a dimensionality reduction process on the value of the bit to 1; if the value of the bit is less than or equal to 0, perform a dimensionality reduction process on the value of the bit to 0, so as to obtain the digital fingerprint corresponding to the segment.
11. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the duplication degree detection method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the duplication degree detection method according to any one of claims 1-5.
Citation Information
Patent Citations
Method for detecting academic document plagiarism based on deep neural networks
CN106095735A
Internet information content similarity definition method
CN106649214A