Article duplicate checking method, device and equipment, readable storage medium and program product

By adjusting the keyword weights in the text to be checked for duplicate content and calculating multi-dimensional similarity, the problem of insufficient adaptability of emerging words in article duplicate checking is solved, and the accuracy of the duplicate checking results is improved.

CN120805881APending Publication Date: 2025-10-17CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510901845.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

Smart Images

  • Figure CN120805881A_ABST
    Figure CN120805881A_ABST
Patent Text Reader

Abstract

The invention provides an article duplicate checking method and device, equipment, a readable storage medium and a program product, and relates to the technical field of natural language process.The method comprises the steps that under the condition that it is determined that a target vocabulary exists, the weight corresponding to each keyword in a text to be subjected to duplicate checking is determined according to the target vocabulary; the target vocabularies are vocabularies of which the occurrence frequency is greater than or equal to a first threshold value and the occurrence frequency is smaller than or equal to a second threshold value in the text retrieval library in the to-be-duplicated text; obtaining a first format similarity, a first statement similarity and a first topic similarity between the to-be-duplicate-checked text and each retrieval text based on each keyword in the to-be-duplicate-checked text and the weight corresponding to the keyword; performing weighted summation on the first format similarity, the first statement similarity and the first topic similarity to obtain first similarities between the text to be subjected to duplicate checking and each retrieval text; and according to the first similarity, screening from the retrieval text set to obtain a duplicate checking result. According to the embodiment of the invention, the duplicate checking result is high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, in particular to an article duplicate detection method, device, equipment, readable storage medium and program product. BACKGROUND

[0002] With the development of technology, the application of natural language processing (NLP) technology in article intelligent duplicate detection is increasingly widespread. Article intelligent duplicate detection is mainly applied in the field of academic education, including but not limited to: graduate thesis duplicate detection, academic journal duplicate detection, conference paper duplicate detection, homework and report duplicate detection, etc.

[0003] The article duplicate detection method in the prior art has the problems of incomplete article understanding and poor adaptability of emerging words when facing articles with long length or insufficient sample quantity in the retrieval library, thereby resulting in poor accuracy of duplicate detection results. SUMMARY

[0004] The purpose of the present application is to provide an article duplicate detection method, device, equipment, readable storage medium and program product, which solves the problem of poor accuracy of duplicate detection results caused by incomplete article understanding and poor adaptability of emerging words in the duplicate detection process.

[0005] In order to solve the above technical problems, the present application provides an article duplicate detection method, comprising:

[0006] In the case where the target word is determined, the weight corresponding to each keyword in the duplicate detection text is determined according to the target word; wherein the target word is a word with a frequency greater than or equal to a first threshold value in the duplicate detection text and a frequency less than or equal to a second threshold value in the text retrieval library, and the first threshold value is greater than the second threshold value;

[0007] Based on each keyword in the duplicate detection text and the weight corresponding to the keyword, the first format similarity, the first sentence similarity and the first topic similarity between the duplicate detection text and each retrieval text in the retrieval text set are obtained; wherein the retrieval text set belongs to the text retrieval library;

[0008] The first format similarity, the first sentence similarity and the first topic similarity are weighted and summed to obtain the first similarity between the duplicate detection text and each retrieval text in the retrieval text set;

[0009] According to the first similarity, the duplicate detection result is obtained from the retrieval text set.

[0010] Optionally, the determining the weight corresponding to each keyword in the text to be checked according to the target vocabulary comprises:

[0011] performing a search in the network article library according to the target vocabulary to obtain a first search result;

[0012] in a case where the first search result indicates that the target article does not exist in the network article library, extracting a target keyword from the keywords of the text to be checked, wherein the keywords of the target article include the target vocabulary, the distance between the target keyword and the target vocabulary in the text to be checked is less than a third threshold, and the occurrence frequency of the target keyword is greater than or equal to a fourth threshold;

[0013] determining the weight corresponding to each keyword in the text to be checked by increasing the weight corresponding to the target keyword based on the weight corresponding to each keyword in the text to be checked preset in advance.

[0014] Optionally, after the performing a search in the network article library according to the target vocabulary to obtain a first search result, the method further comprises:

[0015] in a case where the first search result indicates that the target article exists in the network article library, adding the target article to the text search library to obtain an updated text search library;

[0016] in a case where the first search result indicates that the target article does not exist in the network article library and a target keyword is extracted from the text to be checked, performing a search in the network article library according to the target keyword and adding the search result to the text search library to obtain an updated text search library.

[0017] Optionally, the method further comprises:

[0018] obtaining feature information of the text to be checked, wherein the feature information comprises at least one of the following: content feature, theme feature and author feature;

[0019] determining a second similarity between the text to be checked and each text in the text search library according to the feature information;

[0020] selecting the search text set from the text search library according to the second similarity.

[0021] Optionally, the obtaining the feature information of the text to be checked comprises at least one of the following:

[0022] extracting the content feature of the text to be checked by using a natural language processing technology, wherein the content feature comprises at least one of the following: keyword, text theme and text conclusion.

[0023] identify a research topic and a research direction of the to-be-duplicated text, and obtain a topic feature of the to-be-duplicated text;

[0024] determine an author feature of the to-be-duplicated text according to author background information of the to-be-duplicated text.

[0025] Optionally, the feature information includes a content feature, a topic feature, and an author feature.

[0026] The second similarity between the to-be-duplicated text and each text in the text retrieval library is determined according to the feature information, and the method includes:

[0027] The content feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated according to the content feature.

[0028] The topic feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated according to the topic feature.

[0029] The author feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated according to the author feature.

[0030] The content feature similarity, the topic feature similarity, and the author feature similarity are weighted and summed to obtain the second similarity between the to-be-duplicated text and each text.

[0031] Optionally, the first format similarity between the to-be-duplicated text and each retrieval text in the retrieval text set is obtained based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, and the method includes:

[0032] The to-be-duplicated text is divided into a plurality of first ranges according to the title and the sub-title in the to-be-duplicated text, wherein one first range corresponds to one title or sub-title.

[0033] A semantic network graph model corresponding to the to-be-duplicated text is constructed based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, wherein one first range corresponds to one semantic network graph model, the nodes in the semantic network graph model are text units, the weight of the node corresponds to the weight of the keyword in the corresponding text unit, and the edges in the semantic network graph model are used to represent the relationship between the nodes.

[0034] For each first range, the second format similarity between the to-be-duplicated text and each retrieval text in the retrieval text set is calculated according to the semantic network graph model.

[0035] According to the second format similarity, a first format similarity between the to-be-duplicated text and each search text in the search text set is obtained.

[0036] Optionally, the first theme similarity between the to-be-duplicated text and each search text in the search text set is obtained based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, and the first theme similarity comprises:

[0037] The to-be-duplicated text is input into a pre-trained latent Dirichlet allocation (LDA) model based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, and a theme distribution vector of the to-be-duplicated text output by the LDA model is obtained, wherein the weight value of a keyword in the theme distribution vector is greater than a fourth threshold value.

[0038] According to the theme distribution vector, a first theme similarity between the to-be-duplicated text and each search text in the search text set is calculated.

[0039] The embodiment of the application further provides an article duplication checking device, comprising:

[0040] The first determination module is configured to determine the weight corresponding to each keyword in the to-be-duplicated text according to the target vocabulary when it is determined that the target vocabulary exists, wherein the target vocabulary is a vocabulary with a frequency greater than or equal to a first threshold value in the to-be-duplicated text and a frequency less than or equal to a second threshold value in the text search library, and the first threshold value is greater than the second threshold value.

[0041] The first calculation module is configured to obtain a first format similarity, a first sentence similarity and a first theme similarity between the to-be-duplicated text and each search text in the search text set based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, wherein the search text set belongs to the text search library.

[0042] The second calculation module is configured to perform weighted summation on the first format similarity, the first sentence similarity and the first theme similarity to obtain a first similarity between the to-be-duplicated text and each search text in the search text set.

[0043] The first screening module is configured to screen a duplication checking result from the search text set according to the first similarity.

[0044] The embodiment of the application further provides a network device, comprising a processor, a memory and a program stored in the memory and executable on the processor, wherein the program, when executed on the processor, implements the article duplication checking method according to any one of the above.

[0045] An embodiment of the present invention also provides a readable storage medium, comprising: a program is stored on the readable storage medium, and when the program is executed by a processor, the steps of the article duplication checking method as described in any of the above items are implemented.

[0046] An embodiment of the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the article duplication checking method as described in any of the above items.

[0047] At least one of the above technical solutions of the present invention has the following beneficial effects:

[0048] In the above scheme, during the article duplicate checking process, if the target vocabulary exists in the text to be checked for duplicates, that is, if there is a vocabulary in the text to be checked for duplicates whose frequency of appearance is greater than or equal to the first threshold value and whose frequency of appearance in the text retrieval library is less than or equal to the second threshold value, it is considered that an emerging vocabulary that is relatively new to the vocabulary in the text retrieval library has appeared. Then, based on the target vocabulary, the weight corresponding to each keyword in the text to be checked for duplicates is determined, and based on the weight corresponding to each keyword and keyword in the text to be checked for duplicates, the first similarity between the text to be checked for duplicates and each search text in the search text set is calculated, and the duplicate checking result is obtained by screening based on the first similarity. Compared with the existing technology, this scheme takes into account the influence of emerging vocabulary, enhances the adaptability of emerging vocabulary, and can improve the accuracy of the duplicate checking results.

[0049] Moreover, in the process of calculating the first similarity between the text to be checked for duplicates and each search text in the search text set, this scheme first obtains the first format similarity, first sentence similarity and first topic similarity between the text to be checked for duplicates and each search text in the search text set based on each keyword in the text to be checked for duplicates and the weight corresponding to the keyword, and then performs weighted summation on the first format similarity, first sentence similarity and first topic similarity to obtain the first similarity between the text to be checked for duplicates and each search text in the search text set. Compared with the existing technology, this scheme takes into account the three aspects of format similarity, sentence similarity and topic similarity at the same time, so that the understanding of the article is more comprehensive and the accuracy of the duplicate checking results can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Schematic diagram of the process of checking for duplicate articles according to an embodiment of the present invention;

[0051] Figure 2 Schematic diagram of the training process of the multi-dimensional feature model of the first embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of the construction process of the second similarity model according to the second embodiment of the present invention;

[0053] Figure 4A flowchart of an article duplicate checking method according to an embodiment of the present application is shown in FIG. 1.

[0054] Figure 5 A structural diagram of an article duplicate checking device according to an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0056] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0057] As shown in FIG. 1, an article duplicate checking method according to an embodiment of the present application comprises the following steps. Figure 1

[0058] In step S101, in a case where it is determined that there is a target vocabulary, the weight corresponding to each keyword in the text to be checked is determined according to the target vocabulary; wherein the target vocabulary is a vocabulary in the text to be checked whose frequency of occurrence is greater than or equal to a first threshold value and whose frequency of occurrence in the text retrieval library is less than or equal to a second threshold value, and the first threshold value is greater than the second threshold value.

[0059] Optionally, the target vocabulary can also be referred to as an emerging vocabulary.

[0060] ​In step S101, the frequency of the target vocabulary in the to-be-checked text = the number of times the target vocabulary appears in the to-be-checked text / the total number of times all vocabularies appear in the to-be-checked text, the frequency of the target vocabulary in the text retrieval library = the number of times the target vocabulary appears in the text retrieval library / the total number of times all vocabularies appear in the text retrieval library, if there is a vocabulary in the to-be-checked text whose frequency is greater than or equal to the first threshold value and whose frequency in the text retrieval library is less than or equal to the second threshold value, it is considered that the target vocabulary exists in the to-be-checked text relative to the text retrieval library, and then the weight corresponding to each keyword in the to-be-checked text is determined according to the target vocabulary. Redetermining the weight corresponding to each keyword in the to-be-checked text according to the target vocabulary can enhance the adaptability of emerging vocabularies, and thus improve the accuracy of the duplicate checking result.

[0061] It should be noted that in the case where each keyword in the to-be-checked text corresponds to a preset weight, if it is determined that the target vocabulary exists, the weight corresponding to each keyword in the to-be-checked text is adjusted according to the target vocabulary.

[0062] It should be further noted that in the case where each keyword in the to-be-checked text corresponds to a preset weight, if it is determined that the target vocabulary does not exist, the weight corresponding to each keyword in the to-be-checked text remains unchanged.

[0063] In step S102, based on each keyword in the to-be-checked text and the weight corresponding to the keyword, a first format similarity, a first sentence similarity and a first topic similarity between the to-be-checked text and each retrieval text in a retrieval text set are obtained; wherein the retrieval text set belongs to the text retrieval library.

[0064] In step S102, based on the respective keywords in the to-be-checked duplicate text determined in step S102 and the weights corresponding to the keywords, the similarity between the to-be-checked duplicate text and each search text in the search text set is calculated in three aspects of format, sentence and theme, so that the article understanding is more comprehensive, and the duplicate checking result accuracy is improved. The first format similarity is used to represent the similarity between the to-be-checked duplicate text and the search text in the text format and the text style, and the similarity can be calculated based on a semantic network graph model. The first sentence similarity is used to represent the similarity between the to-be-checked duplicate text and the search text in the sentence, and the similarity can be calculated based on a longest common subsequence (LCS) and a Chinese language model (CLM) sequence (N-gram) overlap method of continuous N words or characters, where N is an integer greater than 1. The first theme similarity is used to represent the similarity between the to-be-checked duplicate text and the search text in the theme distribution, and the similarity can be calculated based on a latent Dirichlet allocation (LDA) model.

[0065] In step S103, the first format similarity, the first sentence similarity and the first theme similarity are weighted and summed to obtain the first similarity between the to-be-checked duplicate text and each search text in the search text set.

[0066] In step S103, the first format similarity, the first sentence similarity and the first theme similarity are weighted and summed to obtain the first similarity between the to-be-checked duplicate text and each search text in the search text set.

[0067] In step S104, the first similarity is used to screen the duplicate checking result from the search text set.

[0068] In an optional embodiment, step S104 includes: sorting the plurality of search texts in the search text set in descending order of the first similarity, and selecting the first n search texts as the duplicate checking result according to the arrangement order, where n is an integer greater than or equal to 1.

[0069] In the article duplication checking process in the embodiment of the application, if there is a target vocabulary in the text to be checked, that is, there is a vocabulary in the text to be checked, the frequency of which is greater than or equal to a first threshold and the frequency of which in the text retrieval library is less than or equal to a second threshold, it is considered that there is a new emerging vocabulary that is relatively novel relative to the vocabulary in the text retrieval library. Then, the weight corresponding to each keyword in the text to be checked is determined according to the target vocabulary, and the first similarity between the text to be checked and each retrieval text in the retrieval text set is calculated according to each keyword in the text to be checked and the weight corresponding to the keyword. The duplication checking result is obtained according to the first similarity. Compared with the prior art, the embodiment of the application considers the influence of the emerging vocabulary, enhances the adaptability of the emerging vocabulary, and can improve the accuracy of the duplication checking result.

[0070] Moreover, in the process of calculating the first similarity between the text to be checked and each retrieval text in the retrieval text set, the first format similarity, the first sentence similarity and the first theme similarity between the text to be checked and each retrieval text in the retrieval text set are obtained based on each keyword in the text to be checked and the weight corresponding to the keyword, and then the first format similarity, the first sentence similarity and the first theme similarity are weighted and summed to obtain the first similarity between the text to be checked and each retrieval text in the retrieval text set. The embodiment of the application considers the format similarity, the sentence similarity and the theme similarity, so that the article understanding is more comprehensive, and the accuracy of the duplication checking result can be improved.

[0071] In some embodiments of the application, the weight corresponding to each keyword in the text to be checked is determined according to the target vocabulary, including:

[0072] According to the target vocabulary, a first retrieval result is obtained by searching in the network article library.

[0073] In the case where the first retrieval result indicates that there is no target article in the network article library, a target keyword is extracted from the keywords of the text to be checked, wherein the keywords of the target article include the target vocabulary, the distance between the target keyword and the target vocabulary in the text to be checked is less than a third threshold, and the frequency of the target keyword is greater than or equal to a fourth threshold.

[0074] Based on the pre-set weight corresponding to each keyword in the text to be checked, the weight corresponding to each keyword in the text to be checked is determined by increasing the weight corresponding to the target keyword.

[0075] In the embodiment of the present application, after determining the target vocabulary, the target vocabulary is searched in the reliable network article library, if the search result indicates that the target article (i.e. the article including the target vocabulary) does not exist in the network article library, the target keyword is extracted from the keywords of the text to be checked, wherein the target keyword is the keyword frequently appearing near the target vocabulary, specifically, when the distance between the target keyword and the target vocabulary in the text to be checked is less than a third threshold value, and the frequency of the target keyword appearing near the target vocabulary is greater than or equal to a fourth threshold value, it is determined that the target keyword is the keyword frequently appearing near the target vocabulary, wherein the distance between two vocabularies can be calculated by the word index difference method or the standardized distance method. After extracting the target keyword, the weight corresponding to the target keyword is increased based on the pre-set weight corresponding to each keyword in the text to be checked, and then the weight corresponding to each keyword in the text to be checked is adjusted and determined.

[0076] In some embodiments of the present application, after the target vocabulary is searched in the network article library to obtain the first search result, the method further comprises:

[0077] In the case that the first search result indicates that the target article exists in the network article library, the target article is added to the text retrieval library to obtain an updated text retrieval library;

[0078] In the case that the first search result indicates that the target article does not exist in the network article library and the target keyword is extracted from the text to be checked, the target keyword is searched in the network article library and the search result is added to the text retrieval library to obtain an updated text retrieval library.

[0079] In the embodiment of the present application, after the target vocabulary is searched in the network article library to obtain the first search result, if the first search result indicates that the target article exists in the reliable network article library, the target article is added to the temporary retrieval library of the text retrieval library to obtain an updated text retrieval library, it should be noted that the temporary retrieval library only plays a role in the current text to be checked.

[0080] If the first search result indicates that the target article does not exist in the reliable network article library, the target keyword is extracted from the keywords of the text to be checked, and then the target keyword is searched in the reliable network article library, and the searched article is added to the temporary retrieval library of the text retrieval library to obtain an updated text retrieval library.

[0081] In some embodiments of the present application, the above article duplication checking method further comprises:

[0082] obtaining feature information of the to-be-duplicated text, the feature information comprising at least one of content feature, theme feature and author feature;

[0083] determining second similarity between the to-be-duplicated text and each text in the text retrieval library according to the feature information;

[0084] obtaining the set of retrieval texts from the text retrieval library according to the second similarity.

[0085] In the embodiment of the application, before step S102, the text retrieval library is filtered by a pre-trained multi-dimensional feature model to obtain the set of retrieval texts used in step S102, wherein the multi-dimensional feature model comprises a plurality of dimensional feature extraction modules and a first similarity model used to calculate similarity. The specific operation is as follows:

[0086] Firstly, the feature information of the to-be-duplicated text is obtained by the feature extraction module, wherein the feature information comprises at least one of content feature, theme feature and author feature; then, the feature information of the to-be-duplicated text and the feature information of each text in the text retrieval library are input into the first similarity model to obtain the second similarity between the to-be-duplicated text and each text in the text retrieval library, wherein the feature information of each text in the text retrieval library is pre-obtained and stored in a database; finally, the texts in the text retrieval library are filtered according to the second similarity to obtain a plurality of retrieval texts, thereby generating the set of retrieval texts. In the process of obtaining the duplication result, the embodiment of the application first performs a first round of filtering on the texts in the text retrieval library based on the feature information of the to-be-duplicated text in multiple dimensions to obtain the set of retrieval texts, and then performs a second round of filtering on the set of retrieval texts based on steps S102-S104 to obtain the duplication result. The embodiment of the application is filtered twice and calculated in multiple dimensions, so that the accuracy of the final duplication result is higher.

[0087] It should be noted that the step of screening the text retrieval library by the pre-trained multi-dimensional feature model can be synchronized with the process of judging whether the to-be-checked duplicate text exists the target word, the process of updating the text retrieval library, and the like. Specifically, after the feature information of the to-be-checked duplicate text is acquired by the feature extraction module, it is judged whether the to-be-checked duplicate text exists the target word based on the feature information. If the to-be-checked duplicate text does not exist the target word, the second similarity between the to-be-checked duplicate text and each text in the current text retrieval library is determined according to the feature information. If the to-be-checked duplicate text exists the target word, and the text retrieval library is updated based on the target word, the second similarity between the to-be-checked duplicate text and each text in the updated text retrieval library is determined according to the feature information. Finally, the retrieval text set is screened from the corresponding text retrieval library according to the second similarity.

[0088] In some embodiments of the present application, the feature information of the to-be-checked duplicate text is acquired by at least one of the following:

[0089] The content features of the to-be-checked duplicate text are extracted by using the natural language processing technology, and the content features include at least one of the following: keywords, text themes and text conclusions.

[0090] The research theme and research direction of the to-be-checked duplicate text are identified to obtain the theme features of the to-be-checked duplicate text.

[0091] The author features of the to-be-checked duplicate text are determined according to the author background information of the to-be-checked duplicate text.

[0092] In the embodiments of the present application, the content features include at least one of the following: keywords, text themes and text conclusions. The content features can be extracted by using the natural language processing technology. For example, the TF-IDF (Term Frequency-Inverse Document Frequency) model is used to identify the keywords in the to-be-checked duplicate text that can best represent the article theme, and the TF-IDF model is used to extract the text theme and text conclusion in the abstract part of the to-be-checked duplicate text.

[0093] For the extraction of the subject feature, a subject modeling method can be used, for example, an LDA subject modeling method, to identify the main research subject and direction of the to-be-duplicated text. For the article with the self-labeled research subject and direction, the main research subject and direction can be directly obtained. In addition, the to-be-duplicated text can be labeled according to the corresponding academic classification (for example, Institute of Electrical and Electronics Engineers (IEEE) and Association for Computing Machinery (ACM)).

[0094] It should be noted that, for the extraction of the content feature and the subject feature, in general, to reduce the use of computing resources, only the abstract part of the article can be extracted, and the content feature and the subject feature can be extracted by means of keyword analysis, subject modeling, etc. on the abstract part.

[0095] The author feature is used to represent the influence of the author. Specifically, the author feature can be set as a single numerical value or a multi-dimensional vector based on requirements. When the author feature is a single numerical value, the comprehensive influence is obtained by weighted summation of the author background and the data in the co-author network of the author through corresponding weights. When the author feature is a multi-dimensional vector, the author background and the data in the co-author network of the author are standardized to form a corresponding vector representation. For the extraction of the author feature, specifically, based on the background information of the author, the previously published papers and citation data of the author are collected to analyze the research field and influence of the author, for example, the awards obtained by the author, the number of citations of the published papers, the level of the published journals, etc. are obtained, and the influence of the author is obtained by weighting through setting coefficients. Based on the co-author network of the author, the co-author relationship of the author is analyzed to understand the position of the author in the academic network. For example, the relationship between the author and the co-authors is obtained by constructing the knowledge graph corresponding to the author, and the influence of the researcher can be improved as the influence of the co-authors is improved.

[0096] In some embodiments of the present application, the feature information includes content features, subject features and author features;

[0097] The second similarity between the to-be-duplicated text and each text in the text retrieval library is determined according to the feature information, including:

[0098] The content feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated according to the content feature;

[0099] The subject feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated according to the subject feature;

[0100] According to the author feature, an author feature similarity between the to-be-duplicated text and each text in the text retrieval library is calculated;

[0101] The content feature similarity, the theme feature similarity and the author feature similarity are weighted and summed to obtain a second similarity between the to-be-duplicated text and each text.

[0102] In the embodiment of the application, after the content feature, the theme feature and the author feature of the to-be-duplicated text are obtained, the feature information is vectorized for convenience of calculation and comparison, and is converted into a numerical form, for example, the feature information of the to-be-duplicated text is converted into a feature vector to represent its attributes in multiple dimensions.

[0103] The feature vector of the to-be-duplicated text and the feature vector of each text in the text retrieval library obtained in advance are input into a first similarity model, the first similarity model is based on a similarity algorithm, for example, a cosine similarity algorithm, a Euclidean distance algorithm, etc., and outputs the content feature similarity, the theme feature similarity and the author feature similarity between the to-be-duplicated text and each text in the text retrieval library, wherein the content feature similarity is calculated according to the content feature of the to-be-duplicated text and the content feature of each text in the text retrieval library, the theme feature similarity is calculated according to the theme feature of the to-be-duplicated text and the theme feature of each text in the text retrieval library, and the author feature similarity is calculated according to the author feature of the to-be-duplicated text and the author feature of each text in the text retrieval library.

[0104] Finally, different weights are given to the content feature similarity, the theme feature similarity and the author feature similarity. Generally, the weights of the content feature and the theme feature are higher, and the two weights can be considered equal, for example, the weights of the two are both allocated as 0.475, at this time, the weight of the author feature is lower, which can be 0.05. However, as the influence of the author increases, it is indicated that the author has strong influence in a certain subdivision field, at this time, the weight of the author feature can be appropriately increased, but generally will not exceed 0.1, but the application does not limit this, and the actual demand can be determined when applied. Based on the weights of the content feature similarity, the theme feature similarity and the author feature similarity, the content feature similarity, the theme feature similarity and the author feature similarity are weighted and summed to obtain a second similarity between the to-be-duplicated text and each text.

[0105] In some embodiments of the present application, the first format similarity between the to-be-duplicated text and each search text in the search text set is obtained based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, and the first format similarity comprises:

[0106] The to-be-duplicated text is divided into a plurality of first ranges according to the title and sub-title in the to-be-duplicated text, wherein one first range corresponds to one title or sub-title;

[0107] A semantic network graph model corresponding to the to-be-duplicated text is constructed based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, wherein one first range corresponds to one semantic network graph model, the node in the semantic network graph model is a text unit, the weight of the node corresponds to the weight of the keyword in the corresponding text unit, and the edge in the semantic network graph model is used to represent the relationship between the nodes.

[0108] For each first range, the second format similarity between the to-be-duplicated text and each search text in the search text set is calculated according to the semantic network graph model.

[0109] The first format similarity between the to-be-duplicated text and each search text in the search text set is obtained according to the second format similarity.

[0110] In the embodiments of the present application, firstly, the use of the title and sub-title of the article is analyzed, so that the article is divided into a plurality of first ranges, and specifically, the to-be-duplicated text is divided into a plurality of first ranges according to the title and sub-title in the to-be-duplicated text, wherein one first range corresponds to one title or sub-title, for example, the to-be-duplicated text is divided into: paper abstract, main text, reference, etc. according to the sub-title, wherein the main text can be further divided into: introduction, research background, method, result, discussion, etc.

[0111] Then, for each first range, a semantic network graph model is constructed in turn, in which the nodes represent the basic units in the text, such as words, phrases, sentences or paragraphs. The edges represent the relationship between the nodes, such as the containing relationship, the reference relationship or the logical connection. Specifically, the nodes include: word nodes representing key words or important words in the article, sentence nodes, each sentence as a node, which can capture the relationship between sentences, paragraph nodes, each paragraph as a node, which is suitable for macro-structure analysis, and the edges include: containing edges representing the words or clauses contained in a sentence or paragraph, reference edges representing that a sentence refers to the content of other sentences, and similarity edges representing the similarity between two sentences or paragraphs based on content similarity calculation. Further, each node has a corresponding weight value, which corresponds to the weight of the key word included in the node. For example, the weight of the word node is determined according to the weight of the corresponding key word, the weight of the sentence node is determined by the weight of the key word contained in the corresponding sentence, and the weight of the paragraph node is determined by the weight of the sentence contained in the corresponding paragraph.

[0112] Secondly, for each first range, the semantic network graph model of the to-be-duplicated text and the semantic network graph model of the search text are compared by using graph isomorphism algorithm, least common ancestor, random walk and other algorithms to capture the deep structure and logical relationship of the text, and the second format similarity between the to-be-duplicated text and each search text in the search text set is obtained. For example, the graph isomorphism algorithm is used to determine whether the structures of two graphs are similar, that is, whether there is a mapping such that the corresponding nodes and edges have the same structure. It can be used to determine the similarity of two articles in logic and structure. The least common ancestor finds the common structural features (such as common nodes) in the two articles, calculates the least common ancestor, and helps to understand their logical similarities.

[0113] Finally, according to the second format similarity, the first format similarity between the to-be-duplicated text and the search text set is obtained. Specifically, there are two methods to determine it. The first method is to set a weight value for each first range, and to weight and sum the second format similarities corresponding to all or part of the first ranges based on the weight values corresponding to the first ranges, to obtain the first format similarity between the to-be-duplicated text and each search text in the search text set. The second method is to select the highest second format similarity as the first format similarity between the to-be-duplicated text and each search text in the search text set from the second format similarities corresponding to all the first ranges.

[0114] In some embodiments of the present application, the first theme similarity between the to-be-duplicated text and each search text in the search text set is obtained based on the weights of the key words in the to-be-duplicated text and the corresponding weights of the key words.

[0115] inputting the to-be-duplicated text into a pre-trained latent Dirichlet allocation (LDA) model based on each keyword in the to-be-duplicated text and the weight corresponding to the keyword, to obtain a topic distribution vector of the to-be-duplicated text output by the LDA model, wherein the weight value of a keyword in the topic distribution vector is greater than a fourth threshold value;

[0116] calculating a first topic similarity between the to-be-duplicated text and each search text in the search text set according to the topic distribution vector.

[0117] In the embodiment of the application, based on the pre-trained LDA model, a topic distribution vector of the to-be-duplicated text is extracted to represent the probability of the to-be-duplicated text on each topic, wherein the topic distribution vector includes multiple topics, topic probabilities, and multiple keywords corresponding to each topic, and the weight of the keyword can directly affect the contribution value of the keyword in the topic distribution vector. For example, the higher the weight of the keyword, the more significant the keyword in the text, and therefore the LDA model is more likely to associate it to a specific topic when generating the topic distribution vector, thus affecting the topic distribution vector.

[0118] Then, based on the cosine similarity algorithm, the first topic similarity between the to-be-duplicated text and each search text in the search text set is calculated according to the topic distribution vector of the to-be-duplicated text and the topic distribution vector of each search text in the search text set.

[0119] The training process of the LDA model is described as follows:

[0120] First, the training data is pre-collected and pre-processed, and then the LDA model is constructed. Then, the parameters are selected, the number of topics is determined, and the LDA model is trained using Python libraries (Gensim) and the like. It should be noted that the number of topics may be different for different research fields.

[0121] In some embodiments of the application, the first sentence similarity between the to-be-duplicated text and each search text in the search text set is obtained, including:

[0122] According to the title and sub-title in the to-be-duplicated article, the to-be-duplicated article is divided into multiple second ranges, wherein one second range corresponds to one title or sub-title.

[0123] For each second range, the second sentence similarity between the to-be-duplicated text and each search text in the search text set is obtained based on the longest common subsequence (LCS) algorithm.

[0124] According to the second sentence similarity, a first sentence similarity between the to-be-checked text and each of the search texts in the search text set is calculated.

[0125] In the embodiment of the present application, a specific method for calculating the first sentence similarity between the to-be-checked text and each of the search texts in the search text set according to the LCS algorithm is illustrated.

[0126] The LCS algorithm is mainly used to measure the similarity between two sequences, and the longest common subsequence refers to the maximum subsequence of characters (or elements) that appear in order but not necessarily continuously in two sequences. For example, for sequences A="ABCBDAB" and B="BDCAB", their LCS is BCAB, and the length is 4. The LCS has two characteristics, namely, order consistency and maximality. The order consistency means that the order of characters in the LCS must be the same as that in the original sequence, but it does not require continuity. The maximality means that among all possible common subsequences, the length of the LCS is the maximum. The calculation of the LCS is as follows.

[0127] First, a dynamic programming table is constructed, and a two-dimensional array dp is created, where dp[i][j] represents the LCS length of the first i characters of sequence A and the first j characters of sequence B. dp[0][*] and dp[*][0] are both initialized to 0, indicating that the LCS length of any number of characters * and an empty sequence is 0. For each pair of characters, if A[i-1]==B[j-1], A[i-1] represents the index starting from 0, indicating the i-th character of sequence A, and B[j-1] represents the i-th character of sequence B, indicating that a common character is found in the first i characters of sequence A and the first j characters of sequence B, and the common character contributes 1, then dp[i][j]=dp[i-1][j-1]+. If A[i-1]!=B[j-1],!= means not equal, then the maximum length of the LCS of the previous characters is taken, dp[i][j]=max(dp[i-1][j],dp[i][j-1]).

[0128] The maximum value of dp[i][j] is obtained, and the LCS of the to-be-checked text and the search text is calculated, so as to evaluate the sentence similarity between them. For example, for the text summary part, the calculated LCS length of the two texts is a, the total length (i.e., the total number of characters) of the A text summary part is b, and the total length of the B text summary part is c. At this time, the sentence similarity between the A text and the B text is For example, the length of the A text summary is 100, the length of the B text summary is 150, and the calculated LCS length of the two texts is 50. At this time, the sentence similarity between the A text and the B text is In addition, the time complexity of the LCS algorithm is 0(mn), where m and n are the lengths of the two sequences, respectively. The space complexity is also 0(mn). When dealing with longer texts, this can consume a lot of time and space, so some optimization methods can be considered in practical applications, such as only saving the current and previous line of state.

[0129] In the computing process of the embodiment of the application, first, according to the title and sub-title in the article to be checked, the article to be checked is divided into a plurality of second ranges, for example, according to the title and sub-title of the article, the article to be checked is divided into a plurality of second ranges such as "paper abstract", "text introduction", "text research background" and the like.

[0130] Then, based on the above-mentioned LCS algorithm, for each second range, the second sentence similarity between the text to be checked and each search text in the search text set is calculated;

[0131] Finally, based on the weight corresponding to each second range, the second sentence similarity of each second range is weighted and summed to obtain the first sentence similarity between the text to be checked and each search text in the search text set.

[0132] Before applying the article duplication checking method provided by the embodiment of the application, a multi-dimensional feature model and a second similarity model need to be trained in advance, wherein in the embodiment of the application, the multi-dimensional feature model is used to calculate the first similarity between the text to be checked and the text in the text search library according to the feature information of the text to be checked and the feature information of the text in the text search library, and the second similarity model is used to calculate the first format similarity, the first sentence similarity and the first theme similarity between the text to be checked and the search text in the search text set. The training process of the multi-dimensional feature model is described in detail in Embodiment One, and the construction process of the second similarity model is described in detail in Embodiment Two:

[0133] Embodiment One: As shown in the figure, the training process of the multi-dimensional feature model is as follows: Figure 2

[0134] Step S201, collect training data, which can be obtained from the following channels: obtain research papers in related fields from academic databases, view the personal homepage of researchers and the articles published by them on online academic social platforms, and obtain research results in a specific field from the paper set of an institution or a conference.

[0135] Step S202, use the feature module to extract the feature information of the training data, and the feature information includes content features, theme features and author features. The specific extraction method is the same as the method of obtaining the feature information of the text to be checked in the embodiment of the application, and will not be described here.

[0136] ​Step S203, the feature information is vectorized to obtain a feature vector.

[0137] Step S204, a second similarity of the text data in the training data is calculated according to the feature vector and using a similarity calculation method (for example, cosine similarity, Euclidean distance, etc.), the training data is trained according to the corresponding feature vector and the second similarity, and a first similarity model is constructed.

[0138] Step S205, the established first similarity model is evaluated to ensure its effectiveness. Through cross-validation, the training data is divided into a training set and a test set, the accuracy of the model is verified, and the actual effect of the model is understood through user feedback, for example: let some scholars use the model, collect their feedback, and understand the actual effect of the model.

[0139] Finally, a multi-dimensional feature model is generated based on the feature module and the first similarity model.

[0140] Embodiment two: as shown in the figure, the construction process of the second similarity model is as follows: Figure 3

[0141] Step S301, training text data is collected from academic databases, online academic social platforms, institutional or conference paper sets, etc., and the training text data is preprocessed, the preprocessing method including: removing useless characters, removing punctuation, filtering stop words, etc., so as to convert the original text data into a cleaner and more structured format.

[0142] Step S302, the first format similarity between each two text data in the training text data is calculated, and the specific calculation method is the same as the first format similarity calculation method between the to-be-checked text and the search text in the embodiment of the application, which will not be repeated here.

[0143] Step S303, the first sentence similarity between each two text data in the training text data is calculated, and the specific calculation method is the same as the first sentence similarity calculation method between the to-be-checked text and the search text in the embodiment of the application, which will not be repeated here.

[0144] Step S304, the first theme similarity between each two text data in the training text data is calculated, and the specific calculation method is the same as the first theme similarity calculation method between the to-be-checked text and the search text in the embodiment of the application, which will not be repeated here.

[0145] Step S305, the first format similarity, the first sentence similarity and the first theme similarity are weighted and summed to obtain the first similarity between the text data.

[0146] ​Based on the steps S301-S305, the model construction is performed to obtain the second similarity model.

[0147] Embodiment three: As shown in the figure, the article duplicate checking method of the embodiment of the application is used to perform the article duplicate checking process as follows: Figure 4

[0148] Step S401, obtaining the to-be-checked text.

[0149] Step S402, performing text preprocessing on the to-be-checked text, and the preprocessing method includes removing useless characters, removing punctuation, filtering stop words, etc.

[0150] Step S403, performing feature extraction on the to-be-checked text according to the feature modules in the multi-dimensional feature model to obtain feature information, and the feature information includes content features, theme features and author features.

[0151] Step S404, judging whether there is a target word in the to-be-checked text based on the content features, if yes, turning to step S405, and if no, turning to step S406.

[0152] Step S405, updating the text retrieval library according to the target word and / or adjusting the weight of the keyword in the second similarity model in the to-be-checked text according to the target word, and the specific steps include:

[0153] performing retrieval on the network article library according to the target word to obtain a first retrieval result;

[0154] in the case where the first retrieval result indicates that there is a target article in the network article library, adding the target article to the text retrieval library to obtain an updated text retrieval library, wherein the keywords of the target article include the target word;

[0155] in the case where the first retrieval result indicates that there is no target article in the network article library, extracting a target keyword from the keywords of the to-be-checked text, wherein the distance between the target keyword and the target word in the to-be-checked text is less than a third threshold value and the appearance frequency of the target keyword is greater than or equal to a fourth threshold value;

[0156] based on the weights of the keywords in the to-be-checked text corresponding to the weights preset in advance, determining the weights of the keywords in the to-be-checked text by increasing the weight corresponding to the target keyword, and / or performing retrieval on the network article library according to the target keyword and adding the retrieval result to the text retrieval library to obtain an updated text retrieval library.

[0157] ​Step S406, according to the first similarity model in the multi-dimension feature model, first, based on the feature information, the content feature similarity, the theme feature similarity and the author feature similarity between the to-be-duplicated text and the texts in the text retrieval library are calculated, and then the content feature similarity, the theme feature similarity and the author feature similarity are weighted and summed to obtain the second similarity between the to-be-duplicated text and the texts in the text retrieval library. It should be noted that if step S405 is transferred to step S406, the text retrieval library here is the updated text retrieval library.

[0158] Step S407, according to the second similarity, the texts in the text retrieval library are screened to obtain a set of retrieval texts.

[0159] Step S408, according to the second similarity model, first, the first format similarity, the first sentence similarity and the first theme similarity between the to-be-duplicated text and the retrieval texts in the set of retrieval texts are calculated, and then the first format similarity, the first sentence similarity and the first theme similarity are weighted and summed to obtain the first similarity between the to-be-duplicated text and each of the retrieval texts in the set of retrieval texts. It should be noted that if step S405 is transferred to step S406, the weight of the keyword of the to-be-duplicated text applied in step S408 is the weight adjusted according to the target vocabulary, that is, the keyword weight obtained in step S405.

[0160] Step S409, according to the first similarity, the retrieval texts in the set of retrieval texts are screened to obtain a duplication result.

[0161] As shown in Figure 5 , the embodiment of the application also provides an article duplication checking device, which comprises:

[0162] The first determination module 501 is configured to, in the case where the target vocabulary exists, determine the weight corresponding to each keyword in the to-be-duplicated text according to the target vocabulary; wherein the target vocabulary is a vocabulary with a frequency greater than or equal to a first threshold value in the to-be-duplicated text and a frequency less than or equal to a second threshold value in the text retrieval library, and the first threshold value is greater than the second threshold value;

[0163] The first calculation module 502 is configured to obtain the first format similarity, the first sentence similarity and the first theme similarity between the to-be-duplicated text and each of the retrieval texts in the set of retrieval texts based on each of the keywords in the to-be-duplicated text and the weight corresponding to the keyword; wherein the set of retrieval texts belongs to the text retrieval library;

[0164] The second calculation module 503 is configured to weight and sum the first format similarity, the first sentence similarity and the first theme similarity to obtain the first similarity between the to-be-duplicated text and each of the retrieval texts in the set of retrieval texts.

[0165] The first screening module 504 is configured to screen the search text set according to the first similarity to obtain a duplicate checking result.

[0166] Optionally, the first determining module 501 comprises:

[0167] The first searching sub-module is configured to search the network article library according to the target vocabulary to obtain a first searching result.

[0168] The first extracting sub-module is configured to extract a target keyword from keywords of the text to be checked in a case where the first searching result indicates that the target article does not exist in the network article library, wherein the keywords of the target article comprise the target vocabulary, the distance between the target keyword and the target vocabulary in the text to be checked is less than a third threshold, and the appearance frequency of the target keyword is greater than or equal to a fourth threshold.

[0169] The first adjusting sub-module is configured to determine the weight of each keyword in the text to be checked by increasing the weight of the target keyword based on the weight of each keyword in the text to be checked.

[0170] Optionally, the apparatus further comprises:

[0171] The first updating module is configured to add the target article to the text search library in a case where the first searching result indicates that the target article exists in the network article library to obtain an updated text search library.

[0172] The second updating module is configured to search the network article library according to the target keyword and add the searching result to the text search library to obtain an updated text search library in a case where the first searching result indicates that the target article does not exist in the network article library and the target keyword is extracted from the text to be checked.

[0173] Optionally, the apparatus further comprises:

[0174] The first obtaining module is configured to obtain feature information of the text to be checked, wherein the feature information comprises at least one of the following: content feature, theme feature and author feature.

[0175] The third calculating module is configured to determine a second similarity between the text to be checked and each text in the text search library according to the feature information.

[0176] The second screening module is configured to screen the text search library according to the second similarity to obtain the search text set.

[0177] Optionally, the first obtaining module comprises at least one of the following:

[0178] The first obtaining sub-module is configured to extract content features of the to-be-duplicated text by using a natural language processing technique, wherein the content features comprise at least one of the following: keywords, a text theme and a text conclusion.

[0179] The first obtaining sub-module is configured to identify a research theme and a research direction of the to-be-duplicated text, and obtain theme features of the to-be-duplicated text.

[0180] The first obtaining sub-module is configured to determine author features of the to-be-duplicated text according to author background information of the to-be-duplicated text.

[0181] Optionally, the feature information in the first obtaining module comprises content features, theme features and author features.

[0182] The third calculating module comprises:

[0183] The first calculating sub-module is configured to calculate content feature similarities between the to-be-duplicated text and each text in the text retrieval library according to the content features.

[0184] The second calculating sub-module is configured to calculate theme feature similarities between the to-be-duplicated text and each text in the text retrieval library according to the theme features.

[0185] The third calculating sub-module is configured to calculate author feature similarities between the to-be-duplicated text and each text in the text retrieval library according to the author features.

[0186] The fourth calculating sub-module is configured to perform weighted summation on the content feature similarities, the theme feature similarities and the author feature similarities, to obtain second similarities between the to-be-duplicated text and each text.

[0187] Optionally, the first calculating module 502 comprises:

[0188] The first dividing sub-module is configured to divide the to-be-duplicated text into a plurality of first ranges according to titles and sub-titles in the to-be-duplicated text, wherein one first range corresponds to one title or sub-title.

[0189] A first construction submodule is configured to construct a semantic network graph model corresponding to the text to be checked for duplicate content based on each of the keywords in the text to be checked for duplicate content and the weights corresponding to the keywords, wherein one of the first ranges corresponds to one of the semantic network graph models, the nodes in the semantic network graph model are text units, the weights of the nodes correspond to the weights of the keywords in the corresponding text units; and the edges in the semantic network graph model are used to represent the relationships between the nodes;

[0190] a fifth calculation submodule, configured to calculate, for each of the first ranges, a second format similarity between the text to be checked for duplicates and each search text in the search text set according to the semantic network graph model;

[0191] The sixth calculation submodule is configured to obtain, based on the second format similarity, a first format similarity between the text to be checked for duplicates and each search text in the search text set.

[0192] Optionally, the first calculation module 502 includes:

[0193] a seventh calculation submodule, configured to input the text to be checked for duplicates into a pre-trained latent Dirichlet allocation (LDA) model based on each of the keywords in the text to be checked for duplicates and the weights corresponding to the keywords, and obtain a topic distribution vector of the text to be checked for duplicates output by the LDA model, wherein the weight values ​​of the keywords in the topic distribution vector are greater than a fourth threshold;

[0194] An eighth calculation submodule is configured to calculate a first topic similarity between the text to be checked for duplicates and each search text in the search text set based on the topic distribution vector.

[0195] It should be noted that the embodiment of the device is a device corresponding to the embodiment of the above method, and all implementation methods in the embodiment of the above method are applicable to the embodiment of the device and can achieve the same technical effect.

[0196] An embodiment of the present invention also provides a network device, comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, it implements the article duplication checking method as described in any of the above items and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0197] The article duplicate checking method is also provided by the embodiment of the application. The article duplicate checking method comprises the steps of: obtaining an article; determining a duplicate article of the article; and determining a duplicate article of the article.

[0198] The article duplicate checking method is also provided by the embodiment of the application. The article duplicate checking method comprises the steps of: obtaining an article; determining a duplicate article of the article; and determining a duplicate article of the article.

[0199] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0200] Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the spirit and scope of the present application. Thus, it is intended that the present application cover the modifications and changes as long as they come within the scope of the appended claims and their equivalents.

Claims

1. A method for checking for duplicate content in an article, characterized in that: include: When it is determined that a target word exists, determining the weight corresponding to each keyword in the text to be checked for duplicates based on the target word; wherein the target word is a word whose appearance frequency in the text to be checked for duplicates is greater than or equal to a first threshold and whose appearance frequency in the text retrieval library is less than or equal to a second threshold, and the first threshold is greater than the second threshold; Based on each of the keywords in the text to be checked for duplicates and the weights corresponding to the keywords, obtaining a first format similarity, a first sentence similarity, and a first topic similarity between the text to be checked for duplicates and each search text in a search text set; wherein the search text set belongs to the text search library; Performing a weighted summation on the first format similarity, the first sentence similarity, and the first topic similarity to obtain first similarities between the text to be checked for duplicates and each search text in the search text set; According to the first similarity, duplicate checking results are obtained by screening from the search text set.

2. The article duplication checking method according to claim 1, characterized in that: Determining the weights corresponding to the keywords in the text to be checked for duplicates based on the target vocabulary includes: Searching the online article library according to the target vocabulary to obtain a first search result; In the case where the first search result indicates that the target article does not exist in the online article library, extracting a target keyword from the keywords of the text to be checked for duplicate content, wherein the keywords of the target article include the target vocabulary, the distance between the target keyword and the target vocabulary in the text to be checked for duplicate content is less than a third threshold, and the occurrence frequency of the target keyword is greater than or equal to a fourth threshold; Based on the preset weights corresponding to the respective keywords in the text to be checked for duplicates, the weights corresponding to the respective keywords in the text to be checked for duplicates are determined by adding the weights corresponding to the target keywords.

3. The article duplication checking method according to claim 2, characterized in that: After searching the online article library according to the target vocabulary and obtaining the first search result, the method further includes: If the first search result indicates that the target article exists in the online article library, adding the target article to the text search library to obtain an updated text search library; When the first search result indicates that the target article does not exist in the online article library and the target keyword is extracted from the text to be checked for duplicates, a search is performed in the online article library based on the target keyword and the search results are added to the text search library to obtain an updated text search library.

4. The article duplication checking method according to claim 1, characterized in that: The method further comprises: Acquiring characteristic information of the text to be checked for duplicates, wherein the characteristic information includes at least one of the following: content characteristics, subject characteristics, and author characteristics; Determining, based on the feature information, a second similarity between the text to be checked for duplicates and each text in the text retrieval library; The search text set is obtained by screening from the text search library according to the second similarity.

5. The article duplication checking method according to claim 4, characterized in that: The step of obtaining characteristic information of the text to be checked for duplicates includes at least one of the following: Extracting content features of the text to be checked for duplicate content using natural language processing technology, wherein the content features include at least one of the following: keywords, text topics, and text conclusions; Identify the research subject and research direction of the text to be checked for duplicates, and obtain the subject features of the text to be checked for duplicates; The author characteristics of the text to be checked for duplicates are determined based on the author background information of the text to be checked for duplicates.

6. The article duplication checking method according to claim 4, characterized in that: The characteristic information includes content characteristics, subject characteristics and author characteristics; Determining the second similarity between the text to be checked for duplicates and each text in the text retrieval library based on the feature information includes: Calculating the content feature similarity between the text to be checked for duplicates and each text in the text retrieval library based on the content features; Calculating the similarity of the subject features between the text to be checked for duplicates and each text in the text retrieval library based on the subject features; Calculating the author feature similarity between the text to be checked for duplicates and each text in the text retrieval library based on the author feature; A weighted sum is performed on the content feature similarity, the subject feature similarity, and the author feature similarity to obtain a second similarity between the text to be checked for duplicates and each of the texts.

7. The article duplication checking method according to claim 1, characterized in that: The obtaining of the first format similarity between the text to be checked for duplicates and each search text in the search text set based on each keyword in the text to be checked for duplicates and the weight corresponding to the keyword includes: Dividing the text to be checked for duplicate content into a plurality of first ranges according to the titles and subtitles in the text to be checked for duplicate content, wherein each first range corresponds to a title or a subtitle; Based on each of the keywords in the text to be checked for duplicate content and the weights corresponding to the keywords, a semantic network graph model corresponding to the text to be checked for duplicate content is constructed, wherein one of the first ranges corresponds to one of the semantic network graph models, the nodes in the semantic network graph model are text units, the weights of the nodes correspond to the weights of the keywords in the corresponding text units; and the edges in the semantic network graph model are used to represent the relationships between the nodes; For each of the first ranges, calculating the second format similarity between the text to be checked for duplicates and each search text in the search text set according to the semantic network graph model; According to the second format similarity, a first format similarity between the text to be checked for duplicates and each search text in the search text set is obtained.

8. The article duplication checking method according to claim 1, characterized in that: The obtaining of the first topic similarity between the text to be checked for duplicates and each search text in the search text set based on each keyword in the text to be checked for duplicates and the weight corresponding to the keyword includes: Based on each of the keywords in the text to be checked for duplicate content and the weights corresponding to the keywords, inputting the text to be checked for duplicate content into a pre-trained latent Dirichlet allocation (LDA) model to obtain a topic distribution vector of the text to be checked for duplicate content output by the LDA model, wherein the weight value of the keyword in the topic distribution vector is greater than a fourth threshold; The first topic similarity between the text to be checked for duplicates and each search text in the search text set is calculated based on the topic distribution vector.

9. An article duplication checking device, characterized in that: include: A first determination module is configured to, upon determining the presence of a target word, determine, based on the target word, a weight corresponding to each keyword in the text to be checked for duplicates; wherein the target word is a word whose appearance frequency in the text to be checked for duplicates is greater than or equal to a first threshold and whose appearance frequency in the text retrieval library is less than or equal to a second threshold, and the first threshold is greater than the second threshold; a first calculation module for obtaining, based on each of the keywords in the text to be checked for duplicates and the weights corresponding to the keywords, a first format similarity, a first sentence similarity, and a first topic similarity between the text to be checked for duplicates and each search text in a search text set; wherein the search text set belongs to the text search library; a second calculation module, configured to perform a weighted summation of the first format similarity, the first sentence similarity, and the first topic similarity to obtain a first similarity between the text to be checked for duplicates and each search text in the search text set; The first screening module is used to screen the search text set to obtain duplicate checking results based on the first similarity.

10. A network device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the article duplication checking method as claimed in any one of claims 1 to 8.

11. A readable storage medium, characterized in that: include: The readable storage medium stores a program, and when the program is executed by the processor, the steps of the article duplication checking method as described in any one of claims 1 to 8 are implemented.

12. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the article duplication checking method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Document duplicate checking method and device and medium

    CN121211029A