A method and device for detecting paper plagiarism based on similarity
By cleaning and screening papers, combining digital fingerprints, word frequency vectors and syntactic similarity methods, the problem of difficult to take into account both detection efficiency and accuracy in the existing technology is solved, and efficient and accurate paper plagiarism detection is achieved.
Patent Information
- Application Number
- CN202411402234.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-09
AI Technical Summary
When the existing paper plagiarism check algorithm processes a large amount of text data, it is difficult to meet the detection accuracy and efficiency at the same time. The word frequency algorithm has high calculation cost, the cosine similarity algorithm is complex, and the SimHash algorithm has low accuracy, so it is impossible to accurately locate the plagiarism part.
A paper plagiarism detection method based on similarity is used. By obtaining the paper to be tested and cleaning, the historical paper database is screened, the digital fingerprint and word frequency vector of the target paper and the paper to be compared, and the plagiarism detection results are judged based on the fingerprint similarity, word frequency similarity and syntactic similarity.
The amount of data calculated in similarity is reduced, the detection efficiency is improved, and the calculation accuracy is ensured through the robustness of digital fingerprints and the accuracy of word frequency algorithms, and the calculation accuracy can be ensured while improving detection efficiency.
Smart Images

Figure CN119357694B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method and device for detecting paper plagiarism based on similarity. Background Art
[0002] With the development of the Internet and the convenient access to information, the problems of plagiarism and piracy have become increasingly prominent. In order to ensure the maintenance of academic atmosphere and research quality, an efficient and accurate paper plagiarism detection algorithm has become an essential tool. The paper plagiarism detection algorithm determines the similarity and duplication rate between multiple text files by analyzing and comparing them. Conventional paper plagiarism detection algorithms mainly include the word frequency algorithm, the cosine similarity algorithm, the SimHash algorithm, etc.
[0003] However, due to the long length of paper texts and the large database resources to be compared, the calculation cost of the word frequency algorithm increases and the efficiency deteriorates. The calculation process of the cosine similarity algorithm is relatively complex, and the efficiency will also be affected due to the increase in data volume. Although the SimHash algorithm has high efficiency in processing large-scale data sets, its accuracy is low and it cannot accurately locate the plagiarized parts. Therefore, the existing paper plagiarism detection algorithms cannot meet both the detection accuracy and the detection efficiency due to the need to process a large amount of data texts. Summary of the Invention
[0004] This application provides a method and device for detecting paper plagiarism based on similarity, which can reduce the amount of data to be processed in similarity calculation, and ensure the calculation accuracy while improving the detection efficiency.
[0005] In a first aspect, an embodiment of this application provides a method for detecting paper plagiarism based on similarity, including:
[0006] Obtain the paper to be detected and clean it to obtain the target paper;
[0007] Obtain the historical paper database and screen to obtain multiple papers to be compared;
[0008] Calculate the digital fingerprints and word frequency vectors of the target paper and each paper to be compared;
[0009] Calculate the fingerprint similarity between the target paper and each paper to be compared based on the digital fingerprints;
[0010] Calculate the word frequency similarity between the target paper and each paper to be compared based on the word frequency vectors;
[0011] Obtain the plagiarism detection result according to the fingerprint similarity and the word frequency similarity.
[0012] Further, the above-mentioned obtaining the paper to be detected and cleaning it to obtain the target paper includes:
[0013] Remove the header, footer, references, and predefined stop words from the paper to be detected;
[0014] Convert all uppercase letters in the paper to be detected to lowercase letters;
[0015] Lemmatize the English text of the paper to be detected to obtain the target paper.
[0016] Further, the above steps of obtaining the historical paper database and screening to obtain multiple papers to be compared include:
[0017] Extract each keyword of the target paper; use the inverted index method to screen the papers to be compared in the historical paper database; at least one keyword exists in the title or abstract of the papers to be compared.
[0018] Further, the above steps of calculating the digital fingerprints and term frequency vectors of the target paper and each paper to be compared include:
[0019] Segment the target paper or the paper to be compared to obtain multiple tokens;
[0020] Process each token with one or more hash functions to obtain digital fingerprints;
[0021] Classify and count each token to obtain multiple tokens of the same type and their corresponding frequencies;
[0022] Construct a term frequency vector based on each token of the same type and its corresponding frequency.
[0023] Further, the above step of processing each token with one or more hash functions to obtain digital fingerprints includes:
[0024] Put each token into different hash functions to map and obtain the corresponding hash values;
[0025] Combine the hash values corresponding to each different hash function into a hash vector;
[0026] Normalize the hash vector to obtain digital fingerprints.
[0027] Further, the above step of processing each token with one or more hash functions to obtain digital fingerprints includes:
[0028] Random step: Randomly arrange each token and put it into a hash function to map and obtain a hash value;
[0029] Repeat the random step multiple times to obtain multiple different hash values;
[0030] Normalize the hash vector composed of each hash value to obtain digital fingerprints.
[0031] Further, the method further includes:
[0032] Use a syntactic analyzer to obtain the syntactic trees of the target paper and each paper to be compared;
[0033] Convert the format of each syntactic tree to obtain each syntactic tree in the treebank format;
[0034] Extract the corresponding syntactic patterns from each syntactic tree in the treebank format;
[0035] Calculate the syntactic similarity between the target paper and each paper to be compared based on the syntactic patterns;
[0036] Obtain the plagiarism detection result according to the syntactic similarity, fingerprint similarity, and word frequency similarity.
[0037] Furthermore, the syntactic pattern is a sub-structure of the treebank, including noun phrases and verb phrases.
[0038] Furthermore, the above-mentioned calculation of the syntactic similarity between the target paper and each paper to be compared based on the syntactic patterns includes:
[0039] Detect the number of overlapping phrases between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared;
[0040] Calculate the edit distance between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared;
[0041] Multiply the reciprocal of the edit distance by the number of overlapping phrases to obtain the syntactic similarity.
[0042] Furthermore, the above-mentioned obtaining of the plagiarism detection result according to the fingerprint similarity and word frequency similarity includes:
[0043] Calculate the first similarity average value of the fingerprint similarity and word frequency similarity with the paper to be compared;
[0044] If the first similarity average value is greater than the first preset threshold, it is determined that there is a plagiarism problem with the target paper.
[0045] Furthermore, the above-mentioned obtaining of the plagiarism detection result according to the syntactic similarity, fingerprint similarity, and word frequency similarity includes:
[0046] Calculate the second similarity average value of the syntactic similarity, fingerprint similarity, and word frequency similarity;
[0047] If the syntactic similarity is greater than the second preset threshold and the first similarity average value is greater than the third preset threshold, or the second similarity average value is greater than the first preset threshold, it is determined that there is a plagiarism problem with the target paper;
[0048] Among them, the first preset threshold is less than the third preset threshold.
[0049] Furthermore, the method further includes:
[0050] Input the target paper and the paper to be compared into the trained first model respectively to obtain the corresponding fine-grained features;
[0051] Input the two obtained fine-grained features into the trained second model to obtain the feature similarity;
[0052] Determine whether the feature similarity is greater than the fingerprint similarity between the target paper and the paper to be compared;
[0053] If so, use the feature similarity as the fingerprint similarity.
[0054] Furthermore, the method further includes:
[0055] Obtain multiple training papers and input them into the word embedding model respectively to obtain the text vectors corresponding to the training papers;
[0056] Use the grid search technique to obtain the optimal convolutional kernel parameters and construct a convolutional deep learning model;
[0057] Input the text vectors of the training papers into the convolutional deep learning model in sequence, and at the same time use the regularization technique to inactivate the random neurons of the convolutional deep learning model to obtain the trained first model.
[0058] In a second aspect, an embodiment of the present application provides a paper plagiarism detection device based on similarity, including:
[0059] A cleaning module for obtaining the paper to be detected and cleaning it to obtain the target paper;
[0060] A screening module for obtaining the historical paper database and screening to obtain multiple papers to be compared;
[0061] A calculation module for calculating the digital fingerprints and word frequency vectors of the target paper and each paper to be compared;
[0062] A fingerprint module for calculating the fingerprint similarity between the target paper and each paper to be compared based on the digital fingerprints;
[0063] A word frequency module for calculating the word frequency similarity between the target paper and each paper to be compared based on the word frequency vectors;
[0064] A judgment module for obtaining the plagiarism detection result according to the fingerprint similarity and the word frequency similarity.
[0065] Furthermore, the above cleaning module includes:
[0066] A removal unit for removing the header, footer, references and preset stop words of the paper to be detected;
[0067] An alphabet conversion unit for converting all uppercase letters of the paper to be detected into lowercase letters;
[0068] A lemmatization unit for lemmatizing the English text of the paper to be detected to obtain a target paper.
[0069] Furthermore, a screening module is used to extract each keyword of the target paper and screen the papers to be compared in the historical paper database by using the inverted index method; at least one keyword exists in the title or abstract of the papers to be compared.
[0070] Furthermore, the above calculation module includes:
[0071] A segmentation unit for segmenting the target paper or the papers to be compared to obtain multiple tokens;
[0072] A hash unit for processing each token by using one or more hash functions to obtain digital fingerprints;
[0073] A statistics unit for classifying and counting each token to obtain multiple homogeneous tokens and corresponding frequencies;
[0074] A vector unit for constructing a term frequency vector according to each homogeneous token and the corresponding frequency.
[0075] Furthermore, the hash unit is used to put each token into different hash functions, map to obtain corresponding hash values; form a hash vector from the hash values corresponding to each different hash function; perform normalization processing on the hash vector to obtain digital fingerprints.
[0076] Furthermore, the hash unit is used to randomly arrange each token and put it into a hash function, map to obtain a hash value; repeat multiple times to obtain multiple different hash values; perform normalization processing on the hash vector composed of each hash value to obtain digital fingerprints.
[0077] Furthermore, the device further includes:
[0078] A syntactic tree module for obtaining the syntactic trees of the target paper and each paper to be compared by using a syntactic analyzer;
[0079] A tree bank module for converting the format of each syntactic tree to obtain each syntactic tree in the tree bank format;
[0080] A syntactic pattern extraction module for extracting corresponding syntactic patterns from each syntactic tree in the tree bank format;
[0081] A syntactic similarity module for calculating the syntactic similarity between the target paper and each paper to be compared based on the syntactic patterns;
[0082] The judgment module is further used to obtain a plagiarism detection result according to the syntactic similarity, fingerprint similarity, and term frequency similarity.
[0083] Furthermore, the above syntactic similarity module includes:
[0084] A coincidence unit for detecting the number of overlapping phrases between the syntactic pattern of a target paper and that of a paper to be compared.
[0085] An edit distance unit for calculating the edit distance between the syntactic pattern of a target paper and that of a paper to be compared.
[0086] A syntactic similarity calculation unit for multiplying the reciprocal of the edit distance by the number of overlapping phrases to obtain the syntactic similarity.
[0087] Furthermore, a judgment module is used to calculate the first similarity average of the fingerprint similarity and the word frequency similarity with the paper to be compared; if the first similarity average is greater than a first preset threshold, it is determined that there is a plagiarism problem with the target paper.
[0088] Furthermore, the judgment module is also used to calculate the second similarity average of the syntactic similarity, the fingerprint similarity, and the word frequency similarity; if the syntactic similarity is greater than a second preset threshold and the first similarity average is greater than a third preset threshold, or the second similarity average is greater than the first preset threshold, it is determined that there is a plagiarism problem with the target paper; where the first preset threshold is less than the third preset threshold.
[0089] Furthermore, the device further includes:
[0090] A refinement module for respectively inputting the target paper and the paper to be compared into a trained first model to obtain corresponding refined features.
[0091] A feature similarity module for inputting the two obtained refined features into a trained second model to obtain the feature similarity.
[0092] A feature update module for determining whether the feature similarity is greater than the fingerprint similarity between the target paper and the paper to be compared; if so, using the feature similarity as the fingerprint similarity.
[0093] Furthermore, the device further includes:
[0094] A word embedding module for obtaining multiple training papers and respectively inputting them into a word embedding model to obtain text vectors corresponding to the respective training papers.
[0095] A parameter module for obtaining optimal convolutional kernel parameters using grid search technology and constructing a convolutional deep learning model.
[0096] A training module for sequentially inputting the text vectors of the respective training papers into the convolutional deep learning model and simultaneously using regularization technology to inactivate random neurons of the convolutional deep learning model to obtain a trained first model.
[0097] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it performs the steps of a method for detecting paper plagiarism based on similarity as described in any one of the above embodiments.
[0098] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a method for detecting paper plagiarism based on similarity as described in any one of the above embodiments.
[0099] In summary, compared with the prior art, the beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include:
[0100] For a method for detecting paper plagiarism based on similarity provided by an embodiment of the present application, first, by screening the historical paper database to determine the papers to be compared, the amount of data required for subsequent processing is reduced, and the overall calculation efficiency of similarity is improved; second, the fingerprint similarity is calculated based on digital fingerprints respectively, and the word frequency similarity is calculated based on word frequency vectors. By combining the two similarities to determine whether there is plagiarism in the paper, the generation of digital fingerprints is not affected by changes in text details and has high robustness, and the results of the word frequency algorithm can accurately locate the parts of the target paper where plagiarism exists, so that while improving the detection efficiency, the calculation accuracy can be guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 It is a flowchart of a method for detecting paper plagiarism based on similarity provided by an exemplary embodiment of the present application.
[0102] Figure 2 It is a flowchart of a cleaning step provided by an exemplary embodiment of the present application.
[0103] Figure 3 It is a flowchart of the steps for calculating word frequency vectors and digital fingerprints provided by an exemplary embodiment of the present application.
[0104] Figure 4 It is a flowchart of the steps for calculating digital fingerprints provided by an exemplary embodiment of the present application.
[0105] Figure 5 It is a flowchart of the steps for calculating digital fingerprints provided by another exemplary embodiment of the present application.
[0106] Figure 6 It is a flowchart of the steps for syntactic tree analysis provided by an exemplary embodiment of the present application.
[0107] Figure 7 It is a flowchart of the steps for calculating syntactic similarity provided by an exemplary embodiment of the present application.
[0108] Figure 8 The flowchart of the fingerprint similarity update step provided for an exemplary embodiment of this application.
[0109] Figure 9 The flowchart of the first model training step provided for an exemplary embodiment of this application.
[0110] Figure 10 The structural diagram of a paper plagiarism detection device based on similarity provided for an exemplary embodiment of this application.
[0111] Figure 11 The structural diagram of the cleaning module provided for an exemplary embodiment of this application.
[0112] Figure 12 The structural diagram of the calculation module provided for an exemplary embodiment of this application. Detailed implementation manners
[0113] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0114] Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0115] Terms in the specification and claims of the present invention and the above accompanying drawings, such as "first", "second", etc. (if any), are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0116] Term explanation:
[0117] 1) Token: Refers to the smallest meaningful unit in the text after word segmentation. In different languages and contexts, tokens may have different forms:
[0118] Word: In languages such as English that use spaces to separate words, tokens are usually individual words.
[0119] Token: In some cases, especially in some Asian languages, a token may refer to a single character.
[0120] Phrase or n-gram: In some text processing tasks, a token can be a combination of multiple consecutive words or characters, such as bigram (two consecutive words), trigram (three consecutive words), etc.
[0121] Root or stem: In text preprocessing using stemming or lemmatization, a token may be the basic or standard form of a word.
[0122] 2) Term Frequency (TF): It refers to the number of times a specific term appears in text data. It is a basic concept in information retrieval and text mining, used to evaluate the importance of a term for a text document.
[0123] First, count each term in the text to calculate the total number of times each term appears in the text. For the term frequency of a specific term, it can be expressed by the following formula:
[0124]
[0125] where t is the term (updated to token in this application), d is the text containing the term (the target paper or the paper to be compared in this application), and the total number of terms refers to the total number of all non-repeating terms in the document.
[0126] In some cases, to eliminate the influence of document length on term frequency, the term frequency is normalized:
[0127]
[0128] The normalized term frequency value usually ranges between 0 and 1.
[0129] It should be noted that term frequency does not consider the importance or semantics of terms in the document. It only reflects the frequency of term occurrences. Therefore, it may overestimate the importance of common terms. So, it is usually combined with the Inverse Document Frequency (IDF) to calculate the TF-IDF weight to more accurately evaluate the significance of terms.
[0130] 3) Hash function: A mathematical function that takes an input (or "key") and computes a fixed-size numerical output, commonly referred to as a "hash value", "hash code", or "digest". A hash function produces the same output whenever it is computed for the same input. Hash functions typically map larger input data to a smaller output space, for example, mapping an input of any length to a hash value of a fixed length. Ideally, a hash function distributes the input evenly across the output space, avoiding clustering in the output. It is almost impossible to reverse-engineer the original input data from the hash value. Hash functions used in cryptography also possess some additional security features, such as resistance to preimage attacks and birthday attacks.
[0131] Common hash functions include MD5, SHA-1, SHA-256, etc., which have different security performances and efficiencies in different application scenarios. However, no hash function is completely collision-resistant, so the selection of an appropriate hash function needs to be determined according to specific application requirements and security requirements.
[0132] 4) Syntactic Tree (or Parse Tree): A graphical representation used to show the grammatical structure of a sentence. In natural language processing (NLP) and computational linguistics, a syntactic tree represents the hierarchical and dependency relationships between words in a sentence through a tree structure. Each node in a syntactic tree represents a grammatical unit, which can be a single word, a phrase, or the entire sentence. The top of the syntactic tree is the root node, representing the whole sentence. Internal nodes usually represent phrases or clauses, which are components of the sentence. Leaf nodes are the individual words in the sentence, located at the bottommost level of the tree. The lines branching out from each internal node represent grammatical relationships, such as subject-predicate structure, verb-object structure, etc.
[0133] The edges in a syntactic tree represent the dependency relationships between words, such as modification relationships, domination relationships, etc. A syntactic tree reflects the hierarchical structure of a sentence, showing how phrases are combined into larger structures. Nodes can also contain grammatical role information, such as noun phrase (NP), verb phrase (VP), etc. In some syntactic trees, leaf nodes also contain the part-of-speech information of the words, such as nouns, verbs, adjectives, etc. Syntactic trees are usually automatically generated by syntactic analyzers (Parsers), which build the tree structure based on certain grammatical rules or statistical models.
[0134] 5) CNN model (Convolutional Neural Network): A deep learning model particularly suitable for processing data with an obvious grid-like topological structure. Its internal structure mainly includes:
[0135] Convolutional Layer: Performs a convolution operation between the convolutional kernel and the input data to extract local features, and introduces non-linearity using an activation function (such as ReLU).
[0136] Activation Function: Such as ReLU, used to increase the non-linear representation ability of the network.
[0137] Pooling Layer: Adopts max pooling or average pooling to reduce the spatial dimension of the data and extract the main features.
[0138] Fully Connected Layer: At the end of the network, converts the feature maps extracted by the convolutional layer and the pooling layer into the final output, such as classification results, feature extraction results, etc.
[0139] The training of CNN is carried out through the backpropagation algorithm and the gradient descent algorithm, and the network weights are updated by calculating the gradient of the loss function. Compared with the method of manually designing features, CNN can automatically learn the features in the data. The advantages of the CNN model are parameter sharing, automatic feature extraction, and good adaptability to the spatial structure of the input data.
[0140] 6) DNN Model (Deep Neural Network): A machine learning model composed of multiple neuron layers, which mimics the structure and working principle of the human brain neural network, and realizes high-performance solutions for complex tasks through hierarchical feature learning and weight adjustment. The basic unit of DNN is the neuron. Each neuron receives inputs from other neurons and changes the influence of the input on the neuron by adjusting the weights. Through multiple non-linear hidden layers, the neural network can approximate complex functions and achieve the effect of universal approximation. The internal structure of DNN includes:
[0141] Input Layer: Receives the original data input.
[0142] Hidden Layers: Composed of multiple neurons. The output of each layer of neurons serves as the input of the next layer, and non-linear transformation is achieved through the activation function.
[0143] Output Layer: Generates the final prediction result according to the task requirements, such as classification or regression.
[0144] In the wave of artificial intelligence and machine learning, DNN has become a powerful tool for solving complex problems with its strong feature learning ability and non-linear processing ability. The advantage of DNN lies in its strong feature learning ability. Compared with the traditional method of manually designing features, DNN can automatically extract useful features from the original data, greatly improving the generalization ability of the model. In addition, the highly non-linear characteristics of DNN enable it to handle complex non-linear relationships.
[0145] Please refer to Figure 1 , the embodiment of this application provides a method for detecting paper plagiarism based on similarity, including:
[0146] Step S1, obtain the paper to be detected and clean it to obtain the target paper.
[0147] Step S2, obtain the historical paper database and screen to obtain multiple papers to be compared.
[0148] Among them, the screening method can adopt a trained clustering analysis model. Input the target paper into it, and the clustering analysis model will perform clustering analysis on the historical paper database, divide the similar papers into the same cluster, and then compare the target paper with the centers or representative papers of these clusters to obtain multiple papers to be compared.
[0149] Specifically, the conventional paper plagiarism detection algorithm usually detects the entire historical paper database and improves the detection accuracy through comprehensive comparison; but for algorithms with relatively complex calculations such as word frequency and cosine similarity, this requires very high computing power or computing time, resulting in low plagiarism detection efficiency. Therefore, this application first screens the historical paper database and then ensures the calculation accuracy through the comprehensive judgment of the two algorithms, so as to obtain accurate detection results while reducing the data volume.
[0150] Step S3, calculate the digital fingerprints and word frequency vectors of the target paper and each paper to be compared.
[0151] Among them, before calculating the digital fingerprints and word frequency vectors of the papers to be compared, the same cleaning process as the target paper can also be performed on each paper to be compared to further improve the efficiency of Step S3.
[0152] Step S4, calculate the fingerprint similarity between the target paper and each paper to be compared based on the digital fingerprints.
[0153] Specifically, the Hamming distance or Jaccard coefficient between the digital fingerprint of the target paper and the digital fingerprint of the paper to be compared can be calculated as the fingerprint similarity of the two digital fingerprints.
[0154] However, the calculation of Hamming distance requires the digital fingerprints of the target paper and the paper to be compared to have the same length, which poses a requirement for the fingerprint algorithm. Although the Jaccard coefficient can handle unequal-length digital fingerprints, since the Jaccard coefficient is calculated based on the union of two digital fingerprints, it is easy to result in a high fingerprint similarity even for a small intersection. Therefore, the performance of the Jaccard coefficient in terms of accuracy is weaker than that of Hamming distance.
[0155] Step S5: Calculate the word frequency similarity between the target paper and each paper to be compared based on the word frequency vectors.
[0156] Specifically, the word frequency similarity can be obtained by calculating the Euclidean distance or Manhattan distance between two word frequency vectors.
[0157] In practical applications, before calculating the word frequency similarity, the word frequency vectors of the target paper and the papers to be compared can be weighted by TF-IDF to reduce the influence of common words, increase the weight of professional terms, and further improve the accuracy of the word frequency similarity.
[0158] Step S6: Obtain the plagiarism detection result based on the fingerprint similarity and the word frequency similarity.
[0159] A paper plagiarism detection method based on similarity provided by the above embodiments. First, by screening the historical paper database to determine the papers to be compared, the amount of data required for subsequent processing is reduced, and the overall calculation efficiency of similarity is improved. Second, the fingerprint similarity is calculated based on digital fingerprints and the word frequency similarity is calculated based on word frequency vectors respectively, and the two similarities are combined to determine whether there is plagiarism in the paper. The generation of digital fingerprints is not affected by changes in text details and has high robustness. The result of the word frequency algorithm can accurately locate the part of the target paper where plagiarism exists, so as to ensure the calculation accuracy while improving the detection efficiency.
[0160] Please refer to Figure 2 , in some embodiments, the above-mentioned obtaining the paper to be detected and cleaning it to obtain the target paper includes:
[0161] Step S11: Remove the header, footer, references, and preset stop words of the paper to be detected.
[0162] Among them, the preset stop words are words such as "of", "already", "and", etc., which are irrelevant to vocabulary and do not contain semantic information. Removing these contents can reduce the sparsity of data and improve the accuracy and efficiency of word segmentation.
[0163] Step S12: Convert all capital letters in the paper to be detected into lowercase letters.
[0164] The case conversion operation here is to avoid judging the same word as different words due to different cases (such as "Apple" and "apple"), thereby affecting the classification of word units and the calculation accuracy of word frequency.
[0165] Step S13, performing morphological restoration on the English text of the paper to be detected to obtain the target paper.
[0166] Among them, word form restoration refers to converting vocabulary into its dictionary form, such as converting "better" to "good", which is equivalent to unifying synonyms and improving the accuracy of word classification.
[0167] In addition to the above, you can also perform stemming on the target paper, for example, restoring vocabulary to its basic form and removing tense descriptions, such as converting "running" to "run".
[0168] The various cleaning processes provided in the above embodiments can reduce the noise in the text, improve the quality of the digital fingerprint and word frequency vector, make them better reflect the semantic content of the text, and thus improve the calculation accuracy of the similarity.
[0169] In some embodiments, the above-mentioned acquisition of the historical paper database and screening to obtain multiple papers to be compared include:
[0170] Extract the keywords of the target paper; use the inverted index method to screen the papers to be compared in the historical paper database; the title or abstract of the paper to be compared contains at least one keyword.
[0171] Specifically, the method of performing inverted indexing in titles and abstracts based on keywords to determine the papers to be compared given in the above embodiment is faster than the aforementioned clustering analysis model in screening, because here only the titles and abstracts of each paper need to be traversed, but the screening accuracy is weaker than the clustering analysis model; in actual application, it can be selected according to the specific technical field of the target paper. For technical fields with more research directions and a wider range, the clustering analysis model can be used for screening first. For some less popular technical fields with fewer related research, the inverted index method can be used first.
[0172] See also Figure 3 In some embodiments, the above calculation of the digital fingerprints and word frequency vectors of the target paper and each paper to be compared may specifically include the following steps:
[0173] Step S31, segment the target paper or the paper to be compared to obtain multiple word-grams.
[0174] Step S32: use one or more hash functions to process each word to obtain a digital fingerprint.
[0175] Step S33: Classify and count each token to obtain multiple homogeneous tokens and their corresponding frequencies.
[0176] Step S34: Construct a term frequency vector based on each homogeneous token and its corresponding frequency.
[0177] Specifically, construct a vocabulary that contains all unique tokens. This vocabulary will be used to map the tokens in the paper to the vector space, and then count the number of occurrences of each token in the vocabulary.
[0178] Convert the vocabulary into vector form, where each dimension of the vector corresponds to a token in the vocabulary. If a token appears in the paper, the corresponding dimension value is the frequency of occurrence. The homogeneous tokens in this application refer to the same tokens, that is, after splitting to obtain multiple tokens, count the same tokens among them, and calculate the frequency of occurrence based on the number of the same tokens, that is, the term frequency.
[0179] Furthermore, in order to eliminate the influence of the length of the paper text on the term frequency, it is possible to standardize the term frequency vectors of the target paper and the paper to be compared (for example, divide each dimension by the total number of words in the paper text).
[0180] Even further, if the dimensional space of the obtained term frequency vector is very large, PCA (Principal Component Analysis) can be used to reduce the dimension of the term frequency vector. This is only applicable to the case where the dimensional space is extremely large; in the specific practice process, to ensure the accuracy of similarity, most of the dimensional spaces of the term frequency vectors do not reach the standard of extremely large dimensional space, and there is no need to reduce the dimension.
[0181] Please refer to Figure 4 , in some embodiments, the above-mentioned process of processing each token with one or more hash functions to obtain digital fingerprints may specifically include the following steps:
[0182] Step S3211: Put each token into different hash functions to map and obtain the corresponding hash values.
[0183] Specifically, assume a set of hash functions as H = {h1, h2, h3, …… h k,}, where h i is a hash function, each hash function is different from each other, and k is the number of hash functions. Input each token into the hash function in turn to obtain the corresponding function value, and take the minimum value among the function values of all tokens as the hash value mapped by this hash function.
[0184] For example, assume that the tokens segmented from the target paper or the paper to be compared are apple, banana, and cherry. They are respectively input into the first hash function: h1(apple) = 17, h1(banana) = 11, h1(cherry) = 5. Among these function values, 5 is the minimum. Then, 5 is taken as the hash value of h1.
[0185] Step S3212: Combine the hash values corresponding to each different hash function to form a hash vector.
[0186] Step S3213: Standardize the hash vector to obtain a digital fingerprint.
[0187] Specifically, to improve the calculation efficiency of fingerprint similarity, the obtained hash vector is standardized so that each value in the digital fingerprint is between 0 and 1, having better generalization ability.
[0188] Please refer to Figure 5 , in some embodiments, the above-mentioned use of one or more hash functions to process each token to obtain a digital fingerprint may specifically include the following steps:
[0189] Step S3221, random step: Randomly arrange each token and put it into the hash function to map and obtain a hash value.
[0190] Step S3222: Repeat the random step multiple times to obtain multiple different hash values.
[0191] Specifically, first perform a random arrangement on each token to shuffle the order of the tokens to reduce the influence of the original data structure on the result. Then, use the shuffled arrangement to cyclically shift the tokens in the arrangement to the right to generate n different arrangement orders, where n is the number of tokens segmented from the target paper or the paper to be compared.
[0192] Assume that the hash function used in this embodiment is h c , input each token in the i-th (1 ≤ i ≤ n) arrangement order into the hash function h c , obtain the corresponding function value, sort the function values according to the positions of their corresponding tokens in the i-th arrangement order, and select the first non-zero function value in the function value arrangement as the hash value corresponding to this arrangement order.
[0193] Execute the above steps for the first k arrangement orders, and take the k obtained hash values as the hash vector; it is equivalent to only repeating the step of "cyclically shifting the tokens in the arrangement to the right" to generate arrangement orders k times.
[0194] Step S3223: Standardize the hash vector composed of each hash value to obtain a digital fingerprint.
[0195] The above embodiments provide a process of constructing digital fingerprints using one hash function. In the specific implementation process, the calculation method using multiple hash functions is relatively simple and can provide accurate fingerprint similarity calculation. However, when dealing with large-scale datasets, especially when the number of papers to be compared is large, the calculation cost is relatively high. While the method of constructing digital fingerprints using a single hash function has high calculation efficiency and is very suitable for use when the number of papers to be compared is large. However, since the estimated variance of the digital fingerprints obtained by a single hash function is smaller than that of multiple hash functions, the accuracy of fingerprint similarity calculation is weaker than that of multiple hash functions.
[0196] Please refer to Figure 6 , in some embodiments, the method further includes:
[0197] Step S01, obtaining the syntactic trees of the target paper and each paper to be compared using a syntactic analyzer.
[0198] Among them, the syntactic analyzer can select analyzers such as Stanford Parser, SpaCy, NLTK, etc.
[0199] Step S02, converting the format of each syntactic tree to obtain each syntactic tree in the treebank format.
[0200] Specifically, the syntactic tree in the treebank format retains the structural information of the sentences in the paper, including the hierarchical structure of phrases, dependency relationships, and syntactic roles, etc. This helps to more deeply understand the semantic and syntactic characteristics of the sentences. Moreover, the treebank format supports the application of more complex syntactic similarity comparison algorithms, such as algorithms based on edit distance, which can more precisely evaluate the differences between sentences. In addition, the treebank format has very good versatility. Not only do many existing NLP tools and libraries support this format, but it can also be adjusted according to multiple languages to achieve cross-language sentence comparison.
[0201] Step S03, extracting the corresponding syntactic patterns from each syntactic tree in the treebank format.
[0202] Among them, the syntactic pattern is a sub-structure of the treebank, including noun phrases and verb phrases.
[0203] Step S04, calculating the syntactic similarity between the target paper and each paper to be compared based on the syntactic patterns.
[0204] Step S05, obtaining the plagiarism detection result according to the syntactic similarity, fingerprint similarity, and word frequency similarity.
[0205] Although the fingerprint similarity mentioned in the foregoing embodiments of the present application can quickly identify the similarity of texts, it does not have a deep enough understanding of semantics. The word frequency similarity focuses on the comparison at the token dimension. Therefore, although the combination of the two has higher accuracy than conventional duplicate checking algorithms, it does not have a good enough understanding of the structural information in the paper, including the components of phrases and sentences.
[0206] Therefore, the present application further adds syntactic similarity comparison to the comparison of fingerprint and word frequency similarities. Syntactic analysis can reveal the dependency relationships and semantic roles between words in the text, reduce misjudgments caused by text editing or format changes, and has good robustness to noise and variations in the text, thereby greatly improving the accuracy of paper duplicate checking.
[0207] Please refer to Figure 7 , in some embodiments, calculating the syntactic similarity of the target paper and each paper to be compared based on the syntactic pattern may specifically include the following steps:
[0208] Step S041, detecting the number of overlapping phrases between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared.
[0209] Among them, the number of overlapping phrases is the number of identical phrases in the syntactic patterns of the two papers.
[0210] Specifically, a Tree-RNN can be used to process whether the two syntactic patterns match.
[0211] Step S042, calculating the edit distance between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared.
[0212] Among them, the edit distance (Levenshtein Distance) is the minimum number of single-character editing (insertion, deletion, or replacement) operations required to convert one string into another string. In the present application, it is the minimum number of editing operations required to convert the syntactic pattern of the target paper into the syntactic pattern of the paper to be compared; when calculating the edit distance, dynamic programming can be introduced, that is, a two-dimensional array (DP table) is established to record the intermediate results, and the array is gradually filled to find the minimum edit distance. In dynamic programming, the solution of each sub-problem depends on the solutions of its adjacent sub-problems, and the state transition equation is usually based on three operations: inserting a node, deleting a node, and replacing a node.
[0213] Step S043, multiplying the reciprocal of the edit distance by the number of overlapping phrases to obtain the syntactic similarity.
[0214] Specifically, the edit distance can reflect the similarity or difference in syntactic structure between two texts. A smaller edit distance means that the two texts are more syntactically similar. Therefore, in this application, the reciprocal of the edit distance is multiplied by the number of overlapping phrases to obtain the syntactic similarity, such that the magnitude of the syntactic similarity is directly proportional to the likelihood of plagiarism.
[0215] The above embodiments calculate the syntactic similarity based on the edit distance and the number of overlapping phrases, fully quantifying the differences between syntactic trees, indirectly reflecting the semantic differences between two thesis texts, and accurately identifying minor changes in sentences.
[0216] In some embodiments, the above-mentioned obtaining the plagiarism detection result based on the fingerprint similarity and the word frequency similarity includes:
[0217] Calculating the first similarity average value of the fingerprint similarity and the word frequency similarity with the thesis to be compared; if the first similarity average value is greater than the first preset threshold, it is determined that there is a plagiarism problem in the target thesis.
[0218] In the prior art, the conventional operation usually is to determine whether a single similarity exceeds the corresponding threshold, but this determination method is overly dependent on the accuracy of the corresponding similarity calculation process and has a poor error tolerance rate; for the two methods of digital fingerprint calculation mentioned in the above embodiments, different calculation methods have a slight impact on the final similarity result; while this application calculates the average value of the two similarities, which is equivalent to combining the similarity degrees of the two theses in terms of text content and word frequency, improving the error tolerance rate, accuracy, and stability of plagiarism determination (if a text is rewritten by changing the word order or using synonyms, the individual word frequency similarity may not be able to detect plagiarism, but the fingerprint similarity may still be able to identify the similarity).
[0219] In one embodiment, the above-mentioned obtaining the plagiarism detection result based on the syntactic similarity, the fingerprint similarity, and the word frequency similarity includes:
[0220] Calculating the second similarity average value of the syntactic similarity, the fingerprint similarity, and the word frequency similarity; if the syntactic similarity is greater than the second preset threshold and the first similarity average value is greater than the third preset threshold, or the second similarity average value is greater than the first preset threshold, it is determined that there is a plagiarism problem in the target thesis; wherein, the first preset threshold is less than the third preset threshold.
[0221] Two judgment methods are given here in this application. One is to calculate the average value of three similarities. This method is more accurate than judging by a single similarity or the average value of two similarities (i.e., the first similarity average value). The other is to judge the syntactic similarity and the first similarity average value separately. At this time, the judgment threshold of the first similarity average value, that is, the third preset threshold, is greater than the first preset threshold. It can be considered that the plagiarism determination of fingerprint similarity and word frequency similarity is more lenient. This method focuses more on the judgment of semantics and sentence structure. Which one to use specifically can be selected according to the needs of researchers.
[0222] Please refer to Figure 8 , in some embodiments, the method further includes:
[0223] Step S71: Input the target paper and the paper to be compared into the trained first model respectively to obtain corresponding fine features.
[0224] Step S72: Input the two obtained fine features into the trained second model to obtain the feature similarity.
[0225] Among them, the first model is a convolutional neural network (CNN model), and the second model is a deep neural network (DNN model).
[0226] Step S73: Judge whether the feature similarity is greater than the fingerprint similarity between the target paper and the paper to be compared.
[0227] Step S74: If so, use the feature similarity as the fingerprint similarity.
[0228] As described in the above embodiments, the calculation of digital fingerprints based on hash functions is affected by the number of hash functions received. The digital fingerprints calculated by single or multiple hash functions are different, resulting in different calculation results of fingerprint similarity, especially when the calculation methods of digital fingerprints used in the target paper and the paper to be compared are different. Therefore, in order to avoid this unexpected situation, this application further adds the calculation of fingerprint similarity by the model. The fingerprint similarity is recalculated through the first model and the second model, and compared with the fingerprint similarity obtained by using the hash function, and the larger value is selected as the fingerprint similarity for subsequent judgment.
[0229] The above embodiments update the fingerprint similarity by adding two new models, realizing "double insurance" for the fingerprint similarity, and avoiding inaccurate calculation of fingerprint similarity caused by inappropriate selection of the number of hash functions.
[0230] Please refer to Figure 9 , in some embodiments, the method further includes:
[0231] Step S81: Obtain multiple training papers and input them into the word embedding model respectively to obtain text vectors corresponding to the respective training papers.
[0232] Among them, the word embedding model can be selected from Word2Vec, GloVe, FastText, etc.
[0233] Specifically, the present application uses text vectors for training to ensure that the dimensions of the training data input into the CNN model are the same, thereby accelerating the training efficiency of the CNN model.
[0234] Step S82: Use the grid search technique to obtain the optimal convolution kernel parameters and construct a convolutional deep learning model.
[0235] Specifically, determining the optimal convolution kernel size and stride is one of the key steps in constructing a CNN model. The selection of these parameters affects the model's ability to capture local features. The grid search system in the present application systematically searches for the optimal convolution kernel size and stride, and can quickly determine the optimal convolution kernel parameters corresponding to the task requirements of the present application. The stride determines the interval at which the convolution kernel slides on the input data. A larger stride can reduce the spatial dimension of the output feature map, helping to reduce the computational amount, but may lose some information. The stride is usually set to 1 or 2.
[0236] Step S83: Input the text vectors of each training paper into the convolutional deep learning model in sequence, and at the same time, use the regularization technique to inactivate the random neurons of the convolutional deep learning model to obtain the trained first model.
[0237] Specifically, during the training process, Dropout randomly inactivates the neurons (or features) in the convolutional deep learning model according to a predetermined probability (usually a hyperparameter, such as 0.5), that is, sets them to zero. This means that in each iteration, the architecture of the network will be slightly different. Since only a part of the neurons are activated in each iteration, this reduces the complex co-adaptation relationship between neurons, that is, prevents the network from relying too much on specific feature combinations.
[0238] In a sense, Dropout is equivalent to training multiple different models (different due to random inactivation in each iteration), and averaging the prediction results of these models during testing. This ensemble method can improve the generalization ability of the model.
[0239] Please refer to Figure 10 , another embodiment of the present application provides a paper plagiarism detection device based on similarity, including:
[0240] The cleaning module 101 is used to obtain the paper to be detected and clean it to obtain the target paper.
[0241] The screening module 102 is used to obtain the historical paper database and screen out multiple papers to be compared.
[0242] A calculation module 103 is configured to calculate the digital fingerprints and term frequency vectors of the target paper and each paper to be compared.
[0243] A fingerprint module 104 is configured to calculate the fingerprint similarity between the target paper and each paper to be compared based on the digital fingerprints.
[0244] A term frequency module 105 is configured to calculate the term frequency similarity between the target paper and each paper to be compared based on the term frequency vectors.
[0245] A judgment module 106 is configured to obtain a plagiarism detection result according to the fingerprint similarity and the term frequency similarity.
[0246] Please refer to Figure 11 , further, the above-mentioned cleaning module 101 includes:
[0247] A removal unit 11 is configured to remove the header, footer, references, and preset stop words of the paper to be detected.
[0248] A letter conversion unit 12 is configured to convert all the capital letters of the paper to be detected into lowercase letters.
[0249] A lemmatization unit 13 is configured to perform lemmatization on the English text of the paper to be detected to obtain the target paper.
[0250] Further, a screening module 102 is configured to extract each keyword of the target paper, and use the inverted index method to screen the papers to be compared in the historical paper database; at least one keyword exists in the title or abstract of the paper to be compared.
[0251] Please refer to Figure 12 , further, the above-mentioned calculation module 103 includes:
[0252] A segmentation unit 31 is configured to segment the target paper or the paper to be compared to obtain a plurality of tokens.
[0253] A hash unit 32 is configured to process each token with one or more hash functions to obtain digital fingerprints.
[0254] A statistics unit 33 is configured to classify and count each token to obtain a plurality of tokens of the same category and the corresponding frequencies.
[0255] A vector unit 34 is configured to construct a term frequency vector according to each token of the same category and the corresponding frequency.
[0256] Further, the hash unit 32 is configured to put each token into different hash functions, map to obtain the corresponding hash values; form a hash vector from the hash values corresponding to each different hash function; perform normalization processing on the hash vector to obtain digital fingerprints.
[0257] Further, the hash unit 32 is used to randomly arrange each token into a hash function to map and obtain a hash value; repeat multiple times to obtain multiple different hash values; and normalize the hash vector composed of each hash value to obtain a digital fingerprint.
[0258] Further, the device further includes:
[0259] A syntactic tree module, which is used to obtain the syntactic trees of the target paper and each paper to be compared by using a syntactic analyzer.
[0260] A tree bank module, which is used to convert the format of each syntactic tree to obtain each syntactic tree in the tree bank format.
[0261] A syntactic extraction module, which is used to extract corresponding syntactic patterns from each syntactic tree in the tree bank format.
[0262] A syntactic similarity module, which is used to calculate the syntactic similarity between the target paper and each paper to be compared based on the syntactic patterns.
[0263] The judgment module is further used to obtain a plagiarism detection result according to the syntactic similarity, fingerprint similarity, and word frequency similarity.
[0264] Further, the above-mentioned syntactic similarity module includes:
[0265] A coincidence unit, which is used to detect the number of overlapping phrases between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared.
[0266] An edit distance unit, which is used to calculate the edit distance between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared.
[0267] A syntactic similarity calculation unit, which is used to multiply the reciprocal of the edit distance by the number of overlapping phrases to obtain the syntactic similarity.
[0268] Further, the judgment module 106 is used to calculate a first similarity average value of the fingerprint similarity and the word frequency similarity with the paper to be compared; if the first similarity average value is greater than a first preset threshold, it is determined that the target paper has a plagiarism problem.
[0269] Further, the judgment module 106 is further used to calculate a second similarity average value of the syntactic similarity, fingerprint similarity, and word frequency similarity; if the syntactic similarity is greater than a second preset threshold and the first similarity average value is greater than a third preset threshold, or the second similarity average value is greater than the first preset threshold, it is determined that the target paper has a plagiarism problem.
[0270] Wherein, the first preset threshold is less than the third preset threshold.
[0271] Further, the device further includes:
[0272] A refinement module for inputting the target paper and the paper to be compared into a trained first model respectively to obtain corresponding refined features.
[0273] A feature similarity module for inputting the two obtained refined features into a trained second model to obtain a feature similarity.
[0274] A feature update module for determining whether the feature similarity is greater than the fingerprint similarity between the target paper and the paper to be compared; if so, using the feature similarity as the fingerprint similarity.
[0275] Further, the apparatus further includes:
[0276] A word embedding module for obtaining multiple training papers and inputting them into a word embedding model respectively to obtain text vectors corresponding to the respective training papers.
[0277] A parameter module for obtaining optimal convolutional kernel parameters by using a grid search technique and constructing a convolutional deep learning model.
[0278] A training module for inputting the text vectors of the respective training papers into the convolutional deep learning model in sequence, and at the same time using a regularization technique to deactivate random neurons in the convolutional deep learning model to obtain a trained first model.
[0279] The specific limitations of the apparatus for detecting paper plagiarism based on similarity provided in this embodiment can be referred to the embodiment of the method for detecting paper plagiarism based on similarity in the foregoing text, and will not be elaborated herein. Each module in the above-mentioned apparatus for detecting paper plagiarism based on similarity can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0280] In the embodiments disclosed in the present application, it should be understood that the disclosed products can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the modules can be in electrical, mechanical or other forms. The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0281] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor executes the steps of a method for detecting paper plagiarism based on similarity as described in any of the above embodiments.
[0282] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference can be made to the embodiments of the method for detecting paper plagiarism based on similarity in the above text, and details will not be repeated here.
[0283] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a method for detecting paper plagiarism based on similarity as described in any of the above embodiments. Among them, the computer-readable storage medium refers to a carrier for storing data, which can but is not limited to including floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks, etc. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0284] For the working process, working details and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the method for detecting paper plagiarism based on similarity in the foregoing text, which will not be elaborated herein.
[0285] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0286] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0287] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A method for detecting paper plagiarism based on similarity, characterized in that Including: Obtain the paper to be detected and clean it to obtain the target paper; Obtain the historical paper database and screen to obtain multiple papers to be compared; Calculate the digital fingerprints and word frequency vectors of the target paper and each of the papers to be compared; Calculate the fingerprint similarity between the target paper and each of the papers to be compared based on the digital fingerprints; Calculate the word frequency similarity between the target paper and each of the papers to be compared based on the word frequency vectors; Use a syntactic analyzer to obtain the syntactic trees of the target paper and each of the papers to be compared; Convert the format of each of the syntactic trees to obtain each of the syntactic trees in the treebank format; Extract the corresponding syntactic patterns from each of the syntactic trees in the treebank format; Calculate the syntactic similarity between the target paper and each of the papers to be compared based on the syntactic patterns; Obtain the plagiarism detection result according to the syntactic similarity, the fingerprint similarity, and the word frequency similarity; 2. The method for detecting paper plagiarism based on similarity according to claim 1, wherein The calculating the digital fingerprints and word frequency vectors of the target paper and each of the papers to be compared includes: Segment the target paper or the paper to be compared to obtain multiple word tokens; Process each of the word tokens with one or more hash functions to obtain the digital fingerprints; Classify and count each of the word tokens to obtain multiple groups of similar word tokens and corresponding frequencies; Construct the word frequency vector according to each of the groups of similar word tokens and the corresponding frequencies; 3. The method for detecting paper plagiarism based on similarity according to claim 2, wherein The processing each of the word tokens with one or more hash functions to obtain the digital fingerprints includes: Put each of the word tokens into different ones of the hash functions to map to the corresponding hash values; Form a hash vector from the hash values corresponding to each of the different hash functions; Perform normalization processing on the hash vector to obtain the digital fingerprints; 4. The method for detecting paper plagiarism based on similarity according to claim 2, characterized in that The processing each of the word tokens with one or more hash functions to obtain the digital fingerprints includes: Random step: Randomly arrange each of the word tokens and put them into the hash functions to map to hash values; Repeat the random step multiple times to obtain multiple different hash values; Perform normalization processing on the hash vector composed of each of the hash values to obtain the digital fingerprints; 5. The method for detecting paper plagiarism based on similarity according to claim 1, wherein The calculating the syntactic similarity between the target paper and each of the papers to be compared based on the syntactic patterns includes: Detect the number of overlapping phrases between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared; Calculate the edit distance between the syntactic pattern of the target paper and the syntactic pattern of the paper to be compared; Multiply the reciprocal of the edit distance by the number of overlapping phrases to obtain the syntactic similarity; 6. The method for detecting paper plagiarism based on similarity according to claim 1, wherein Also including: Input the target paper and the papers to be compared into the trained first model respectively to obtain the corresponding fine features; Input the two obtained fine features into the trained second model to obtain the feature similarity; Judge whether the feature similarity is greater than the fingerprint similarity between the target paper and the paper to be compared; If so, use the feature similarity as the fingerprint similarity; 7. The method for detecting paper plagiarism based on similarity according to claim 6, characterized in that, Also including: Obtain multiple training papers and input them into the word embedding model respectively to obtain the text vectors corresponding to each of the training papers; Use the grid search technique to obtain the optimal convolutional kernel parameters and construct a convolutional deep learning model; The text vectors of the respective training papers are sequentially input into the convolutional deep learning model, and at the same time, a regularization technique is used to inactivate random neurons in the convolutional deep learning model, thereby obtaining the trained first model.
8. A paper plagiarism detection device based on similarity, characterized in that, It includes: A cleaning module for obtaining a paper to be detected and cleaning it to obtain a target paper; A screening module for obtaining a historical paper database and screening out a plurality of papers to be compared; A calculation module for calculating the digital fingerprints and word frequency vectors of the target paper and each of the papers to be compared; A fingerprint module for calculating the fingerprint similarity between the target paper and each of the papers to be compared based on the digital fingerprints; A word frequency module for calculating the word frequency similarity between the target paper and each of the papers to be compared based on the word frequency vectors; A syntactic tree module for obtaining the syntactic trees of the target paper and each of the papers to be compared by using a syntactic analyzer; A tree bank module for converting the formats of the respective syntactic trees to obtain the syntactic trees in the tree bank format; A syntactic pattern extraction module for extracting corresponding syntactic patterns from the syntactic trees in the tree bank format; A syntactic similarity module for calculating the syntactic similarity between the target paper and each of the papers to be compared based on the syntactic patterns; A judgment module for obtaining a plagiarism detection result according to the syntactic similarity, fingerprint similarity, and word frequency similarity.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the similarity-based paper plagiarism detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fingerprint feature-based text copy detection system and method
CN105912514A