Article duplicate checking method for preventing generation of duplicate

Through the methods of textual paragraphs and textual intention vectors, combined with attention mechanisms and supervised learning, the problems that are difficult to identify rewrites and retelling in the existing technology are solved, and the reliability and anti-cracking ability of article plagiarism checking are improved.

CN120429422AActive Publication Date: 2025-08-05HUAZHONG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510483573.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-05
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing plagiarism checking technology is difficult to effectively identify, rewritten and retell, and is easily cracked by artificial intelligence, resulting in a decrease in the reliability of plagiarism checking.

Method used

The method of text-information paragraphs and text-information vectors is used to simulate human reading comprehension through the attention mechanism AM algorithm, combined with the keyword co-occurrence matrix and disambiguation parameters, the text-information paragraphs of the article are identified, and ANN is used for supervised learning division to enhance the effectiveness of plagiarism checking.

Benefits of technology

Effectively identifying AI recapitalization reduces the effect of AI through keyword substitution and grammatical transformation, improves the reliability of article plagiarism checking, and enhances the ability to combat cracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429422A_ABST
    Figure CN120429422A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field related to artificial intelligence, and particularly relates to an article duplicate checking method for preventing duplicate generation, which comprises the following steps that: duplicate checking is carried out by using text meaning segments and text meaning vectors, and the text meaning vector of each text meaning segment is obtained by inputting an attention mechanism AM algorithm through all vector groups corresponding to all keywords of the text meaning segment, obtaining a text meaning vector of the text meaning segment; the vector group is a vector group formed by word vectors of the keywords and disambiguation parameter vectors, the disambiguation parameter vector of each keyword is a vector formed by disambiguation parameters between the keyword and various effective words in the corpus, and the dimension of the disambiguation parameter vector is the same as that of the word vector of the keyword; word vectors of the keywords are formed by connecting co-occurrence matrixes of the keywords to symbolized normalized vectors of the co-occurrence matrixes. According to the method, a duplicate generation algorithm is counteracted, so that the AI cannot effectively reduce the weight of the article, and the reliability of the current article duplicate checking technology is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence related technologies, and more specifically, relates to a method for checking for duplicate articles to prevent repetitive generation. Background Art

[0002] Article plagiarism detection technology compares similarities between articles to identify duplicate and plagiarized content. It has been widely used in the review process of journals. However, current plagiarism detection technology faces serious challenges from artificial intelligence (AI), and its reliability is declining. The latest generative AI, such as the ChatGPT4 model, can achieve deep plagiarism reduction. Articles modified with this model can reduce the duplication rate to 5-10% while maintaining the basic meaning. The resulting plagiarism detection results meet the plagiarism detection requirements of most journals. This poses new challenges to article plagiarism detection technology.

[0003] Current article duplication detection technologies have developed limited preventative measures against artificial intelligence, such as: 1. Patent number CN119292658A, based on structural embedding vectors and sequential embedding representations, fuses the structure and sequence of the source code through an attention mechanism to generate a global vector; based on this global vector, the similarity between the source code to be detected and the target source code is calculated. 2. Patent number CN119271528A, drawing on the Reflexion framework and creating an agent based on the ReAct mechanism, uses a reflection mechanism to evaluate current results and conduct in-depth analysis, resulting in more comprehensive duplication detection results.

[0004] However, the current article duplication detection technology still has two difficulties that cannot be solved: 1. Article duplication detection has poor recognition ability for rewriting and restatement. For example, even if the meanings of two sentences are exactly the same, if all words are replaced with synonyms, the duplication rate obtained by the duplication check will still decrease. For another example, if an English document is translated into Chinese and then translated back into English, the duplication rate obtained by the duplication check will also decrease, but the actual meaning has not changed. 2. The algorithm for article duplication detection is easy to crack and can be easily attacked by artificial intelligence in a targeted manner. Taking the article fingerprint-based duplication detection as an example, the code for generating the hash value can be calculated by the duplication rate of a specific sample text segment. It only needs to use the sample text segment to perform a few dozen duplication checks to basically crack the duplication detection code, which has poor security. Summary of the Invention

[0005] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a method for checking for duplicate articles to prevent rewriting, which aims to simulate the human reading comprehension process, deeply learn the expression of text, complete the extraction and comparison of text meaning, and efficiently realize the identification of rewriting or rewriting.

[0006] To achieve the above objectives, according to one aspect of the present invention, a method for checking for duplicate articles to prevent repetition is provided, comprising:

[0007] The article to be checked for duplicate content is divided into words; valid non-connective words are identified from all words, and a co-occurrence matrix is generated for each valid word; keywords are determined based on the weights of each valid word in the article; the article is divided into paragraphs based on the keywords to obtain multiple paragraphs; and the disambiguation parameters between each valid word in the article and various valid words in the corpus are solved;

[0008] Calculate the r of each keyword in each text segment i To the keyword r immediately following it i+1 The middle text placeholder parameter L(r i ), the value is determined in advance based on the linguistic characteristics of the stop words corresponding to the text based on experiments; construct the keyword r i The symbolic normalized vector of Among them, the vector The positive and negative signs of L(r i ) is the same, vector The dimension of is determined in advance through experiments; if L(r i ) is greater than the dimension, then the vector Each dimension of is equal to one-half of the square root of the dimension, otherwise, the vector V ri The first dimension value of The square root of L(r) is 1 / 2, and the other dimensions are 0. The dimensions of the first dimension are the same as L(r i ) value; if The sign is negative, then the keyword r i With the keyword r i+1 Exchange positions, otherwise do not exchange;

[0009] The co-occurrence matrix of each keyword in each semantic segment is concatenated to the back of its symbolic normalized vector to form the word vector of the keyword; the word vector and disambiguation parameter vector of each keyword in each semantic segment form a vector group, and all the vector groups corresponding to all keywords in the semantic segment are input into the attention mechanism AM algorithm to obtain the semantic vector of the semantic segment; wherein the disambiguation parameter vector of each keyword is a vector composed of the disambiguation parameters between the keyword and various valid words in the corpus, and its dimension is the same as the word vector dimension of the keyword;

[0010] The duplicate checking rate of each semantic segment is calculated based on the semantic vector of the semantic segment, and the duplicate checking rate of the article to be checked is calculated based on the duplicate checking rates of each semantic segment of the article to be checked.

[0011] Furthermore, we use adhesion judgment to implement word segmentation through ANN, and the segmentation method is as follows:

[0012] For the three characters A, B, and C, the probability of the word composed of A and B appearing in the entire corpus is P AB , the probability of the word composed of B and C appearing in the entire corpus is P BC , if P AB >>P BC , then B and A form a word; if P AB >>P BC , then B and C form a word; if P AB ≈P BC , then B, A and C do not constitute a word.

[0013] Furthermore, the method of determining keywords is:

[0014] Assign a weight to each valid word according to the importance of the text;

[0015] Determine the maximum weight T among all valid word weights in the article to be checked for duplicates max , the weight is Valid words within the range are used as keywords; determine whether the number of keywords meets the preset number, if not, take the relative weight Valid words with a low range and close to the range until the number of keywords reaches the preset number.

[0016] Furthermore, the way to assign weights to each word is as follows:

[0017]

[0018]

[0019] TF-IDF(t,d)=TF(t,d)×IDF(t)

[0020] T(t)=TF-IDF(t,d)

[0021] Where TF(t,d) represents the frequency parameter of word t in the article d to be checked for duplicates; n t represents the number of times word t appears in the article d to be checked for duplicates; n0 represents the total number of words in the article d to be checked for duplicates; IDF(t) represents the inverse document frequency of word t; N t It represents the number of documents in the corpus where word t appears, N represents the total number of documents in the corpus; TF-IDF(t,d) represents the term frequency-inverse document frequency of word t in the article d to be checked for duplicates, and T(t) represents the weight of word t.

[0022] Furthermore, the maximum weight T among all the valid word weights in the article to be checked is determined. max Previously, it also included: updating the weight of each valid word in the following way:

[0023] Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a Wc represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, i Represents the known valid words c in the corpus i The co-occurrence matrix of , i = 1, 2, ..., m; m represents the number of all valid words in the corpus; if the disambiguation parameters between the valid word a and the m valid words in the corpus are not greater than the preset value Δ, then the weight of the valid word a does not need to be modified; otherwise, the weight of the valid word a is updated according to the following formula:

[0024]

[0025] Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ; T′(a) is the weight of the updated effective word a; TF(c′ i ,d) represents the valid word c′ in the article d to be checked for duplicates i The word frequency parameter, IDF(c′ i ) represents a valid word c′ i The inverse document frequency of

[0026] Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final weight of the valid word a in the article to be checked for duplicates.

[0027] Further, solve the co-occurrence matrix W of each valid word a in the article to be checked for duplicates a The implementation is:

[0028] In the article to be checked for duplicates, determine the n valid words before and after the valid word a, and record them as b respectively. -n ,b -n+1 ,……,b n ;

[0029] Determine b i The probability of occurrence of a P a,bi : In all documents, search for b in the range -n to n before and after each valid word a i Words, can find bi The statistical probability of the word, as P a,bi , i=-n,……,n;

[0030] All P a,bi Construct a vector and use it as the Y of the HMM algorithm 1,2n Matrix, Y 1,2n The matrix transposed as [P(Y|X)] of the HMM algorithm 2n,1 Matrix, input HMM algorithm, output result X 1,m That is the co-occurrence matrix W of the effective word a a .

[0031] Furthermore, before forming the word vector of the keyword, the method also includes updating the co-occurrence matrix of each valid word a, specifically:

[0032] Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a Wc represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, i Represents the known valid words c in the corpus i The co-occurrence matrix of the effective word a is , i = 1, 2, ..., m; m represents the number of all valid words in the corpus; if the disambiguation parameters between the effective word a and the m valid words in the corpus are not greater than the preset value Δ, then the co-occurrence matrix of the effective word a does not need to be modified; otherwise, the co-occurrence matrix of the effective word a is updated according to the following formula:

[0033]

[0034] Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ;W a ′ is the co-occurrence matrix of the effective word a after update; W c′i Represents the known valid words c′ in the corpus i The co-occurrence matrix of

[0035] Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final co-occurrence matrix of the valid words a in the article to be checked for duplicates.

[0036] Furthermore, the method of dividing the text into paragraphs according to keywords is as follows: the positions of all keywords in the article to be checked for duplicates, the positions of periods in each sentence of the article to be checked for duplicates and the positions of paragraph divisions are input into ANN to obtain the division results.

[0037] Furthermore, the supervised learning conditions when supervising ANN learning include:

[0038] (1) The division of the paragraph is located at the period or paragraph division;

[0039] (2) Let P λ is the ratio of the number of keywords in the first paragraph to the number of all keywords in this paragraph, then among all possible divisions, the division P λ The average value should be the highest;

[0040] (3) The number of text paragraphs should not be less than half of the number of natural paragraphs, and should not be more than the number of natural paragraphs.

[0041] According to another aspect of the present invention, a computer-readable storage medium is provided, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to execute the steps of the above-mentioned method.

[0042] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:

[0043] 1. The present invention proposes a method for checking duplicate articles to prevent paraphrase generation. It innovatively uses the concept of semantic segments and the parameter semantic vector, so that the duplicate check does not only stop at literal comparison, but also can understand the meaning of key words, so that the effect of AI paraphrase generation using keyword replacement is further reduced. Among them, the method of the present invention proposes that the semantic vector of each semantic segment is obtained by inputting all vector groups corresponding to all keywords in the semantic segment into the attention mechanism AM algorithm to obtain the semantic vector of the semantic segment, thereby identifying the AI paraphrase for the entire paragraph; the vector group is a vector group composed of the keyword's word vector and the disambiguation parameter vector, and the disambiguation parameter vector of each keyword is a vector composed of the disambiguation parameters between the keyword and various valid words in the corpus, and its dimension is the same as the keyword's word vector dimension, thereby identifying the AI paraphrase performed by the keyword synonym replacement method; the keyword's word vector is formed by concatenating the keyword's co-occurrence matrix to its symbolized normalized vector, while taking into account the information of both the article's language structure and the meaning of the words, thereby enhancing the effectiveness of counter-paraphrase generation. Therefore, the present invention can effectively identify rewriting and restatement, effectively counter the AI restatement generation algorithm, making it impossible for AI to effectively reduce the plagiarism of articles, increasing the reliability of current article plagiarism detection technology, and can solve the problem that plagiarism detection technology is difficult to identify AI-generated and plagiarized articles, that is, plagiarism detection is ineffective for rewritten and restated articles, and the plagiarism detection algorithm is easily cracked.

[0044] 2. The present invention also proposes updating keyword weights and / or co-occurrence matrices to achieve keyword disambiguation. This, on the one hand, eliminates the need for rigid word-to-word comparisons in duplicate checking, significantly reducing the effectiveness of AI restatement generation algorithms such as keyword substitution and grammatical transformation. On the other hand, disambiguation can increase information entropy, complicating AI inverse deduction formulas and significantly increasing the difficulty of AI cracking.

[0045] 3. The present invention further proposes implementing segmentation via an artificial neural network (ANN). Using supervised learning ANN to generate relevant parameters, AI counters AI, enabling the present invention to rapidly iterate despite the threat of cracking by restatement-generated AI, further enhancing the present invention's advantage in countering cracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flowchart of a method for checking for duplicate articles to prevent repetition provided by an embodiment of the present invention;

[0047] Figure 2 A flowchart of another method for checking for duplicate articles to prevent repetition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0049] Example 1

[0050] A method for checking duplicate articles to prevent repetition, such as Figure 1 Shown, including:

[0051] The article to be checked for duplicate content is divided into words; valid non-connective words are identified from all words, and a co-occurrence matrix is generated for each valid word; keywords are determined based on the weights of each valid word in the article; the article is divided into paragraphs based on the keywords to obtain multiple paragraphs; and the disambiguation parameters between each valid word in the article and various valid words in the corpus are solved;

[0052] Calculate the r of each keyword in each text segment i To the keyword r immediately following it i+1 The middle text placeholder parameter L(r i ), the value is determined in advance based on the linguistic characteristics of the stop words corresponding to the text based on experiments; construct the keyword r i The symbolic normalized vector of Among them, the vector The positive and negative signs of L(r i ) is the same, vector The dimension of is determined in advance through experiments; if L(r i ) is greater than the dimension, then the vector Each dimension of takes the value of one-half of the square root of the dimension, otherwise, the vector The value of the first dimension is L(r i ) is one-tenth of the square root of L(r), and the values of the other dimensions are 0. The dimensions of the first dimension are the same as L(r i ) value; if The sign is negative, then the keyword r i With the keyword r i+1 Exchange positions, otherwise do not exchange;

[0053] Concatenate the co-occurrence matrix of each keyword in each semantic paragraph behind its symbolized normalized vector to form the word vector of the keyword; form a vector group with the word vectors of each keyword and the disambiguation parameter vector in each semantic paragraph, and input all the said vector groups corresponding to all keywords in this semantic paragraph into the attention mechanism AM algorithm to obtain the semantic vector of this semantic paragraph; wherein, the disambiguation parameter vector of each keyword is a vector composed of the disambiguation parameters between this keyword and various valid words in the corpus, and its dimension is the same as the dimension of the word vector of this keyword;

[0054] Calculate the duplicate rate of each semantic paragraph according to the semantic vector of each semantic paragraph, and calculate the duplicate rate of the article to be checked based on the duplicate rates of each semantic paragraph of the article to be checked.

[0055] The method of this embodiment needs to first identify valid words that are not conjunctions from all words. The definition of valid words is as follows: Words other than stop words (conjunctions, auxiliary words, pronouns, etc.) in the sense of computer science are all valid words. All valid words can be divided by directly excluding stop words.

[0056] In addition, regarding the construction of the symbolized normalized vector. Suppose there are p keywords in a semantic paragraph, i = 1, 2,..., p, a keyword r i to the keyword r i+1 immediately following it has a placeholder parameter L(r i ), and this parameter is determined by all the stop words between the two keywords. Extract the stop words among them, and the direct sum of the linguistic features of all stop words in terms of data is L(r i ). The relevant parameters for converting the linguistic features of stop words into data need to be obtained through experiments. For example, "de" is -1, "le" is 0, etc.

[0057] Set the parameter χ, and the symbolized normalized vector of the keyword r i can be expressed as:

[0058]

[0059] where V ri is the symbolized normalized vector of the keyword r i , and it is a χ-dimensional column vector. When L(r i ) < χ, the first [L(r ri )] (the square brackets indicate rounding down) dimensions of V i are and the subsequent dimensions are 0. When L(r i ) ≥ χ, all dimensions are As a preferred solution, χ can take 12, 13 or 14, which is determined by experiments.

[0060] If V ri sgn[L(r i )] is greater than zero, then the keyword r i If it is less than 0, this keyword will be swapped with the keyword immediately following it.

[0061] The matrix concatenation method is: generate a column vector U with dimension χ+m r As the keyword r i The word vector of is (m represents the number of valid words in the corpus), the first χ dimension is equal to its symbolic normalized vector, and the last m dimension is equal to its co-occurrence matrix.

[0062] The simplified workflow of the AM algorithm is as follows:

[0063] The keyword with the largest number in a certain paragraph is called the semantic keyword of this paragraph. Let the semantic keyword of a certain paragraph be r′0, and all the keywords are rearranged into r′1, r′2, ... r′ p , whose word vector is U r′1 , U r′2 , ...U r′p , then the semantic vector U of this semantic segment P for:

[0064]

[0065] Among them, S r′0,r′i Represents the contextual keyword r′0 and any keyword r′ i The disambiguation parameter between .

[0066] As an example, the formula for checking the duplicate rate of a paragraph can be:

[0067]

[0068] Where λ is the order of the text segments, U i Represents the semantic vector of each semantic segment in the corpus.

[0069] As an example, the formula for the full text duplication rate can be:

[0070]

[0071] where λ t is the number of Chinese semantic segments in the article to be checked for plagiarism, and Ψ0 is the full-text plagiarism rate.

[0072] As a preferred implementation, the stickiness judgment is adopted to implement word segmentation through ANN, and the segmentation method is as follows:

[0073] For the three characters A, B, and C, the probability of the word composed of A and B appearing in the entire corpus is PAB , the probability of the word composed of B and C appearing in the entire corpus is P BC , if P AB >>P BC , then B and A form a word; if P AB >>P BC , then B and C form a word; if P AB ≈P BC , then B, A and C do not constitute a word.

[0074] We can use ANN with adhesion criteria to improve word segmentation accuracy. The database is a corpus and idioms, and machine learning is performed on the probability that a certain word and its preceding and following words form a word.

[0075] As a preferred implementation, the above method of determining keywords is:

[0076] Assign a weight to each valid word according to the importance of the text;

[0077] Determine the maximum weight T among all valid word weights in the article to be checked for duplicates max , the weight is Valid words within the range are used as keywords; determine whether the number of keywords meets the preset number, if not, take the relative weight Valid words with a low range and close to the range until the number of keywords reaches the preset number.

[0078] For example, for the weights of all valid words in the article, find its maximum value T max , all values are in All words in the range are keywords. Then, let the number of words in the article be N d (Total number of words divided in S1), if the number of keywords at this time (Square brackets indicate rounding down), then continue to take words with lower weights until the number of keywords reaches until.

[0079] As a further preferred embodiment, the way to assign weights to each word is as follows:

[0080]

[0081]

[0082] TF-IDF(t,d)=TF(t,d)×IDF(t)

[0083] T(t)=TF-IDF(t,d)

[0084] Where TF(t,d) represents the frequency parameter of word t in the article d to be checked for duplicates, which is usually a number greater than 0 and much less than 1; n t Indicates the number of times word t appears in the article d to be checked for duplicates; n0 indicates the total number of words in the article d to be checked for duplicates; IDF(t) indicates the inverse document frequency of word t, which usually has three types of values: type 1 values are between 0-0.001, type 2 values are between 2.3-3.5, and type 3 values are greater than 5; N t It represents the number of documents in the corpus in which word t appears, and N represents the total number of documents in the corpus. TF-IDF(t,d) represents the term frequency minus the inverse document frequency of word t in the article d to be checked for duplicates, and T(t) represents the weight of word t, which is equal to TF-IDF(t,d).

[0085] As a further preferred embodiment, the maximum weight T among all the valid word weights in the article to be checked for duplicates is determined. max Previously, it also included: updating the weight of each valid word in the following way:

[0086] Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, W ci Represents the known valid words c in the corpus i The co-occurrence matrix is i = 1, 2, ..., m; m represents the number of all valid words in the corpus (including the articles to be checked for duplicates); if the disambiguation parameters between the valid word a and the m valid words in the corpus are not greater than the preset value Δ, then the weight of the valid word a does not need to be modified; otherwise, the weight of the valid word a is updated according to the following formula:

[0087]

[0088] Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ; T′(a) is the weight of the updated effective word a; TF(c′ i ,d) represents the valid word c′ in the article d to be checked for duplicates i The word frequency parameter, IDF(c′ i ) represents a valid word c′ i The inverse document frequency of

[0089] Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final weight of the valid word a in the article to be checked for duplicates.

[0090] As a preferred implementation, the co-occurrence matrix W of each valid word a in the article to be checked for duplicates can be solved. a The implementation is:

[0091] In the article to be checked for duplicates, determine the n valid words before and after the valid word a, and record them as b respectively. -n ,b -n+1 ,……,b n ;

[0092] Determine b i The probability of occurrence of a P a,bi : In all documents, search for b in the range -n to n before and after each valid word a i Words, can find b i The statistical probability of the word, as P a,bi , i=-n,……,n;

[0093] All P a,bi Construct a vector and use it as the Y of the HMM algorithm 1,2n Matrix, Y 1,2n The matrix transposed as [P(Y|X)] of the HMM algorithm 2n,1 Matrix, input HMM algorithm, output result X 1,m That is the co-occurrence matrix W of the effective word a a .

[0094] That is, first set the parameter n to represent the position of the valid word, and stipulate that the first valid word before a valid word is position -1, and so on, the nth valid word before a valid word is position -n, and the nth valid word after a valid word is position n. Suppose all the valid words in the interval from -n to n in the article to be checked for duplicates are b -n ,b -n+1 ,……,b n .

[0095] In the article to be checked for duplicates, for a valid word a, solve its co-occurrence matrix W a :

[0096] For a and each b in the article to be checked for plagiarism i (i=-n,...,n), solve b i Regarding the probability of occurrence of a; all P corresponding to a a,bi Constitute a vector, which is the Y of the HMM algorithm 1,2nMatrix; Y 1,2n The matrix transposed as [P(Y|X)] of the HMM algorithm 2n,1 Matrix, input HMM algorithm, output result X 1,m This is the co-occurrence matrix W of the word a .

[0097] As a further preferred embodiment, before forming the word vector of the keyword, the method further includes updating the co-occurrence matrix of each valid word a, specifically:

[0098] Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a Wc represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, i Represents the known valid words c in the corpus i The co-occurrence matrix of the valid word a is , i = 1, 2, ..., m; m represents the number of all valid words in the corpus (including the current article to be checked for duplicates); if the disambiguation parameters between the valid word a and the m valid words in the corpus are not greater than the preset value Δ, then there is no need to modify the co-occurrence matrix of the valid word a; otherwise, the co-occurrence matrix of the valid word a is updated according to the following formula:

[0099]

[0100] Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ;W a ′ is the co-occurrence matrix of the updated valid word a; W c′i Represents the known valid words c′ in the corpus i The co-occurrence matrix of

[0101] Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final co-occurrence matrix of the valid words a in the article to be checked for duplicates.

[0102] It should be noted that updating the weights and updating the co-occurrence matrix are both disambiguation operations. Both updates can be performed simultaneously, or one or the other can be selected for disambiguation.

[0103] After completing the update of weights and co-occurrence matrix, a, c′1, c′2, ..., c′ in the article to be checked for duplicates m′These words are replaced with a. The purpose of this operation is to prevent repetition. In the process of AI searching for synonyms of words in the repetition process, synonyms are merged to minimize the impact of the AI repetition process on the article's duplication rate.

[0104] As an example, the default value Δ is set to approximately 0.937, determined experimentally. After disambiguation, repeat the above steps and set the parameter Δ to approximately 0.894. Experiments have shown that articles using AI deep deduplication are prone to multiple repetitions, and a single disambiguation cannot completely eliminate this effect. Multiple disambiguations can be performed, and the parameter Δ can be slightly increased to offset the effect of AI deep deduplication.

[0105] Furthermore, the method of dividing the text into paragraphs according to keywords is as follows: the positions of all keywords in the article to be checked for duplicates, the positions of periods in each sentence of the article to be checked for duplicates and the positions of paragraph divisions are input into ANN to obtain the division results.

[0106] Furthermore, the supervised learning conditions when supervising ANN learning include:

[0107] (1) The division of the text paragraph is located at the period or paragraph division;

[0108] (2) Let P λ is the ratio of the number of keywords in the first paragraph to the number of all keywords in this paragraph, then among all possible divisions, the division P λ The average value should be the highest;

[0109] (3) The number of text paragraphs should not be less than half of the number of natural paragraphs, and should not be more than the number of natural paragraphs.

[0110] The method of this embodiment contains artificial intelligence and deep learning modules, whose parameters can be adjusted automatically through training, without the need to manually adjust the parameters one by one. Therefore, the introduction of neural networks and hidden Markov models only leaves input and output, and the working process and the specific values of various parameters do not need to be elaborated in the above description. The method of this embodiment is particularly suitable for checking Chinese papers for plagiarism, because Chinese stop words are more directional and linguistic features are easy to grasp, and the resulting word vectors are more effective.

[0111] Based on the content of the above embodiment 1, Figure 2 A block diagram of one of the duplicate checking methods is shown.

[0112] Example 2

[0113] The present application also relates to a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the computer program is executed by a processor.

[0114] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0115] The relevant technical solutions are the same as above and will not be repeated here.

[0116] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for checking for duplicate articles to prevent repetition, characterized in that: include: Divide the articles to be checked for plagiarism into words; Identify valid words that are non-connective words from all words, and generate a co-occurrence matrix for each valid word; Determine keywords based on the weights of each valid word in the article; divide the article into paragraphs based on the keywords to obtain multiple paragraphs; solve the disambiguation parameters between each valid word in the article and various valid words in the corpus; Calculate the r of each keyword in each text segment i To the keyword r immediately following it i+1 The middle text placeholder parameter L(r i ), the value is determined in advance based on the linguistic characteristics of the stop words corresponding to the text based on experiments; construct the keyword r i The symbolic normalized vector of Among them, the vector The positive and negative signs of L(r i ) is the same, vector The dimension of is determined in advance through experiments; if L(r i ) is greater than the dimension, then the vector Each dimension of takes the value of one-half of the square root of the dimension, otherwise, the vector The value of the first dimension is L(r i ), the other dimensions are 0, and the dimensions of the first dimension are the same as L(r i ) value; if The sign is negative, then the keyword r i With the keyword r i+1 Exchange positions, otherwise do not exchange; The co-occurrence matrix of each keyword in each semantic segment is concatenated to the back of its symbolic normalized vector to form the word vector of the keyword; the word vector and disambiguation parameter vector of each keyword in each semantic segment form a vector group, and all the vector groups corresponding to all keywords in the semantic segment are input into the attention mechanism AM algorithm to obtain the semantic vector of the semantic segment; wherein the disambiguation parameter vector of each keyword is a vector composed of the disambiguation parameters between the keyword and various valid words in the corpus, and its dimension is the same as the word vector dimension of the keyword; The duplicate checking rate of each semantic segment is calculated based on the semantic vector of the semantic segment, and the duplicate checking rate of the article to be checked is calculated based on the duplicate checking rates of each semantic segment of the article to be checked.

2. The article duplication checking method according to claim 1, characterized in that: Adhesion judgment is used to implement word segmentation through ANN. The segmentation method is as follows: For the three characters A, B, and C, the probability of the word composed of A and B appearing in the entire corpus is P AB , the probability of the word composed of B and C appearing in the entire corpus is P BC , if P AB >>P BC , then B and A form a word; if P AB >>P BC , then B and C form a word; if P AB ≈P BC , then B, A and C do not constitute a word.

3. The article duplication checking method according to claim 1, wherein: The method of determining keywords is: Assign a weight to each valid word according to the importance of the text; Determine the maximum weight T among all valid word weights in the article to be checked for duplicates max , the weight is Valid words within the range are used as keywords; determine whether the number of keywords meets the preset number, if not, take the relative weight Valid words with a low range and close to the range until the number of keywords reaches the preset number.

4. The article duplication checking method according to claim 3, wherein: The way to assign weights to each word is: TF-IDF(t,d)=TF(t,d)×IDF(t) T(t)=TF-IDF(t,d) Where TF(t,d) represents the frequency parameter of word t in the article d to be checked for duplicates; n t represents the number of times word t appears in the article d to be checked for duplicates; n0 represents the total number of words in the article d to be checked for duplicates; IDF(t) represents the inverse document frequency of word t; N t It represents the number of documents in the corpus where word t appears, N represents the total number of documents in the corpus; TF-IDF(t,d) represents the term frequency-inverse document frequency of word t in the article d to be checked for duplicates, and T(t) represents the weight of word t.

5. The article duplication checking method according to claim 3, characterized in that: Determine the maximum weight T among all the valid word weights in the article to be checked for duplicates max Previously, it also included: updating the weight of each valid word in the following way: Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a Wc represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, i Represents the known valid words c in the corpus i The co-occurrence matrix of , i = 1, 2, ..., m; m represents the number of all valid words in the corpus; if the disambiguation parameters between the valid word a and the m valid words in the corpus are not greater than the preset value Δ, then the weight of the valid word a does not need to be modified; otherwise, the weight of the valid word a is updated according to the following formula: Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ; T′(a) is the weight of the updated effective word a; TF(c′ i ,d) represents the valid word c′ in the article d to be checked for duplicates i The word frequency parameter, IDF(c′ i ) represents a valid word c′ i The inverse document frequency of Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final weight of the valid word a in the article to be checked for duplicates.

6. The article duplication checking method according to claim 1, wherein: Solve the co-occurrence matrix W of each valid word a in the article to be checked for duplicates a The implementation is: In the article to be checked for duplicates, determine the n valid words before and after the valid word a, and record them as b respectively. -n ,b -n+1 ,……,b n ; Determine b i The probability of occurrence of a P a,bi : In all documents, search for b in the range -n to n before and after each valid word a i Words, can find b i The statistical probability of the word, as P a,bi , i=-n,……,n; All P a,bi Construct a vector and use it as the Y of the HMM algorithm 1,2n Matrix, Y 1,2n The matrix transposed as [P(Y|X)] of the HMM algorithm 2n,1 Matrix, input HMM algorithm, output result X 1,m That is the co-occurrence matrix W of the effective word a a .

7. The article duplication checking method according to claim 1, characterized in that: Before forming the word vector of the keyword, the method also includes updating the co-occurrence matrix of each valid word a, specifically: Solve the effective word a in the article to be checked and the i-th effective word c in the corpus i The disambiguation parameter S between a,ci , the larger the value of the disambiguation parameter, the more likely a and c are to be different. i The more likely they are to become synonyms; among them, W a Wc represents the co-occurrence matrix of valid words a in the article to be checked for duplicates, i Represents the known valid words c in the corpus i The co-occurrence matrix of the effective word a is , i = 1, 2, ..., m; m represents the number of all valid words in the corpus; if the disambiguation parameters between the effective word a and the m valid words in the corpus are not greater than the preset value Δ, then the co-occurrence matrix of the effective word a does not need to be modified; otherwise, the co-occurrence matrix of the effective word a is updated according to the following formula: Where m' represents the number of valid word types in the corpus corresponding to the disambiguation parameter being greater than the preset value Δ, and the corresponding valid word types are recorded as c'1, c'2, ..., c' m′ ;W a ′ is the co-occurrence matrix of the effective word a after update; W c′i Represents the known valid words c′ in the corpus i The co-occurrence matrix of Increase the value of the preset value Δ, repeat the above updating operation for a preset number of times, and obtain the final co-occurrence matrix of the valid words a in the article to be checked for duplicates.

8. The article duplication checking method according to claim 1, wherein: The method of dividing the text into paragraphs according to keywords in the article to be checked for duplicate content is as follows: the positions of all keywords in the article to be checked for duplicate content, the positions of periods in each sentence of the article to be checked for duplicate content and the positions of paragraph divisions are input into ANN to obtain the division results.

9. The article duplication checking method according to claim 8, characterized in that: The supervised learning conditions in supervised learning ANN include: (1) The division of the paragraph is located at the period or paragraph division; (2) Let P λ is the ratio of the number of keywords in the first paragraph to the number of all keywords in this paragraph, then among all possible divisions, the division P λ The average value should be the highest; (3) The number of text paragraphs should not be less than half of the number of natural paragraphs, and should not be more than the number of natural paragraphs.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to perform the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Text duplicate checking method and device based on attention mechanism, equipment and storage medium

    CN110347790A

  • Domain-oriented science and technology project duplicate checking method and system

    CN116431763A

  • Automatic tagging method and apparatus, and computer device and storage medium

    WO2019153552A1