Method and device for detecting academic disused literature
By acquiring and matching the semantic features and layout features of academic literature, and using a multi-dimensional semantic feature index library for searching, the problem of in-depth rewriting and plagiarism in the existing technology is solved, and a more accurate and efficient detection of academic misconduct is achieved.
Patent Information
- Application Number
- CN202510558613.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing academic misconduct literature detection technology cannot identify deeply rewritten plagiarism and distinguish between key content and non-critical content in duplicate texts.
By obtaining the semantic features and layout features of the literature to be analyzed, the preset multi-dimensional semantic feature index library is used to search semantic similarity, and combining semantic features and layout features for matching and rearranging, to obtain detection results.
It improves the ability to identify plagiarism after deep rewrites and duplication of key and non-critical content, and enhances the accuracy and recall of the detection results.
Smart Images

Figure CN120086379A_ABST
Abstract
Description
Background Art
[0002] Plagiarism is a typical manifestation of academic misconduct, with various forms, including simple verbatim copying of text, shallow rewriting such as word substitution and sentence structure adjustment, and deep rewriting that combines methods such as expansion, abbreviation, narrative style and narrative logic transformation. At present, the detection of academic text plagiarism is mainly achieved through text content similarity retrieval, and there are mainly three technical routes for similarity retrieval: full-text inverted index based on characters, words, phrases and N-grams; semantic hash index; vector index based on representation learning. The three indexes are suitable for different services and have different applicable scenarios. The generation quality and retrieval effect of the indexes also have their own advantages and disadvantages. Due to the significant improvement in vector representation learning ability in recent years, the vector index has obvious advantages in the recall rate at the semantic level, and the vector index has gradually become the mainstream technical method for content retrieval, or is combined with other traditional index methods to jointly improve the similarity retrieval effect.
[0003] In the prior art, various academic misconduct literature detection technologies can effectively detect plagiarism in paper content, including plagiarism after shallow semantic rewriting, by means of vector representation and retrieval technologies. However, there are still obvious deficiencies: First, in the face of deep rewriting (abbreviation, expansion, rewriting, etc.) by humans or large AI models (such as DeepSeek, ChatGPT, Wenxin Yiyan, etc.), since the rewritten content varies greatly in terms of word usage, language expression narrative style and content logic organization method, although the semantics are the same, existing detection software is difficult to identify such plagiarism; Second, there is a certain error rate in the detection results. When the language text expression method is highly repetitive with non-critical content but the key core objects are different, the text is easily misjudged as repetitive.
[0004] The fundamental reason is that the current technologies based on vector representation or based on word, n-gram repeated retrieval cannot accurately express the complete semantics of the text. For example, when the text is long, the vector representation ability decreases, and it is difficult for vector-based methods to completely express the semantics; when generating vectors by dividing sentences or fragments, due to the large semantic span of the core key content, the generated vectors cannot accurately represent the text semantics. And currently, whether it is based on vector semantic representation or based on words or n-gram methods, it is impossible to distinguish the repetition of key and non-key content in the repeated text.
[0005] In summary, the existing academic misconduct literature detection technologies have deficiencies in dealing with the recognition of plagiarism in deep rewriting and distinguishing the repetition of key and non-key content in the text, and urgent improvement is needed. Summary of the Invention
[0006] To solve the above problems in the prior art, that is, the existing academic misconduct literature detection technology cannot identify plagiarism detection with deep rewriting and cannot distinguish the repetition of key content and non-key content in repeated texts, the first aspect of the present invention proposes a method for detecting academic misconduct literature, and the method includes the following steps: S100. Obtain the literature to be analyzed, and use a pre-trained feature extraction model to extract features from the literature to be analyzed to obtain query semantic features and layout features; the query semantic features include text body semantic features, title keyword fusion semantic features, and abstract semantic features; S200. Perform semantic similarity retrieval in a preset multi-dimensional semantic feature index library based on the query semantic features to obtain similar semantic features and corresponding similar paragraph IDs and similar literature IDs; the multi-dimensional semantic feature index library includes various literature sets and semantic feature sets, and any semantic feature in the semantic feature set is associated with an original literature paragraph ID, an original literature ID, and an original text position in various literature sets through an index; S300. Perform weighted counting on the similar paragraph IDs and the similar literature IDs respectively, and perform a descending order sorting based on the obtained count values to obtain a similar sorting result; S400. Match the literature to be analyzed and the query semantic features with the similar literature IDs and the corresponding similar semantic features in the similar sorting result to determine the semantic feature similarity; S500. Match the layout features with a preset layout feature library to determine the layout similarity; S600. Rearrange the similar sorting result based on the semantic feature similarity and / or layout similarity to obtain a detection result.
[0007] In some preferred embodiments, the method for using a pre-trained feature extraction model to extract features from the literature to be analyzed is as follows: Perform a structural analysis on the literature to be analyzed, extract the title, abstract, natural paragraphs of the text body, figures and sub-figures, tables, formulas, and references of the literature to be analyzed, and record the layout order and layout area between the structural contents; Divide the extracted abstract into five parts: background, purpose, method, result, and conclusion, and perform vectorization on each part respectively to obtain abstract semantic features with different dimensions and structures; Sort the natural paragraphs of the extracted text body and extract the content logical link, where the content logical link includes main logical link sentences and side branch logical link sentences, perform vectorization on the main logical link sentences and side branch logical link sentences respectively to obtain text body semantic features, and record the main logical link sentences, side branch logical link sentences, and original text paragraphs corresponding to the semantic features; Vectorize the extracted title and keywords respectively, and then calculate the weighted integrated semantic features of the title keywords. Parse the extracted structural content to obtain layout features centered on figures, tables, and formulas.
[0008] In some preferred embodiments, semantic similarity retrieval is performed in a preset multi-dimensional semantic feature index library based on the query semantic features. The method is as follows: Based on the query semantic features, perform similarity retrieval recall in the same type of semantic feature indexes in the multi-dimensional semantic feature index library to obtain several alternative similar semantic features for each query semantic feature; wherein, the types include text semantic features, integrated semantic features of title keywords, and abstract semantic features, and similarity thresholds are respectively set for the main logical link sentences and side branch logical link sentences in the text semantic features. Select the alternative similar semantic features with similarity greater than the preset similarity threshold as similar semantic features. Obtain the similar semantic features, the similarity of the similar semantic features, and the original document paragraph ID and original document ID pointed to by the vector index of the similar semantic feature, and use the original document paragraph ID and original document ID as the corresponding similar paragraph ID and similar document ID.
[0009] In some preferred embodiments, the layout features centered on figures, tables, and formulas are obtained. The method is as follows: S110: Starting from figures, tables, and formulas respectively, perform pre-order traversal to obtain the area of the pre-order sequence whose first is text as the pre-order text area. S120: Starting from figures, tables, and formulas respectively, perform post-order traversal to obtain the area of the post-order sequence whose first is text as the post-order text area. S130: Repeat steps S110 - S120 until the layout feature sequences of all figures, tables, and formulas are obtained: FigureID i ={ Figure , SubFigure , PreTxt , NextTxt}; TableID i ={ Table , CellNum , PreTxt , NextTxt}; FormulaID i ={ Formula , LatexNum , PreTxt , NextTxt}; Among them, FigureID i is the layout feature sequence of the i th figure, Figure , SubFigure , PreTxt , NextTxt respectively represent the figure area, the average area of sub - figures, the area of the preceding text, and the area of the following text; TableID i is the layout feature sequence of the i th table, Table , CellNum , PreTxt , NextTxt respectively represent the surface area, the number of table cells, the area of the preceding text, and the area of the following text; FormulaID i is the layout feature sequence of the i th table, Formula , LatexNum , PreTxt , NextTxt respectively represent the formula area, the number of formula latex commands, the area of the preceding text, and the area of the following text.
[0010] In some preferred embodiments, weighted counting is respectively performed on the similar paragraph IDs and the similar literature IDs, and the method is as follows: According to the types of query semantic features based on which the similar semantic features are retrieved and recalled, different weights are assigned to the similar paragraph IDs or similar literature IDs; Among them, the types include the main text semantic features, the fused semantic features of title keywords, and the abstract semantic features, and different weights are respectively set for the main logical link sentences and the side - branch logical link sentences in the main text semantic features; The number of times each similar paragraph ID or similar literature ID is recalled is calculated according to the weights.
[0011] In some preferred embodiments, the semantic feature similarity includes any one or more of the main logical link sentence similarity, the side - branch logical link sentence similarity, the fused semantic feature similarity of title keywords, and the abstract semantic feature similarity.
[0012] In some preferred embodiments, the layout similarity is determined, and the method is as follows: The layout features are respectively tested by the difference ratio, and it is judged whether the difference ratios in four dimensions between the layout features of the literature to be analyzed and the layout features of the same type in the literature to be compared simultaneously satisfy the preset parameter set K; If the condition is met, the matching degree count of the document layout features corresponding to the current layout features is incremented by 1, where the same figure, table, and formula features that meet the condition are not cumulatively counted; otherwise, no count is made. Until all layout features of the document to be analyzed are verified, the layout similarity is obtained.
[0013] In some preferred embodiments, the similar sorting results are rearranged based on the semantic feature similarity, and the method is as follows: The similar sorting results are rearranged with the semantic feature similarity as the weight, and the repetition situation between paragraphs is considered; among them, the repetition weight of the main logical link sentences is greater than that of the side branch logical link sentences, the repetition weight within the same paragraph is greater than that across paragraphs, and the continuous repetition weight is greater than the scattered repetition weight.
[0014] In some preferred embodiments, if there is only an independent layout feature library generated from the factory thesis literature collection, after the document to be analyzed obtains the layout features, it jumps to step S500.
[0015] In a second aspect of the present invention, an apparatus for detecting academic misconduct documents is proposed, and the apparatus includes: An index construction module configured to construct a multi-dimensional semantic feature index library by using a trained feature extraction model, where the multi-dimensional semantic feature index library includes various literature collections and semantic feature collections, and any semantic feature in the semantic feature collection is associated with an original document paragraph ID, an original document ID, and an original text position in various literature collections through an index; A feature extraction module configured to obtain a document to be analyzed and extract features from the document to be analyzed by using a trained feature extraction model to obtain query semantic features and layout features; A semantic similarity retrieval module configured to perform semantic similarity retrieval in a preset multi-dimensional semantic feature index library based on the query semantic features to obtain similar semantic features and corresponding similar paragraph IDs and similar document IDs; A semantic similarity calculation module configured to match the document to be analyzed and the query semantic features with the similar document IDs and corresponding similar semantic features in the similar sorting results to determine the semantic feature similarity; A layout similarity calculation module configured to match the layout features with a preset layout feature library to determine the layout similarity; A result sorting module configured to respectively perform weighted counting on the similar paragraph IDs and the similar document IDs, and perform a descending order sorting based on the obtained count values to obtain similar sorting results; it is also configured to rearrange the similar sorting results based on the semantic feature similarity and / or layout similarity to obtain detection results.
[0016] Advantages of the present invention: (1) The method of the present invention extracts semantic features of the main text, fuses semantic features of title keywords, abstract semantic features and layout features through structuring. According to the retrieval results of different semantic features, a variety of fusion methods are adopted to calculate the similarity for semantic feature matching, obtain the original text content, and conduct comparison calculations together. It can improve the effect of semantic repetition detection after deep modification of the text by various methods such as abbreviation, expansion, fusion of narrative styles and narrative logics, improve the recall rate, and at the same time improve the accuracy of the detection results and can distinguish semantic repetitions between key content and non-key content. (2) For the text content of existing literature, the text is decomposed by designing the main logical link and the branch logical link. After the main logical link and the branch logical link are transformed into their respective logical link sentences, vectors are generated as semantic feature vectors. The main logical link sentences are used as the input for extracting key semantic features, and the branch logical link sentences are used as the input for non-key semantic features. It can ensure the integrity of semantic segmentation of the text content, and at the same time can also distinguish the key part and the non-key part in the text content. (3) In view of the structural characteristics of academic literature and the importance of different parts, for the non-main text content, according to the structural characteristics and importance of different parts of the literature, a weighted mean vector representation method is designed, and the weights of title keywords are fused as independent semantic features. After structuring the abstract, semantic features of "background", "purpose", "method", "result", and "conclusion" are established respectively, improving the semantic feature representation ability and refining the granularity of semantic feature representation. (4) A layout structure feature representation method with figures, tables, and formulas as the core content and a literature layout structure matching method based on this method are proposed, which is a general auxiliary method for comparing the layout of various literatures. And it can achieve rapid matching of factory papers. Even under the premise of only a small amount of factory paper data, it is still possible to compare the layout features of factory papers. Description of the Drawings
[0017] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more obvious: Figure 1 is a flowchart of the method for detecting academic misconduct literature in an embodiment of the present invention; Figure 2 is a structural schematic diagram of a semantic feature set in a multi-dimensional semantic feature index library in an embodiment of the present invention; Figure 3 is a flowchart of extracting features and analyzing the content of each structure in an embodiment of the present invention. Detailed Embodiments
[0018] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than limiting the invention. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings.
[0019] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and embodiments.
[0020] The present invention improves the situation of repetition where it is impossible to distinguish key content from non - key content in repetitive texts. It extracts the logical chain of the main text content, performs structured conversion on the abstract and title respectively, then performs vector conversion respectively, conducts similarity retrieval and recall through a pre - constructed multi - dimensional semantic feature index library, fuses and calculates the matching results, realizes in - depth retrieval of the input literature, and can determine whether it is a repetition of the main logical link sentences or the side - branch logical link sentences.
[0021] To more clearly illustrate the method of the present invention for improving the detection effect of academic misconduct literature, the following combines Figure 1 to elaborate on each step in the embodiments of the present invention.
[0022] A method for detecting academic misconduct literature according to the first embodiment of the present invention includes steps S100 - S600, and each step is described in detail as follows: S100. Obtain the literature to be analyzed, and use a pre - trained feature extraction model to extract features from the literature to be analyzed, obtaining query semantic features and layout features; the query semantic features include main text semantic features, title keyword fusion semantic features, and abstract semantic features.
[0023] Preferably, the method of using a pre - trained feature extraction model to extract features from the literature to be analyzed is as follows: Conduct a structured analysis on the literature to be analyzed, extract the title, abstract, natural paragraphs of the main text, figures and sub - figures, tables, formulas, and references of the literature to be analyzed, and record the layout order and layout area between the structural contents. Segment the extracted abstract into five parts: background, purpose, method, result, and conclusion, and perform vectorization on each part respectively to obtain abstract semantic features with different dimensions and structures. Sort the natural paragraphs of the extracted main text and extract the content logical chain. The content logical chain includes main logical link sentences and side - branch logical link sentences. Perform vectorization on the main logical link sentences and side - branch logical link sentences respectively to obtain main text semantic features, and record the main logical link sentences, side - branch logical link sentences, and original main text paragraphs corresponding to the semantic features. Vectorize the extracted title and keywords respectively, and then calculate the weighted combined semantic features of the title and keywords. Parse the extracted structural contents to obtain the layout features centered on figures, tables, and formulas.
[0024] In this embodiment, if the extracted abstract is not a structured abstract, first perform a structured conversion, and then extract it into five parts: "background", "purpose", "method", "result", and "conclusion".
[0025] According to the structural characteristics of academic documents and the importance of different parts, design different vectorization representation methods. The method of representing the combination of the document title and keywords based on the weighted mean improves the accuracy of semantic feature representation, and the method of first structuring the abstract and then regarding it as an independent semantic feature respectively refines the granularity of semantic feature representation.
[0026] In this embodiment, dependency syntactic analysis or an AI large model is used to extract the content logical link, but those skilled in the art can also use other methods for logical link extraction. The content logical link is a chain relationship among various content elements such as various entities, concepts, terms, actions, etc., and is a formal semantic representation of the text content.
[0027] The content logical link mentioned in this application includes the main logical link sentence and the collateral logical link sentences. There is only one main logical link sentence, which reflects the key elements of the core content of the text and the key semantics of the text content. There can be multiple collateral logical link sentences, which serve as supplements to the content on the main link.
[0028] By designing the main logical link sentence and the collateral logical link sentences to decompose the text, the main logical link sentence is used as the input for extracting key semantic features, and the collateral logical link sentences are used as the input for non-key semantic features. Compared with the existing text segmentation methods based on the complete text, sentence segmentation, fixed length, sliding window segmentation, etc., the way of logical link can ensure the integrity of semantic segmentation of the text content, and at the same time distinguish the key content and non-key content of the academic document text.
[0029] In this embodiment, for the extraction of the logical link, based on the method of the present invention, extract the logical link of a certain paragraph in a certain paper, specifically: Text natural paragraph: {Biodiversity is the integration of the ecological complex formed by organisms and their environment, as well as various ecological processes related thereto, including animals, plants, microorganisms, the genes they possess, and the complex ecosystems they form with their living environment. It generally includes three components: genetic diversity, species diversity, and ecosystem diversity. The survival of humans depends on other organisms. The rich and diverse organisms and ecosystems provide the material basis and environmental basis for human survival. In recent years, with the continuous growth of the world's population and the increasing scope and intensity of human activities, how to protect and rationally utilize biodiversity is a challenge faced by countries around the world.} The logical link of the extracted content is: Main link sentence: Biodiversity → Ecological complex and ecological processes → Dependence of human survival Side branch link sentence: 1. Biodiversity → Genetic diversity, species diversity, ecosystem diversity 2. Biodiversity → Influence of human activities → Challenges of protection and utilization}
[0030] Further preferably, the extracted title and keywords are respectively vectorized, and then the semantic features of the title-keyword fusion are calculated by weighted calculation. The method is as follows: 1) Vectorize the obtained title. The vectorization method is not limited, and TV is obtained, representing the semantic features of the title; 2) Vectorize one or more obtained keywords. The same method as in step 1) is used for vectorization, and {KV i} is obtained, representing the semantic features of the keywords, i representing the keyword serial number; 3) Calculate the semantic features of the title-keyword fusion. The calculation method is as follows: ; where TKV represents the semantic feature vector of the title-keyword fusion, represents the weight of the title, n is the number of keywords, W i represents the i th keyword weight.
[0031] In this embodiment, the weights of n keywords can be the same, or can decrease according to the keyword serial number, or can partially decrease and partially be the same.
[0032] The representation method of the layout structure features with figures, tables, and formulas as the core content and the literature layout structure matching method based on this method can be used as a general auxiliary method for various literature layout comparisons.
[0033] Further preferably, obtain the layout features centered on figures, tables, and formulas, and the method is as follows: S110: Respectively starting from figures, tables, and formulas, perform a preorder sequence traversal, and obtain the area of the preorder sequence whose first is the body text as the preorder text area; S120: Respectively starting from figures, tables, and formulas, perform a postorder sequence traversal, and obtain the area of the postorder sequence whose first is the body text as the postorder text area; S130: Repeat steps S110 - S120 until the layout feature sequences of all figures, tables, and formulas are obtained: FigureID i ={ Figure , SubFigure , PreTxt , NextTxt}; TableID i ={ Table , CellNum , PreTxt , NextTxt}; FormulaID i ={ Formula , LatexNum , PreTxt , NextTxt}; Among them, FigureID i is the layout feature sequence of the i th figure, Figure , SubFigure , PreTxt , NextTxt respectively represent the figure area, average sub - figure area, preorder text area, and postorder text area; TableID i is the layout feature sequence of the i th table, Table , CellNum , PreTxt , NextTxt respectively represent the surface area, number of table cells, preorder text area, and postorder text area; FormulaID i is the layout feature sequence of the i th table, Formula , LatexNum , PreTxt , NextTxt respectively represent the formula area, number of formula latex commands, preorder text area, and postorder text area.
[0034] In this embodiment, further explanations are made for the preorder text area and the postorder text area: In step S100, a pre-trained feature extraction model is used to extract features from the literature to be analyzed, extracting the title, abstract, natural paragraphs of the body text, figures and sub-figures, tables, formulas, and references of the literature to be analyzed, and recording the layout order and layout area between the structural contents; assuming the extraction order is: Title, abstract, [Body text paragraph P1, Body text paragraph P2, Body text paragraph P3], Figure 1 , Figure 2 , Formula 1, [Body text paragraph P4, Body text paragraph P5], Formula 2, Figure 3 , Figure 4, [Body text paragraph P6, Body text paragraph P7] Table 1 [Body text paragraph P8, Body text paragraph P9] [References]; Then, Figure 3 The corresponding area of the previous text is: the area of [Body text paragraph P4, Body text paragraph P5], and the area of the subsequent text is the area of [Body text paragraph P6, Body text paragraph P7].
[0035] As an option, in this embodiment, the text area can also be replaced by the text length.
[0036] In this embodiment, structuring is extracted using a visual model. As an option, other common structuring methods in the art can also be used.
[0037] S200. Semantic similarity retrieval is performed one by one in a preset multi-dimensional semantic feature index library based on each query semantic feature to obtain a retrieval result set for each semantic feature. The retrieval result set includes similar semantic features and corresponding similar paragraph IDs and similar document IDs; the multi-dimensional semantic feature index library includes various types of document sets and semantic feature sets. Any semantic feature in the semantic feature set is associated with an original document paragraph ID, an original document ID, and an original text location through indexing in various types of document sets.
[0038] Preferably, the multi-dimensional semantic feature index library is obtained from various types of documents obtained. By using a pre-trained feature extraction model to extract features from various types of document sets, semantic features corresponding to various types of documents are obtained, including body text semantic features, title keyword fusion semantic features, and abstract semantic features. At the same time, layout features are obtained, and index construction is performed based on the semantic features of various types of documents obtained. Each semantic feature is linked with an identifier pointing to the original text content, including a document ID and an original text location.
[0039] Preferably, the various types of literature collections include academic literature collections and factory paper collections, and the layout feature library includes a factory paper layout feature library and an academic literature collection layout feature library. The factory paper layout feature library and the academic literature collection layout feature library are extracted based on the academic literature collection and the factory paper collection respectively.
[0040] Preferably, if there is only an independent factory paper layout feature library generated from the factory paper literature collection, after the literature to be analyzed obtains the layout features, it jumps to step S500.
[0041] More preferably, the index construction method can adopt various existing vector index methods according to the scale of the data volume; the semantic features generated by the same vector generation method can also be merged into the same index when constructing the index.
[0042] Specifically, in this embodiment, the structured content obtained in step S100 can be stored in blocks or as a whole. The original text position (various types of literature collections) here can be either the position of the whole storage or the position of the block storage, which is not limited.
[0043] Preferably, the query semantic features are respectively subjected to similarity retrieval in the same type of semantic feature index in the multi-dimensional semantic feature index library to obtain a retrieval result set for each semantic feature. The method is as follows: First, based on the query semantic features, perform similarity retrieval recall of the same type of semantic features in the multi-dimensional semantic feature index library to obtain several alternative similar semantic features for each query semantic feature, and obtain the title + keyword retrieval result set T = { t i}, the structured abstract retrieval results {A = { a i}, B = { b i}, C = { c i}, D = { d i}, E = { e i}}, the main link retrieval results {M 1 = { m 1i}, M 2 = { m 2i}, …} and the side branch link retrieval results {S 1 = { s 1i}, S 2 = { s 2i}, …}; Among them, A, B, C, D, and E respectively represent the retrieval result sets of the contents of "background", "purpose", "method", "result", and "conclusion"; the types include the main text semantic features, the fused semantic features of the title keywords, and the abstract semantic features, and similarity thresholds are respectively set for the main logical link sentences and the collateral logical link sentences in the main text semantic features; Then, the alternative similar semantic features with similarity greater than the preset similarity threshold are screened as the similar semantic features; Finally, the similar semantic features, the similarity of the similar semantic features, the original document paragraph ID and the original document ID pointed to by the vector index of the similar semantic features are obtained, and the original document paragraph ID and the original document ID are used as the corresponding similar paragraph ID and similar document ID.
[0044] In this embodiment, the number of elements in each set is flexibly controlled according to the similarity threshold or a fixed number is set.
[0045] S300. Respectively perform weighted counting on the similar paragraph ID and the similar document ID, and perform a descending order sorting based on the obtained count values to obtain a similar sorting result. The method is as follows: According to the types of the query semantic features based on which the similar semantic features are retrieved and recalled, different weights are assigned to the similar paragraph ID or the similar document ID; Among them, the types include the main text semantic features, the fused semantic features of the title keywords, and the abstract semantic features, but different weights are respectively set for the main logical link sentences and the collateral logical link sentences in the main text semantic features; Calculate the number of times each similar paragraph ID or similar document ID is recalled according to the weights; According to the above counting results, perform a descending order sorting on each similar paragraph ID or similar document ID respectively.
[0046] Preferably, in this embodiment, the retrieval result of the main text semantic features is the main factor for grouped counting, and the count value of the similar document ID retrieved and recalled by the main link can be set to have a weight equal to or greater than that of the collateral link.
[0047] Further preferably, in this embodiment, weighted counting is performed based on the above method, and the result is: A certain similar document ID is retrieved and recalled 5 times, among which 2 times are recalled through the main logical link sentence and 3 times are recalled through the secondary logical link sentence; then the weighted count value corresponding to the similar document ID is 3*MW + 2*SW; Suppose that a total of 2 similar passage IDs are hit in the 5 retrievals of the similar literature ID. One natural passage ID is retrieved by 2 main logical link sentences and 1 side branch logical link sentence. The weighted count value of this similar passage ID is 2*MW + 1*SW; the other similar passage ID is retrieved by 1 main logical link sentence and 1 side branch logical link sentence. The weighted count value of this similar passage ID is 1*MW + 1*SW.
[0048] Among them, MW is the weight of retrieval and recall of main logical link sentences, and SW is the weight of recall of side branch logical link sentences.
[0049] Further preferably, in this embodiment, when performing weighted counting, the retrieval results of the title + keyword and the structured abstract features can be considered, or not.
[0050] Further preferably, when the number of various types of IDs obtained is extremely high and some results need to be filtered, only the IDs with high count values can be selected according to needs, or the similarity threshold can be further increased to continue screening and filtering IDs, so as to reduce the subsequent comparison process and improve the detection speed.
[0051] S400. Match the literature to be analyzed and the query semantic features with the similar literature IDs and the corresponding similar semantic features in the similar sorting results to determine the semantic feature similarity; Among them, the semantic feature similarity includes any one or more of the main logical link sentence similarity, the side branch logical link sentence similarity, the title keyword fusion semantic feature similarity, and the abstract semantic feature similarity.
[0052] Preferably, when calculating the similarity, a vector similarity calculation method is adopted, including but not limited to various methods such as vector distance and dot product.
[0053] S500. Match the layout features with a preset layout feature library to determine the layout similarity. The method is as follows: Perform tests on the layout features using the difference ratio respectively to determine whether the difference ratios in four dimensions between the layout features of the literature to be analyzed and the layout features of the same type of the literature to be compared simultaneously satisfy the preset parameter set K; The test formula is: , where i = 1, 2, 3, 4; Further, different preset parameter sets can be used for the three layout types of figures, tables, and formulas k i , a represents any one of the figure, table, or formula layout features in the literature input by the user, b represents a figure, table, or formula layout feature of the same type in the literature to be compared; If it is satisfied, the matching degree count of the document layout feature corresponding to the current layout feature is incremented by 1, and the same figure, table, and formula features that meet the conditions are not cumulatively counted; otherwise, no count is made. Until all layout features of the document to be analyzed are inspected, the layout similarity is obtained.
[0054] Preferably, if the document to be analyzed obtained is a factory paper, after obtaining the layout features, jump to step S500, directly perform layout feature matching between the layout features of the paper input by the user and the factory papers in the current existing document set, and the papers with high matching degrees can be fused and calculated with other arbitrary factory paper feature detection methods, and a factory paper risk may be prompted.
[0055] S600. Rearrange the similar sorting results based on the semantic feature similarity and / or layout similarity to obtain the detection result.
[0056] Preferably, different sorting rules are selected according to the instructions input by the user. In this embodiment, the semantic feature similarity of the body text semantic features is used as the main judgment basis: Use the semantic feature similarity as the weight to rearrange the similar sorting results, calculate the repetition ratio of the main logical link sentences and the side branch logical link sentences that are repeated between the natural paragraphs of the document input by the user and the natural paragraphs in the result document, and consider the repetition situation between paragraphs; among them, the weight of the main logical link sentence repetition is greater than the weight of the side branch logical link sentence repetition, the weight of the same paragraph repetition is greater than the weight of the cross-paragraph repetition, and the weight of the continuous repetition is greater than the weight of the scattered repetition.
[0057] Further preferably, when calculating the repetition ratio, one or more methods such as the text ratio, the number of words and numbers, the link repetition ratio, and the link repetition quantity can be used.
[0058] Preferably, the multiple semantic features and layout features constructed by the above methods can be independently applied, or according to the retrieval results of different semantic features, multiple fusion methods can be adopted for semantic feature matching. Further, encapsulate the identified repeated content together with the calculation results, and integrate it into existing other various academic misconduct detection software as the detection result of possible plagiarism behaviors such as rewriting, abbreviating, expanding, and manuscript washing. The content prompted as repeated in the detection result belongs to the key content repetition of the main logical link, or the non-key content repetition of the side branch logical link, or both are repeated, thereby improving the detection effect of existing various academic misconduct detection software.
[0059] In the above embodiments, although the various steps are described in the above order, those skilled in the art can understand that, in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0060] The device for improving the detection effect of academic misconduct documents according to the second embodiment of the present invention, the device includes: An index construction module, configured to construct a multi-dimensional semantic feature index library by using a trained feature extraction model. The multi-dimensional semantic feature index library includes various types of document sets and semantic feature sets. Any semantic feature in the semantic feature set is associated with an original document paragraph ID, an original document ID, and an original text position in various types of document sets through an index; A feature extraction module, configured to obtain a document to be analyzed, and use a trained feature extraction model to extract features from the document to be analyzed, so as to obtain a query semantic feature and a layout feature; A semantic similarity retrieval module, configured to perform semantic similarity retrieval in a preset multi-dimensional semantic feature index library based on the query semantic feature, so as to obtain a similar semantic feature and corresponding similar paragraph IDs and similar document IDs; A semantic similarity calculation module, configured to match the document to be analyzed and the query semantic feature with the similar document IDs and corresponding similar semantic features in the similar sorting result to determine the semantic feature similarity; A layout similarity calculation module, configured to match the layout feature with a preset layout feature library to determine the layout similarity; A result sorting module, configured to respectively perform weighted counting on the similar paragraph IDs and the similar document IDs, and perform a descending order sorting based on the obtained count values to obtain a similar sorting result; and is also configured to rearrange the similar sorting result based on the semantic feature similarity and / or layout similarity to obtain a detection result.
[0061] Those skilled in the art of the present technology can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein.
[0062] It should be noted that the system for detecting academic misconduct documents provided in the above embodiments is only illustrated by dividing the above functional modules. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as improper limitations of the present invention.
[0063] An electronic device according to a third embodiment of the present invention includes: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above method for detecting academic misconduct documents.
[0064] A computer-readable storage medium according to a fourth embodiment of the present invention stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above method for detecting academic misconduct documents.
[0065] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described electronic device and computer-readable storage medium can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0066] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0067] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0069] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or represent a specific order or sequence.
[0070] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or device / apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to those processes, methods, articles, or devices / apparatus.
[0071] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A method for detecting academic misconduct documents, characterized in that: The following steps are involved: S100, obtaining a document to be analyzed, and using a pre-trained feature extraction model to extract features from the document to be analyzed, to obtain query semantic features and format features; The query semantic features include text semantic features, title keyword fusion semantic features and abstract semantic features; S200, based on the query semantic feature, a semantic similarity search is performed in a preset multi-dimensional semantic feature index library to obtain similar semantic features and corresponding similar paragraph IDs and similar document IDs; the multi-dimensional semantic feature index library includes a document set and a semantic feature set, and any semantic feature in the semantic feature set is associated with an original document paragraph ID, an original document ID and an original text position in the document set through an index; S300, performing weighted counting on the similar paragraph IDs and the similar document IDs respectively, and sorting them in descending order based on the obtained count values to obtain a similarity sorting result; S400, matching the document to be analyzed and the query semantic features with the similar document IDs and the corresponding similar semantic features in the similarity ranking results to determine the semantic feature similarity; S500, matching the layout features with a preset layout feature library to determine layout similarity; S600: Rearrange the similarity sorting results based on the semantic feature similarity and / or layout similarity to obtain a detection result.
2. A method for detecting academic misconduct according to claim 1, characterized in that: The pre-trained feature extraction model is used to extract features from the documents to be analyzed. The method is as follows: Conduct structured analysis on the documents to be analyzed, extract the title, abstract, natural paragraphs of the main text, figures and sub-figures, tables, formulas, and references of the documents to be analyzed, and record the layout order and layout area between each structural content; The extracted abstract is divided into five parts: background, purpose, method, result, and conclusion, and vectorized respectively to obtain structured abstract semantic features of different dimensions; Sort the extracted natural paragraphs of the main text and extract the content logic links, the content logic links include main logic link sentences and side logic link sentences, respectively vectorize the main logic link sentences and side logic link sentences to obtain the semantic features of the main text, and record the main logic link sentences, side logic link sentences and original paragraphs of the main text corresponding to the semantic features; The extracted titles and keywords are vectorized separately, and then the title-keyword fusion semantic features are weighted and calculated; The extracted structural contents are parsed to obtain layout features centered on diagrams, tables, and formulas.
3. A method for detecting academic misconduct according to claim 2, characterized in that: Based on the query semantic features, semantic similarity retrieval is performed in the preset multi-dimensional semantic feature index library. The method is as follows: Based on the query semantic feature, similar retrieval and recall are performed in the semantic feature index of the same category in the multi-dimensional semantic feature index library to obtain a number of candidate similar semantic features for each query semantic feature; wherein the categories include text semantic features, title keyword fusion semantic features and abstract semantic features, and similarity thresholds are set for the main logical link sentences and the side branch logical link sentences in the text semantic features respectively; Selecting candidate similar semantic features whose similarity is greater than a preset similarity threshold as similar semantic features; The similar semantic features, the similarity of the similar semantic features, and the original document paragraph ID and the original document ID pointed to by the vector index of the similar semantic features are obtained, and the original document paragraph ID and the original document ID are used as the corresponding similar paragraph ID and similar document ID.
4. A method for detecting academic misconduct according to claim 2, characterized in that: Get the layout features centered on figures, tables, and formulas by: S110, starting from the graph, table, and formula, respectively, traverse the pre-order sequence, and obtain the area of the first pre-order sequence that is the body text as the pre-order text area; S120, starting from the graph, table, and formula, traversing the subsequent sequences respectively, and obtaining the area of the first subsequent sequence that is the main text as the subsequent text area; S130, repeat steps S110-S120 until the layout feature sequences of all figures, tables, and formulas are obtained: < i>FigureID i ={ Figure , SubFigure , PreTxt , NextTxt }; TableID i ={ Table , CellNum , PreTxt , NextTxt }; FormulaID i ={ Formula , LatexNum , PreTxt , NextTxt }; in, FigureID i For the i The layout feature sequence of the graph, Figure , SubFigure , PreTxt , NextTxt Respectively represent the area of the graph, the average area of the subgraph, the area of the preceding text, and the area of the succeeding text; TableID i For the i The layout feature sequence of a table, Table , CellNum , PreTxt , NextTxt They represent the surface area, the number of table cells, the area of the preceding text, and the area of the succeeding text respectively; FormulaID i For the i The layout feature sequence of a table, Formula , LatexNum , PreTxt , NextTxt Respectively represent the formula area, formula latex Number of commands, area of preceding text, area of following text.
5. A method for detecting academic misconduct documents according to claim 3, characterized in that: The similar paragraph IDs and the similar document IDs are weighted and counted respectively, and the method is as follows: Assigning different weights to similar paragraph IDs or similar document IDs according to the type of query semantic features based on which the similar semantic features are retrieved; The categories include text semantic features, title keyword fusion semantic features and abstract semantic features. The main logical link sentences and side branch logical link sentences in the text semantic features are weighted respectively. The number of times each similar paragraph ID or similar document ID is recalled is calculated according to the above weights.
6. A method for detecting academic misconduct documents according to claim 5, characterized in that: The semantic feature similarity includes one or more of the similarity of main logical link sentences, similarity of side logical link sentences, similarity of title keyword fusion semantic features and similarity of abstract semantic features.
7. A method for detecting academic misconduct according to claim 2, characterized in that: The method for determining the layout similarity is: The format features are tested respectively by using difference ratios to determine whether the difference ratios in four dimensions between the format features of the document to be analyzed and the format features of the same type of document to be compared satisfy a preset parameter set at the same time; If the condition is met, the document layout feature matching degree count corresponding to the current layout feature is increased by 1, where the same figure, table, and formula features that meet the condition are not counted; otherwise, they are not counted; Until all the layout features of the document to be analyzed are checked, the layout similarity is obtained.
8. A method for detecting academic misconduct according to claim 6 or 7, characterized in that: The similarity ranking results are rearranged based on the semantic feature similarity, and the method is as follows: The similarity ranking results are rearranged using the semantic feature similarity as the weight, and the repetition between paragraphs is taken into account; among them, the repetition weight of the main logical link sentence is greater than the repetition weight of the branch logical link sentence, the repetition weight in the same paragraph is greater than the repetition weight across paragraphs, and the continuous repetition weight is greater than the dispersed repetition weight.
9. A method for detecting academic misconduct documents according to claim 1, characterized in that: If there is only an independent format feature library generated by the factory paper document collection, after analyzing the documents to obtain the format features, jump to step S500.
10. A device for detecting academic misconduct documents, according to a method for detecting academic misconduct documents according to any one of claims 1 to 9, characterized in that: The device includes: An index construction module is configured to construct a multi-dimensional semantic feature index library using the trained feature extraction model, wherein the multi-dimensional semantic feature index library includes various document sets and semantic feature sets, and any semantic feature in the semantic feature set is associated with an original document paragraph ID, an original document ID, and an original text position in the various document sets through an index; A feature extraction module is configured to obtain the document to be analyzed, and use the trained feature extraction model to extract features from the document to be analyzed, so as to obtain query semantic features and format features; A semantic similarity retrieval module is configured to perform semantic similarity retrieval in a preset multi-dimensional semantic feature index library based on the query semantic feature, and obtain similar semantic features and corresponding similar paragraph IDs and similar document IDs; A semantic similarity calculation module is configured to match the document to be analyzed and the query semantic features with the similar document IDs and the corresponding similar semantic features in the similarity ranking results to determine the semantic feature similarity; A layout similarity calculation module, configured to determine layout similarity based on matching the layout features with a preset layout feature library; The result sorting module is configured to perform weighted counting on the similar paragraph IDs and the similar document IDs respectively, and to perform descending sorting based on the obtained count values to obtain similarity sorting results; and is also configured to rearrange the similarity sorting results based on semantic feature similarity and / or layout similarity to obtain detection results.
Citation Information
Patent Citations
Method and device for determining literature similarity based on semantic analysis
CN114580557A
Long text retrieval method and device based on knowledge system
CN115248839A
Anti-plagiarism method and system based on AI
CN118333032A
Pathological literature searching and dialogue system based on large language model and RAG technology
CN118643128A
Article recall sorting method, intelligent question and answer processing method and computing equipment
CN119441417A
Cited By
Multi-mode image searching method based on images
CN120804355A
Factory paper detection method, system and equipment based on migration characteristics
CN121189303A
A plant paper detection method, system and device based on migration features
CN121189303B
Scientific research subject similarity examination method, device, equipment and medium
CN121859025A