Document plagiarism judgment method and system based on semantic vector library and large language model

By building a semantic vector library and using a large language model, combining semantic similarity and text overlap to judge document plagiarism, the problem of difficult to identify deep plagiarism in the existing technology is solved, and more accurate and intelligent plagiarism detection is achieved.

CN120234420AInactive Publication Date: 2025-07-01TRS INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510703998.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify deep-level document plagiarism, especially to circumvent detection by replacing synonyms and adjusting sentence structure.

Method used

Using a method based on semantic vector library and large language model, we use the semantic vector library to build a semantic vector, obtain the semantic vectors of the search text, filter similar documents, and combine semantic similarity, text overlap and large language model to determine whether there is plagiarism and generate a judgment basis.

Benefits of technology

It improves the accuracy and intelligence of document comparison, provides objective and interpretable basis for plagiarism judgment, and reduces misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234420A_ABST
    Figure CN120234420A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, provides a document plagiarism judgment method and system based on a semantic vector library and a large language model, and aims to map each split retrieval text into semantic vectors through a semantic vectorization model, effectively capture deep semantic information of the text and provide a data basis for the subsequent process. The comparison semantic vector matched with the retrieval semantic vector of each retrieval text is retrieved from the semantic vector library, similar texts are accurately screened, and the document comparison range is narrowed. And enabling each retrieval text to be in one-to-one correspondence with each comparison text of the similar document to form each similar text pair. Whether plagiarism exists between similar text pairs or not is comprehensively judged on the basis of three dimensions of semantic similarity, text coincidence degree and large language model judgment, judgment basis is generated, semantic comparison accuracy is improved, objective and explainable basis is provided for judgment, and the intelligent level of scientific and technological novelty checking and thesis duplicate checking is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method and system for judging document plagiarism based on a semantic vector library and a large language model. Background Art

[0002] Sci-tech novelty search is an important link in scientific and technological research and academic research, aiming to judge the novelty and innovation of technologies or achievements.

[0003] Currently, it mainly relies on the text overlap degree to judge the similarity between documents, and can perform duplicate checking through means such as keyword matching and sentence structure analysis. However, document writers can avoid detection by replacing synonyms, adjusting sentence structures, etc., making it difficult to effectively identify deep-level plagiarism behaviors. Therefore, how to perform document semantic comparison based on a semantic vector library and a large language model to detect whether there is plagiarism by the writer is a very important issue.

[0004] Based on this, this specification provides a method and system for judging document plagiarism based on a semantic vector library and a large language model. Summary of the Invention

[0005] To solve the problem that traditional duplicate checking methods are difficult to effectively identify deep-level plagiarism behaviors, the present invention proposes a method and system for judging document plagiarism based on a semantic vector library and a large language model. By retrieving comparison documents similar to the retrieved document from the semantic vector library and comprehensively judging whether there is plagiarism in the retrieved document based on three dimensions: semantic similarity, text overlap degree, and large language model judgment, it not only improves the accuracy of semantic comparison, but also provides an objective and interpretable basis for the determination of academic misconduct, significantly improving the intelligent level of sci-tech novelty search and paper duplicate checking.

[0006] This specification provides a method for judging document plagiarism based on a semantic vector library and a large language model, including: S1: Construct a semantic vector library: Preprocess a preset set of comparison documents to obtain comparison semantic vectors of each comparison text, and construct a semantic vector library; S2: Construct retrieval semantic vectors based on the retrieved document: Obtain the retrieved document and split it into each retrieval text; Use a semantic vectorization model to determine the retrieval semantic vector of each retrieval text; S3: Screen similar documents: From the semantic vector library, determine the comparison semantic vectors that match each retrieval semantic vector to generate the retrieval results of each retrieval text, and sort and screen each comparison document in the comparison document set according to the retrieval results to obtain similar documents; S4: Determine whether there is plagiarism based on a large language model and generate a judgment basis: Corresponding each of the retrieved texts with each comparison text of the similar document one by one to form each pair of similar texts; Determine whether the semantic similarity between the retrieved text and the comparison text in each pair of similar texts exceeds a first threshold. If it exceeds the first threshold, continue to determine whether the text overlap between the retrieved text and the comparison text in each pair of similar texts exceeds a second threshold. If it does not exceed the second threshold, splice the retrieved text and the comparison text in each pair of similar texts and supplement them into the prompt template. After obtaining the complete prompt, input it into the large language model to determine whether there is a suspicion of plagiarism and generate a judgment basis.

[0007] Optionally, the method further includes: If it does not exceed the first threshold, determine that the judgment result of each pair of similar texts is not suspected of plagiarism, and determine that the judgment basis of each pair of similar texts is that the semantic similarity does not exceed the first threshold; If it exceeds the second threshold, determine that the judgment result of each pair of similar texts is suspected of plagiarism, and determine that the judgment basis of each pair of similar texts is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.

[0008] Optionally, the method further includes: Generate a retrieval report corresponding to the retrieved document and the similar document according to the judgment result and judgment basis of whether each pair of similar texts is suspected of plagiarism; Determine a final retrieval report according to the retrieval report and display it.

[0009] Optionally, the specific steps of splitting the retrieved document into each retrieved text in S2 include: Perform text extraction on the retrieved document to obtain the text data corresponding to the retrieved document; According to the chapter hierarchy structure of the retrieved document, split the text data to obtain each chapter data; Filter each chapter data according to the preset filtering rules to obtain each target chapter data; According to the granularity requirement corresponding to the retrieved document, use natural language processing technology to split each target chapter data to obtain each retrieved text.

[0010] Optionally, S1 specifically includes: Obtain a set of comparison documents including each comparison document uploaded, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; Use a semantic vectorization model to determine the comparison semantic vectors corresponding to each comparison text of each comparison document; Generate labels corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein, the label includes comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector.

[0011] Optionally, in step S4, the one-to-one correspondence between each retrieved text and each comparison text of the similar document to form each similar text pair specifically includes: According to the retrieval semantic vector of each retrieved text and the comparison semantic vector of each comparison text of the similar document, determine the comparison text that matches each retrieved text from each comparison text of the similar document; Form each similar text pair by combining each retrieved text with the determined comparison text.

[0012] Optionally, in step S4, splicing the retrieved text and the comparison text in each similar text pair and supplementing them into the prompt template to obtain a complete prompt specifically includes: Through a preset prompt template, splice the preset comparison rules, the retrieved text and the comparison text in each similar text pair to obtain a complete prompt.

[0013] Optionally, the large language model in step S4 is one of a general large language model and a fine-tuned general large language model, and the fine-tuned general large language model is obtained by fine-tuning the general large language model based on training text pairs and label results.

[0014] Optionally, in step S3, determining the comparison semantic vector that matches each retrieval semantic vector from the semantic vector library specifically includes: According to a preset relationship field, determine the retrieval information corresponding to each retrieval semantic vector, and filter the comparison semantic vectors in the semantic vector library according to the retrieval information; wherein, the relationship field is at least one of a document title, a document type, and a chapter type; Determine the comparison semantic vector that matches each retrieval semantic vector from the filtered comparison semantic vectors.

[0015] This specification also provides a document plagiarism judgment system based on a semantic vector library and a large language model. The system includes a data storage module, a data preprocessing module, a semantic vectorization module, a vector retrieval module, and a document comparison module, wherein: The data preprocessing module is used to preprocess the comparison document set to obtain the comparison semantic vectors of each comparison text, and split the retrieved document into each retrieved text; The semantic vectorization module is used to call a semantic vectorization model to determine the retrieval semantic vector of each retrieved text; The data storage module is used to construct and store a semantic vector library based on the comparison semantic vectors; The vector retrieval module is used to determine, from the semantic vector library, the comparison semantic vectors that match each retrieval semantic vector, so as to generate each retrieval result, and sort and filter each comparison document in the comparison document set according to the retrieval results to obtain similar documents; The document comparison module is used to pair each retrieval text with each comparison text of the similar documents one by one to form each similar text pair; determine whether the semantic similarity between the texts in each similar text pair exceeds a first threshold. If it exceeds the first threshold, determine whether the text overlap degree between the texts in each similar text pair exceeds a second threshold. If it does not exceed the second threshold, splice the texts in each similar text pair and supplement them into the prompt template, and after obtaining the completed prompt, input it into the large language model to determine whether there is a suspicion of plagiarism and generate a judgment basis.

[0016] Optionally, the document comparison module is further used to, if it does not exceed the first threshold, determine that the judgment result of each similar text pair is not suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold; if it exceeds the second threshold, determine that the judgment result of each similar text pair is suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold.

[0017] Optionally, the system further includes a report generation module; The report generation module is used to generate a retrieval report corresponding to the retrieval document and the similar documents according to the judgment results and judgment bases of whether each similar text pair is suspected of plagiarism; determine a final retrieval report according to the retrieval report and display it.

[0018] Optionally, the data preprocessing module is specifically used to perform text extraction on the retrieval document to obtain the text data corresponding to the retrieval document; split the text data according to the chapter hierarchy structure of the retrieval document to obtain each chapter data; filter each chapter data according to preset filtering rules to obtain each target chapter data; split each target chapter data through natural language processing technology according to the granularity requirement corresponding to the retrieval document to obtain each retrieval text.

[0019] Optionally, the data preprocessing module is further used to store the semantic vectorization model, the large language model, the retrieval document, and the comparison document set; The data preprocessing module is specifically configured to obtain a set of comparison documents including each comparison document uploaded, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; The semantic vectorization module is further configured to use the semantic vectorization model to determine comparison semantic vectors respectively corresponding to each comparison text of each comparison document; The data storage module is specifically configured to generate labels corresponding to each comparison semantic vector, and construct a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein, the label includes comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector.

[0020] Optionally, the document comparison module is specifically configured to determine, from each comparison text of the similar documents, a comparison text that matches each retrieval text according to the retrieval semantic vectors of the retrieval texts and the comparison semantic vectors of each comparison text of the similar documents; and form each retrieval text and the determined comparison text into each similar text pair.

[0021] Optionally, the document comparison module is specifically configured to splice a preset comparison rule, the retrieval text and the comparison text in each similar text pair through a preset prompt template to obtain a complete prompt word.

[0022] Optionally, the large language model is one of a general large language model and a fine-tuned general large language model, and the fine-tuned general large language model is obtained by fine-tuning the general large language model based on training text pairs and label results.

[0023] Optionally, the vector retrieval module is specifically configured to determine retrieval information corresponding to each retrieval semantic vector according to a preset relationship field, and screen the comparison semantic vectors in the semantic vector library according to the retrieval information; wherein, the relationship field is at least one of a document title, a document type, and a chapter type; and determine, from the screened comparison semantic vectors, comparison semantic vectors that match each retrieval semantic vector.

[0024] The above at least one technical solution adopted in this specification can achieve the following beneficial effects: The document plagiarism judgment method based on the semantic vector library and the large language model provided in this specification first preprocesses the preset comparison document set to obtain the comparison semantic vectors of each comparison text, and constructs a semantic vector library. The retrieved document is obtained and split into each retrieved text. A semantic vectorization model is used to determine the retrieved semantic vector of each retrieved text. By mapping each retrieved text split out by the semantic vectorization model into a high-dimensional semantic vector, the deep semantic information of the text can be effectively captured, providing a data basis for subsequent vector retrieval and semantic-level similarity detection.

[0025] Then, from the semantic vector library, the comparison semantic vectors that match each retrieved semantic vector are determined to generate the retrieval results of each retrieved text, and based on each retrieval result, each comparison document in the comparison document set is sorted and screened to obtain similar documents. By retrieving the comparison semantic vectors that match the retrieved semantic vectors of the retrieved text from the semantic vector library, and sorting and screening each comparison document according to the retrieval results to obtain similar documents, similar texts similar to the retrieved document can be accurately screened out, narrowing the scope of document comparison and improving the speed of document comparison.

[0026] After that, each retrieved text is corresponded to each comparison text of the similar document one by one to form each similar text pair. It is judged whether the semantic similarity between the retrieved text and the comparison text in each similar text pair exceeds the first threshold. If the semantic similarity exceeds the first threshold, it is continued to judge whether the text overlap degree between the retrieved text and the comparison text in each similar text pair exceeds the second threshold. If the text overlap degree does not exceed the second threshold, the retrieved text and the comparison text in each similar text pair are spliced and supplemented into the prompt template. After obtaining the complete prompt, it is input into the large language model to determine whether there is suspicion of plagiarism and generate a judgment basis. Judging whether the retrieved document has plagiarism based on three dimensions: semantic similarity, text overlap degree, and large language model judgment not only improves the accuracy of semantic comparison, but also provides an objective and interpretable basis for the determination of academic misconduct, significantly improving the intelligent level of novelty search and paper plagiarism checking.

[0027] In the present invention, data preprocessing is performed on the retrieved document and each comparison document to perform text extraction, chapter extraction, chapter filtering, and splitting according to the granularity requirement, filter out useless data, and eliminate irrelevant content, improving the processing efficiency of subsequent document comparison. Moreover, the retrieved text is split into each retrieved text to respectively retrieve the comparison semantic vectors that match the retrieved semantic vectors of each retrieved text, and the granularity of the retrieved text and the comparison text is the same, avoiding the problem of window sliding denoising due to inconsistent granularity. When performing vector retrieval, the vector retrieval method of the semantic vector library can be relied on, combined with relationship fields, that is, document title, document type, and chapter type, to retrieve each retrieved text to obtain the corresponding retrieval result.

[0028] In the present invention, during document comparison, first, according to each retrieval semantic vector and the comparison semantic vectors of each comparison text of the similar documents, the comparison text that matches each retrieval text is determined from each comparison text, and each retrieval text and the determined comparison text of each are combined to form each similar text pair. Then, the semantic similarity between the similar text pairs is compared, and when the semantic similarity does not exceed the first threshold, the judgment result is directly determined to be not suspected of plagiarism and the judgment basis is generated. When the semantic similarity exceeds the first threshold, the text overlap degree between the similar text pairs is compared, and when the text overlap degree exceeds the second threshold, the judgment result is determined to be suspected of plagiarism and the judgment basis is generated. Additionally, when the semantic similarity exceeds the first threshold and the text overlap degree does not exceed the second threshold, the large language model is used to judge whether the similar text pairs with high semantic similarity and low text overlap degree are suspected of plagiarism, so as to solve the problem that it is difficult to intuitively judge whether the text pairs with high semantic similarity but lack significant overlap in literal meaning are suspected of plagiarism. Compared with the traditional method of only judging suspected plagiarism based on high semantic similarity, the present invention can continue to judge based on the large language model when the semantic similarity is high, so as to judge the situation where the semantic similarity is high but it is judged not to be suspected of plagiarism, thereby avoiding misjudgment due to only judging based on semantic similarity.

[0029] In the present invention, through the prompt template, the comparison rule and the similar text are spliced and input into the large language model to determine the judgment result and judgment basis of the similar text pair. Based on the semantic analysis and understanding ability of the large language model and its own massive knowledge, the large language model is driven by the prompt to judge whether it is suspected of plagiarism and generate the judgment basis, and the comparison rule is input into the large language model together as knowledge data, so that the large language model can judge and generate the judgment basis according to the comparison rule and its own massive knowledge, improving the accuracy of the output result and simplifying the subsequent manual work.

[0030] In the present invention, based on the decision tree, it is judged whether there is plagiarism in the retrieved document in turn according to three dimensions: semantic similarity, text overlap degree, and large language model judgment, so as to avoid the problem of misjudgment caused by using a single dimension for judgment, and generate the judgment basis corresponding to the judgment result to provide the judgment basis for the novelty search personnel and writing personnel to determine. This not only improves the accuracy of document comparison but also increases the interpretability. Additionally, according to the judgment result and judgment basis of whether each similar text pair is suspected of plagiarism, a retrieval report can be generated, and based on the retrieval report, the final retrieval report can be determined and displayed to provide it for the novelty search personnel and writing personnel to view.

[0031] The large language model in the present invention can be a general large language model or a fine-tuned general large language model. The fine-tuned general large language model is a large language model obtained by fine-tuning the general large language model. Compared with the traditional training logic, methods such as fine-tuning the general large language model make the fine-tuned large language model more suitable for the semantic comparison application scenario, and the output judgment results and judgment bases are more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings described herein are used to provide a further understanding of the present specification and form a part of the present specification. The illustrative embodiments of the present specification and their descriptions are used to explain the present specification and do not constitute an improper limitation of the present specification. In the drawings: Figure 1 is a schematic flow chart of a document plagiarism judgment method based on a semantic vector library and a large language model provided in the present specification; Figure 2 is a schematic flow chart of a semantic vector library construction process provided in the present specification; Figure 3 is a schematic diagram of a document semantic retrieval and comparison process provided in the present specification; Figure 4 is a schematic diagram of a document comparison process assisted by a large language model provided in the present specification; Figure 5 is a schematic diagram of a plagiarism determination process assisted by a large language model provided in the present specification Figure 6 is a schematic diagram of a document plagiarism judgment system based on a semantic vector library and a large language model provided in the present specification; Figure 7 is a schematic diagram of a document plagiarism judgment system architecture based on a semantic vector library and a large language model provided in the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The present specification provides a document plagiarism judgment method and system based on a semantic vector library and a large language model. The following describes in detail the technical solutions provided by each embodiment of the present specification with reference to the drawings.

[0034] Figure 1 is a schematic flow chart of a document plagiarism judgment method based on a semantic vector library and a large language model provided in the present specification. As Figure 1 shown, it specifically includes the following steps: S1: Construct a semantic vector library: Preprocess a preset set of comparison documents to obtain comparison semantic vectors of each comparison text, and construct a semantic vector library.

[0035] In this specification, the device for document comparison can build a semantic vector library, that is, preprocess a preset set of comparison documents to obtain the comparison semantic vectors of each comparison text, and build a semantic vector library. Among them, the device for document comparison can be a server, or a system or a certain module in the system, or an electronic device such as a desktop computer or a laptop computer. For the convenience of description, hereinafter, only the server is used as the execution subject to illustrate the document plagiarism determination method provided in this specification based on the semantic vector library and the large language model. Specifically, the server can preprocess the comparison documents in the set of comparison documents compared with the retrieved document to obtain the comparison semantic vectors of each comparison text, and build a semantic vector library. The above set of comparison documents includes each comparison document, each comparison document corresponds to each comparison text, and each comparison text has a corresponding comparison semantic vector. Each comparison text can be a document uploaded by the user, and the comparison document is a document used to detect whether the retrieved document has plagiarism. The comparison document can be in any format, such as documents in PDF, Word and other formats. The comparison document can be a thesis, a project report, etc. The above comparison semantic vector is a high-dimensional vector obtained by mapping the comparison text using a semantic vectorization model. The above preprocessing includes text extraction, chapter extraction, chapter filtering, and splitting according to granularity requirements. The above semantic vector library at least includes the comparison semantic vectors corresponding to each comparison text of each pre-uploaded comparison document.

[0036] Furthermore, the server can obtain a set of comparison documents including each comparison document, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document. Using a semantic vectorization model, determine the comparison semantic vectors respectively corresponding to each comparison text of each comparison document. Generate labels corresponding to each comparison semantic vector, and build a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector. Among them, the process of performing data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document is similar to the process of performing data preprocessing on the retrieved document to obtain each retrieved text corresponding to the retrieved document, which will not be elaborated here. When using the semantic vectorization model to determine the comparison semantic vectors respectively corresponding to each comparison text of each comparison document, the server can, for each comparison document and for each comparison text of the comparison document, input the comparison text into the semantic vectorization model to output the comparison semantic vector corresponding to the comparison text through the semantic vectorization model. The comparison semantic vector is a high-dimensional dense vector.

[0037] In addition to including each comparison semantic vector, the above semantic vector library also includes the tags corresponding to each comparison semantic vector. The tag includes the comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector. The comparison document information can be the name or identifier of the comparison document to which the comparison semantic vector belongs, etc. The chapter information can be the information of the chapter to which the comparison text corresponding to the comparison semantic vector belongs. The chapter information includes the chapter name, identifier, keywords, etc. The comparison text information can include the comparison text corresponding to the comparison semantic vector and the page number of the page where the comparison text is located. By setting tags for each comparison semantic vector, it is possible to quickly and accurately match the comparison semantic vector during subsequent vector retrieval. For example, there are comparison documents 1 to 3. Among them, comparison document 1 includes comparison texts A to D. The tag of the comparison semantic vector a corresponding to comparison text A can include comparison document information, chapter information, and comparison text information. Assuming that the comparison document information is the name of comparison document 1, that is, comparison document 1, the chapter information is references, and the comparison text information is comparison text A and page 3. Therefore, the tag of comparison semantic vector a can be expressed in the form of "comparison document information - chapter information - comparison text information", that is, "comparison document 1 - references - comparison text A and page 3".

[0038] S2: Construct retrieval semantic vectors based on the retrieval document: Obtain the retrieval document and split it into each retrieval text; Use the semantic vectorization model to determine the retrieval semantic vector of each retrieval text.

[0039] In this specification, the server can construct retrieval semantic vectors based on the retrieval document, that is, obtain the retrieval document and split it into each retrieval text. Then use the semantic vectorization model to determine the retrieval semantic vector of each retrieval text. Among them, the above retrieval document can be a document uploaded by the user, and the retrieval document is a document that needs to be detected for plagiarism. The retrieval document can be in any format, such as documents in PDF, Word, etc. The retrieval document can also be a thesis, a project report, etc. The above splitting of the retrieval document into each retrieval document is actually data preprocessing of the retrieval document to obtain each retrieval text corresponding to the retrieval document. Specifically, the server can perform text extraction on the retrieval document to obtain the text data corresponding to the retrieval document. According to the chapter hierarchy structure of the retrieval document, the text data is split to obtain each chapter data. According to the preset filtering rules, each chapter data is filtered to obtain each target chapter data. According to the granularity requirement corresponding to the retrieval document, through natural language processing technology, each target chapter data is split to obtain each retrieval text.

[0040] Among them, since the above-mentioned retrieval document can be in any format, the above text extraction from the retrieval document is to extract the pure text content to remove format interference. The extracted text data includes the chapter hierarchy structure, header page numbers, and all text content of the retrieval document. The method of text extraction from the retrieval document can be any existing algorithm or model, which is not specifically limited in this specification. According to the chapter hierarchy structure of the retrieval document, the above text data is split to split the text data into a tree structure of chapters and sub-chapters, and each chapter data is obtained. At this time, each chapter data no longer includes the above header page numbers and other data, but only includes the data corresponding to each chapter in the retrieval document to avoid the interference of useless header page numbers and other data. The above filtering rules are preset. Through these filtering rules, the chapter data that meets the filtering rules is determined and filtered out to obtain the target chapter data. Specifically, the filtering rules can be used to identify the chapter data in the retrieval document through preset keywords, titles, or format features, and filter out the chapter data that meets the above-set keywords, titles, or format features to obtain the target chapter data. For example, the above keywords can be "papers" and "literatures", the title can be "References", and the format features can be "bold" and "font size 2". Through the above keywords, titles, or format features, paragraphs that do not participate in the retrieval, such as journal information and organizational structures, can be filtered out, and the chapters related to the retrieval target can be screened out, that is, the target chapter data, to eliminate irrelevant content and improve the efficiency of subsequent processing. It should be noted that the above filtering rules are only an example, and this specification does not limit the specific content of the filtering rules, which can be set according to your own needs.

[0041] The above granularity requirements can be preset. The granularity requirements can be paragraphs or sentences. The server can identify the paragraph boundaries and sentence segmentation points in each target chapter data through natural language processing technology according to the granularity requirements corresponding to the retrieval document, so as to split each target chapter data to obtain each retrieval text, ensuring that each retrieved text split out is independent and complete in terms of paragraphs or sentences, which is convenient for subsequent semantic analysis.

[0042] The above semantic vectorization model is a dense vectorization model. The above retrieval semantic vectors are high-dimensional dense vectors. This semantic vectorization model captures the semantic information of the text by mapping the text into high-dimensional vectors. This semantic vectorization model can effectively represent the semantic features of the text and provides a data basis for semantic similarity detection and document comparison. The semantic vectorization model can be a general model such as BGE, BERT, GPT, etc. for mapping text into high-dimensional vectors. Of course, it can also be a model fine-tuned from the above general models, which is not specifically limited in this specification. Specifically, the server can input each retrieval text into the semantic vectorization model for each retrieval text to output the retrieval semantic vector corresponding to the retrieval text through the semantic vectorization model.

[0043] S3: Screening Similar Documents: From the semantic vector library, determine the comparison semantic vectors that match each retrieved semantic vector to generate the retrieval results for each retrieved text, and sort and screen each comparison document in the comparison document set according to the respective retrieval results to obtain similar documents.

[0044] In this specification, the server can screen similar documents, that is, from the semantic vector library, determine the comparison semantic vectors that match each retrieved semantic vector to generate the retrieval results for each retrieved text, and sort and screen each comparison document in the comparison document set according to the respective retrieval results to obtain similar documents. Among them, the above retrieval results can be the similarity between the retrieved semantic vector and the determined comparison semantic vector that matches the retrieved semantic vector, can also be the comparison semantic vector that matches the retrieved semantic vector, or can be the comparison semantic vector that matches the retrieved semantic vector and the similarity between the retrieved semantic vector and the comparison semantic vector that matches the retrieved semantic vector.

[0045] When determining the comparison semantic vectors that match each retrieved semantic vector from the semantic vector library to generate the retrieval results corresponding to each retrieved text, the server can calculate the similarity between each retrieved semantic vector and each comparison semantic vector in the semantic vector library respectively, and use the comparison semantic vector corresponding to the maximum similarity as the comparison semantic vector that matches the retrieved semantic vector. The server can use the maximum similarity as the retrieval result of the retrieved text corresponding to the retrieved semantic vector, can also use the comparison semantic vector corresponding to the maximum similarity as the retrieval result of the retrieved text corresponding to the retrieved semantic vector, or can use the maximum similarity and its corresponding comparison semantic vector as the retrieval result of the retrieved text corresponding to the retrieved semantic vector. When calculating the similarity between the retrieved semantic vector and each comparison semantic vector in the semantic vector library respectively, the cosine similarity or Euclidean distance between the retrieved semantic vector and each comparison semantic vector in the pre-constructed semantic vector library can be calculated respectively and used as the similarity.

[0046] In addition, when determining the comparison semantic vectors that match each retrieved semantic vector from the semantic vector library, the server can use the vector space model or indexing technology to determine the comparison semantic vectors that match each retrieved semantic vector from the semantic vector library to quickly match the comparison semantic vectors that match each retrieved semantic vector.

[0047] When sorting and filtering each comparison document in the comparison document set according to each retrieval result to obtain similar documents, the server can determine the similarity score corresponding to each comparison document in the comparison document set according to each similarity included in each retrieval result. Sort and filter each comparison document according to each similarity score to determine similar documents. Among them, the retrieval result is the similarity between the retrieval semantic vector of the retrieval text and the comparison semantic vector matching the retrieval semantic vector of the retrieval document. The above similarity score is the sum of the similarities corresponding to each comparison text in the comparison document. The number of the above similar texts can be preset.

[0048] Specifically, the server can, according to each similarity included in each retrieval result, for each comparison document, determine the first score corresponding to each comparison text in the comparison document, and determine the sum of each first score, and use it as the similarity score of the comparison document. Sort each comparison document in descending order according to each similarity score to obtain a document sequence. Determine a specified number of comparison documents in the order of the document sequence and use them as similar documents. Among them, the above first score is determined according to each similarity included in each of the above retrieval results. When the comparison text matches the retrieval text, there is a similarity between the retrieval result corresponding to the retrieval text and the comparison text, so this similarity is directly used as the first score of the comparison text. However, when the comparison text does not match any of the above retrieval texts, there is no similarity between any of the above retrieval results and the comparison text, so the first score of the comparison text can be directly zero. The above specified number is a preset value, such as 1 or 10. The above specified number is the number of determined similar documents.

[0049] S4: Judging whether there is plagiarism based on the large language model and generating a judgment basis: Corresponding each of the above retrieval texts to each comparison text of the similar documents one by one to form each similar text pair; judging whether the semantic similarity between the retrieval text and the comparison text in each similar text pair exceeds a first threshold. If it exceeds the first threshold, continue to judge whether the text overlap degree between the retrieval text and the comparison text in each similar text pair exceeds a second threshold. If it does not exceed the second threshold, splice the retrieval text and the comparison text in each similar text pair and supplement them into the prompt template. After obtaining the complete prompt, input it into the large language model to determine whether there is suspected plagiarism and generate a judgment basis.

[0050] In this specification, the server can determine whether there is plagiarism based on a large language model and generate a basis for judgment, that is, each retrieved text is corresponded to each comparison text of the similar documents one by one to form each similar text pair. It is determined whether the semantic similarity between the retrieved text and the comparison text in each similar text pair exceeds a first threshold. If the semantic similarity exceeds the first threshold, it is further determined whether the text overlap between the retrieved text and the comparison text in each similar text pair exceeds a second threshold. If the text overlap does not exceed the second threshold, the retrieved text and the comparison text in each similar text pair are spliced and supplemented into the prompt template. After obtaining the complete prompt, it is input into the large language model to determine whether there is a suspicion of plagiarism and generate a basis for judgment. Among them, each similar text pair can be determined from both the retrieved document and the similar documents. Each similar text pair includes a retrieved text and a comparison text. The retrieved text is all from the retrieved document, and the comparison text is all from the similar document. In addition, when there are multiple similar documents, the server can, for each similar document, determine each similar text pair between the similar text and the retrieved document according to the respective comparison semantic vectors of the similar document and the respective retrieved semantic vectors of the retrieved document.

[0051] When the above-mentioned each retrieved text is corresponded to each comparison text of the similar documents one by one to form each similar text pair, the server can, according to the retrieved semantic vectors of each retrieved text and the comparison semantic vectors of each comparison text of the similar document, determine the comparison text that matches each retrieved text from each comparison text of the similar document, and form each similar text pair by combining each retrieved text with each determined comparison text.

[0052] Further, the server may determine a first vector set composed of the respective retrieval semantic vectors of the retrieved documents, and determine a second vector set composed of the respective comparison semantic vectors of the similar documents. Determine the Cartesian product between the first vector set and each second vector set. Determine the similarity between the vectors in each vector combination in the Cartesian product. According to the similarity between the vectors in each vector combination in the Cartesian product, from the respective comparison texts of the similar documents, determine the comparison text similar to each retrieval text, and use each retrieval text and each determined comparison text as a similar text pair. Among them, the above Cartesian product includes all possible vector combinations between the first vector set and the second vector set. For example, if the first vector set includes retrieval semantic vectors 1 to 3 and the second vector set includes comparison semantic vectors 4 to 5, then all possible vector combinations between the first vector set and the second vector set, that is, the Cartesian product, are retrieval semantic vector 1 and comparison semantic vector 4, retrieval semantic vector 1 and comparison semantic vector 5, retrieval semantic vector 2 and comparison semantic vector 4, retrieval semantic vector 2 and comparison semantic vector 5, retrieval semantic vector 3 and comparison semantic vector 4, and retrieval semantic vector 3 and comparison semantic vector 5. The above similarity may be cosine similarity or Euclidean distance. The comparison text in each of the above similar text pairs is the comparison text in the similar document with the greatest similarity to the retrieval text in each similar text pair. That is, the similarity between the retrieval text and the comparison text in each similar text pair in the similar document is the greatest, and the comparison text is also one of the texts in the similar document.

[0053] In addition, when there are multiple similar documents, each similar document has a corresponding second vector set, and each second vector set and the first vector set have a corresponding Cartesian product, which includes all possible vector combinations between each second vector set and the first vector set. Specifically, the server may determine a first vector set composed of the respective retrieval semantic vectors of the retrieved documents. For each similar document, determine a second vector set composed of the respective comparison semantic vectors of the similar document. Determine the Cartesian product between the first vector set and the second vector set of the similar document. Determine the similarity between the vectors in each vector combination in the Cartesian product corresponding to the similar document. According to the similarity between the vectors in each vector combination in the Cartesian product corresponding to the similar document, from the respective comparison texts of the similar document, determine the comparison text similar to each retrieval text, and use each retrieval text and each determined comparison text as the similar text pair corresponding to the similar document.

[0054] In this specification, in addition to constructing a first vector set based on each retrieval semantic vector, constructing a second vector set based on each comparison semantic vector, and then determining similar text pairs based on the first vector set and the second vector set, the server can also construct a first vector library based on each retrieval semantic vector, construct a second vector library based on each comparison semantic vector, and then determine similar text pairs based on the first vector library and the second vector library to improve efficiency.

[0055] In this specification, there is a risk of misjudgment when solely relying on the semantic vector similarity for plagiarism judgment. For example, there are a large number of fixed sentence paragraphs (such as experimental steps, background introductions) in the document, making it difficult to distinguish between structural similarity and plagiarism, or the semantically similar texts may lack significant literal overlap, making it difficult to intuitively judge whether there is a plagiarism relationship. Therefore, to overcome the above technical defects, the server can determine the semantic similarity between the retrieval text and the comparison text in each similar text pair. When the semantic similarity of each similar text pair exceeds the first threshold, determine the text overlap degree between the retrieval text and the comparison text in each similar text pair. When the text overlap degree of each similar text pair does not exceed the second threshold, splice the retrieval text and the comparison text in each similar text pair and supplement them into the prompt template, and after obtaining the complete prompt, input it into the large language model to obtain the judgment result and judgment basis of each similar text pair output by the large language model. Among them, the semantic similarity is the cosine similarity or Euclidean distance between the retrieval semantic vector corresponding to the retrieval text and the comparison semantic vector corresponding to the comparison text in the similar text pair. The first threshold and the second threshold are respectively preset values. The above text overlap degree is the literal overlap degree between the retrieval text and the comparison text in the similar text pair, and this text overlap degree can be determined by a preset algorithm or model, which is not specifically limited in this specification.

[0056] When splicing the retrieved text and the comparison text in each similar text pair and supplementing them into the prompt template to obtain a complete prompt, the server can splice the preset comparison rules, the retrieved text and the comparison text in each similar text pair through the preset prompt template to obtain a complete prompt. Among them, the above-mentioned prompt template is pre-set. The prompt template includes fixed text and each card slot. Each card slot is used to fill in the retrieved text, the comparison text and the comparison rules in the similar text pair. The comparison rules are the principles of scientific and technological novelty search and document comparison, and can also be the knowledge data of several judgment bases for plagiarism and judgment cases. The comparison rules can be pre-set. The fixed text can be pre-set content used to make the comparison rules, the retrieved text and the comparison text in the similar text pair continuous. The above-mentioned prompt template is used to splice the comparison rules, the retrieved text and the comparison text in the similar text pair. The above-mentioned prompt template can be "Adopt [Card Slot 1] to determine whether [Card Slot 2] is suspected of plagiarizing [Card Slot 3], and generate a judgment result representing whether it is suspected of plagiarism and the judgment basis corresponding to the judgment result:". In the prompt template, "[Card Slot 1]" is used to fill in the comparison rules, "[Card Slot 2]" is used to fill in the retrieved text, and "[Card Slot 3]" is used to fill in the comparison text. Therefore, the above-mentioned complete prompt can be "Adopt [Comparison Rule] to determine whether [Retrieved Text] is suspected of plagiarizing [Comparison Text], and generate a judgment result representing whether it is suspected of plagiarism and the judgment basis corresponding to the judgment result:".

[0057] The above-mentioned large language model (LLM) is a neural network model based on deep learning, with powerful semantic understanding and analysis capabilities. The above-mentioned large language model can be any existing general large language model, or a large language model obtained by fine-tuning a general large language model, that is, a fine-tuned general large language model. This specification does not make specific limitations. The above-mentioned large language model is used to determine whether the retrieved text in the complete prompt word is suspected of plagiarizing the comparison text based on the complete prompt word, and further generate a judgment basis. The judgment result determined by the above-mentioned large language model can be one of suspected plagiarism and not suspected plagiarism, and the judgment basis is generated by the large language model. The above-mentioned fine-tuned general large language model is obtained by fine-tuning and training the general large language model based on the training text pairs and label results. During fine-tuning training, the server can first obtain the general large language model, then obtain the training text pairs used to train the general large language model, splice the first text and the second text in the training text pairs to obtain a third text, input the third text into the general large language model to obtain an output result, and the output result includes a judgment result and a judgment basis. According to the label result and the output result of the training text pair, the general large language model is fine-tuned to use the trained general large language model (i.e., the fine-tuned general large language model) as the above-mentioned large language model. Among them, the label result of the above-mentioned training text pair is pre-manually marked and includes a judgment result and a judgment basis. The first text and the second text in the training text pair respectively represent the comparison text and the retrieved text. The training text pair can be a text pair for semantic comparison in history.

[0058] In addition, when there are multiple similar documents, the server can determine the semantic similarity between the retrieved text and the comparison text in each similar text pair for each similar document and for each similar text pair between the similar document and the retrieved document. When the semantic similarity of the similar text pair exceeds the first threshold, the text overlap between the retrieved text and the comparison text in the similar text pair is determined. When the text overlap of the similar text pair does not exceed the second threshold, through a preset prompt word template, the preset comparison rule, the retrieved text and the comparison text in the similar text pair are spliced to obtain a complete prompt word, and the complete prompt word is input into the large language model to obtain the judgment result and the judgment basis of the similar text pair output by the large language model. For those similar text pairs with relatively close semantics but relatively low text overlap, intelligent judgment needs to be combined with the large language model. Through prompt engineering (Prompt Engineering), that is, the prompt word template, a complete prompt word is constructed, so that the large language model judges whether there is a suspicion of plagiarism between the similar text pairs based on the principles of novelty search and document comparison, and further generates a judgment basis for subsequent verification.

[0059] In this specification, after determining the semantic similarity between the retrieved text and the comparison text in each similar text pair, when the semantic similarity of each similar text pair does not exceed the first threshold, that is, if the above semantic similarity does not exceed the first threshold, the server can determine that the judgment result of each similar text pair is not suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold. Among them, since the semantic similarity does not exceed the first threshold, it indicates that there are significant differences in semantics between the retrieved text and the comparison text, so the retrieved text is not suspected of plagiarism, and the corresponding judgment basis is that the semantic similarity does not exceed the first threshold. In addition, when there are multiple similar documents, when the semantic similarity of the similar text pair does not exceed the first threshold, the server can determine that the judgment result of the similar text pair is not suspected of plagiarism, and determine that the judgment basis of the similar text pair is that the semantic similarity does not exceed the first threshold.

[0060] In addition, after determining that the semantic similarity exceeds the first threshold, when the text overlap degree of each similar text pair exceeds the second threshold (that is, the text overlap degree exceeds the second threshold), the server determines that the judgment result of each similar text pair is suspected of plagiarism, and determines that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold. Among them, since the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold, it indicates that the semantics between the retrieved text and the comparison text are similar and the text overlap degree is high, so the retrieved text is suspected of plagiarism, and the corresponding judgment basis is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold. In addition, when there are multiple similar documents, after the semantic similarity of the similar text pair exceeds the first threshold, the server can determine that the judgment result of the similar text pair is suspected of plagiarism when the text overlap degree of the similar text pair exceeds the second threshold, and determine that the judgment basis of the similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold.

[0061] In some embodiments of this specification, the server can generate a final retrieval report according to the judgment results and judgment bases of whether each similar text pair is suspected of plagiarism in the above S4 step. Specifically, the server can first generate a retrieval report corresponding to the retrieved document and the similar document according to the judgment results and judgment bases of whether each similar text pair is suspected of plagiarism, and determine the final retrieval report according to the retrieval report and display it. Among them, the above retrieval report can be a summary of each judgment result and each judgment basis. In addition, the above retrieval report can also include the names or identifiers of similar text pairs and similar documents. The above final retrieval report can be the retrieval report. This final retrieval report can be displayed to users, that is, science and technology novelty search personnel or writing personnel.

[0062] In addition, when there are multiple similar documents, the server can generate a retrieval document and a retrieval report corresponding to each similar document based on the judgment results and bases for whether each pair of similar texts is suspected of plagiarism, and determine the final retrieval report according to each retrieval report. Among them, each pair of similar texts determined in S4 above is determined based on each similar document and the retrieval document, that is, multiple pairs of similar texts can be determined from one similar document and the retrieval document. Therefore, the server can generate a retrieval document and a retrieval report corresponding to each similar document based on the judgment results and bases for each pair of similar texts between the retrieval document and each similar document. Specifically, the server can generate a retrieval report corresponding to the similar document and the retrieval document for each similar document according to the judgment results and bases corresponding to each pair of similar texts between the similar document and the retrieval text.

[0063] The above final retrieval report can be a summary of each retrieval report, that is, directly use each retrieval report as the final retrieval report. Of course, the above final retrieval report can also be one of each retrieval report, that is, when determining the final retrieval report according to each retrieval report, the server can determine the plagiarism scores corresponding to each similar document according to each retrieval report, and then determine the final retrieval report from each retrieval report according to each plagiarism score. Among them, the plagiarism score can be the number of judgments of suspected plagiarism. Specifically, the server can determine the number of judgments of suspected plagiarism in the retrieval report corresponding to the similar document and the retrieval document for each similar document, and use it as the plagiarism score corresponding to the similar document. Determine the retrieval report between the similar document corresponding to the maximum plagiarism score and the retrieval document according to the plagiarism scores corresponding to each similar document, and use it as the final retrieval report.

[0064] In some embodiments of this specification, when performing vector retrieval, a preset vector retrieval strategy may be adopted to determine, from a semantic vector library, comparison semantic vectors that match each retrieval semantic vector. This vector retrieval strategy may retrieve from comparison texts that meet the relationship fields according to the pre-set relationship fields. The relationship fields may be at least one of a document title, a document type, and a chapter type, or may be other fields, which are not specifically limited in this specification. The relationship fields are pre-set by the user. Taking the chapter type as the relationship field as an example, the server may retrieve from comparison texts with the same chapter type as the retrieval text to perform semantic matching between the retrieval text corresponding to the same type of chapter and the comparison text, improving the accuracy of semantic matching. Based on this, when determining, in the above S3, comparison semantic vectors that match each retrieval semantic vector from the semantic vector library, the server may determine the retrieval information corresponding to each retrieval semantic vector according to the preset relationship fields, and screen the comparison semantic vectors in the semantic vector library according to the retrieval information. From the screened comparison semantic vectors, determine comparison semantic vectors that match each retrieval semantic vector. Among them, the retrieval information is the field value of the preset relationship field of the retrieval text corresponding to the retrieval semantic vector. Taking the chapter type as the relationship field as an example, this retrieval information is the chapter type of the retrieval text corresponding to the retrieval semantics. Taking the retrieval document or the comparison document as a thesis as an example, the chapter type may include a title, an abstract, an introduction, a literature review, a method, a conclusion, an acknowledgement, references, and an appendix, etc. Each comparison text corresponding to each of the above-screened comparison semantic vectors is a vector that meets the retrieval information of each retrieval semantic vector. When screening the comparison semantic vectors in the semantic vector library according to the retrieval information, for each comparison semantic vector in the semantic vector library, when the field value corresponding to the relationship field of the comparison text of this comparison semantic vector is consistent with the retrieval information, then this comparison semantic vector is used as the screened comparison semantic vector.

[0065] In addition, for each retrieval semantic vector, the server may also, according to the preset relationship fields, use the retrieval information of the retrieval text corresponding to this retrieval semantic vector, and then determine, from the semantic vector library, each comparison semantic vector that meets the retrieval information corresponding to this retrieval semantic vector, and determine, from the determined comparison semantic vectors, a comparison semantic vector that matches this retrieval semantic vector. Among them, the process of determining, from the determined comparison semantic vectors, a comparison semantic vector that matches this retrieval semantic vector is similar to the process of determining, in the above S3, comparison semantic vectors that match each retrieval semantic vector from the semantic vector library, and will not be elaborated here.

[0066] In some embodiments of this specification, there may be multiple comparison texts that match each retrieval text. Therefore, the above retrieval results may include at least one of multiple comparison semantic vectors that match the retrieval semantic vector of the retrieval document, and the similarity between the retrieval semantic vector and the multiple comparison semantic vectors that match the retrieval semantic vector. The number of comparison texts that match each retrieval text can be preset.

[0067] In some embodiments of this specification, when generating a retrieval report corresponding to the retrieval document and each similar document based on the judgment results and bases of each similar text pair between the retrieval document and each similar document, the server can generate a first report corresponding to the retrieval document and each similar text based on the judgment results and bases of each similar text pair between the retrieval document and each similar document. Determine similar text pairs that meet the preset target conditions based on the judgment results and bases of each similar text pair between the retrieval document and each similar document, and generate a second report between the retrieval document and each similar document based on the judgment results and bases of the determined similar text pairs. Use the first report and the second report as the retrieval reports between the retrieval text and each similar document. Among them, the above first report is a summary of all judgment results and bases, and the second report is a summary of the judgment results and bases that meet the preset target conditions. The target conditions are preset, and the target conditions include semantic similarity detection items (that is, the semantic similarity exceeds the first threshold, the text overlap does not exceed the second threshold, and the judgment result output by the large language model is suspected of plagiarism), large language model judgment of non-suspicious plagiarism items (that is, the semantic similarity exceeds the first threshold, the text overlap does not exceed the second threshold, and the judgment result output by the large language model is not suspected of plagiarism), and items with highly similar semantic texts (that is, the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold). The above semantic similarity detection items and large language model judgment of non-suspicious plagiarism items are mainly for text pairs with relatively high semantic similarity and relatively low text overlap (that is, the semantic similarity exceeds the first threshold and the text overlap does not exceed the threshold), and the judgment of whether there is suspected plagiarism is obtained through the large language model. By additionally generating the second report, it helps subsequent users verify the judgment results and bases. The specific implementation process of generating a retrieval report corresponding to the retrieval document and the similar document based on the judgment results and bases of whether each similar text pair between the retrieval document and the similar document is suspected of plagiarism is similar to the above process and will not be elaborated here.

[0068] In some embodiments of this specification, taking the description text of global environmental changes as an example, assume Text 1: "With the intensification of global climate change, the rise in sea level has become a problem that cannot be ignored. Scientists warn that if no measures are taken, many coastal cities may be flooded." If only the order of the text is simply reversed, the original text (i.e., Text 1) can be transformed into Text 2, that is, Text 2: "With the intensification of global climate change, scientists warn that the rise in sea level has become a problem that cannot be ignored. If no measures are taken, many coastal cities may be flooded." Since the text itself has hardly changed, the literal overlap between Text 2 and Text 1 is nearly 100%. Through the traditional method of comparing literal overlap (i.e., text overlap), this form of plagiarism can be easily detected.

[0069] However, if the key words in the original text are replaced with synonyms to transform it into Text 3, that is, Text 3: "Given the continuous deterioration of global warming, the rise in sea level has become a serious threat. Researchers warn that if no countermeasures are taken, many coastal cities may be devastated." Then, from the perspective of text overlap, the overlap between Text 3 and Text 1 is less than 28.7%, and the repeated words are relatively common and non-consecutive non-academic words. The method of text overlap detection is difficult to detect this form of plagiarism. But since the replaced words are all synonyms with similar semantics, from the direction of semantic vector comparison, the similarity is 86.9%. The semantics between Text 1 and Text 3 are highly similar, and this pair of texts suspected of plagiarism can be found through the method of semantic similarity retrieval and comparison.

[0070] The following lists the comparison results of these two situations of text overlap and semantic similarity in the form of a table, and adds the judgment results and judgment bases based on the large language model, that is, the judgment results and judgment bases obtained by comparing in the Figure 1 described manner, as shown in Table 1 below. Table 1 is a text comparison case. This Table 1 includes the retrieved text (i.e., Text 1), the compared texts (i.e., Text 2 and Text 3). The text similarity between Text 1 and Text 2 is 100%, and the semantic vector similarity (i.e., semantic similarity) is 99.4%. The judgment result of the large language model is suspected of plagiarism, and the judgment basis is "the two texts are basically the same". The text similarity between Text 1 and Text 3 is 28.7%, and the semantic vector similarity (i.e., semantic similarity) is 86.9%. The judgment result of the large language model is suspected of plagiarism, and the judgment basis is "the key content and expressions are highly consistent, and the semantic repetition is obvious".

[0071]

[0072] Figure 1 ​The method described above mainly consists of three parts. The first part is to construct a semantic vector library, that is, to construct a semantic vector library according to the comparison semantic vectors corresponding to each comparison text of each comparison document. The specific process is the process of constructing the semantic vector library in S1 above. As Figure 2 shown Figure 2 is a schematic flowchart of a process for constructing a semantic vector library provided in this specification. Figure 2 Starting from the "start" node, first obtain each uploaded comparison document, and perform text extraction on each comparison document to obtain the comparison text data corresponding to each comparison document, that is Figure 2 "document data extraction" in. Then, according to the chapter hierarchy structure of each comparison document, split the comparison text data to obtain each comparison chapter data. According to the pre-set filtering rules, filter each comparison chapter data to obtain each comparison target chapter data, that is Figure 2 "document chapter splitting and screening" in. According to the granularity requirements corresponding to each comparison document, use natural language processing technology to split each comparison target chapter data to obtain each comparison text, that is Figure 2 "paragraph and sentence splitting" in. Use a semantic vectorization model to determine the comparison semantic vectors corresponding to each comparison text of each comparison document, that is Figure 2 "text semantic vectorization" in. Generate labels corresponding to each comparison semantic vector, and construct a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector, that is Figure 2 "data storage" in. Finally, it reaches the "end" node.

[0073] The second part is to perform similar document retrieval based on the semantic vector library, that is, to preprocess the retrieval document and perform retrieval in the semantic vector library. Integrate the retrieval results corresponding to each retrieval text, and narrow the range of similar documents through weighted ranking to obtain each similar text. The specific process is the process of S2~S3 above. As Figure 3 shown Figure 3 is a schematic diagram of a document semantic retrieval and comparison process provided in this specification. Figure 3 Starting from the "start" node, first obtain the retrieval document, and perform data preprocessing on the retrieval document to obtain each retrieval text corresponding to the retrieval document, that is Figure 3 "data preprocessing" in. Use a semantic vectorization model to determine the retrieval semantic vectors corresponding to each retrieval text, that is Figure 3 "semantic vectorization" in. Determine the comparison semantic vectors matching each retrieval semantic vector from the semantic vector library to generate the retrieval results corresponding to each retrieval text, that is Figure 3"Vector library retrieval" in. According to the similarities included in each retrieval result, determine the similarity scores corresponding to each comparison document. According to each similarity score, sort and screen each comparison document to determine similar documents, that is Figure 3 "Similar document screening and sorting" in. Finally, until the "End" node.

[0074] The third part, document comparison assisted by large language models, that is, based on methods such as semantic vector comparison, text comparison, and large language model judgment, compare the retrieved documents and similar documents, combine the large language model to generate the basis for whether there is plagiarism, and integrate all comparison results to generate a retrieval report, as Figure 4 shown, Figure 4 is a schematic diagram of a document comparison process assisted by a large language model provided in this specification, Figure 4 starting from the "Start" node in, first determine each similar text pair between the retrieved document and the similar document according to the retrieval semantic vectors of the retrieved document and the comparison semantic vectors of the similar document, that is Figure 4 "Semantic retrieval matching" in. According to the semantic similarity and text overlap degree between the retrieved text and the comparison text in each similar text pair, use a large language model to determine the judgment result and judgment basis of each similar text pair, that is Figure 4 "Judge whether there is suspected plagiarism and generate judgment basis" in. According to the judgment results and judgment bases of each similar text pair between the retrieved document and the similar document, generate a retrieval report corresponding to the retrieved document and the similar document, and according to the retrieval report, determine the final retrieval report, that is Figure 4 "Summarize results to generate a report" in. Finally, until the "End" node.

[0075] When performing the above "Judge whether there is suspected plagiarism and generate judgment basis" and "Summarize results to generate a report", as Figure 5 shown, Figure 5 is a schematic diagram of a plagiarism determination process assisted by a large language model provided in this specification, Figure 5 starting from the "Start" node in, first determine the semantic similarity between the retrieved text and the comparison text in each similar text pair, and judge whether the semantic similarity of each similar text pair exceeds the first threshold, that is Figure 5 "Judge whether the semantic similarity exceeds the first threshold" in. If not (that is, the semantic similarity of each similar text pair does not exceed the first threshold), the server can determine that the judgment result of each similar text pair is not suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold, that is Figure 5 "Judge as not suspected of plagiarism" in.

[0076] If so (i.e., the semantic similarity of each similar text pair exceeds the first threshold), continue to determine the text overlap degree between the retrieved text and the comparison text in each similar text pair, and judge whether the text overlap degree of each similar text pair exceeds the second threshold, that is Figure 5 In "<judging whether the text overlap degree exceeds the second threshold>", if so (i.e., the text overlap degree of each similar text pair exceeds the second threshold), the server can determine that the judgment result of each similar text pair is suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold, that is Figure 5 In "<judged as suspected of plagiarism>". If not (i.e., the text overlap degree of each similar text pair does not exceed the second threshold), through a preset prompt word template, splice the preset comparison rules, the retrieved text and the comparison text in each similar text pair to obtain a complete prompt word, and input the complete prompt word into the large language model to obtain the judgment result and judgment basis of each similar text pair output by the large language model, that is Figure 5 In "<the large language model judges whether it is suspected of plagiarism and generates a judgment basis>". If the large language model judges that it is suspected of plagiarism (i.e., if so), the judgment result is suspected of plagiarism, that is Figure 5 In "<judged as suspected of plagiarism>". If the large language model judges that it is not suspected of plagiarism (i.e., if not), the judgment result is not suspected of plagiarism, that is Figure 5 In "<judged as not suspected of plagiarism>". After that, according to the judgment result and judgment basis of each similar text pair, generate a retrieval report corresponding to the retrieved document and the similar document, and generate a final retrieval report according to the retrieval report, that is Figure 5 In "<summarize the results and generate a report>". Finally, until the "end" node.

[0077] Based on the above Figure 1 shown method, this specification provides a document plagiarism judgment system based on a semantic vector library and a large language model, as Figure 6 shown, Figure 6 is a schematic diagram of a document plagiarism judgment system provided in this specification. The system includes a data storage module 100, a data preprocessing module 101, a semantic vectorization module 102, a vector retrieval module 103, and a document comparison module 104.

[0078] The above data preprocessing module 101 can preprocess the comparison document set to obtain the comparison semantic vectors of each comparison text, and split the retrieved document into each retrieved text. The specific implementation process of the data preprocessing module 101 is the same as the above Figure 1The implementation processes of splitting the retrieved document into each retrieved text in S1 and S2 are the same. When splitting the retrieved document into each retrieved text as described above, the data preprocessing module 101 can perform text extraction on the retrieved document to obtain the text data corresponding to the retrieved document. According to the chapter hierarchical structure of the retrieved document, the text data is split to obtain each chapter data. According to the preset filtering rules, each chapter data is filtered to obtain each target chapter data. According to the granularity requirement corresponding to the retrieved document, each target chapter data is split through natural language processing technology to obtain each retrieved text.

[0079] The above semantic vectorization module 102 can call the semantic vectorization model to determine the retrieval semantic vector of each retrieved text. The specific implementation process of the semantic vectorization module 102 is the same as that in Figure 1 S2 in the above to determine the retrieval semantic vector of each retrieved text using the semantic vectorization model, and will not be elaborated here.

[0080] The above data storage module 100 can construct and store a semantic vector library based on each comparison semantic vector. The data storage module 100 can also store the semantic vectorization model, the large language model, the retrieved document, and the comparison document set.

[0081] The above vector retrieval module 103 can determine the comparison semantic vector matching each retrieval semantic vector from the semantic vector library to generate the retrieval result corresponding to each retrieved text, and sort and filter each comparison document in the comparison document set according to each retrieval result to obtain similar documents. The specific implementation process of the vector retrieval module 103 is the same as that in Figure 1 S3 in the above, and will not be elaborated here.

[0082] In addition to including each comparison semantic vector, the above semantic vector library can also include the label corresponding to each comparison semantic vector. The label can include the comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector. The above retrieval result can be the similarity between the retrieval semantic vector of the retrieved text and the comparison semantic vector matching the retrieval semantic vector of the retrieved document, or can be the comparison semantic vector matching the retrieval semantic vector of the retrieved document, or can be the comparison semantic vector matching the retrieval semantic vector of the retrieved document and the similarity between the retrieval semantic vector and the comparison semantic vector matching the retrieval semantic vector.

[0083] Based on this, when retrieving the similarity between the retrieval semantic vector of the retrieval text and the comparison semantic vector that matches the retrieval semantic vector of the retrieval document, when the vector retrieval module 103 sorts and filters each comparison document in the comparison document set according to each retrieval result to obtain similar documents, it can determine the similarity score corresponding to each comparison document according to each similarity included in each retrieval result. Sort and filter each comparison document according to each similarity score to determine similar documents.

[0084] The above-mentioned document comparison module 104 can correspond each retrieval text to each comparison text of the similar document one by one to form each similar text pair. Determine whether the semantic similarity between the texts (i.e., the retrieval text and the comparison text) in each similar text pair exceeds the first threshold. If it exceeds the first threshold, continue to determine whether the text overlap between the texts in each similar text pair exceeds the second threshold. If it does not exceed the second threshold, splice the texts (i.e., the retrieval text and the comparison text) in each similar text pair and supplement them to the prompt template. After obtaining the complete prompt, input it into the large language model to determine whether it is suspected of plagiarism and generate a judgment basis. The specific implementation process of the document comparison module 104 is the same as that of S5 in the above Figure 1 and will not be elaborated here.

[0085] Specifically, the document comparison module 104 can first determine, from each comparison text of the similar document, the comparison text that matches each retrieval text according to the retrieval semantic vector of each retrieval text and the comparison semantic vector of each comparison text of the similar document. Combine each retrieval text with each determined comparison text to form each similar text pair. Further, it can first determine the first vector set composed of the retrieval semantic vectors of the retrieval documents, and determine the second vector set composed of the comparison semantic vectors of the similar documents. Determine the Cartesian product between the first vector set and the second vector set. Determine the similarity between the vectors in each vector combination in the Cartesian product. According to the similarity between the vectors in each vector combination in the Cartesian product, determine, from each comparison text of the similar document, the comparison text that is similar to each retrieval text, and use each retrieval text and each determined comparison text as a similar text pair.

[0086] After that, the above-mentioned document comparison module 104 determines the semantic similarity between the retrieved text and the compared text in each similar text pair. When the semantic similarity of each similar text pair exceeds the first threshold, the text overlap degree between the retrieved text and the compared text in each similar text pair is determined. When the text overlap degree of each similar text pair does not exceed the second threshold, the retrieved text and the compared text in each similar text pair are concatenated and supplemented into the prompt template to obtain a complete prompt, which is then input into the large language model to obtain the judgment result and judgment basis of each similar text pair output by the large language model. Among them, when obtaining the complete prompt, the above-mentioned document comparison module 104 can concatenate the preset comparison rules, the retrieved text and the compared text in each similar text pair through the preset prompt template to obtain a complete prompt.

[0087] In addition, when the semantic similarity of each similar text pair does not exceed the first threshold, that is, if the semantic similarity does not exceed the first threshold, the above-mentioned document comparison module 104 can determine that the judgment result of each similar text pair is not suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold. In addition, when the text overlap degree of each similar text pair exceeds the second threshold (that is, the text overlap degree exceeds the second threshold), it can be determined that the judgment result of each similar text pair is suspected of plagiarism, and it is determined that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold.

[0088] In some embodiments of this specification, the above-mentioned system may further include a report generation module 105, and the report generation module 105 can generate a final retrieval report according to the judgment results and judgment bases of whether each similar text pair in the document comparison module 104 is suspected of plagiarism. Specifically, the above-mentioned report generation module 105 can generate a retrieval report corresponding to the retrieved document and the similar document according to the judgment results and judgment bases of whether each similar text pair is suspected of plagiarism, and determine the final retrieval report according to the retrieval report and display it. The specific implementation process of the report generation module 105 is the same as that in the above Figure 1 The implementation process of generating the final retrieval report according to the judgment results and judgment bases of whether each similar text pair is suspected of plagiarism in step S4 above will not be repeated here.

[0089] In some embodiments of this specification, the above-mentioned data storage module 100 can also perform addition, deletion, modification, and query operations on the data it stores, such as the semantic vector library.

[0090] In some embodiments of this specification, the above-mentioned data preprocessing module 101 may further obtain a set of comparison documents including each comparison document uploaded, and perform data preprocessing on each comparison document to obtain respective comparison texts corresponding to each comparison document. The semantic vectorization module 102 may further use a semantic vectorization model to determine respective comparison semantic vectors corresponding to the respective comparison texts of each comparison document. The data storage module 100 may generate labels corresponding to the respective comparison semantic vectors, and construct a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector.

[0091] In some embodiments of this specification, the above-mentioned large language model is one of a general large language model and a fine-tuned general large language model, and the fine-tuned general large language model is obtained by fine-tuning the general large language model based on training text pairs and label results. The above-mentioned data preprocessing module 101 may further obtain the general large language model; obtain training text pairs for training the general large language model, and splice the first text and the second text in the training text pairs to obtain a third text. Input the third text into the general large language model to obtain an output result. Fine-tune the general large language model according to the label result and the output result of the training text pair, so as to use the trained general large language model (i.e., the above-mentioned fine-tuned general large language model) as the large language model. Among them, the output result includes a judgment result and a judgment basis.

[0092] In some embodiments of this specification, when the above-mentioned vector retrieval module 103 determines a comparison semantic vector matching each retrieval semantic vector from the semantic vector library, it may determine retrieval information corresponding to each retrieval semantic vector according to a preset relationship field, and screen the comparison semantic vectors in the semantic vector library according to the retrieval information. Determine a comparison semantic vector matching each retrieval semantic vector from the screened comparison semantic vectors. Among them, the relationship field is at least one of a document title, a document type, and a chapter type.

[0093] In some embodiments of this specification, when the above-mentioned report generation module 105 generates a retrieval report corresponding to the retrieval document and the similar document based on the judgment results and judgment bases of each similar text pair, it may generate a first report corresponding to the retrieval document and the similar text according to the judgment results and judgment bases of each similar text pair of the retrieval document and the similar document. Determine similar text pairs that meet the preset target conditions according to the judgment results and judgment bases of each similar text pair of the retrieval document and the similar document, and generate a second report corresponding to the retrieval document and the similar document according to the judgment results and judgment bases of the determined similar text pairs. Use the first report and the second report as the retrieval report between the retrieval text and the similar document.

[0094] In some embodiments of this specification, the above system mainly provides services externally in the form of B / S architecture software services. The overall architecture of this system is divided into a data layer, an application layer, an interaction layer, and external data, such as Figure 7 shown Figure 7 is a schematic diagram of a document plagiarism judgment system architecture based on a semantic vector library and a large language model provided in this specification. The data layer mainly includes a model and a semantic vector library. The model includes a semantic vectorization model and a large language model. The application layer mainly includes a data storage module 100, a data preprocessing module 101, a semantic vectorization module 102, a vector retrieval module 103, a document comparison module 104, and a report generation module 105. The above interaction layer mainly provides functions for interacting with users, mainly including providing a UI interface to interact with users, processing user input and output (such as form submission, page rendering), and displaying the results processed by the application layer. The interaction layer provides functions for document warehousing, document retrieval, and report generation, that is, a document warehousing page, a document retrieval page, and a report generation page. Users can batch upload each comparison document or upload a single comparison document through the document warehousing page, and the document warehousing page sends each comparison document to the data storage module 100 for storage. Users can upload a retrieval document through the document retrieval page, and the document retrieval page sends the retrieval document to the data storage module 100 for storage. In addition, users can also set a vector retrieval strategy through the document retrieval page, and the document retrieval page sends the vector retrieval strategy to the vector retrieval module 103 for subsequent use in vector retrieval. The above report generation module 105 can display the generated final retrieval report through the above report generation page, and the report generation page provides functions for previewing the final retrieval report and downloading the report. In addition, users can also edit and modify the final retrieval report on this report generation page. The above external data mainly includes two parts: each comparison document to be retrieved (i.e., a comparison document set) and the retrieval document that needs to be retrieved.

[0095] It should be noted that the above specific embodiments can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the present invention can still be modified or equivalently replaced. In short, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the patent of the present invention.

Claims

1. A method for judging document plagiarism based on a semantic vector library and a large language model, characterized in that, Including: S1: Construct a semantic vector library: Preprocess a preset set of comparison documents to obtain the comparison semantic vectors of each comparison text, and construct a semantic vector library; S2: Construct retrieval semantic vectors based on the retrieval document: Obtain the retrieval document and split it into each retrieval text; Use a semantic vectorization model to determine the retrieval semantic vector of each retrieval text; S3: Screen similar documents: From the semantic vector library, determine the comparison semantic vectors that match each retrieval semantic vector to generate the retrieval results of each retrieval text, and sort and screen each comparison document in the comparison document set according to the retrieval results to obtain similar documents; S4: Judge whether there is plagiarism based on a large language model and generate a judgment basis: Correspond the retrieval texts with the comparison texts of the similar documents one by one to form each similar text pair; Judge whether the semantic similarity between the retrieval text and the comparison text in each similar text pair exceeds a first threshold. If it exceeds the first threshold, continue to judge whether the text overlap degree between the retrieval text and the comparison text in each similar text pair exceeds a second threshold. If it does not exceed the second threshold, splice the retrieval text and the comparison text in each similar text pair and supplement them to the prompt template to obtain a complete prompt, and then input it into the large language model to determine whether there is suspected plagiarism and generate a judgment basis.

2. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, characterized in that The method further includes: If it does not exceed the first threshold, determine that the judgment result of each similar text pair is not suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold; If it exceeds the second threshold, determine that the judgment result of each similar text pair is suspected of plagiarism, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold.

3. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, characterized in that, The method further includes: Generate a retrieval report corresponding to the retrieval document and the similar document according to the judgment result and judgment basis of whether each similar text pair is suspected of plagiarism; Determine and display the final retrieval report according to the retrieval report.

4. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, wherein, Specifically, in S2, splitting the retrieval document into each retrieval text includes: Perform text extraction on the retrieval document to obtain the text data corresponding to the retrieval document; According to the chapter hierarchy structure of the retrieval document, split the text data to obtain each chapter data; Filter each chapter data according to preset filtering rules to obtain each target chapter data; According to the granularity requirement corresponding to the retrieval document, use natural language processing technology to split each target chapter data to obtain each retrieval text.

5. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, characterized in that, Specifically, S1 includes: Obtain a comparison document set including each comparison document uploaded, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; Use a semantic vectorization model to determine the comparison semantic vectors respectively corresponding to each comparison text of each comparison document; Generate labels corresponding to each comparison semantic vector, and construct a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein, the label includes comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector.

6. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, wherein, In step S4, the one-to-one correspondence between each retrieved text and each comparison text of the similar document to form each similar text pair specifically includes: According to the retrieval semantic vector of each retrieved text and the comparison semantic vector of each comparison text of the similar document, determine the comparison text that matches each retrieved text from each comparison text of the similar document; Form each similar text pair by combining each retrieved text with the determined comparison text.

7. A method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, characterized in that, In step S4, splicing the retrieved text and the comparison text in each similar text pair and supplementing them into the prompt template to obtain a complete prompt specifically includes: Through a preset prompt template, splice the preset comparison rules, the retrieved text and the comparison text in each similar text pair to obtain a complete prompt.

8. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, wherein, The large language model in step S4 is one of a general large language model and a fine-tuned general large language model, and the fine-tuned general large language model is obtained by fine-tuning the general large language model based on training text pairs and label results.

9. The method for judging document plagiarism based on a semantic vector library and a large language model according to claim 1, wherein, In step S3, determining the comparison semantic vector that matches each retrieval semantic vector from the semantic vector library specifically includes: According to a preset relationship field, determine the retrieval information corresponding to each retrieval semantic vector, and screen the comparison semantic vectors in the semantic vector library according to the retrieval information; wherein, the relationship field is at least one of a document title, a document type, and a chapter type; From the screened comparison semantic vectors, determine the comparison semantic vector that matches each retrieval semantic vector.

10. A document plagiarism judgment system based on a semantic vector library and a large language model for the method according to any one of claims 1 to 9, characterized in that, The system includes a data storage module, a data preprocessing module, a semantic vectorization module, a vector retrieval module, and a document comparison module, wherein: The data preprocessing module is used to preprocess the comparison document set to obtain the comparison semantic vectors of each comparison text, and split the retrieved document into each retrieved text; The semantic vectorization module is used to call a semantic vectorization model to determine the retrieval semantic vector of each retrieved text; The data storage module is used to construct and store a semantic vector library based on the comparison semantic vectors; The vector retrieval module is used to determine the comparison semantic vector that matches each retrieval semantic vector from the semantic vector library to generate each retrieval result, and sort and screen each comparison document in the comparison document set according to the retrieval results to obtain similar documents; The document comparison module is used to correspond each of the retrieved texts to each comparison text of the similar documents one by one to form each pair of similar texts; determine whether the semantic similarity between the texts in each pair of similar texts exceeds a first threshold. If it exceeds the first threshold, determine whether the text overlap degree between the texts in each pair of similar texts exceeds a second threshold. If it does not exceed the second threshold, splice the texts in each pair of similar texts and supplement them to the prompt template, and after obtaining the complete prompt, input it into the large language model to determine whether there is a suspicion of plagiarism and generate a judgment basis.

11. The document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that The document comparison module is further used to, if it does not exceed the first threshold, determine that the judgment result of each pair of similar texts is not suspected of plagiarism, and determine that the judgment basis of each pair of similar texts is that the semantic similarity does not exceed the first threshold; If it exceeds the second threshold, determine that the judgment result of each pair of similar texts is suspected of plagiarism, and determine that the judgment basis of each pair of similar texts is that the semantic similarity exceeds the first threshold and the text overlap degree exceeds the second threshold.

12. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The system further includes a report generation module; The report generation module is used to generate a retrieval report corresponding to the retrieved document and the similar document according to the judgment result and judgment basis of whether each pair of similar texts is suspected of plagiarism; determine the final retrieval report according to the retrieval report and display it.

13. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The data preprocessing module is specifically used to perform text extraction on the retrieved document to obtain the text data corresponding to the retrieved document; split the text data according to the chapter level structure of the retrieved document to obtain each chapter data; filter each chapter data according to the preset filtering rules to obtain each target chapter data; split each target chapter data through natural language processing technology according to the granularity requirement corresponding to the retrieved document to obtain each retrieved text.

14. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The data preprocessing module is further used to store the semantic vectorization model, the large language model, the retrieved document, and the set of comparison documents; The data preprocessing module is specifically used to obtain a set of comparison documents including each comparison document uploaded, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; The semantic vectorization module is further used to use the semantic vectorization model to determine the comparison semantic vectors corresponding to each comparison text of each comparison document; The data storage module is specifically used to generate labels corresponding to each comparison semantic vector, and construct a semantic vector library according to each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein, the label includes the comparison document information, chapter information, and comparison text information corresponding to the comparison semantic vector.

15. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The document comparison module is specifically configured to determine, from each comparison text of the similar documents, the comparison text that matches each retrieval text according to the retrieval semantic vectors of the respective retrieval texts and the comparison semantic vectors of each comparison text of the similar documents; and form each similar text pair by combining each retrieval text with the determined comparison text for each retrieval text.

16. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The document comparison module is specifically configured to splice the preset comparison rules, the retrieval text and the comparison text in each similar text pair through a preset prompt template to obtain a complete prompt.

17. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The large language model is one of a general large language model and a fine-tuned general large language model, and the fine-tuned general large language model is obtained by fine-tuning the general large language model based on training text pairs and label results.

18. A document plagiarism judgment system based on a semantic vector library and a large language model according to claim 10, characterized in that, The vector retrieval module is specifically configured to determine the retrieval information corresponding to each retrieval semantic vector according to a preset relationship field, and screen the comparison semantic vectors in the semantic vector library according to the retrieval information; wherein, the relationship field is at least one of a document title, a document type, and a chapter type; and determine the comparison semantic vectors that match each retrieval semantic vector from the screened comparison semantic vectors.

Citation Information

Patent Citations

  • Chinese and English cross-language plagiarism recognition method based on characteristics and content of translations

    CN103544326A

  • Text comparison method and device

    CN111967243A

  • Tibetan composition plagiarism detection prototype system

    CN112131859A

  • Text comparison method and device, equipment and medium

    CN113821601A

  • Plagiarism detecting method and plagiarism detecting apparatus

    KR100863943B1

Cited By

  • Code adoption condition determination method and device, electronic equipment and storage medium

    CN120508837A

  • Code adoption situation determination method and apparatus, electronic device, and storage medium

    CN120508837B