Science and technology project declaration duplicate checking method and system based on large language model and Jaccard similarity coefficient

The method leverages a large language model to extract and vectorize document segments for precise similarity checking, addressing inefficiencies and inaccuracies in existing systems by using Jaccard similarity scores and re-ranking, thereby improving the efficiency and accuracy of project proposal document similarity analysis.

CN120316243APending Publication Date: 2025-07-15INSPUR SOFTWARE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510443948.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing scientific and technological project application application methods are inefficient, unable to accurately locate core research content, poor recognition of variants and synonyms, and lack of industry-orientedness, resulting in inaccurate and inefficient plagiarism results.

Method used

The core content of the document is extracted and split into fragments, converted into vectors through the text embedding model, calculated the Euclidean distance or cosine similarity, combined with the Jaccard similarity coefficient to evaluate the similarity, and introduced the re-rank model optimization results to generate a plagiarism check report.

Benefits of technology

It improves the accuracy and efficiency of plagiarism checking, can accurately identify duplicate or highly similar projects, reduce irrelevant information, build an efficient and intelligent scientific and technological management system, and avoid duplicate project establishment and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316243A_ABST
    Figure CN120316243A_ABST
Patent Text Reader

Abstract

The invention discloses a science and technology project declaration duplicate checking method and system based on a large language model and a Jaccard similarity coefficient, belongs to the technical field of natural language processing, and aims to solve the technical problem of how to improve the accuracy and efficiency of science and technology project declaration duplicate checking. According to the technical scheme, the method comprises the steps that document core content is extracted, wherein the core content of a science and technology project document to be subjected to duplicate checking is extracted through a large language model subjected to data training of a plurality of science and technology projects; splitting document segments: splitting the document core content into a plurality of document segments based on a natural language processing technology; document vectorization storage: converting document fragments into vectors through a text embedding model, and storing the vectors in a vector database; calculating a vector distance and retrieving historical items: calculating the Euclidean distance or cosine similarity between the vector of the item document fragment to be subjected to duplicate checking and the vector of the historical item document fragment, and extracting topK historical items with the closest distance; calculating a Jaccard similarity coefficient; aggregating the document similarity; and generating a duplicate checking result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and specifically to a method and system for detecting duplicate submissions of science and technology project application forms based on large language models and Jaccard similarity coefficients. Background Art

[0002] With the rapid development of technology, the number of science and technology project application forms has been increasing continuously, and the importance of duplicate submission detection has become increasingly prominent. Existing duplicate submission detection methods mainly rely on traditional text comparison technologies, and there are many problems and deficiencies in dealing with large amounts of data, as follows:

[0003] ① Low duplicate submission detection efficiency: When existing duplicate submission detection systems process a large number of science and technology project application forms, they often consume a large amount of time and computing resources. For example, some systems need to compare documents word by word, resulting in an overly long duplicate submission detection time and unable to meet the requirements of efficient management.

[0004] ② Inability to accurately locate the core research content: Existing duplicate submission detection methods usually can only perform surface text comparison and cannot deeply understand the semantics and core research content of the documents. This results in a large amount of irrelevant information in the duplicate submission detection results, reducing the accuracy and reliability of duplicate submission detection.

[0005] ③ Poor ability to identify variants and synonyms: The expressions in science and technology project application forms are diverse, and the same concept may have multiple different expressions. Existing duplicate submission detection methods are difficult to identify these variants and synonyms, resulting in frequent missed detections.

[0006] ④ Lack of industry specificity: Existing duplicate submission detection systems are mostly general-purpose and lack an understanding of the industry characteristics and professional terms of science and technology project application forms, and are unable to accurately identify and process industry-specific content.

[0007] In summary, how to improve the accuracy and efficiency of duplicate submission detection of science and technology project application forms is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0008] The technical task of the present invention is to provide a method and system for detecting duplicate submissions of science and technology project application forms based on large language models and Jaccard similarity coefficients to solve the problem of how to improve the accuracy and efficiency of duplicate submission detection of science and technology project application forms.

[0009] The technical task of the present invention is realized in the following way. A method for detecting duplicate submissions of science and technology project application forms based on large language models and Jaccard similarity coefficients is as follows:

[0010] Extract the core content of the document: Extract the core content of the science and technology project document to be detected for duplicate submissions through a large language model that has been trained with a number of science and technology project data.

[0011] Split document fragments: Based on natural language processing technology, split the core content of the document into multiple document fragments; among them, each fragment includes one or more key sentences for subsequent vectorization processing;

[0012] Vectorized storage of documents: Convert document fragments into vectors through a text embedding model and store the vectors in a vector database;

[0013] Calculate vector distance and retrieve historical items: Calculate the Euclidean distance or cosine similarity between the vectors of the document fragments of the item to be checked for duplication and the vectors of the document fragments of historical items, and extract the top K historical items with the closest distance;

[0014] Calculate Jaccard similarity coefficient: Recalculate the Jaccard similarity coefficient between the document fragments of the item to be checked for duplication and the document fragments of historical items to evaluate the similarity between the document fragments of the item to be checked for duplication and the document fragments of historical items;

[0015] Aggregate document similarity: Aggregate the Jaccard similarity coefficients of all fragments through weighted average to form the overall similarity of the document;

[0016] Generate duplicate check results: If the overall similarity of the document exceeds the set threshold, it is determined that there is duplication or high similarity between the document of the item to be checked for duplication and the item documents in the database, and a project duplicate check report is generated; among them, the project duplicate check report includes a list of similar items and similarity score information.

[0017] Preferably, the core content of the document is extracted as follows:

[0018] Input the scientific and technological project document to be checked for duplication into a large language model;

[0019] The large language model performs semantic analysis on the scientific and technological project document to be checked for duplication, and extracts core keywords and key sentences;

[0020] Generate a core content summary of the document based on the core keywords and key sentences.

[0021] Preferably, the document fragments are split as follows:

[0022] Split the core content summary of the document into multiple document fragments;

[0023] Perform preprocessing operations of word segmentation and stop word removal on each document fragment.

[0024] Preferably, the Jaccard similarity coefficient is calculated as follows:

[0025] For each document fragment of the item to be checked for duplication and the document fragment of the historical item, extract the keyword set;

[0026] Calculate the intersection and union of the two sets;

[0027] Calculate the Jaccard similarity coefficient. The formula is: Jaccard coefficient = the number of keywords in the intersection / the number of keywords in the union.

[0028] Preferably, optimize the overall similarity of the documents as follows:

[0029] Adjust the embedding threshold through multiple rounds of experiments and data analysis;

[0030] Evaluate the duplicate-checking efficiency and accuracy under different embedding thresholds to obtain the evaluation results;

[0031] Select the optimal embedding threshold according to the evaluation results.

[0032] More preferably, after calculating the vector distance, introduce a re-rank model to re-rank the retrieval results to further accurately locate the items with high similarity as follows:

[0033] Use the re-rank model to re-rank the top K historical project document fragments;

[0034] Recalculate the Jaccard similarity coefficient to ensure the accuracy of the duplicate-checking results.

[0035] Among them, the Jaccard similarity coefficient is an index widely used in calculating the similarity of sets. It measures the similarity between two sets by comparing their intersection and union.

[0036] A duplicate-checking system for science and technology project application forms based on a large language model and the Jaccard similarity coefficient. The system includes:

[0037] A data preprocessing module for extracting the core content of the science and technology project document to be checked for duplicates through a large language model that has been trained with several science and technology project data, splitting the document core content into multiple document fragments based on natural language processing technology; and converting the document fragments into vectors through a text embedding model (embedding model);

[0038] A data storage module for storing the vectors in a vector database;

[0039] A comparison module for calculating the Euclidean distance or cosine similarity between the vectors of the document fragments to be checked for duplicates and the vectors of the historical project document fragments, extracting the top K historical projects with the closest distance; recalculating the Jaccard similarity coefficient between the document fragments to be checked for duplicates and the historical project document fragments to evaluate the similarity between the document fragments to be checked for duplicates and the historical project document fragments; and aggregating the Jaccard similarity coefficients of all fragments through weighted averaging to form the overall similarity of the document;

[0040] The report output module is used to generate the duplicate check result. Specifically, if the overall similarity of the document exceeds the set threshold, it is determined that there is duplication or high similarity between the document to be checked for the project and the project documents in the database, and a project duplicate check report is generated. Among them, the project duplicate check report includes a list of similar projects and similarity score information.

[0041] Preferably, the system further includes an optimization module for optimizing the overall similarity of the document. Specifically, first, through multiple rounds of experiments and data analysis, the embedding threshold is adjusted, then the duplicate check efficiency and accuracy under different embedding thresholds are evaluated to obtain the evaluation result, and then according to the evaluation result, the optimal embedding threshold is selected;

[0042] After the vector distance calculation, the comparison module introduces a re-rank model to re-rank the retrieval results to further accurately locate the projects with high similarity. Specifically, first, the re-rank model is used to re-rank the topK historical project document segments; then the Jaccard similarity coefficient is recalculated to ensure the accuracy of the duplicate check result.

[0043] An electronic device, characterized in that it includes: a memory and at least one processor;

[0044] Wherein, a computer program is stored on the memory;

[0045] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the method for checking the duplicate of the science and technology project application form based on the large language model and the Jaccard similarity coefficient as described above.

[0046] A computer-readable storage medium, in which a computer program is stored, and the computer program can be executed by a processor to implement the method for checking the duplicate of the science and technology project application form based on the large language model and the Jaccard similarity coefficient as described above.

[0047] The method and system for checking the duplicate of the science and technology project application form based on the large language model and the Jaccard similarity coefficient of the present invention have the following advantages:

[0048] (1) The present invention adopts text similarity calculation and document core content extraction and noise reduction technologies based on the application of the large language model, which can efficiently identify and process key information in the document, providing a solid foundation for subsequent duplicate check work;

[0049] (2) The document processing and analysis technologies of the present invention, such as document content extraction, document chunking, and document vectorized storage, also involve the system architecture design and implementation of the duplicate checking system from document preprocessing, storage, document duplicate checking to report output, which not only improves the duplicate checking efficiency, but also realizes the dual advantages of high efficiency and sustainable development by reducing unnecessary calculation processes;

[0050] (3) The present invention provides a project management auxiliary tool for science and technology management workers, focusing on the duplicate checking of science and technology project application forms; through the application of the present invention, an efficient science and technology project management working mechanism can be effectively established, and strong technical support can be provided for the construction of the academic integrity system;

[0051] (4) The present invention introduces the Jaccard similarity coefficient into the field of science and technology project duplicate checking. By accurately calculating the similarity of the keyword sets of project documents, it can effectively identify potential duplicate or highly similar projects, thereby improving the accuracy and efficiency of duplicate checking;

[0052] (5) The present invention can accurately analyze science and technology project documents, effectively remove irrelevant information, and significantly improve the accuracy of duplicate checking;

[0053] (6) The present invention improves the duplicate checking efficiency: facing the duplicate checking scenario of a large number of science and technology projects, it greatly improves the duplicate checking efficiency;

[0054] (7) The present invention improves the accuracy of duplicate checking: the large language model in the present invention has been trained with a large amount of industry data, and can accurately extract the core information in the project application form; greatly increasing the accuracy and reliability of the duplicate checking results;

[0055] (8) The present invention realizes intelligent management: the present invention brings an intelligent duplicate checking method to science and technology management work, reduces manual intervention, improves the intelligent level of science and technology management, and helps to build a more efficient and standardized science and technology management system;

[0056] (9) Economic benefits: Through an efficient duplicate checking mechanism, duplicate project establishment and resource waste are avoided, and the overall efficiency of science and technology projects is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The present invention will be further described below with reference to the accompanying drawings.

[0058] Appendix Figure 1 is a structural block diagram of a science and technology project application form duplicate checking system based on a large language model and the Jaccard similarity coefficient. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0059] The method and system for checking the duplicates of science and technology project application forms based on a large language model and the Jaccard similarity coefficient of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.

[0060] Example 1:

[0061] This example provides a method for detecting duplicate submissions of scientific and technological project applications based on a large language model and the Jaccard similarity coefficient. The method is as follows:

[0062] S1. Extract the core content of the document: Use a large language model that has been trained with data from a number of scientific and technological projects to extract the core content of the scientific and technological project document to be checked for duplicates.

[0063] S2. Split the document into fragments: Based on natural language processing techniques, split the core content of the document into multiple document fragments; each fragment includes one or more key sentences for subsequent vectorization processing.

[0064] S3. Vectorize and store the document: Convert the document fragments into vectors through a text embedding model and store the vectors in a vector database, significantly reducing the consumption of comparison computing resources for irrelevant documents.

[0065] S4. Calculate the vector distance and retrieve historical projects: Calculate the Euclidean distance or cosine similarity between the vectors of the document fragments of the project to be checked for duplicates and the vectors of the document fragments of historical projects, and extract the top K historical projects with the closest distance.

[0066] S5. Calculate the Jaccard similarity coefficient: Recalculate the Jaccard similarity coefficient between the document fragments of the project to be checked for duplicates and the document fragments of historical projects to evaluate the similarity between them.

[0067] S6. Aggregate the document similarity: Aggregate the Jaccard similarity coefficients of all fragments through weighted averaging to form the overall similarity of the document, improving the reliability of the results.

[0068] S7. Generate the duplicate check result: If the overall similarity of the document exceeds the set threshold, it is determined that there are duplicates or high similarities between the scientific and technological project document to be checked for duplicates and the project documents in the database, and a project duplicate check report is generated; the project duplicate check report includes a list of similar projects and similarity score information.

[0069] The specific process of extracting the core content of the document in step S1 of this example is as follows:

[0070] S101. Input the scientific and technological project document to be checked for duplicates into the large language model.

[0071] S102. The large language model performs semantic analysis on the scientific and technological project document to be checked for duplicates, extracts core keywords and key sentences.

[0072] S103. Generate a core content summary of the document based on the core keywords and key sentences.

[0073] The splitting of the document fragments in step S2 of this embodiment is specifically as follows:

[0074] S201. Split the abstract of the core content of the document into multiple document fragments;

[0075] S202. Perform preprocessing operations of word segmentation and stop word removal on each document fragment.

[0076] The calculation of the Jaccard similarity coefficient in step S5 of this embodiment is specifically as follows:

[0077] S501. Extract the keyword sets for each document fragment of the item to be checked for duplication and the historical item document fragment;

[0078] S502. Calculate the intersection and union of the two sets;

[0079] S503. Calculate the Jaccard similarity coefficient, and the formula is: Jaccard coefficient = the number of keywords in the intersection / the number of keywords in the union.

[0080] In this embodiment, the overall similarity of the optimized document is specifically as follows:

[0081] ① Adjust the embedding threshold through multiple rounds of experiments and data analysis;

[0082] ② Evaluate the duplicate checking efficiency and accuracy under different embedding thresholds to obtain the evaluation results;

[0083] ③ Select the optimal embedding threshold according to the evaluation results.

[0084] In this embodiment, after calculating the vector distance, a re-rank model is introduced to re-rank the retrieval results to further accurately locate the items with high similarity, specifically as follows:

[0085] ① Use the re-rank model to re-rank the top K historical item document fragments;

[0086] ② Recalculate the Jaccard similarity coefficient to ensure the accuracy of the duplicate checking results.

[0087] Among them, the Jaccard similarity coefficient is an index widely used in the calculation of set similarity, and it measures the similarity between two sets by comparing their intersection and union.

[0088] Embodiment 2:

[0089] As shown in the appendix Figure 1 This embodiment provides a duplicate checking system for science and technology project application forms based on a large language model and the Jaccard similarity coefficient. The system includes:

[0090] A data preprocessing module, which is used to extract the core content of the scientific and technological project document to be checked for duplication through a large language model that has been trained with the data of several scientific and technological projects, split the core content of the document into multiple document fragments based on natural language processing technology; and convert the document fragments into vectors through a text embedding model (embedding model);

[0091] A data storage module, which is used to store the vectors in a vector database;

[0092] A comparison module, which is used to calculate the Euclidean distance or cosine similarity between the vectors of the document fragments of the project to be checked for duplication and the vectors of the document fragments of historical projects, and extract the top K historical projects with the closest distance; recalculate the Jaccard similarity coefficient between the document fragments of the project to be checked for duplication and the document fragments of historical projects, and evaluate the similarity between the document fragments of the project to be checked for duplication and the document fragments of historical projects; and aggregate the Jaccard similarity coefficients of all fragments by weighted average to form the overall similarity of the document;

[0093] A report output module is used to generate a duplicate check result. Specifically, if the overall similarity of the document exceeds the set threshold, it is determined that there is duplication or high similarity between the project document to be checked for duplication and the project documents in the database, and a project duplicate check report is generated; among them, the project duplicate check report includes a list of similar projects and similarity score information.

[0094] This embodiment further includes an optimization module, which is used to optimize the overall similarity of the document. Specifically, first, through multiple rounds of experiments and data analysis, the embedding threshold is adjusted, then the duplicate check efficiency and accuracy under different embedding thresholds are evaluated to obtain evaluation results, and then according to the evaluation results, the optimal embedding threshold is selected.

[0095] After calculating the vector distance, the comparison module in this embodiment introduces a re-rank model to re-rank the retrieval results to further accurately locate projects with high similarity. Specifically, first, use the re-rank model to re-rank the document fragments of the top K historical projects; then recalculate the Jaccard similarity coefficient to ensure the accuracy of the duplicate check result.

[0096] Embodiment 3:

[0097] This embodiment also provides an electronic device, including: a memory and a processor;

[0098] Wherein, the memory stores computer execution instructions;

[0099] The processor executes the computer execution instructions stored in the memory, so that the processor executes the method for checking the duplication of scientific and technological project application forms based on a large language model and Jaccard similarity coefficient in any embodiment of the present invention.

[0100] The processor can be a central processing unit (CPU), or it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0101] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, at least one magnetic disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0102] Embodiment 4:

[0103] This embodiment also provides a computer-readable storage medium, in which multiple instructions are stored. The instructions are loaded by the processor to cause the processor to execute the method for checking the duplication of science and technology project application forms based on the large language model and the Jaccard similarity coefficient in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.

[0104] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0105] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0106] In addition, it should be clear that not only can the partial or all of the actual operations be completed by executing the program code read by a computer, but also by the operating system operating on the computer and other components based on the instructions of the program code, so as to implement the functions of any one of the above embodiments.

[0107] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer. Subsequently, based on the instructions of the program code, the CPU and other components installed on the expansion board or the expansion unit execute the partial and all of the actual operations, so as to implement the functions of any one of the above embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for checking the similarity of science and technology project application forms based on large language models and Jaccard similarity coefficients, characterized in that, The method is as follows: Extract the core content of the document: Use a large language model that has been trained with data from several scientific and technological projects to extract the core content of the scientific and technological project document to be checked for duplication; Split the document fragments: Based on natural language processing technology, split the core content of the document into multiple document fragments; each fragment includes one or more key sentences; Vectorize and store the document: Convert the document fragments into vectors through a text embedding model and store the vectors in a vector database; Calculate the vector distance and retrieve historical projects: Calculate the Euclidean distance or cosine similarity between the vectors of the document fragments of the project to be checked for duplication and the vectors of the historical project document fragments, and extract the top K historical projects with the closest distance; Calculate the Jaccard similarity coefficient: Recalculate the Jaccard similarity coefficient between the document fragments of the project to be checked for duplication and the historical project document fragments to evaluate the similarity between them; Aggregate the document similarity: Aggregate the Jaccard similarity coefficients of all fragments using weighted average to form the overall similarity of the document; Generate the duplication check result: If the overall similarity of the document exceeds the set threshold, it is determined that there is duplication or high similarity between the scientific and technological project document to be checked for duplication and the project documents in the database, and a project duplication check report is generated; the project duplication check report includes a list of similar projects and similarity score information.

2. The method for checking the duplication of scientific and technological project application forms based on large language models and Jaccard similarity coefficients according to claim 1, characterized in that, The specific process of extracting the core content of the document is as follows: Input the scientific and technological project document to be checked for duplication into the large language model; The large language model performs semantic analysis on the scientific and technological project document to be checked for duplication, extracts core keywords and key sentences; Generate a core content summary of the document based on the core keywords and key sentences.

3. The method for checking the similarity of scientific and technological project application forms based on large language models and Jaccard similarity coefficients according to claim 1, wherein, The specific process of splitting the document fragments is as follows: Split the core content summary of the document into multiple document fragments; Perform preprocessing operations of word segmentation and stop word removal on each document fragment.

4. The method for checking the similarity of scientific and technological project application forms based on large language models and Jaccard similarity coefficients according to claim 1, characterized in that, The specific process of calculating the Jaccard similarity coefficient is as follows: For each document fragment of the project to be checked for duplication and the historical project document fragment, extract the keyword sets; Calculate the intersection and union of the two sets; Calculate the Jaccard similarity coefficient, and the formula is: Jaccard coefficient = the number of keywords in the intersection / the number of keywords in the union.

5. The method for detecting duplicate submissions of science and technology project application forms based on large language models and Jaccard similarity coefficients according to claim 1, wherein Optimize the overall similarity of the document, specifically as follows: Adjust the embedding threshold through multiple rounds of experiments and data analysis; Evaluate the duplication check efficiency and accuracy under different embedding thresholds to obtain the evaluation results; Select the optimal embedding threshold according to the evaluation results.

6. The method for checking the duplication of science and technology project application forms based on large language models and Jaccard similarity coefficients according to any one of claims 1-5, characterized in that, After calculating the vector distance, introduce a re-rank model to re-rank the retrieval results to further accurately locate projects with high similarity, specifically as follows: Use the re-rank model to re-rank the top K historical project document fragments; Recalculate the Jaccard similarity coefficient to ensure the accuracy of the duplication check results.

7. A plagiarism detection system for science and technology project application forms based on large language models and Jaccard similarity coefficients, characterized in that, The system includes: A data preprocessing module, which is used to extract the core content of the scientific and technological project document to be checked for duplication through a large language model that has been trained with data from several scientific and technological projects, split the core content of the document into multiple document fragments based on natural language processing technology; and convert the document fragments into vectors through a text embedding model; A data storage module for storing vectors in a vector database; A comparison module for calculating the Euclidean distance or cosine similarity between the vectors of the document fragments of the project to be checked for duplication and the vectors of the document fragments of historical projects, and extracting the top K historical projects with the closest distance; recalculating the Jaccard similarity coefficient between the document fragments of the project to be checked for duplication and the document fragments of historical projects to evaluate the similarity between the document fragments of the project to be checked for duplication and the document fragments of historical projects; and aggregating the Jaccard similarity coefficients of all fragments by weighted average to form the overall similarity of the document; A report output module for generating a duplicate check result, specifically: if the overall similarity of the document exceeds a set threshold, it is determined that there is duplication or high similarity between the document of the project to be checked for duplication and the project documents in the database, and a project duplicate check report is generated; wherein, the project duplicate check report includes a list of similar projects and similarity score information.

8. The technology project application form duplicate checking system based on the large language model and Jaccard similarity coefficient according to claim 7, characterized in that, The system further includes an optimization module for optimizing the overall similarity of the document, specifically: first adjusting the embedding threshold through multiple rounds of experiments and data analysis, then evaluating the duplicate check efficiency and accuracy under different embedding thresholds to obtain an evaluation result, and then selecting the optimal embedding threshold according to the evaluation result; After calculating the vector distance, the comparison module introduces a re-rank model to re-rank the retrieval results to further accurately locate projects with high similarity, specifically: first using the re-rank model to re-rank the document fragments of the top K historical projects; then recalculating the Jaccard similarity coefficient to ensure the accuracy of the duplicate check result.

9. An electronic device, characterized in that, Including: A memory and at least one processor; Wherein, a computer program is stored on the memory; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the method for checking the duplicate of the science and technology project application form based on the large language model and the Jaccard similarity coefficient as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program can be executed by the processor to implement the method for checking the duplicate of the science and technology project application form based on the large language model and the Jaccard similarity coefficient as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Report filling-in-place unit identification method and device based on multi-level clustering

    CN121117212A