Science and technology project document duplicate checking method and device based on multi-level abstract generation and medium

By constructing a multi-level abstract generation model and vector database, and combining large language models to conduct plagiarism checking for scientific and technological project documents, the problem of lack of deep semantic understanding of the plagiarism checking methods in the existing technology is solved, and higher plagiarism checking accuracy and reliability are achieved, and are suitable for scientific and technological project management.

CN120509394APending Publication Date: 2025-08-19STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1

Patent Information

Application Number
CN202510548709.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing scientific and technological project document plagiarism checking methods lack deep semantic understanding, resulting in insufficient reliability and accuracy of plagiarism checking results and limited application scope.

Method used

By constructing a multi-level abstract generation model, using large language models and manual rules to generate sample pairs, fine-tune them in combination with text feature extraction models, constructing vector databases for historical scientific and technological projects, performing structured analysis and similarity search, and generating plagiarism check reports.

Benefits of technology

It improves the accuracy and reliability of scientific and technological project document plagiarism checking, can accurately capture deep similarities, expands the scope of application of plagiarism checking methods, ensures project originality, and avoids duplicate funding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509394A_ABST
    Figure CN120509394A_ABST
Patent Text Reader

Abstract

The invention discloses a science and technology project document duplicate checking method and device based on multi-level abstract generation and a medium, belongs to the technical field of document duplicate checking, and solves the problem of how to improve the reliability and accuracy of science and technology project document duplicate checking. The method comprises the following steps: firstly, carrying out structured analysis on a project document to be subjected to duplicate checking, introducing a fine-tuned abstract generation model to obtain a multi-level abstract of a project to be subjected to duplicate checking, and extracting a feature vector of the multi-level abstract through a fine-tuned and trained text feature extraction model; secondly, performing similarity search in a vector database based on feature vectors corresponding to the dimensions of the main abstract, further calculating cosine similarities based on the dimensions of the subordinate abstract, and sequencing the weighted similarities to obtain a similarity rank; and finally, according to the structured analysis information and the similarity ranking result, LLM similarity analysis is performed by using a large language model, a document duplicate checking report is formed, the originality of the science and technology project is ensured, repeated subsidization is avoided, and the method has important practical application value for science and technology project management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of document duplication checking, and relates to a method, device and medium for checking duplicate content in scientific and technological project documents based on multi-level summary generation. Background Art

[0002] In science and technology project management, checking for duplicate content in project applications is a crucial step to ensure originality and avoid duplicate funding. Existing methods for checking for duplicate content rely primarily on keyword matching and simple text similarity calculations, often failing to accurately capture the deeper similarities within project content, resulting in insufficient reliability and accuracy in the results.

[0003] Prior art, such as the invention patent application with publication number CN118036586A, discloses a document duplication detection method for an intelligent document review system. This method uses the TF-IDF model to perform vectorization processing based on optimized word segmentation results, obtaining document feature vectors for the document to be checked for duplicates and the target document to be checked for duplicates. The degree of duplication in the document to be checked for duplicates is then determined by measuring the cosine similarity between the document feature vectors. However, TF-IDF calculates word importance based on word frequency and does not consider the semantic relationship between words. For example, "car" and "sedan," while having similar meanings, may be considered completely different words in the TF-IDF model. Furthermore, because TF-IDF focuses only on the frequency of occurrence of individual words, it cannot effectively handle words whose meaning depends on context. For example, "apple" may refer to fruit or the technology company Apple in different sentences.

[0004] For example, the invention patent application with the publication number CN118313365A discloses a method for checking duplicate content in distribution network projects based on natural language processing. The method uses a word segmentation tool to segment the project name and check for duplicate content. If the similarity shown in the check result is greater than the set threshold, the project book or research report is further checked for duplicate content. However, this method relies on the construction of duplicate content checking logic rules. It is necessary to aggregate and classify the keywords based on the importance of each keyword and set different weights. For example, this method provides a duplicate content checking logic rule for a distribution network scenario. When it is needed for other application scenarios, the duplicate content checking logic rule needs to be reset. The scope of application of this method is limited and it does not have good promotion significance.

[0005] For example, the invention patent application CN117951256A discloses a document duplication detection method based on hierarchical feature vector search. This method sequentially extracts the first-level word frequency feature vector, the second-level context semantic feature vector, and the third-level topic distribution feature vector of the document to be checked for duplicates, constructs an inverted index, and calculates the similarity between the document to be checked for duplicates and each document in the index. However, while this method can capture the basic semantic information of the document and the semantic associations between words, it does not consider the multi-level content of the project document. Extracting multi-level feature vectors and calculating similarity for the entire document requires a large amount of computing resources, and the granularity of the duplicate detection is relatively coarse.

[0006] Therefore, it is particularly important to overcome the technical difficulties of existing scientific and technological project document duplication detection methods, such as the lack of deep semantic understanding, duplication detection accuracy, and poor promotion capabilities. Summary of the Invention

[0007] The technical problem to be solved by the present invention is how to improve the reliability and accuracy of duplicate checking of scientific and technological project documents.

[0008] The present invention solves the above technical problems through the following technical solutions:

[0009] The method for checking duplicate documents of scientific and technological projects based on multi-level summary generation includes:

[0010] S1. Preprocess the scientific and technological project documents, divide the scientific and technological project documents into multiple text blocks according to the chapter structure, and divide each text block into project-level and topic-level content;

[0011] S2. Construct a text pair dataset based on a large language model and manual rules, build and train a text feature extraction model for scientific and technological projects;

[0012] S3. Construct a multi-dimensional dataset dedicated to summary generation, build a summary generation model, and fine-tune it. Combined with the text feature extraction model, the summary generation model is constrained for consistency from both intra-document and cross-document perspectives.

[0013] S4. Build a vector database of historical scientific and technological projects, perform structured analysis on scientific and technological project documents to be checked for duplicates, pass the results into a summary generation model to obtain multi-level summaries, perform similarity searches in the vector database, and calculate similarity rankings;

[0014] S5. The structured information and similarity rankings of similar scientific and technological projects and the scientific and technological projects to be checked for duplicate are passed into the large language model for similarity analysis to generate a duplicate checking report.

[0015] Furthermore, the step S2 includes the following steps:

[0016] S21. Generate sample pairs as training data using a large language model and manual rules, where each piece of training data includes a text pair and a similarity score, wherein the text pairs include related text pairs extracted from item-level and topic-level texts in the same document, and text pairs where text items in different documents are paired with each other;

[0017] S22, after the text pair is constructed, use the preset instruction prompt words to let multiple large language models generate similarity scores S = {s1, s2, s3, ..., s p×k}, obtain high-quality annotated text pair dataset Among them, p is the number of instruction prompt words, and k is the number of large language models used;

[0018] S23. Using the text pair dataset to pre-train the text semantic feature extraction model Fine-tune and obtain a text semantic feature extraction model specifically for scientific and technological projects The following logic is used to express the pre-trained text semantic feature extraction model using the text pair dataset Make fine adjustments:

[0019]

[0020] Among them, N is the number of text pairs required to calculate a loss, x n , x′ n Represents text pair datasets Two texts of a single training data in , n∈[1,N] and is an integer, sim(·) is the similarity of the text pair calculated according to the model, y n Similarity scores for large language model annotations.

[0021] Furthermore, the pre-trained text semantic feature extraction model Select the BGE-large-zh-v1.5 model.

[0022] Furthermore, the S3 includes:

[0023] S31. Divide the text blocks of each part from the scientific and technological project documents according to the chapter structure, and construct a multi-dimensional summary generation special dataset through expert annotation;

[0024] S32. Build a summary generation model remember Among them, A is the main summary dimension, P is the instruction prompt word corresponding to A, and A u is the subordinate summary dimension, P u A u The corresponding command prompt word, doc, is the pre-processed scientific and technological project document content;

[0025] S33. Calculate the supervised fine-tuning loss L of the summary generation model CrossEntropy ;

[0026] S34, calculating the intra-document consistency constraint of the fusion text semantic feature extraction model;

[0027] S35. Calculate cross-document consistency constraints of the fusion text semantic feature extraction model;

[0028] S36. Calculate the total loss function L all , fine-tune the summary generation model.

[0029] Furthermore, the S34 includes:

[0030] S341, generate feature vectors of each summary dimension in the document, record Among them, V A is the eigenvector of A, A u The eigenvector of Represents the fine-tuned text semantic feature extraction model;

[0031] S342, generate feature vectors of all original text entries in the document, record Among them, t m For the lowest level single text entry, V m The lowest level single text entry t m The corresponding feature vector, M is the number of single text entries at the bottom level;

[0032] S343. Constructing feature consistency constraints L within the document intro :

[0033]

[0034] Among them, v1 is the total project overview vector, v2 is the average value of the summary vectors of different dimensions, and v3 is the feature vector of all original text entries in the document. Take v1 = V A , Represents the square of the Euclidean distance.

[0035] Furthermore, the S35 includes:

[0036] S351, generate cross-document summary, record Among them, doc j For the jth document, A j For doc j The main summary dimension of P j Aj The corresponding command prompt words, For doc j The subordinate summary dimension of for Corresponding command prompt words;

[0037] S352, calculating the similarity of the summary text features, wherein the similarity of the summary text features includes the similarity sim1 of the main summary dimension and the similarity sim2 of the subordinate summary dimension;

[0038] S353, calculating the similarity sim3 of the bottom-up text features;

[0039] S354. Construct cross-document feature consistency constraint L inter :

[0040]

[0041] Furthermore, the summary generation model The model structure and initialization parameters adopt Qwen2.5-7B-Instruct. The model input is the scientific and technological project text, and the output is the summarized project text.

[0042] Furthermore, the S4 includes:

[0043] S41. Use existing scientific and technological project documents to build a vector database of historical scientific and technological projects

[0044] S42, perform structural analysis on the scientific and technological project documents to be checked for duplicate content, recorded as doc * , pass the parsing results into the fine-tuned summary generation model to obtain the multi-level summary of the model, which includes the multi-dimensional project summary text A * , and project structured text

[0045] S43, calculate the feature vector of the summary text corresponding to the main summary dimension through the text semantic feature extraction model, record in, Abstract text A * The eigenvector of Abstract text ; Search in the vector database to find the top c items with the highest semantic content similarity, and obtain the summary text of the subordinate summary dimension; wherein, the following logic is used to express the execution in the vector database Search in the query to find the top c items with the highest semantic content similarity in the summary

[0046]

[0047] in, Indicates that from the vector database according to The first c items after similarity sorting, Indicates that the vector database in the previous operation is

[0048] S44. Calculate the cosine similarity of the summary texts corresponding to the subordinate summary dimension, sort the weighted similarities, and obtain a similarity ranking. The following logic is used to express the calculation of the cosine similarity of the summary texts corresponding to the subordinate summary dimension:

[0049]

[0050] Among them, w u is the weight of the u-th subordinate summary dimension, doc for the project document to be checked for duplicates * The inner product of the feature vector of the u-th summary dimension of the query entry z.

[0051] An electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above-mentioned method for checking duplicate technology project documents based on multi-level summary generation, and the processor is configured to execute the program stored in the memory.

[0052] A storage medium stores a computer program, which, when run by a processor, executes the steps of the above-mentioned method for checking duplicate technology project documents based on multi-level summary generation.

[0053] The advantages of the present invention are:

[0054] (1) The present invention constructs a summary text dataset containing multiple dimensions in project documents, and fine-tunes the summary generation model by jointly supervising fine-tuning loss, consistency constraints within the same document, and consistency constraints across documents. By combining intra-document feature alignment and cross-document feature alignment, the deviation of the model output can be reduced.

[0055] (2) The present invention first performs a structured analysis on the project documents to be checked for duplicate content, and inputs them into a fine-tuned summary generation model to obtain a multi-level summary of the project to be checked for duplicate content, and extracts the feature vectors of the multi-level summary through a fine-tuned and trained text feature extraction model; secondly, a similarity search is performed in the vector database based on the feature vectors corresponding to the main summary dimensions, and an initial screening is performed when checking for duplicate content in the project to narrow the scope of the check, and then the cosine similarity based on the subordinate summary dimension is further calculated, and the weighted similarity is sorted to obtain a similarity ranking; finally, based on the structured analysis information and the similarity ranking results, a large language model is used to perform LLM similarity analysis to form a document duplication check report, thereby ensuring the originality of the scientific and technological projects and avoiding duplicate funding, which has important practical application value for scientific and technological project management.

[0056] (3) The present invention can more accurately identify and compare the similarities of scientific and technological project documents, accurately capture the deep similarities of project contents, overcome the technical difficulties of existing document duplication detection methods, such as lack of deep semantic understanding, poor duplication detection accuracy, and poor promotion capabilities, effectively improve the duplication detection depth and versatility of the method, and ensure the reliability and accuracy of the duplication detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a flow chart of a method for checking duplicate technology project documents based on multi-level summary generation according to a first embodiment of the present invention;

[0058] Figure 2 This is an example diagram of the original scientific and technological project document structure of the first embodiment of the present invention;

[0059] Figure 3 This is an example diagram of the scientific and technological project documents after pre-processing and division according to the first embodiment of the present invention;

[0060] Figure 4 This is an example diagram of a single piece of training data according to the first embodiment of the present invention;

[0061] Figure 5 This is an example diagram of the preset instruction prompt words in the first embodiment of the present invention;

[0062] Figure 6 Schematic diagram of the summary generation model fine-tuning process according to the first embodiment of the present invention;

[0063] Figure 7 Schematic diagram of the consistency constraint of the summary generation model within the same document according to the first embodiment of the present invention;

[0064] Figure 8 Schematic diagram of consistency constraints of the inter-document summary generation model according to the first embodiment of the present invention;

[0065] Figure 9This is an example diagram of the vector database organizational structure of the first embodiment of the present invention;

[0066] Figure 10 This is a flowchart of the scientific and technological project similarity calculation according to the first embodiment of the present invention;

[0067] Figure 11 Schematic diagram of the steps for checking duplicate technology project documents to be checked for duplicate technology according to the first embodiment of the present invention;

[0068] Figure 12 Schematic diagram of instruction prompt words used in similarity analysis according to the first embodiment of the present invention. DETAILED DESCRIPTION

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0070] The technical solution of the present invention is further described below with reference to the accompanying drawings and specific embodiments:

[0071] Example 1

[0072] like Figure 1 Specifically, a method for in-depth duplicate checking of scientific and technological projects based on multi-level summary generation is disclosed, comprising the following steps:

[0073] S1. Preprocessing stage: divide the scientific and technological project documents into multiple text blocks according to the chapter structure, and divide each text block into project-level and topic-level content;

[0074] Furthermore, the scientific and technological project documents include but are not limited to project task orders, project requirements tables, project feasibility study reports, project work reports, project acceptance reports, etc.

[0075] In this embodiment, a project task book is taken as an example for explanation. The project task book document generally includes project-level content and its corresponding subject-level content. Document processing technologies such as Python-Docx or OCR are used to divide the scientific and technological project document into chapter structures, obtain the text blocks of each part, and then divide each text block into project-level and subject-level content.

[0076] This embodiment is described by taking the project task book as an example. Figure 2 The structure of the original scientific and technological project document is shown in the figure. Since many contents in the original document are irrelevant to the requirements of duplicate checking, this embodiment processes the scientific and technological project document in the pre-processing stage as follows: Figure 3 The document structure shown is divided into key contents at the project level and subject level, formatted, and the parts required for duplicate checking are extracted. Figure 2 Extracted from the structure Figure 3 part.

[0077] S2. Construct a text pair dataset based on a large language model and manual rules, build a text feature extraction model for scientific and technological projects, and train it. This specifically includes the following steps:

[0078] S21. Generate sample pairs as training data using a large language model and manual rules, where a single piece of training data includes a text pair and a similarity score.

[0079] Furthermore, the text pairs include related text pairs extracted from project-level and topic-level texts in the same document, as well as text pairs in which text items in different documents are paired with each other.

[0080] In this embodiment, in order to construct a high-quality data set, on the one hand, it is necessary to ensure that the similarity scores in the data set are relatively evenly distributed, and on the other hand, it is necessary to ensure that the similarity scores are consistent with reality. Therefore, it is necessary to include training data with low similarity scores as well as training data with high similarity scores. Since most of the existing scientific and technological project documents are not related, a large amount of data with low similarity scores can be constructed by pairing the text entries in the documents with each other. For projects that have been approved, they can be naturally considered to be different contents, so it is easy to obtain text pairs with low similarity. The key to obtaining data with relatively high similarity scores by pairing the document contents between different documents is how to construct data with relatively high similarity scores. By observing the document structure, it can be found that the project-level content and the subject-level content of the same document correspond to each other or the relevant part of the text can be used to meet this demand.

[0081] like Figure 4 The figure shows a single piece of training data. A single piece of data consists of a text pair and a similarity score. Text1 and text2 represent text pairs, and sim_score represents the similarity score.

[0082] S22, after the text pair is constructed, use the preset instruction prompt word Prompt to let multiple large language models generate similarity scores S = {s1, s2, s3, ..., s p×k}, obtain high-quality annotated text pair dataset Where p is the number of instruction prompt words, k is the number of large language models used, and to eliminate uncertainty, the final similarity score is obtained by removing extreme values and taking the average. The following logic is used to express the final similarity score s:

[0083]

[0084] Among them, s i is the i-th similarity score, i∈[1,p×k] and is an integer, max(S) is the maximum similarity score, and min(S) is the minimum similarity score. The high-quality annotated text pair dataset is finally obtained and recorded as Among them, T1 and T2 are the two texts in the text pair.

[0085] like Figure 5 The figure shows an example of a specific instruction prompt word Prompt. The instruction prompt word Prompt is a text fragment or question input to the large language model to guide the model to generate corresponding output. Prompt can be a complete sentence, a question, a descriptive beginning, or even a keyword.

[0086] S23. Using text to compare datasets Pre-trained text semantic feature extraction model Fine-tune and obtain a text semantic feature extraction model specifically for scientific and technological projects

[0087] This embodiment uses the loss function L MSE Model Fine-tune the pre-trained text semantic feature extraction model using the following logic Make fine adjustments:

[0088]

[0089] Among them, N is the number of text pairs required to calculate a loss, x n , x′ n Represents text pair datasets Two texts of a single training data in n∈[1,N] are integers, sim(·) is the similarity of the text pair calculated according to the model, y n Similarity scores for large language model annotations.

[0090] In this embodiment, the pre-trained text semantic feature extraction model Select the BGE-large-zh-v1.5 model to generate text features.

[0091] S3. Construct a multi-dimensional dataset dedicated to summary generation, build a summary generation model, and fine-tune it. Combined with the text feature extraction model, constrain the summary generation model's consistency from both within-document and across-document perspectives. Specifically, the following steps are included:

[0092] S31. Divide the text blocks of each part from the scientific and technological project documents according to the chapter structure, and construct a multi-dimensional summary generation-specific dataset through expert annotation to fine-tune the instructions of the summary generation model.

[0093] The multi-dimensional summary generation dedicated data set includes multiple summary dimensions and instruction prompt words corresponding to the multiple summary dimensions. In this embodiment, the multiple summary dimensions include the following six dimensions: project overview, research objectives, research content and technical route, innovation points and technical advantages, expected results and indicators, application scenarios and market prospects. Each summary dimension is represented by A, A1, A2, A3, A4, and A5 respectively; for each summary dimension, the instruction prompt words used to generate the summary are represented by P, P1, P2, P3, P4, and P5 respectively. The instruction prompt words are used to fine-tune the instructions of the summary generation model.

[0094] S32. Build a summary generation model The following logic is used:

[0095]

[0096] Among them, doc is the pre-processed scientific and technological project document content, A is the main summary dimension, which in this embodiment represents the summary dimension corresponding to the project overview, A u is the subordinate summary dimension, P is the instruction prompt word corresponding to A, P u A u The corresponding instruction prompt words in this embodiment correspond to research objectives, research content and technical routes, innovations and technical advantages, expected results and indicators, application scenarios and market prospects.

[0097] In this embodiment, the summary generation model The model structure and initialization parameters adopt Qwen2.5-7B-Instruct. The model input is the scientific project text, and the output is the summarized project text. Qwen2.5-7B-Instruct extends the context of the open source Qwen model to 1M length, making the summary generation model The output summary text is shorter and more concise.

[0098] S33. Computational summary generation model The supervised fine-tuning loss L CrossEntropy .

[0099] In this embodiment, the supervised fine-tuning loss represents the distribution of model output when fine-tuning the model based on the dedicated dataset generated by the summary. With the target distribution y u This embodiment uses cross entropy loss and uses the following logic to calculate the supervised fine-tuning loss:

[0100]

[0101] Among them, L CrossEntropy To supervise the fine-tuning loss, is the model output distribution, y u is the target distribution.

[0102] like Figure 6 As shown in the figure, the scientific and technological project documents are divided into multiple text blocks according to the chapter structure. A multi-dimensional summary generation dataset is constructed by expert annotation, and the summary generation model is trained by supervised fine-tuning.

[0103] S34. Computational fusion text semantic feature extraction model Consistency constraints within the document.

[0104] In this embodiment, if Figure 7 As shown, the S34 specifically includes the following steps:

[0105] S341. Generate feature vectors for each summary dimension in the document, using the following logic:

[0106]

[0107] Among them, V A is the eigenvector of A, represents the fine-tuned text semantic feature extraction model, A u The eigenvector of .

[0108] S342. Generate feature vectors for all original text entries in the document using the following logic:

[0109]

[0110] Among them, V m The lowest level single text entry t m The corresponding eigenvector, t m is the lowest level single text entry, and M is the number of the lowest level single text entries.

[0111] S343. Constructing feature consistency constraints L within the document intro , using the following logic:

[0112]

[0113]

[0114] Among them, v1 is the total project overview vector. In this embodiment, the feature vector of the main summary dimension is taken. v2 is the average value of the summary vectors of different dimensions. In this embodiment, the average value of the feature vectors of the subordinate summary dimension is taken. v3 is the feature vector of all original text items in the document. center is the average value of v1, v2, and v3, v w is the w-th eigenvector, w=1,2,3, Represents the square of the Euclidean distance. In this embodiment, v1, v2, and v3 are used to represent bottom-up text features.

[0115] S35. Computational fusion text semantic feature extraction model Cross-document consistency constraints.

[0116] In this embodiment, the summary text features and bottom-up text features between documents are obtained respectively to ensure that the similarity calculated using the summary text features is consistent with the similarity calculated using the bottom-up text features. Figure 8 As shown, taking two documents doc1 and doc2 as an example, the following steps are specifically included:

[0117] S351. Generate a cross-document summary using the following logic:

[0118]

[0119] Among them, doc j For the jth document, A j For doc j The main summary dimension of P j A j The corresponding command prompt words, For doc j The subordinate summary dimension of for The corresponding command prompt word.

[0120] S352. Calculate the similarity of the summary text features using the following logic:

[0121]

[0122] Among them, sim1 is the similarity of the main summary dimension, sim2 is the similarity of the subordinate summary dimension, and A 1 is the main summary dimension of doc1, A 2 is the main summary dimension of doc2, is the subordinate summary dimension of doc1, It is the subordinate summary dimension of doc2.

[0123] S353, calculating the similarity of bottom-up text features, including the following steps:

[0124] In this embodiment, the documents are divided into blocks according to the structure of the key content extracted in the preprocessing stage, and then the bottom-up similarity between the two documents is calculated. Taking the subject-level content "subject name" as an example, assuming that the research content of documents doc1 and doc2 has H and Q items respectively, their similarities can form an H×Q similarity matrix. The element in the hth row and qth column of the matrix is a hq , then the similarity of the two documents in terms of research content is expressed using the following logic:

[0125] Among them, sim 课题名称 is the similarity between two documents with the content of “topic name” at the topic level, max q (a hq ) is the maximum value in the hth row.

[0126] The final bottom-up similarity is the average of all sub-block similarities, expressed using the following logic:

[0127]

[0128] Among them, sim3 is the similarity of bottom-up text features, L is the number of elements in Ω, L = |Ω|, Ω is the topic-level content of the text block, Ω = {Ω1, Ω2, ...}, in this embodiment, Ω = {topic name, research content, ...}, sim l It is actually the similarity of the text content of the lth topic level content, expressed as sim 课题名称 For example, the actual calculation is as follows Figure 3 The text content after the "Topic Name" shown.

[0129] S354. Construct cross-document feature consistency constraint L inter , using the following logic:

[0130]

[0131] Among them, sim w For the wth similarity, take sim1, sim2 and sim3 in steps S352-S353, sim mean It is the mean of the similarity sim1 of the main summary dimension, the similarity sim2 of the subordinate summary dimension, and the similarity sim3 of the bottom-up text features.

[0132] S36. Calculate the total loss function L all , for the summary generation model Fine-tune training, using the following logic to express the total loss function L all :

[0133] L all =λ1L CrossEntropy +λ2l intro +λ3l inter

[0134] Among them, λ1, λ2, and λ3 correspond to the weight coefficients of the loss function respectively. They are adjusted and optimized during model fine-tuning training to ensure that λ1, λ2, and λ3 are all greater than 0 and their sum is 1.

[0135] S4. Build a vector database of historical scientific and technological projects, perform structured analysis on the scientific and technological project documents to be checked for duplicates, pass the results into the summary generation model to obtain multi-level summaries, perform similarity search in the vector database, and calculate the similarity ranking.

[0136] S41. Use existing scientific and technological project documents to build a vector database of historical scientific and technological projects.

[0137] After completing the summary generation model After fine-tuning training, it is necessary to use existing scientific and technological project documents to build a historical scientific and technological project vector library for scientific and technological project duplication checking and analysis. The vector database should not only store the basic information of the original document 1,2,3 ,…, each summary dimension and the corresponding feature vector need to be stored. A single entry in the vector database is recorded as The vector database is denoted as like Figure 9 Shown is the specific organizational structure of the vector database.

[0138] S42, perform structured analysis on the scientific and technological project documents to be checked for duplicates, and record the analysis results as doc * , pass the parsing results into the fine-tuned summary generation model Get a multi-level summary of the model summary, the multi-level summary includes multi-dimensional project summary text A * , and project structured text

[0139] In this embodiment, a preliminary feasibility study report of a project is selected as an example of a scientific and technological project document to be checked for duplicates. The following logic is used to represent a multi-dimensional project summary text:

[0140]

[0141] Among them, A * is the summary text of the main summary dimension, The summary text for the subordinate summary dimension.

[0142] S43. Summary text A corresponding to the main summary dimension * Through text semantic feature extraction model Calculate the feature vector, search in the vector database, find the top c items with the highest summary semantic content similarity, and obtain the summary text of the subordinate summary dimension.

[0143] In this embodiment, the following logic is used to represent the summary text A corresponding to the main summary dimension: * Through text semantic feature extraction model Compute the eigenvectors:

[0144]

[0145] in, Abstract text A * The eigenvector of Abstract text The eigenvector of .

[0146] The following logic is used to express the execution in the vector database Search in the query to find the top c items with the highest semantic content similarity in the summary

[0147]

[0148] in, Indicates that from the vector database according to The first c items after similarity sorting, Indicates that the vector database in the previous operation is

[0149] S44: Calculate the cosine similarity of the summary texts in the subordinate summary dimension, sort the weighted similarities, and obtain a similarity ranking.

[0150] Based on the multi-dimensional summary text of the suspected project retrieved in step S43, the cosine similarity is calculated from the five subordinate summary dimensions: research objectives, research content and technical routes, innovation points and technical advantages, expected results and indicators, and application scenarios and market prospects. The weighted similarities are sorted to obtain the similarity ranking. The cosine similarity of the summary text corresponding to the subordinate summary dimensions is calculated using the following logic:

[0151]

[0152] Among them, w u is the weight of the u-th subordinate summary dimension, doc for the project document to be checked for duplicates *The inner product of the feature vector of the u-th summary dimension of the queried entry z. Since the feature vectors have been normalized within the model, this embodiment uses the inner product of the above two to represent the cosine similarity.

[0153] S5. The structured information and similarity calculation results of similar scientific and technological projects and the scientific and technological projects to be checked for duplicate are passed into the large language model for similarity analysis to generate a duplicate checking report.

[0154] In this embodiment, the ranking results of project document structured information and hierarchical aggregation similarity are utilized, combined with the instruction prompt words used in similarity analysis, and a large language model technology is used to form a document duplication report, giving the project similarity analysis results to help scientific and technological project managers understand the project duplication situation.

[0155] Furthermore, the structured information is the project structured text obtained after the structured parsing in step S42. This embodiment parses the original document (such as a document in word format) into Figure 3 The structure shown is represented in json format in the computer, and can only be used for subsequent feature vector calculation operations based on this format. The similarity calculation result is a similarity ranking, and the duplicate check report includes a conclusion on whether the scientific and technological projects to be checked are similar. If the conclusion is similar, the duplicate check report also includes an analysis of the similarity reasons. Figure 12 Shown are examples of instruction prompt words used in similarity analysis.

[0156] like Figure 10-11 As shown, the present invention first performs structured analysis on the project documents to be checked for duplicate content, and passes them into the fine-tuned summary generation model to obtain a multi-level summary of the project to be checked for duplicate content (stage 1). The feature vector extracted by the text feature extraction model is first queried in the vector database based on the feature vector corresponding to the main summary dimension to find similar projects with high similarity ranking (stage 2). This is used for initial screening when checking for duplicate content. Since the main summary dimension only represents a description of the project as a whole, it can only narrow the scope of the check for duplicate content, while the subordinate summary dimension has a focus on its own dimension compared to the main summary dimension. Therefore, the cosine similarity based on the subordinate summary dimension is further calculated, the weighted similarity is sorted, and the similarity ranking is obtained (stage 3). Finally, based on the ranking results of the project document structured information and the hierarchical aggregation similarity, the large language model is used to perform LLM similarity analysis to form a document duplicate check report (stage 4). This ensures the originality of scientific and technological projects and avoids duplicate funding, which has important practical application value for scientific and technological project management.

[0157] In order to make the summary generation model Adapting to the content of scientific and technological project texts, the model is fine-tuned and trained in scientific and technological project texts through joint supervision fine-tuning loss, intra-document consistency constraints of the fusion feature extraction model, and cross-document consistency constraints of the fusion feature extraction model. Combined with intra-document feature alignment and cross-document feature alignment, the deviation of the model output is reduced.

[0158] The present invention can more accurately identify and compare the similarities of scientific and technological project documents, accurately capture the deep similarities of project contents, overcome the technical difficulties of existing document duplication detection methods such as lack of deep semantic understanding, duplication detection accuracy and poor promotion capabilities, effectively improve the duplication detection depth and method versatility of the duplication detection method, and ensure the reliability and accuracy of the duplication detection results.

[0159] Example 2

[0160] A device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the method for checking duplicate technology project documents based on multi-level summary generation in Example 1, and the processor is configured to execute the program stored in the memory.

[0161] Example 3

[0162] A storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for checking duplicate technology project documents based on multi-level summary generation in embodiment 1.

[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for checking duplicate technology project documents based on multi-level summary generation, characterized in that: include: S1. Preprocess the scientific and technological project documents, divide the scientific and technological project documents into multiple text blocks according to the chapter structure, and divide each text block into project-level and topic-level content; S2. Construct a text pair dataset based on a large language model and manual rules, build and train a text feature extraction model for scientific and technological projects; S3. Construct a multi-dimensional dataset dedicated to summary generation, build a summary generation model, and fine-tune it. Combined with the text feature extraction model, the summary generation model is constrained for consistency from both intra-document and cross-document perspectives. S4. Build a vector database of historical scientific and technological projects, perform structured analysis on scientific and technological project documents to be checked for duplicates, pass the results into a summary generation model to obtain multi-level summaries, perform similarity searches in the vector database, and calculate similarity rankings; S5. The structured information and similarity rankings of similar scientific and technological projects and the scientific and technological projects to be checked for duplicate are passed into the large language model for similarity analysis to generate a duplicate checking report.

2. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 1 is characterized in that: The S2 comprises the following steps: S21. Generate sample pairs as training data using a large language model and manual rules, where each piece of training data includes a text pair and a similarity score, wherein the text pairs include related text pairs extracted from item-level and topic-level texts in the same document, and text pairs where text items in different documents are paired with each other; S22, after the text pair is constructed, use the preset instruction prompt words to let multiple large language models generate similarity scores S = {s1, s2, s3, ..., s p×k }, obtain high-quality annotated text pair dataset Among them, p is the number of instruction prompt words, and k is the number of large language models used; S23. Using the text pair dataset to pre-train the text semantic feature extraction model Fine-tune and obtain a text semantic feature extraction model specifically for scientific and technological projects The following logic is used to express the pre-trained text semantic feature extraction model using the text pair dataset Make fine adjustments: Among them, N is the number of text pairs required to calculate a loss, x n , x′ n Represents text pair datasets Two texts of a single training data in n∈[1,N] are integers, sim(·) is the similarity of the text pair calculated according to the model, y n Similarity scores for large language model annotations.

3. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 2 is characterized in that: The pre-trained text semantic feature extraction model Select the BGE-large-zh-v1.5 model.

4. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 2 is characterized in that: The S3 includes: S31. Divide the text blocks of each part from the scientific and technological project documents according to the chapter structure, and construct a multi-dimensional summary generation special dataset through expert annotation; S32. Build summary generation model remember Among them, A is the main summary dimension, P is the instruction prompt word corresponding to A, and A u is the subordinate summary dimension, P u A u The corresponding command prompt word, doc, is the pre-processed scientific and technological project document content; S33. Calculate the supervised fine-tuning loss L of the summary generation model CrossEntropy ; S34, calculating the intra-document consistency constraint of the fusion text semantic feature extraction model; S35. Calculate cross-document consistency constraints of the fusion text semantic feature extraction model; S36. Calculate the total loss function L all , fine-tune the summary generation model.

5. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 4 is characterized in that: The S34 includes: S341, generate feature vectors of each summary dimension in the document, record Among them, V A is the eigenvector of A, A u The eigenvector of Represents the fine-tuned text semantic feature extraction model; S342, generate feature vectors of all original text entries in the document, record Among them, t m For the lowest level single text entry, V m The lowest level single text entry t m The corresponding feature vector, M is the number of single text entries at the bottom level; S343. Constructing feature consistency constraints L within the document intro : Among them, v1 is the total project overview vector, v2 is the average value of the summary vectors of different dimensions, and v3 is the feature vector of all original text entries in the document. Take v1 = V A , Represents the square of the Euclidean distance.

6. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 5 is characterized in that: The S35 includes: S351, generate cross-document summary, record Among them, doc j For the jth document, A j For doc j The main summary dimension of P j A j The corresponding command prompt words, For doc j The subordinate summary dimension of for Corresponding command prompt words; S352, calculating the similarity of the summary text features, wherein the similarity of the summary text features includes the similarity sim1 of the main summary dimension and the similarity sim2 of the subordinate summary dimension; S353, calculating the similarity sim3 of the bottom-up text features; S354. Construct cross-document feature consistency constraint L inter :

7. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 4 is characterized in that: The summary generation model The model structure and initialization parameters adopt Qwen2.5-7B-Instruct. The model input is the scientific and technological project text, and the output is the summarized project text.

8. The method for checking duplicate technology project documents based on multi-level summary generation according to claim 4 is characterized in that: The S4 includes: S41. Use existing scientific and technological project documents to build a vector database of historical scientific and technological projects S42, perform structural analysis on the scientific and technological project documents to be checked for duplicate content, recorded as doc * , pass the parsing results into the fine-tuned summary generation model to obtain the multi-level summary of the model, which includes the multi-dimensional project summary text A * , and project structured text S43, calculate the feature vector of the summary text corresponding to the main summary dimension through the text semantic feature extraction model, record in, Abstract text A * The eigenvector of Abstract text ; Search in the vector database to find the top c items with the highest semantic content similarity, and obtain the summary text of the subordinate summary dimension; wherein, the following logic is used to express the execution in the vector database Search in the query to find the top c items with the highest semantic content similarity in the summary in, Indicates that from the vector database according to The first c items after similarity sorting, Indicates that the vector database in the previous operation is S44. Calculate the cosine similarity of the summary texts corresponding to the subordinate summary dimension, sort the weighted similarities, and obtain a similarity ranking. The following logic is used to express the calculation of the cosine similarity of the summary texts corresponding to the subordinate summary dimension: Among them, w u is the weight of the u-th subordinate summary dimension, doc for the project document to be checked for duplicates * The inner product of the feature vector of the u-th summary dimension of the query entry z.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the method for checking duplicate technology project documents based on multi-level summary generation as described in any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is run by a processor, the steps of the method for checking duplicate technology project documents based on multi-level summary generation according to any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Document duplicate checking method based on hierarchical feature vector search

    CN117951256A

  • Document duplicate checking method for intelligent document review system and storage medium

    CN118036586A

  • Distribution network project duplicate checking method based on natural language processing

    CN118313365A

Cited By

  • Text similarity calculation method and system based on multi-algorithm fusion

    CN121279321A