Scientific and technological document duplicate checking method and device, equipment and medium
By constructing the target scientific and technological knowledge graph and fine-tuning the large model, the problems of low efficiency and accuracy in scientific and technological document duplication checking were solved, achieving more efficient and accurate duplication checking results.
Patent Information
- Application Number
- CN202510417917.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-09-23
AI Technical Summary
The efficiency and accuracy of scientific document duplication checking in existing technologies are low, making it difficult to meet the needs of speed and accuracy.
By constructing the target scientific and technological knowledge graph, fine-tuning the large model based on the sample scientific and technological document set, obtaining a similar document set, and using the fine-tuned target large model to check for duplicates, the model's adaptability is improved.
The efficiency and accuracy of checking for duplicate content in scientific and technological documents have been improved, the checking method has been optimized, and the scope of checking for duplicate content has been narrowed.
Smart Images

Figure CN120688473A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of clean energy technology, and in particular to a method, device, equipment and medium for checking for duplicates in scientific and technological documents. Background Art
[0002] With the development of technology, duplicate checking of scientific and technological documents in the industry is very important for people's work. In related technologies, duplicate checking of scientific and technological documents can be carried out through string comparison, but the efficiency is poor.
[0003] Therefore, it is very important to improve the efficiency and accuracy of duplicate checking of scientific and technological documents. Summary of the Invention
[0004] The purpose of this application is to solve one of the technical problems in the above technology at least to a certain extent.
[0005] The first aspect of the present application provides a method for checking for duplicates in scientific and technological documents, comprising: obtaining a fine-tuning sample set based on an acquired sample scientific and technological document set to fine-tune the candidate large model to obtain a fine-tuned target large model; constructing a corresponding target scientific and technological knowledge graph, and based on the target scientific and technological knowledge graph, obtaining a similar document set of the document object to be checked for duplicates; calling the model capability of the target large model, and obtaining the target duplicate checking result of the document object based on the similar document set through the model capability.
[0006] The second aspect of the present application provides a device for checking for duplicates in scientific and technological documents, comprising: a fine-tuning module for obtaining a fine-tuning sample set based on an acquired sample scientific and technological document set, so as to fine-tune the candidate large model and obtain a fine-tuned target large model; a first acquisition module for constructing a corresponding target scientific and technological knowledge graph, and obtaining a similar document set of the document object to be checked for duplicates based on the target scientific and technological knowledge graph; a second acquisition module for calling the model capability of the target large model, and obtaining the target duplicate checking result of the document object based on the similar document set through the model capability.
[0007] The third aspect embodiment of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the scientific and technological document duplication checking method provided in the first aspect of the present application.
[0008] The fourth aspect of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method for checking for duplicate scientific and technological documents provided in the first aspect of the present application.
[0009] The fifth aspect of the present application provides a computer program product. When the instruction processor in the computer program product is executed, the method for checking for duplicate technology documents provided in the first aspect of the present application is executed.
[0010] The method and device for checking for duplicates in scientific and technological documents proposed in this application obtain a fine-tuned target macro-model based on a sample set of scientific and technological documents, obtain a similar document set of the document object to be checked for duplicates based on a target scientific and technological knowledge graph, and then obtain a target duplicate checking result of the document object based on the similar document set based on the model capability of the fine-tuned target macro-model. In this application, the scope of documents for duplicate checking of the document object is narrowed by the target scientific and technological knowledge graph, and the adaptability of the target macro-model to the industry to which the scientific and technological documents belong is improved by fine-tuning the macro-model. The target duplicate checking result of the scientific and technological documents is obtained based on the target macro-model, thereby improving the efficiency and accuracy of the duplicate checking of scientific and technological documents and optimizing the method for checking for duplicates in scientific and technological documents.
[0011] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0013] Figure 1 A flowchart of a method for checking duplicate technology documents according to an embodiment of the present application is provided;
[0014] Figure 2 This is a flow chart of a method for checking for duplicate technology documents according to another embodiment of the present application;
[0015] Figure 3 This is a flow chart of a method for checking for duplicate technology documents according to another embodiment of the present application;
[0016] Figure 4 This is a flow chart of a method for checking for duplicate technology documents according to another embodiment of the present application;
[0017] Figure 5 This is a schematic structural diagram of a scientific document duplicate checking device according to an embodiment of the present application;
[0018] Figure 6 It is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0020] The following describes the scientific and technological document duplication checking method, device, equipment and medium of the embodiments of the present application with reference to the accompanying drawings.
[0021] Figure 1 This is a flow chart of a method for checking duplicate technology documents according to an embodiment of the present application. Figure 1 As shown, the method includes:
[0022] S101, obtaining a fine-tuning sample set based on the obtained sample scientific and technological document set to fine-tune the candidate large model to obtain a fine-tuned target large model.
[0023] In the daily work process, it may be necessary to check for duplicates in scientific and technological documents. In this scenario, the duplicate checking of scientific and technological documents can be achieved by calling the model capabilities of the big model.
[0024] Optionally, before calling the model capabilities of the large model to check for duplicate scientific and technological documents, the large model can be fine-tuned, wherein the large model that needs to be fine-tuned can be marked as a candidate large model, and the samples used when fine-tuning the candidate large model can be determined as sample scientific and technological documents, and the set consisting of the sample scientific and technological documents can be determined as the fine-tuning sample set used when fine-tuning the large model.
[0025] Optionally, screening can be performed from an open source scientific document database to screen out some scientific documents that meet the set conditions from the open source database, and a sample scientific document set can be constructed based on the part of scientific documents. Furthermore, a sample construction method in related technology can be used to construct samples of each sample scientific document in the screened sample scientific document set, thereby obtaining a fine-tuned sample set based on the sample scientific document set.
[0026] In this scenario, the candidate large model can be fine-tuned based on the fine-tuning samples included in the fine-tuning sample set according to the large model fine-tuning method in the relevant technology, and the large model obtained after fine-tuning is determined as the target large model after fine-tuning.
[0027] S102, constructing a corresponding target scientific and technological knowledge graph, and obtaining a similar document set of the document object to be checked for duplicates based on the target scientific and technological knowledge graph.
[0028] In an embodiment of the present application, before checking for duplicates on document objects through the target large model, historical scientific documents within a set historical time range can be screened based on the similarity of the document objects, and the document objects can be further checked for duplicates from some of the screened historical scientific documents.
[0029] Optionally, the historical scientific documents within a set historical time range can be integrated by constructing a knowledge graph, and the scientific documents that need to be checked for duplication can be checked for duplication based on the historical scientific documents included in the integrated knowledge graph.
[0030] Optionally, historical scientific and technological documents within a set historical time range can be obtained from an open source database, and a knowledge graph can be drawn based on this part of historical scientific and technological documents through the knowledge graph construction method in the relevant technology. The knowledge graph obtained by the drawing can then be used as the knowledge graph used for duplicate checking and determined as the target scientific and technological knowledge graph.
[0031] In an embodiment of the present application, the scientific and technological documents that need to be checked for duplicates can be marked as document objects. In this scenario, the content included in the document object can be checked for duplicates based on the content of the various scientific and technological documents included in the target scientific and technological knowledge graph.
[0032] Among them, based on the similarity acquisition algorithm in the relevant technology, the various text contents in the target scientific and technological knowledge graph and the various text contents included in the document object can be algorithmically processed to obtain the similarity parameters between each text content included in the document object and the text contents in each graph in the target scientific and technological knowledge graph, and then based on the similarity parameters, some scientific and technological documents similar to the document object are screened out from the various scientific and technological documents included in the target scientific and technological knowledge graph as similar documents of the document object, and then a similar document set composed of the similar documents is obtained.
[0033] S103, calling the model capability of the target large model, and obtaining the target duplicate checking result of the document object based on the similar document set through the model capability.
[0034] In an embodiment of the present application, the model capabilities of the target large model obtained after fine-tuning can be called, and the document objects can be checked for duplicates for a second time based on a similar document set through the called model capabilities of the target large model.
[0035] Optionally, by calling the model capabilities of the target large model, the document content of each similar document in the similar document set and the document content included in the document object can be obtained. Furthermore, for any document content included in the document object, based on calling the model capabilities of the target large model, some document content whose similarity with the document content meets preset conditions is filtered out from the document content included in each similar document, and based on the screening results, a duplicate checking result of the document object based on the target scientific and technological knowledge graph is generated, wherein the duplicate checking result can be determined as the target duplicate checking result.
[0036] The method for checking for duplicates in scientific and technological documents proposed in this application obtains a fine-tuned target macro-model based on a sample set of scientific and technological documents, obtains a similar document set of the document object to be checked for duplicates based on a target scientific and technological knowledge graph, and then obtains a target duplicate checking result of the document object based on the similar document set based on the model capability of the fine-tuned target macro-model. In this application, the scope of documents for duplicate checking of the document object is narrowed by the target scientific and technological knowledge graph, and the adaptability of the target macro-model to the industry to which the scientific and technological documents belong is improved by fine-tuning the macro-model. The target duplicate checking result of the scientific and technological documents is obtained based on the target macro-model, thereby improving the efficiency and accuracy of the duplicate checking of scientific and technological documents and optimizing the method for checking for duplicates in scientific and technological documents.
[0037] In the above embodiment, the fine-tuning of the candidate large model and the acquisition of the target duplicate checking results can be combined with Figure 2 Further understanding, Figure 2 This is a flow chart of a method for checking duplicate technology documents according to another embodiment of the present application. Figure 2 As shown, the method includes:
[0038] S201, constructing a target scientific and technological knowledge graph corresponding to the sample scientific and technological document set.
[0039] Optionally, an initial set of scientific and technological documents required for knowledge graph construction is obtained from a source document library, and content is extracted from each initial scientific and technological document to pre-process the extracted content to obtain candidate scientific and technological documents.
[0040] In an embodiment of the present application, an open source scientific and technological document storage database can be used as a source document library. In this scenario, the scientific and technological documents required to construct the knowledge graph can be obtained from the source document library as initial scientific and technological documents, and the collection of the initial scientific and technological documents can be determined as the initial scientific and technological document set.
[0041] Optionally, the scientific and technological industry to which the scientific and technological documents to be checked for duplicates belong can be obtained, and a corresponding scientific and technological document screening strategy can be constructed based on the scientific and technological industry. Furthermore, scientific and technological documents can be retrieved and screened from the source document library based on the strategy, and the initial scientific and technological documents that match the scientific and technological document screening strategy can be obtained, thereby forming the initial set of scientific and technological documents required to construct the scientific and technological knowledge graph.
[0042] Among them, the source document library can be crawled by crawling technology in related technologies to obtain the initial set of scientific and technological documents, or the initial set of scientific and technological documents can be obtained based on other methods that can retrieve and screen scientific and technological documents from the source document library. No specific restrictions are made here.
[0043] In an embodiment of the present application, content extraction can be performed on the initial set of scientific and technological documents based on the elements required to construct a knowledge graph, and the extracted content can be preprocessed based on a preset data preprocessing strategy, and the documents obtained after preprocessing the content extracted from each initial scientific and technological document can be determined as candidate scientific and technological documents.
[0044] As an example, Figure 3 As shown, it is possible to Figure 3 The initial technical documents shown are structurally disassembled and Figure 3 The content extraction tool shown extracts content from the initial scientific and technological document after structural decomposition.
[0045] It should be noted that Figure 3 The content extraction tool shown can be constructed based on a document to text (Docx2Txt) algorithm in the related art, or can be constructed using other algorithms capable of implementing content extraction, which is not specifically limited here.
[0046] As an example, the document-to-text algorithm can be understood in conjunction with the following:
[0047] Install the toolkit (python-docx), confirm that the Python version and python-docx version are consistent, use the python-docx toolkit to read the document in the specified path, obtain the corresponding format (doc) document, and traverse all block-level elements in the doc document. Block-level elements can include paragraphs, tables, and images in the doc document.
[0048] Furthermore, a block-level element recognition tool is developed to identify elements in documents, including pictures, tables, and paragraphs. The python-docx toolkit can use XML elements to represent Word documents. By parsing XML, the current element block type is identified, all element blocks in the document are traversed, all element block types are identified, and content is extracted for each element block type. The extracted content can be saved in a python list to provide data support for downstream tasks.
[0049] like Figure 3 As shown, the extracted content can be pre-processed, wherein the Figure 3The regular expression 1 shown matches the characters in the extracted content and passes Figure 3 The data cleaning module shown cleans the matched symbols, and can also be used to Figure 3 The data cleaning module shown cleans the headers, footers, some redundant punctuation marks, stop words and other related content included in the extracted content. Furthermore, the document obtained after the above-mentioned preprocessing of the content extracted from the initial scientific document is used as a candidate scientific document corresponding to the initial scientific document.
[0050] The regular expression 1 may be constructed based on the symbols [, . , . ? ! ; : ""''''@#$%^&*()_+{}|<>\[\]\\-=`~\'""''""…], or based on other elements that meet preset conditions, which are not specifically limited here.
[0051] Optionally, the document structure of each candidate scientific document is decomposed to obtain a set of decomposed candidate scientific document text blocks, and knowledge extraction is performed on each candidate scientific document text block to obtain a corresponding knowledge graph construction file.
[0052] In an embodiment of the present application, each candidate scientific document can be subjected to a decomposition algorithm based on a preset document structure decomposition algorithm, and then, based on the results of the algorithm processing, each candidate scientific document text block after the decomposition of each candidate scientific document can be obtained, and the set composed of each candidate scientific document text block included in any candidate scientific document can be determined as the candidate scientific document text block set under the candidate scientific document.
[0053] The candidate scientific and technological document text blocks may include text blocks composed of characters, or text blocks composed of information in formats such as tables, which are not specifically limited here.
[0054] Furthermore, based on the knowledge extraction algorithm used in constructing the knowledge graph in the relevant technology, knowledge extraction is performed on the text block level of the candidate scientific document under each candidate scientific document, and then based on the extracted knowledge, the construction file required to construct the knowledge graph is obtained and determined as the knowledge graph construction file.
[0055] As an example, the formats of the candidate scientific and technological documents may need to be adjusted so that the document formats of the candidate scientific and technological documents are consistent, such as Figure 3 As shown, it can be achieved through Figure 3 The format normalization shown is used to adjust the consistency of each candidate scientific document, wherein the Figure 3 The structural disassembly module shown is based on Figure 3 The regular expression 2 and the document structure information shown are used to decompose the candidate scientific and technological documents after consistency adjustment in terms of document structure dimension, thereby obtaining a set of candidate scientific and technological document text blocks included in each candidate scientific and technological document.
[0056] It should be noted that the format of the candidate scientific and technological documents may be
[0057] Among them, for any candidate scientific document, regular expression 2 and document structure information (xml) can be used to disassemble it, analyze the structural characteristics of the document, design regular expressions for document structure markers, and respectively identify the start and end positions of text blocks. In text blocks without special marks, the text is divided based on the style information of xml, and finally a text block with complete semantic content is obtained. The table information is read through the style of xml, and the read table content is saved, and then the candidate scientific document text blocks included in the candidate scientific document are obtained, thereby forming a candidate scientific document text block set under the candidate scientific document.
[0058] like Figure 3 As shown, for each candidate scientific document, the candidate scientific document text block set can be obtained by Figure 3 The content disassembly module shown is based on Figure 3 The model capability of the large model shown is to decompose the content dimension of each candidate scientific document text block, thereby realizing the extraction of knowledge carried in each candidate scientific document text block. Furthermore, based on the file integration method of construction in related technologies, the extracted knowledge is integrated, and the integrated files that can meet the knowledge graph construction requirements are determined as knowledge graph construction files.
[0059] in, Figure 3 The model capability of the large model shown may be a model capability based on a generative algorithm or other types of model capabilities, which are not specifically limited here.
[0060] In this example, based on the natural language processing algorithms and large language model capabilities in related technologies, the content of text blocks composed of words and text blocks composed of tables can be decomposed and analyzed respectively, and knowledge extraction can be performed on the text content after structural decomposition, so as to summarize the content and extract keywords of each text block. It can be understood that based on natural language processing and the model capabilities of large models, semantic understanding can be performed on each candidate scientific and technological document text block, so as to generate content summary information and keyword information, and then obtain a knowledge graph construction file.
[0061] Optionally, a file is constructed based on the knowledge graph to obtain the target scientific and technological knowledge graph.
[0062] Among them, based on the knowledge graph construction method in the relevant technology, the knowledge graph can be drawn and constructed based on the knowledge graph construction file, so as to obtain the constructed target scientific and technological knowledge graph.
[0063] In addition, the knowledge graph construction file can be converted into a file in a format (xlxs, json, dict), and the converted file can be input into the pre-acquired knowledge graph construction tool to obtain the constructed target science and technology knowledge graph.
[0064] S202 : extracting sample question text blocks, sample answer text blocks, and probability distribution of sample categories to which the sample question text blocks belong from a sample scientific document collection.
[0065] Optionally, data cleaning and data preprocessing are performed on each historical scientific and technological document in the source document library to obtain a sample scientific and technological document set.
[0066] In an embodiment of the present application, historical scientific documents that can be used for fine-tuning the large model can be screened from the source document library based on the scientific and technological industry to which the fine-tuned large model is applicable, and based on the historical scientific and technological documents obtained by screening, sample scientific and technological documents belonging to the training samples used when fine-tuning the large model are obtained, and the collection composed of the sample scientific and technological documents is determined as a sample scientific and technological document set.
[0067] Optionally, each historical scientific and technological document may be cleaned and preprocessed based on a data cleaning and preprocessing method in related technologies, and each scientific and technological document obtained after cleaning and preprocessing may be determined as each sample scientific and technological document.
[0068] As an example, redundant characters and page numbers and other redundant symbols in each historical scientific document can be deleted, as well as redundant formatting information such as extra spaces and tabs in each historical scientific document. Redundant text content such as footnotes and text descriptions of charts in each historical scientific document can also be deleted, thereby achieving data cleaning and preprocessing of each historical scientific document, and obtaining sample scientific documents after the historical scientific documents have been cleaned and preprocessed.
[0069] Optionally, semantic decomposition is performed on each sample scientific and technological document to obtain a sample text block of each sample scientific and technological document.
[0070] In an embodiment of the present application, the title information of each sample scientific document can be obtained, and the text content corresponding to the different semantic paragraph levels included in each sample scientific document can be disassembled according to the title information, and the text blocks obtained by the disassembly can be determined as the sample text blocks included in each sample scientific document.
[0071] As an example, for any sample scientific document T clean , according to the sample technology document T clean The title information included in the sample scientific document T clean Divide the text content into different semantic paragraph levels. cleanThe specific situation defines the rules that the title complies with R H , using Python character matching and regular expressions and other related methods, in the sample technology document T clean Search for the rule R H All title instances H1, H2,…, Hn.
[0072] In this example, based on the extracted title example Hi, the sample scientific document T clean Divide into n semantic blocks B1, B2, ..., Bn. Each semantic block Bi starts with the title Hi and contains all the text content between the title Hi and the next title H(i+1) (or the end of the document). In this example, the sample scientific document T clean It is effectively split into multiple text blocks with clear semantic block structures, and the multiple text blocks are sample scientific and technological documents T clean Various blocks of sample text are included.
[0073] Optionally, the logical relationship between each sample text block is obtained, and based on the logical relationship, a sample question text block and a sample answer text block of the sample question text block are obtained from each sample text block.
[0074] In an embodiment of the present application, there is a certain degree of logical relationship between the sample text blocks. In this scenario, the contextual relationship between the sample text blocks can be determined based on the contextual relationship in the logical relationship. Specifically, for each sample text block included in any contextual logical relationship, the sample text block in the front position of the sample text blocks in the contextual logical relationship can be obtained as a sample question text block, and the sample text block in the back position in the contextual logical relationship can be obtained as a sample answer text block corresponding to the above-mentioned sample question text block.
[0075] As an example, for any title set S extracted from a sample scientific document H , S H ={H1, H2, ..., Hn}, where the text contents of Hq and Ha in the title set are assumed to have a causal logical relationship, wherein the text content Bq corresponding to Hq is the cause, and the text content Ba corresponding to Ha is the result caused by the text content corresponding to Hq. It can be understood that the text content Bq corresponding to Hq is the question, and the text content Ba corresponding to Ha is the answer to the question.
[0076] Among them, the text block corresponding to the cause in the causal relationship can be determined as the text block in the front position in the causal relationship, and the text block corresponding to the effect in the causal relationship can be determined as the text block in the back position in the causal relationship. In this example, the text block Bq corresponding to the title Hq in the front position in the title pair {Hq, Ha} can be used as the sample question text block in this example, and the text block Ba corresponding to the title Ha in the back position can be used as the sample answer text block in this example.
[0077] Optionally, a document category set corresponding to the historical science and technology document set is obtained, and each sample document category to which the sample question text block belongs and the category probability of each sample document category are obtained from the document category set to obtain the sample category probability distribution of the sample question text block.
[0078] In an embodiment of the present application, it is necessary to fine-tune the large model used in checking for duplicates according to the scientific and technological industry to which the scientific and technological documents to be checked for duplicates belong. In this scenario, it is necessary to divide the scientific and technological industries to which the sample scientific and technological text blocks included in each sample scientific and technological document belong.
[0079] Among them, the industry category to which each historical scientific and technological document in the historical scientific and technological document collection belongs can be obtained to obtain a document category set corresponding to the historical scientific and technological document collection.
[0080] Optionally, a division strategy for industry classification of each document category in the document category set is obtained, and each sample question text block is divided into industry categories based on the division strategy. For any sample question text block, a category analysis can be performed on the sample question text block based on the division strategy, and then the category to which the sample question text block belongs is determined based on the result of the category analysis, and the category can be marked as the sample document category.
[0081] Among them, in the scenario where the sample text category includes multiple categories, corresponding algorithm processing can also be performed based on the category probability algorithm in the relevant technology to obtain the probability that the sample question text block belongs to each sample text category, and determine the category probability of each sample document category.
[0082] Optionally, an algorithm model can be obtained that can obtain the sample document category to which the sample question text block belongs and the category probability of each sample document category, and each sample question text block and each document category in the document category set are input into the algorithm model. The sample document category to which each sample question text block belongs and the category probability corresponding to each sample document category are obtained through the model capability of the algorithm model.
[0083] Furthermore, based on the sample document categories to which any sample question text block belongs and the category probabilities of the sample document categories, a sample category probability distribution corresponding to the sample question text block is generated.
[0084] S203: Based on the sample question text blocks, the sample answer text blocks and the sample category probability distribution, a sample question and answer group is constructed to obtain a fine-tuning sample set for the candidate large model.
[0085] In an embodiment of the present application, based on the question and answer group construction method in the relevant technology, a question and answer group can be constructed for the sample question text block, the sample answer text block and the corresponding sample category probability distribution, thereby obtaining a question and answer group consisting of the sample question text block, the sample answer text block and the sample category probability distribution as a sample question and answer group.
[0086] Optionally, a first candidate question-answer pair is obtained based on the sample question text block and the sample answer text block.
[0087] In the embodiment of the present application, the first candidate question-answer pair can be understood in conjunction with the following formula:
[0088] (Q,A)1=L((B q ,B a )P1)
[0089] In the above formula, (Q, A)1 represents the first candidate question-answer pair, L represents the construction model for obtaining the first candidate question-answer pair, P1 represents the prompt word input into the model when constructing the first candidate question-answer pair, and B q represents a sample question text block, B a Represents a block of sample answer text.
[0090] Optionally, a second candidate question-answer pair is obtained based on the sample question text block and the probability distribution of the sample list corresponding to the sample question text block.
[0091] In the embodiment of the present application, the second candidate question-answer pair can be understood in conjunction with the following formula:
[0092] (Q,A)2=L(B i ,P c (B i )P2)
[0093] In the above formula, (Q,A)2 represents the sample question and answer group, B i represents a sample question text block, P c (B i ) represents the probability distribution of the sample list corresponding to the sample question text block, P2 represents the prompt word input into the model when constructing the second candidate question-answer pair, and L represents the construction model for obtaining the second candidate question-answer pair.
[0094] Optionally, a sample question and answer group is obtained based on the first candidate question and answer pair and the second candidate question and answer pair.
[0095] In an embodiment of the present application, for any sample question text block, the first candidate question and answer pair and the corresponding second candidate question and answer pair corresponding to the sample question text block can be combined, and the question and answer group obtained by the combination can be used as the sample question and answer group.
[0096] Furthermore, based on the sample construction algorithm in the relevant technology, the input data and label data and other relevant information when fine-tuning the large model are extracted from any sample question and answer group, and the extracted relevant information is processed by the sample construction algorithm. Then, based on the result of the algorithm processing, the samples used when fine-tuning the large model corresponding to the sample question and answer group are obtained as fine-tuning samples, and then the set composed of the fine-tuning samples is determined as the fine-tuning sample set.
[0097] In step S204, the fine-tuning samples in the fine-tuning sample set are input into the candidate large model to obtain the output result of the candidate large model, and based on the output result and the label of the fine-tuning sample, the fine-tuning loss of the candidate large model is obtained to fine-tune the candidate large model to obtain the fine-tuned target large model.
[0098] In an embodiment of the present application, fine-tuning samples used in the fine-tuning round of the current model can be obtained from the fine-tuning sample set and input into the candidate large model to be fine-tuned to obtain the output result of the candidate large model.
[0099] Optionally, based on the loss value algorithm in the relevant technology, the output results of the candidate large model and the label information of the fine-tuning sample can be algorithmically processed, and then the loss value of the candidate large model in the current fine-tuning round can be obtained based on the result of the algorithm processing as the fine-tuning loss.
[0100] In this scenario, the model parameters of the candidate large model can be adjusted based on the fine-tuning loss, and the next fine-tuning sample can be returned to continue fine-tuning the adjusted candidate large model until the end conditions of the large model fine-tuning are met. The fine-tuning of the candidate large model can be ended, and the fine-tuned candidate large model can be determined as the target large model.
[0101] Among them, the fine-tuning conditions of the large model can be set based on the fine-tuning rounds or based on the output results of the candidate large model, and are not specifically limited here.
[0102] As an example, Figure 4 As shown, it can be achieved by Figure 4 The method shown obtains the fine-tuning sample set used when fine-tuning the large model, such as Figure 4 As shown, it can be achieved by Figure 4The manual labeling method shown obtains the data required for sample construction through the data acquisition module, preprocesses the data, and then manually labels the preprocessed data. Furthermore, the data set composed of the manually labeled data is evaluated and verified to identify whether the data set meets the construction conditions of the fine-tuning sample.
[0103] Furthermore, when it is identified that the data set meets the conditions for constructing fine-tuning samples, a fine-tuning sample set obtained by manual annotation can be obtained based on the data set.
[0104] like Figure 4 As shown, you can also Figure 4 The knowledge distillation method shown in the figure obtains a fine-tuning sample set, wherein the data required for sample construction can be obtained through the data acquisition module, and the data is cleaned. After semantic splitting of the cleaned data, the Figure 4 The selection of the question-answer pair generation model shown determines the acquisition model for the first candidate question-answer pair and the second candidate question-answer pair proposed in the above embodiment.
[0105] Further, through Figure 4 The analysis task dataset construction module shown in FIG is based on the obtained question-answer pair generation model to obtain the first candidate question-answer pair, and Figure 4 The classification task dataset construction module shown obtains a second candidate question-answer pair based on the acquired question-answer pair model, and then obtains a question-answer group dataset based on the first candidate question-answer pair and the second candidate question-answer pair.
[0106] like Figure 4 As shown, through Figure 4 The dataset quality evaluation module shown performs quality evaluation on the question-answer group dataset to identify whether the dataset meets the construction conditions of fine-tuning samples. Furthermore, a fine-tuning sample set obtained based on the knowledge distillation method is obtained based on the question-answer group dataset that meets the construction conditions.
[0107] like Figure 4 As shown, you can also Figure 4 The data capture method shown in the figure is used to obtain a fine-tuning sample set, wherein the open source data can be captured by a data capture program, and the captured data can be processed accordingly to obtain the corresponding data set, and then the fine-tuning sample set can be obtained by the data capture program. Figure 4 The dataset quality evaluation module shown identifies whether the dataset meets the construction conditions of fine-tuning samples. Further, when it is identified that the dataset meets the construction conditions, a fine-tuning sample set obtained by data crawling can be obtained based on the dataset.
[0108] Further, based on Figure 4 The fine-tuning sample sets obtained in the three ways shown are used to fine-tune the large model, thereby obtaining the fine-tuned target large model.
[0109] S205 , obtaining a document splitting strategy for the document object, and decomposing the document object based on the document splitting strategy to obtain a first text vector for each text block object in the document object.
[0110] In an embodiment of the present application, the text that needs to be checked for plagiarism can be determined as the text object to be checked for plagiarism. In the scenario of checking for plagiarism on the text object, the text object can be split into multiple text blocks. By checking for plagiarism on the multiple text blocks separately, the plagiarism checking result of the text object can be obtained.
[0111] The strategy for splitting the text object to be checked for duplicates may be determined as the document splitting strategy for the text object, and the multiple text blocks obtained after splitting the text object based on the document splitting strategy may be determined as the multiple text block objects in the document object.
[0112] It should be noted that the document splitting strategy of the text object can be determined based on the decomposition method used to decompose the document structure of each candidate scientific and technological document when constructing the target scientific and technological map, or it can be determined based on other types of document decomposition methods, and no specific limitation is made here.
[0113] Optionally, each text block object may be processed based on a feature vector extraction algorithm in related technologies, and then a feature vector of each text block object may be obtained according to the result of the algorithm processing and determined as the first text vector of each text block object.
[0114] S206, obtaining the second text vector of each piece of scientific and technological knowledge in the target scientific and technological knowledge graph, so as to obtain a candidate similar document set of the document object based on each first text vector and each second text vector.
[0115] In an embodiment of the present application, feature vectors may be extracted for each piece of scientific and technological knowledge in the target scientific and technological knowledge graph, and the extracted vectors of each piece of scientific and technological knowledge may be determined as the second text vector of each piece of scientific and technological knowledge.
[0116] Optionally, for any first text vector, the vector cosine similarity between the first text vector and each second text vector is calculated.
[0117] In an embodiment of the present application, for any first text vector, the similarity between the first text vector and each second text vector can be calculated, wherein, for any second text vector, the first text vector and the second text vector can be respectively processed based on the cosine similarity algorithm in the relevant technology, and then the cosine similarity between the first text vector and the second text vector is obtained according to the result of the algorithm processing as the vector cosine similarity between the two.
[0118] Optionally, for any second text vector, in response to a vector cosine similarity between the second text vector and the first text vector being greater than or equal to a preset cosine similarity threshold, the second text vector is determined to be a similar text vector to the first text vector.
[0119] It can be understood that, for any second text vector, when the vector cosine similarity between the second text vector and the first text vector is greater than or equal to a preset cosine similarity threshold, it can be determined that the second text vector may be the reference text of the first text vector, that is, there is part of the text content corresponding to the first text vector, and the similarity between the text content corresponding to the second text vector meets the preset conditions of the reference text.
[0120] In this scenario, the second text vector may be determined as a similar text vector to the first text vector.
[0121] Optionally, documents are recalled from the source document library based on similar text vectors to obtain candidate similar documents of the document object to construct a candidate similar document set.
[0122] In an embodiment of the present application, similar text content corresponding to similar text vectors can be obtained, and through the document recall method in the relevant technology, similar documents can be recalled from the source document library based on the similar text content, and the recalled documents can be determined as candidate similar documents of the text object.
[0123] Furthermore, based on the set of candidate similar documents, a candidate similar document set is obtained.
[0124] S207 , using the model capability of the target large model, respectively compare the similarity between the document object and each candidate similar document to obtain the document similarity between the document object and each candidate similar document.
[0125] In an embodiment of the present application, the text object and each candidate similar document can be input into the fine-tuned target large model, and the similarity between the document object and each candidate similar document can be obtained through the model capabilities of the target large model, and the similarity between the document object and each candidate similar document can be determined as the document similarity of each candidate similar document.
[0126] S208, based on the similarity of each document, obtain the target duplicate checking result of the document object based on the target scientific and technological knowledge graph.
[0127] Optionally, for any candidate similar document, in response to the document similarity between the document object and the candidate similar document being greater than or equal to a preset document similarity threshold, the candidate similar document is determined to be the target similar document of the document object.
[0128] Among them, when the document similarity between any candidate similarity document and the document object is greater than or equal to the preset document similarity threshold, it can be determined that the candidate similarity document is a reference document that has a certain degree of similarity with the document object. In this scenario, the candidate similarity document can be determined as the target similarity document of the document object.
[0129] Optionally, in response to the existence of at least one target similar document in the candidate similar document set, failure of duplicate checking is determined as a target duplicate checking result.
[0130] It can be understood that when there is at least one target similar document in the candidate similar document set, it can be determined that the text content of the document object is constructed based on the text content in the at least one target similar document as a reference. In this scenario, it can be determined that the document object fails the duplicate check, and the failed duplicate check is determined as the target duplicate check result of the document object.
[0131] Optionally, in response to the target similar document not existing in the candidate similar document set, determining that the duplicate check is qualified is the target duplicate check result.
[0132] It can be understood that when the target similar document does not exist in the candidate similar document set, it can be determined that the degree of similarity between the text content of the document object and the text content of each candidate similar document does not meet the preset duplication checking conditions. In this scenario, it can be determined that the document object has passed the duplication check, and the qualified duplication check will be determined as the target duplication checking result of the document object.
[0133] The method for checking for duplicates in scientific and technological documents proposed in this application obtains a fine-tuned target macro-model based on a sample set of scientific and technological documents, obtains a similar document set of the document object to be checked for duplicates based on a target scientific and technological knowledge graph, and then obtains a target duplicate checking result of the document object based on the similar document set based on the model capability of the fine-tuned target macro-model. In this application, the scope of documents for duplicate checking of the document object is narrowed by the target scientific and technological knowledge graph, and the adaptability of the target macro-model to the industry to which the scientific and technological documents belong is improved by fine-tuning the macro-model. The target duplicate checking result of the scientific and technological documents is obtained based on the target macro-model, thereby improving the efficiency and accuracy of the duplicate checking of scientific and technological documents and optimizing the method for checking for duplicates in scientific and technological documents.
[0134] Corresponding to the method for checking for duplicate scientific documents proposed in the above-mentioned embodiments, an embodiment of the present application also proposes a device for checking for duplicate scientific documents. Since the device for checking for duplicate scientific documents proposed in the embodiment of the present application corresponds to the method for checking for duplicate scientific documents proposed in the above-mentioned embodiments, the implementation method of the above-mentioned method for checking for duplicate scientific documents is also applicable to the device for checking for duplicate scientific documents proposed in the embodiment of the present application, and will not be described in detail in the following embodiments.
[0135] Figure 5 This is a schematic diagram of the structure of a scientific document duplicate checking device according to an embodiment of the present application. Figure 5 As shown, the scientific document duplicate checking device 500 includes a fine-tuning module 51, a first acquisition module 52 and a second acquisition module 53, wherein:
[0136] A fine-tuning module 51 is configured to obtain a fine-tuning sample set based on the obtained sample scientific and technological document set, so as to fine-tune the candidate large model and obtain a fine-tuned target large model;
[0137] The first acquisition module 52 is used to construct a corresponding target scientific and technological knowledge graph, and obtain a similar document set of the document object to be checked for duplicates based on the target scientific and technological knowledge graph;
[0138] The second acquisition module 53 is used to call the model capability of the target large model, and obtain the target duplicate checking result of the document object based on the similar document set through the model capability.
[0139] In an embodiment of the present application, the fine-tuning module 51 is also used to: extract sample question text blocks, sample answer text blocks, and the probability distribution of sample categories to which the sample question text blocks belong from a sample scientific document set; construct a sample question and answer group based on the probability distribution of the sample question text blocks, sample answer text blocks, and sample lists to obtain a fine-tuning sample set for the candidate large model; input the fine-tuning samples in the fine-tuning sample set into the candidate large model to obtain the output result of the candidate large model, and obtain the fine-tuning loss of the candidate large model based on the output result and the label of the fine-tuning sample to fine-tune the candidate large model to obtain the fine-tuned target large model.
[0140] In an embodiment of the present application, the fine-tuning module 51 is also used to: perform data cleaning and data preprocessing on each historical scientific document in the source document library to obtain a sample scientific document set; perform semantic decomposition on each sample scientific document to obtain a sample text block of each sample scientific document; obtain the logical relationship between each sample text block, and obtain a sample question text block and a sample answer text block of the sample question text block from each sample text block based on the logical relationship; obtain a document category set corresponding to the historical scientific document set, and obtain each sample document category to which the sample question text block belongs and the category probability of each sample document category from the document category set to obtain a sample category probability distribution of the sample question text block.
[0141] In an embodiment of the present application, the fine-tuning module 51 is further used to: obtain a first candidate question-answer pair based on a sample question text block and a sample answer text block; obtain a second candidate question-answer pair based on a sample question text block and a probability distribution of a sample list corresponding to the sample question text block; and obtain a sample question-answer group based on the first candidate question-answer pair and the second candidate question-answer pair.
[0142] In an embodiment of the present application, the first acquisition module 52 is also used to: obtain the initial set of scientific and technological documents required for knowledge graph construction from the source document library, and extract content from each initial scientific and technological document to pre-process the extracted content to obtain each candidate scientific and technological document; decompose the document structure of each candidate scientific and technological document to obtain a set of decomposed candidate scientific and technological document text blocks, and extract knowledge from each candidate scientific and technological document text block to obtain a corresponding knowledge graph construction file; and obtain a target scientific and technological knowledge graph based on the knowledge graph construction file.
[0143] In an embodiment of the present application, the second acquisition module 53 is also used to: obtain the document splitting strategy of the document object, and decompose the document object based on the document splitting strategy to obtain the first text vector of each text block object in the document object; obtain the second text vector of each scientific and technological knowledge in the target scientific and technological knowledge graph, and obtain a candidate similar document set of the document object based on each first text vector and each second text vector; through the model capability of the target large model, perform a similarity comparison on the document object and each candidate similar document respectively to obtain the document similarity between the document object and each candidate similar document; based on the similarity of each document, obtain the target duplication checking result of the document object based on the target scientific and technological knowledge graph.
[0144] In an embodiment of the present application, the second acquisition module 53 is further used to: calculate, for any first text vector, the vector cosine similarity between the first text vector and each second text vector; for any second text vector, in response to the vector cosine similarity between the second text vector and the first text vector being greater than or equal to a preset cosine similarity threshold, determine that the second text vector is a similar text vector of the first text vector; and recall documents from the source document library based on the similar text vectors to obtain candidate similar documents of the document object to construct a candidate similar document set.
[0145] In an embodiment of the present application, the second acquisition module 53 is also used to: for any candidate similar document, in response to the document similarity between the document object and the candidate similar document being greater than or equal to a preset document similarity threshold, determine the candidate similarity document as the target similar document of the document object; in response to the presence of at least one target similar document in the candidate similar document set, determine that the duplicate check fails to pass as the target duplicate check result; in response to the absence of the target similar document in the candidate similar document set, determine that the duplicate check passes as the target duplicate check result.
[0146] The scientific and technological document duplication checking device proposed in this application obtains a fine-tuned target macro-model based on a sample scientific and technological document set, obtains a similar document set of the document object to be checked for duplicates based on a target scientific and technological knowledge graph, and then obtains a target duplication checking result of the document object based on the similar document set based on the model capability of the fine-tuned target macro-model. In this application, the scope of documents for duplicate checking of the document object is narrowed by the target scientific and technological knowledge graph, and the adaptability of the target macro-model to the industry to which the scientific and technological document belongs is improved by fine-tuning the macro-model. The target duplication checking result of the scientific and technological document is obtained based on the target macro-model, thereby improving the efficiency and accuracy of the duplicate checking of the scientific and technological document and optimizing the duplicate checking method of the scientific and technological document.
[0147] To achieve the above embodiments, the present application also provides an electronic device, a computer-readable storage medium, and a computer program product.
[0148] Figure 6 This is a block diagram of an electronic device according to an embodiment of the present application, such as Figure 6 As shown, the device 600 includes a memory 61, a processor 62, and a computer program stored in the memory 61 and executable on the processor 62. When the processor 62 executes the program instructions, the execution is realized. Figures 1 to 4 The embodiment of the present invention provides a method for checking for duplicate scientific and technological documents.
[0149] In order to implement the above embodiment, the present application also provides a non-transitory computer readable storage medium storing computer instructions, which is used to enable the computer to execute Figures 1 to 4 The embodiment of the present invention provides a method for checking for duplicate scientific and technological documents.
[0150] In order to implement the above embodiment, the present application also provides a computer program product, when the instruction processor in the computer program product executes Figures 1 to 4 The embodiment of the present invention provides a method for checking for duplicate scientific and technological documents.
[0151] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0152] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0153] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0154] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0155] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0156] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0157] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0158] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for checking duplicate content in scientific and technological documents, characterized in that: The method comprises: Obtaining a fine-tuning sample set based on the obtained sample scientific and technological document set to fine-tune the candidate large model to obtain a fine-tuned target large model; Construct a corresponding target scientific and technological knowledge graph, and based on the target scientific and technological knowledge graph, obtain a similar document set of the document object to be checked for duplicates; The model capability of the target large model is called, and the target duplicate checking result of the document object is obtained based on the similar document set through the model capability.
2. The method according to claim 1, characterized in that The method of obtaining a fine-tuning sample set based on the obtained sample scientific and technological document set to fine-tune the candidate large model to obtain a fine-tuned target large model includes: Extracting sample question text blocks, sample answer text blocks, and probability distribution of sample categories to which the sample question text blocks belong from a sample scientific and technological document collection; Based on the sample question text block, the sample answer text block and the sample list probability distribution, a sample question and answer group is constructed to obtain the fine-tuning sample set of the candidate large model; The fine-tuning samples in the fine-tuning sample set are input into the candidate large model to obtain the output result of the candidate large model, and based on the output result and the label of the fine-tuning sample, the fine-tuning loss of the candidate large model is obtained to fine-tune the candidate large model to obtain the fine-tuned target large model.
3. The method according to claim 2, characterized in that The step of extracting sample question text blocks, sample answer text blocks, and probability distribution of sample categories to which the sample question text blocks belong from the sample scientific and technological document collection includes: Performing data cleaning and data preprocessing on each historical scientific and technological document in the source document library to obtain the sample scientific and technological document set; Semantically decomposing each sample scientific and technological document to obtain a sample text block of each sample scientific and technological document; Acquiring logical relationships between the sample text blocks, and obtaining the sample question text block and the sample answer text block of the sample question text block from the sample text blocks based on the logical relationships; The document category set corresponding to the historical science and technology document set is obtained, and each sample document category to which the sample question text block belongs and the category probability of each sample document category are obtained from the document category set to obtain the sample category probability distribution of the sample question text block.
4. The method according to claim 3, characterized in that The constructing of a sample question-answer group based on the sample question text block, the sample answer text block, and the sample list probability distribution includes: Obtaining a first candidate question-answer pair based on the sample question text block and the sample answer text block; Obtaining a second candidate question-answer pair based on the sample question text block and a probability distribution of a sample list corresponding to the sample question text block; The sample question and answer group is obtained based on the first candidate question and answer pair and the second candidate question and answer pair.
5. The method according to claim 1, wherein The step of constructing a target scientific and technological knowledge graph corresponding to the sample scientific and technological document set includes: Obtain the initial set of scientific and technological documents required for knowledge graph construction from the source document library, and extract the content of each initial scientific and technological document to pre-process the extracted content to obtain candidate scientific and technological documents; Decompose the document structure of each candidate scientific document to obtain a set of decomposed candidate scientific document text blocks, and extract knowledge from each candidate scientific document text block to obtain a corresponding knowledge graph construction file; A file is constructed based on the knowledge graph to obtain the target scientific and technological knowledge graph.
6. The method according to claim 1, wherein The method of obtaining a candidate similar document set of the document object to be checked for duplicates based on the target scientific and technological knowledge graph, and obtaining a target duplicate checking result of the document object based on the candidate similar document set by using the model capability of the target large model, includes: Obtaining a document splitting strategy for the document object, and decomposing the document object based on the document splitting strategy to obtain a first text vector for each text block object in the document object; Obtaining the second text vector of each piece of scientific knowledge in the target scientific knowledge graph, so as to obtain the candidate similar document set of the document object based on each first text vector and each second text vector; By using the model capability of the target large model, the similarity between the document object and each candidate similar document is compared respectively to obtain the document similarity between the document object and each candidate similar document; Based on the similarity of each document, the target duplication checking result of the document object based on the target scientific and technological knowledge graph is obtained.
7. The method according to claim 6, characterized in that The base obtains the second text vector of each scientific knowledge in the target scientific knowledge graph to obtain the candidate similar document set of the document object based on each first text vector and each second text vector, including: For any first text vector, calculating the vector cosine similarity between the first text vector and each second text vector; For any second text vector, in response to a vector cosine similarity between the second text vector and the first text vector being greater than or equal to a preset cosine similarity threshold, determining that the second text vector is a similar text vector to the first text vector; Documents are recalled from the source document library according to the similar text vectors to obtain candidate similar documents of the document object to construct the candidate similar document set.
8. The method according to claim 6, characterized in that The step of obtaining the target duplicate checking result of the document object based on the target scientific and technological knowledge graph based on the similarity of each document includes: For any candidate similar document, in response to the document similarity between the document object and the candidate similar document being greater than or equal to a preset document similarity threshold, determining the candidate similar document as a target similar document for the document object; In response to at least one target similar document existing in the candidate similar document set, determining a duplicate check failure as the target duplicate check result; In response to the target similar document not existing in the candidate similar document set, determining that the duplicate check is qualified is the target duplicate check result.
9. A device for checking duplicates in scientific and technological documents, characterized in that: The device comprises: A fine-tuning module, configured to obtain a fine-tuning sample set based on the obtained sample scientific and technological document set, so as to fine-tune the candidate large model and obtain a fine-tuned target large model; The first acquisition module is used to construct a corresponding target scientific and technological knowledge graph, and obtain a similar document set of the document object to be checked for duplicates based on the target scientific and technological knowledge graph; The second acquisition module is used to call the model capability of the target large model, and obtain the target duplicate checking result of the document object based on the similar document set through the model capability.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.