Document semantic comparison method and system based on large language model agent
Through the document semantic comparison method of the large language model intelligent agent, combined with the semantic vector library and document comparison intelligent agent, the problem of inaccurate semantic relevance identification in existing scientific and technological novelty retrieval is solved, and high-accuracy document comparison and intelligent novelty retrieval are achieved.
Patent Information
- Application Number
- CN202510775582.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technology novelty search methods are unable to effectively identify the deep semantic correlations between documents, resulting in inaccurate detection results.
A document semantic comparison method based on a large language model agent is adopted to retrieve similar documents through the semantic vector library. The semantic similarity, text overlap and document comparison agent are combined to judge the relevance between documents and generate an objective and explainable retrieval report.
It improves the accuracy and intelligence of scientific and technological novelty retrieval, can accurately identify the semantic relevance between documents, provide objective judgment basis, and simplify manual work.
Smart Images

Figure CN120671678A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a document semantic comparison method and system based on a large language model intelligent agent. Background Art
[0002] Scientific and technological novelty search is an important part of scientific and technological research and development and academic research, aiming to determine the novelty and innovation of technology or results.
[0003] Currently, the similarity between documents is mainly determined by the degree of text overlap. Keyword matching and sentence structure analysis can be used to determine whether there is correlation between documents. However, different writers may use very different expressions to describe semantically similar content. Furthermore, document writers may plagiarize ideas from other documents by replacing synonyms, adjusting sentence structures, and other means to evade detection. Therefore, methods based on text duplication comparison cannot detect semantic correlation between documents. Therefore, how to perform semantic correlation comparisons between documents based on semantic vector libraries and large language models to improve the accuracy of scientific and technological novelty searches has become a very important issue.
[0004] Based on this, this specification provides a document semantic comparison method and system based on a large language model agent. Summary of the Invention
[0005] In order to solve the problem that traditional duplicate checking methods are difficult to effectively identify deep-level semantic relevance, the present invention proposes a document semantic comparison method and system based on a large language model intelligent agent. It retrieves comparison documents that are similar to the retrieved document through a semantic vector library, and comprehensively judges whether the retrieved document and the comparison document are related based on three dimensions: semantic similarity, text overlap, and document comparison intelligent agent judgment. This not only improves the accuracy of semantic comparison, but also improves the accuracy of scientific and technological novelty retrieval, and provides an objective and explainable basis for the judgment of scientific and technological novelty retrieval, significantly improving the intelligence level of scientific and technological novelty retrieval.
[0006] This specification provides a document semantic comparison method based on a large language model agent, including:
[0007] S1: Split the search document and construct a search semantic vector: Obtain a search document and perform data preprocessing on the search document to obtain search texts corresponding to the search document; use a semantic vectorization model to determine the search semantic vectors corresponding to the search texts;
[0008] S2: Retrieve similar documents: retrieve the comparison semantic vectors that match each retrieval semantic vector from the pre-built semantic vector library, and determine similar documents based on the comparison text corresponding to the label of each retrieved comparison semantic vector;
[0009] S3: Generate similar text pairs: match each search text of the search document with each comparison text of the similar document one by one to obtain similar text pairs;
[0010] S4: Fine-tuning the large language model in the agent and constructing a document comparison agent: Based on the retrieval scenario information, in the agent, a document comparison knowledge base is selected from a plurality of knowledge bases by calling a tool component in a tool component library, and the general large language model in the agent is fine-tuned based on the training data and the document comparison knowledge base to obtain a document comparison agent. The knowledge base includes a knowledge base constructed from local data and a knowledge base constructed from data acquired from the network.
[0011] S5: Constructing prompt words and judging semantic relevance based on the document comparison agent: Based on the semantic similarity and text overlap of each similar text pair, the texts in each similar text pair are spliced and added to the prompt words, the prompt words are input into the document comparison agent, and the document comparison agent determines knowledge data matching the prompt words from the document comparison knowledge base, and determines the judgment result and judgment basis of whether each similar text pair is relevant based on the knowledge data and the prompt words and the document comparison large language model in the document comparison agent;
[0012] S6: Generate a final search report: Generate a search report corresponding to the search document and the similar document based on the judgment result and judgment basis of each similar text pair, and determine a final search report based on the search report.
[0013] Optionally, the method for constructing the prompt word in S5 is:
[0014] Calculate the semantic similarity of each similar text pair and determine whether it exceeds a first threshold; if it exceeds the first threshold, continue to calculate the text overlap in each similar text pair and determine whether it exceeds a second threshold; if it does not exceed the second threshold, splice the search text and comparison text in each similar text pair and add them to the prompt word.
[0015] Optionally, the method further includes:
[0016] If the semantic similarity does not exceed the first threshold, determining that the judgment result of each similar text pair is irrelevant, and determining that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold;
[0017] If the second threshold is exceeded, the judgment result of each similar text pair is determined to be relevant, and the judgment basis of each similar text pair is determined to be that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
[0018] Optionally, performing data preprocessing on the search document in S1 to obtain each search text corresponding to the search document specifically includes:
[0019] Performing text extraction on the search document to obtain text data corresponding to the search document;
[0020] Splitting the text data according to the chapter hierarchical structure of the search document to obtain data of each chapter;
[0021] Filtering the chapter data according to pre-set filtering rules to obtain target chapter data;
[0022] According to the granularity requirements corresponding to the search documents, each target chapter data is split using natural language processing technology to obtain various search texts.
[0023] Optionally, the pre-constructing of the semantic vector library in S2 specifically includes:
[0024] Obtaining an uploaded comparison document set including each comparison document, and performing data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document;
[0025] Using the semantic vectorization model, determining the comparison semantic vectors corresponding to the comparison texts of each comparison document;
[0026] Generate a label corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector.
[0027] Optionally, the S2 specifically includes:
[0028] Determining, from a pre-built semantic vector library, a comparison semantic vector that matches each search semantic vector to generate a search result corresponding to each search text; wherein the search result includes a label of the retrieved comparison semantic vector;
[0029] According to the comparison texts corresponding to the respective search results, the comparison documents in the comparison document set are sorted and screened to determine similar documents; wherein the semantic vector library at least includes the comparison semantic vectors of the comparison texts corresponding to the respective comparison documents in the comparison document set uploaded in advance.
[0030] Optionally, the retrieving a comparison semantic vector matching each search semantic vector from a pre-built semantic vector library in S2 specifically includes:
[0031] Determining the search information corresponding to each search semantic vector based on a preset relationship field, and screening the comparison semantic vectors in the pre-built semantic vector library based on the search information; wherein the relationship field is at least one of a document title, a document type, and a section type;
[0032] From the screened comparison semantic vectors, a comparison semantic vector that matches each search semantic vector is determined.
[0033] Optionally, in S4, fine-tuning the general language model in the agent based on the training data and the document comparison knowledge base to obtain the document comparison agent specifically includes:
[0034] Retrieving knowledge data matching the training data from the document comparison knowledge base using retrieval enhancement generation technology;
[0035] Inputting the field values corresponding to the instruction field and the input field in the knowledge data, the training data, and the general language model in the agent to obtain an output result;
[0036] Fine-tuning the universal large language model according to the output result and the field value corresponding to the label field in the training data;
[0037] Obtaining validation data and evaluating the fine-tuned general large language model based on the validation data;
[0038] After the evaluation is passed, the fine-tuned universal large language model is used as the document comparison large language model to obtain a document comparison agent.
[0039] Optionally, the tool component library includes tool components for knowledge base query, network query, calling interfaces of natural language processing technology and calling large language models.
[0040] This specification also provides a document semantic comparison system based on a large language model agent, the system comprising a data storage module, a data preprocessing module, a semantic vectorization module, a vector retrieval module, a document comparison module, a report generation module, and an agent construction module, wherein:
[0041] The data storage module is used to store the model, the pre-built semantic vector library, the search documents and the comparison documents in the pre-uploaded comparison document set, wherein the model includes a semantic vectorization model and a document comparison agent;
[0042] The agent construction module is configured to select a document comparison knowledge base from a plurality of knowledge bases by calling a tool component in a tool component library in the agent based on the retrieval scenario information, and fine-tune a general large language model in the agent based on the training data and the document comparison knowledge base to obtain a document comparison agent, wherein the knowledge base includes a knowledge base constructed from local data and a knowledge base constructed from data acquired from the network;
[0043] The data preprocessing module is used to perform data preprocessing on the search document to obtain each search text corresponding to the search document;
[0044] The semantic vectorization module is used to call the semantic vectorization model to determine the search semantic vectors corresponding to the respective search texts;
[0045] The vector retrieval module is configured to retrieve a comparison semantic vector that matches each search semantic vector from the semantic vector library, and determine similar documents based on a comparison text corresponding to a label of each retrieved comparison semantic vector;
[0046] The document comparison module is configured to match each of the search texts with each of the comparison texts of the similar documents one by one to obtain similar text pairs; based on the semantic similarity and text overlap of each similar text pair, the texts in each similar text pair are spliced and added to the prompt word; the prompt word is input into the document comparison agent; the document comparison agent determines knowledge data matching the prompt word from the document comparison knowledge base; and based on the knowledge data and the prompt word, the document comparison large language model in the document comparison agent determines the judgment result and judgment basis of whether each similar text pair is relevant;
[0047] The report generating module is used to generate a search report corresponding to the search document and the similar document according to the judgment result and judgment basis of each similar text pair, and determine a final search report based on the search report.
[0048] Optionally, the document comparison module is specifically used to calculate the semantic similarity of each similar text pair and determine whether it exceeds a first threshold. If it exceeds the first threshold, continue to calculate the text overlap in each similar text pair and determine whether it exceeds a second threshold. If it does not exceed the second threshold, splice the search text and comparison text in each similar text pair and add them to the prompt word.
[0049] Optionally, the document comparison module is also used to, if it does not exceed the first threshold, determine that the judgment result of each similar text pair is irrelevant, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold; if it exceeds the second threshold, determine that the judgment result of each similar text pair is relevant, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
[0050] Optionally, the data preprocessing module is specifically used to perform text extraction on the search document to obtain text data corresponding to the search document; split the text data according to the chapter hierarchy structure of the search document to obtain data of each chapter; filter the data of each chapter according to pre-set filtering rules to obtain data of each target chapter; and split each target chapter data according to the granularity requirements corresponding to the search document through natural language processing technology to obtain each search text.
[0051] Optionally, the data preprocessing module is further configured to obtain an uploaded comparison document set including each comparison document, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document;
[0052] The semantic vectorization module is further configured to determine, using the semantic vectorization model, the comparison semantic vectors corresponding to the comparison texts of each comparison document;
[0053] The data storage module is specifically used to generate a label corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector.
[0054] Optionally, the vector retrieval module is specifically used to determine, from a pre-constructed semantic vector library, a comparison semantic vector that matches each search semantic vector, to generate a search result corresponding to each search text; wherein the search result includes a label of the retrieved comparison semantic vector; according to the comparison text corresponding to each search result, each comparison document in the comparison document collection is sorted and screened to determine similar documents; wherein the semantic vector library includes at least the comparison semantic vector of each comparison text corresponding to each comparison document in the comparison document collection uploaded in advance.
[0055] Optionally, the vector retrieval module is specifically used to determine the retrieval information corresponding to each retrieval semantic vector based on a preset relationship field, and to filter the comparison semantic vectors in a pre-built semantic vector library based on the retrieval information; wherein the relationship field is at least one of the document title, document type and chapter type; and from the filtered comparison semantic vectors, determine the comparison semantic vector that matches each retrieval semantic vector.
[0056] Optionally, the document comparison agent construction module is specifically used to use retrieval enhancement generation technology to retrieve knowledge data matching the training data from the document comparison knowledge base; input the field values corresponding to the knowledge data, the instruction field and the input field in the training data into the general large language model in the agent to obtain an output result; fine-tune the general large language model according to the output result and the field value corresponding to the label field in the training data; obtain verification data, and evaluate the fine-tuned general large language model according to the verification data; after the evaluation passes, use the fine-tuned general large language model as the document comparison large language model to obtain a document comparison agent.
[0057] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0058] The document semantic comparison method based on a large language model agent provided in this specification first obtains the search document and performs data preprocessing on the search document to obtain the search texts corresponding to the search document. A semantic vectorization model is used to determine the search semantic vectors corresponding to each search text. The semantic vectorization model maps the split search texts into high-dimensional semantic vectors, which can effectively capture the deep semantic information of the text and provide a data basis for subsequent vector retrieval and semantic-level correlation detection.
[0059] Next, a matching semantic vector matching each search semantic vector is retrieved from a pre-built semantic vector library, and similar documents are identified based on the matching text corresponding to the label of each retrieved matching semantic vector. The semantic vector library contains at least the matching semantic vector corresponding to each matching text in each matching document in the pre-uploaded matching document collection. By retrieving matching semantic vectors from the semantic vector library that match the search semantic vector of the search text, similar documents similar to the search document are accurately screened, narrowing the scope of document comparison and increasing the speed of document comparison.
[0060] Afterwards, each search text is matched one-to-one with each comparison text of a similar document to obtain each similar text pair. Based on the semantic similarity and text overlap of each similar text pair, the text in each similar text pair is spliced and supplemented into the prompt word, and the prompt word is input into the document comparison agent. The document comparison agent determines the knowledge data that matches the prompt word from the document comparison knowledge base, and based on the knowledge data and the prompt word, the document comparison large language model in the document comparison agent determines the judgment result and judgment basis of whether each similar text pair is relevant. Based on the judgment result and judgment basis of each similar text pair between the search document and the similar document, a search report corresponding to the search document and the similar document is generated, and the final search report is determined based on the search report. Comprehensively judging whether the search document and the similar document are relevant based on the three dimensions of semantic similarity, text overlap and document comparison agent judgment not only improves the accuracy of semantic comparison, but also provides an objective and explainable basis for the judgment of scientific and technological novelty retrieval, significantly improving the intelligent level of scientific and technological novelty retrieval. In addition, the document comparison agent automatically determines knowledge data from the document comparison knowledge base, and introduces knowledge data including professional knowledge when generating judgment results and judgment basis. Based on the semantic analysis, understanding ability and massive knowledge of the document comparison large language model in the document comparison agent, the document comparison agent is driven by prompt words to make judgments on relevance and generate judgment basis, and the knowledge data is input into the document comparison large language model in the document comparison agent as expert knowledge, so that the document comparison large language model can make judgments and generate judgment basis based on the knowledge data and its own massive knowledge, thereby improving the accuracy of the output results and simplifying subsequent manual work.
[0061] When constructing a document comparison agent, the present invention first selects a document comparison knowledge base from several knowledge bases in the agent by calling a tool component in a tool component library based on retrieval scenario information, and fine-tunes the general large language model in the agent based on the training data and the document comparison knowledge base to obtain a document comparison agent. Compared with traditional training logic, methods such as fine-tuning and training the general large language model make the large language model obtained after fine-tuning more suitable for semantic comparison application scenarios. By selecting a knowledge base suitable for the retrieval scenario information, and based on a document comparison knowledge base including massive professional knowledge and training data including field values corresponding to instruction fields, input fields, and label fields, the general large language model is fine-tuned, so that the general large language model can learn professional knowledge and apply the professional knowledge to document comparison, thereby making the judgment results and judgment basis output by the document comparison agent based on the fine-tuned general large language model more accurate. In addition, in addition to the large language model, the above-mentioned document comparison agent also includes a tool component library, which not only expands the capabilities of the large language model, but also makes the document comparison agent more adaptable to complex document comparison processing scenarios.
[0062] The present invention performs data preprocessing on the search document and each comparison document to perform text extraction, chapter extraction, chapter filtering, and splitting according to granularity requirements, thereby filtering out useless data, eliminating irrelevant content, and improving the processing efficiency of subsequent document comparison. In addition, the search text is split into individual search texts to retrieve comparison semantic vectors that match the search semantic vectors of each search text, and the granularity of the search text and the comparison text is consistent, avoiding the problem of sliding window denoising due to inconsistent granularity. During vector retrieval, the vector retrieval method of the semantic vector library can be relied upon, combined with the relationship fields, namely the document title, document type, and chapter type, to search each search text to obtain the corresponding search results.
[0063] In the present invention, when comparing documents, the semantic similarity between similar text pairs is first compared, and when the semantic similarity exceeds a first threshold, the text overlap between similar text pairs is compared, and when the text overlap is not greater than a second threshold, the knowledge data and prompt words are input into the document comparison large language model in the document comparison agent, and the judgment result and judgment basis of the similar text pairs are determined by the document comparison large language model, so that the document comparison agent can be used to judge whether similar text pairs with high semantic similarity and low text overlap are related, so as to solve the problem that it is difficult to intuitively judge whether text pairs that are semantically similar but lack significant overlap in the literal sense are related. Compared with the traditional method of judging relevance based only on high semantic similarity, the present invention can continue to judge based on the document comparison agent when the semantic similarity is high, so as to judge the situation where the semantic similarity is high but the text is not related, thereby avoiding the situation where misjudgment occurs based only on semantic similarity judgment.
[0064] The present invention determines the relevance of a retrieved document and similar documents based on a decision tree, sequentially taking into account semantic similarity, text overlap, and document comparison agent judgment. This avoids the problem of misjudgment caused by using a single dimensional judgment, and generates a corresponding judgment basis for the judgment result. This not only improves the accuracy of document comparison but also increases interpretability. Furthermore, a search report can be generated based on the relevance judgment results and judgment basis for each similar text pair. Based on the search report, a final search report is determined and displayed for review by novelty searchers and writers.
[0065] The present invention first retrieves knowledge data that matches the training data from the document comparison knowledge base through retrieval enhancement generation technology, and inputs the field values corresponding to the instruction field and input field in the knowledge data and training data into the general large language model in the intelligent agent to obtain the output result. According to the output result and the field value corresponding to the label field in the training data, the general large language model is fine-tuned. The fine-tuned general large language model is evaluated based on the verification data. The fine-tuned general large language model is evaluated by using independent verification data to ensure that the fine-tuned general large language model can also perform well on unseen data. The quality of the fine-tuned general large language model is judged by evaluation, and after passing the evaluation, the fine-tuned general large language model is used as the document comparison large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0067] Figure 1 A flowchart of a document semantic comparison method based on a large language model agent provided in this specification;
[0068] Figure 2 A flowchart of a semantic vector library construction process provided in this specification;
[0069] Figure 3 A schematic diagram of a document semantic retrieval and comparison process provided in this specification;
[0070] Figure 4 A schematic diagram of a document comparison process assisted by a document comparison agent provided in this specification;
[0071] Figure 5 A schematic diagram of a document comparison agent-assisted related determination process provided in this specification;
[0072] Figure 6 A schematic diagram of a document semantic comparison system based on a large language model agent provided in this specification;
[0073] Figure 7 This is a schematic diagram of a document semantic comparison system architecture based on a large language model agent provided in this specification. DETAILED DESCRIPTION
[0074] This specification provides a document semantic comparison method and system based on a large language model intelligent agent. The technical solutions provided by each embodiment of this specification are described in detail below in conjunction with the accompanying drawings.
[0075] Figure 1 This is a flowchart of a document semantic comparison method based on a large language model agent provided in this specification, such as Figure 1 As shown, the specific steps include:
[0076] S1: Split the search document and construct a search semantic vector: obtain the search document, and perform data preprocessing on the search document to obtain the search texts corresponding to the search document; use the semantic vectorization model to determine the search semantic vectors corresponding to the search texts.
[0077] In this specification, the device used for document comparison can split the retrieval document and construct a retrieval semantic vector, that is, obtain the retrieval document, and perform data preprocessing on the retrieval document to obtain each retrieval text corresponding to the retrieval document. A semantic vectorization model is used to determine the retrieval semantic vector corresponding to each retrieval text. Among them, the device used for document comparison can be a server, a system or a module in the system, or an electronic device such as a desktop computer or a laptop computer. For the sake of ease of description, the document semantic comparison method based on a large language model intelligent agent provided in this specification is described below with the server as the execution subject.
[0078] The above-mentioned search document can be a document uploaded by the user, and the search document is a document that needs to be tested for relevance. The search document can be in any format, such as PDF, Word and other formats. The search document can also be a thesis, project report, etc. The above-mentioned preprocessing includes text extraction, chapter extraction, chapter filtering, and splitting according to granularity requirements. Specifically, when the above-mentioned data preprocessing is performed on the search document to obtain the search texts corresponding to the search document, the server can perform text extraction on the search document to obtain the text data corresponding to the search document. According to the chapter hierarchical structure of the search document, the text data is split to obtain the data of each chapter. According to the pre-set filtering rules, the data of each chapter is filtered to obtain the data of each target chapter. According to the granularity requirements corresponding to the search document, each target chapter data is split through natural language processing technology to obtain each search text.
[0079] Wherein, since the above-mentioned search document can be in any format, the above-mentioned text extraction of the search document is to extract the plain text content to remove format interference. The text data obtained by extraction includes the chapter hierarchy structure, header page number and all text content of the search document. The method of performing text extraction on the search document can be any existing algorithm or model, and this specification does not make specific restrictions. The above-mentioned text data is split according to the chapter hierarchy structure of the search document to split the text data into a tree structure of chapters and sub-chapter to obtain each chapter data. At this time, the each chapter data no longer includes data such as the above-mentioned header page number, but only includes data corresponding to each chapter in the search document to avoid interference from useless data such as header page number. The above-mentioned filtering rules are pre-set, and the chapter data that meets the filtering rules are determined by the filtering rules and filtered out to obtain the target chapter data. Specifically, the filtering rules can be to identify the chapter data in the search document through pre-set keywords, titles or format features to filter the chapter data that meet the keywords, titles or format features set above to obtain the target chapter data. For example, the above keywords can be "papers" or "documents," the title can be "references," and the format features can be "bold" and "font size 2." Using these keywords, titles, or format features, you can filter out paragraphs that are not part of the search, such as journal information and organizational structure, to select chapters related to the search target, i.e., target chapter data, and eliminate irrelevant content, thereby improving subsequent processing efficiency. It should be noted that the above filtering rules are only an example, and this specification does not limit the specific content of the filtering rules. These rules can be set according to your needs.
[0080] The above-mentioned granularity requirements can be pre-set, and the granularity requirements can be paragraphs or sentences. The server can identify the paragraph boundaries and sentence segmentation points in each target chapter data through natural language processing technology based on the granularity requirements corresponding to the retrieval document, so as to split each target chapter data and obtain each retrieval text, ensuring that the paragraphs or sentences of each split retrieval text are independent and complete, so as to facilitate subsequent semantic analysis.
[0081] When the semantic vectorization model is used to determine the retrieval semantic vectors corresponding to each retrieval text, the server can input the retrieval text into the semantic vectorization model for each retrieval text, so as to output the retrieval semantic vector corresponding to the retrieval text through the semantic vectorization model. Among them, the semantic vectorization model is a dense vectorization model, and the above-mentioned retrieval semantic vector is a high-dimensional dense vector. The semantic vectorization model captures the text semantic information by mapping the text into a high-dimensional vector. The semantic vectorization model can effectively characterize the semantic features of the text and provide a data basis for semantic similarity detection and document comparison. The semantic vectorization model can be a general model for mapping text into a high-dimensional vector, such as BGE, BERT, GPT, etc., and of course it can also be a model after fine-tuning the above-mentioned general model. This specification does not make specific limitations.
[0082] S2: Retrieve similar texts: retrieve comparison semantic vectors that match each retrieval semantic vector from a pre-built semantic vector library, and determine similar documents based on the comparison text corresponding to the label of each retrieved comparison semantic vector.
[0083] In this specification, the server can retrieve similar documents, that is, retrieve the comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library, and determine similar documents based on the comparison text corresponding to the label of each comparison semantic vector retrieved. Among them, the semantic vector library at least includes the comparison semantic vectors of each comparison text corresponding to each comparison document in the comparison document set uploaded in advance. The above-mentioned comparison document set includes each comparison document, each comparison document corresponds to each comparison text, and each comparison text has a corresponding comparison semantic vector. Each comparison document can be a document uploaded by the user, and the comparison document is a document used to detect whether it is related to the search document. The comparison document can be in any format, such as PDF, Word and other formats. The comparison document can be a thesis, project report, etc. The above-mentioned comparison semantic vector is a high-dimensional vector obtained by mapping the comparison text using a semantic vectorization model. The above-mentioned search results may include the label of the retrieved comparison semantic vector. Specifically, the server can determine the comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library to generate a search result corresponding to each search text, and determine each similar document from each comparison document based on the comparison text corresponding to each search result.
[0084] When the semantic vector library is pre-built as described above, the server can obtain the uploaded comparison document set including each comparison document, and perform data pre-processing on each comparison document to obtain each comparison text corresponding to each comparison document. A semantic vectorization model is used to determine the comparison semantic vectors corresponding to each comparison text of each comparison document. A label corresponding to each comparison semantic vector is generated, and a semantic vector library is constructed based on each comparison semantic vector and the label corresponding to each comparison semantic vector. The above-mentioned process of performing data pre-processing on each comparison document to obtain each comparison text corresponding to each comparison document is similar to the above-mentioned process of performing data pre-processing on the retrieval document to obtain each retrieval text corresponding to the retrieval document, and will not be repeated here. When the semantic vectorization model is used as described above to determine the comparison semantic vectors corresponding to each comparison text of each comparison document, the server can input the comparison text into the semantic vectorization model for each comparison document and for each comparison text of the comparison document, so as to output the comparison semantic vector corresponding to the comparison text through the semantic vectorization model. The comparison semantic vector is a high-dimensional dense vector.
[0085] In addition to the comparison semantic vectors, the above-mentioned semantic vector library also includes a label corresponding to each comparison semantic vector, and the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector. The comparison document information may be the name or identifier of the comparison document to which the comparison semantic vector belongs, and the chapter information may be the information of the chapter to which the comparison text corresponding to the comparison semantic vector belongs, and the chapter information may be the chapter name, identifier, keyword, etc. The comparison text information may include the comparison text corresponding to the comparison semantic vector and the page number of the page where the comparison text is located. By setting a label for each comparison semantic vector, the comparison semantic vector can be quickly and accurately matched during subsequent vector retrieval. For example, there are comparison documents 1 to 3, where comparison document 1 includes comparison texts A to D. The label of the comparison semantic vector a corresponding to comparison text A may include comparison document information, chapter information and comparison text information. Assume that the comparison document information is the name of comparison document 1, that is, comparison document 1, the chapter information is the reference, and the comparison text information is the comparison text A and page 3. Therefore, the label of the comparison semantic vector a can be expressed in the form of "comparison document information-chapter information-comparison text information", that is, "comparison document 1-reference-comparison text A and page 3".
[0086] In addition to including the label of the retrieved comparison semantic vector, the above-mentioned search results may also include the similarity between the retrieval semantic vector of the retrieved text and the comparison semantic vector matching the retrieval semantic vector of the retrieved document (i.e., the retrieved comparison semantic vector), and may also include the comparison semantic vector matching the retrieval semantic vector of the retrieved document. This specification does not make specific restrictions.
[0087] When determining the comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library to generate the search results corresponding to each search text, the server can calculate the similarity between the search semantic vector and each comparison semantic vector in the pre-built semantic vector library for each search semantic vector, and use the comparison semantic vector corresponding to the maximum similarity as the comparison semantic vector that matches the search semantic vector. The server can use the label corresponding to the comparison semantic vector corresponding to the maximum similarity (i.e., the retrieved comparison semantic vector) and the maximum similarity as the search result of the search text corresponding to the search semantic vector, or can use the maximum similarity, the label corresponding to the retrieved comparison semantic vector, and the retrieved comparison semantic vector as the search result of the search text corresponding to the search semantic vector. This specification does not make specific restrictions. When calculating the similarity between the search semantic vector and each comparison semantic vector in the pre-built semantic vector library, the cosine similarity or Euclidean distance between the search semantic vector and each comparison semantic vector in the pre-built semantic vector library can be calculated and used as the similarity.
[0088] In addition, when determining the comparison semantic vector that matches each retrieval semantic vector from the pre-built semantic vector library, the server can use a vector space model or indexing technology to determine the comparison semantic vector that matches each retrieval semantic vector from the pre-built semantic vector library, so as to quickly match the comparison semantic vector that matches each retrieval semantic vector.
[0089] Furthermore, when searching for similar documents, the server can also determine, from a pre-built semantic vector library, a matching semantic vector that matches each search semantic vector, to generate search results corresponding to each search text. Based on the matching text corresponding to each search result, the matching documents in the matching document collection are sorted and screened to determine similar documents.
[0090] When determining each similar document from each comparison document based on the comparison text corresponding to each search result, or when sorting and screening each comparison document in the comparison document set based on the comparison text corresponding to each search result to determine similar documents, the server can determine the similarity score corresponding to each comparison document based on each similarity included in each search result and the comparison text corresponding to each label. Based on each similarity score, each comparison document is screened to determine each similar document. The search result can be the similarity between the search semantic vector of the search text and the comparison semantic vector that matches the search semantic vector of the search document, as well as the label of the comparison semantic vector that matches the search semantic vector of the search document. The similarity score is the sum of the similarities corresponding to each comparison text in the comparison document.
[0091] Specifically, the server can determine the first score corresponding to each comparison text in the comparison document for each comparison document based on the similarities included in each search result and the comparison text corresponding to each label, and determine the sum of the first scores, and use it as the similarity score of the comparison document. According to the similarity scores, the comparison documents are sorted in descending order to obtain a document sequence. According to the order of the document sequence, a specified number of comparison documents are determined and used as similar documents. Among them, the above-mentioned first score is determined based on the similarities included in each search result and the comparison text corresponding to each label. When the comparison text matches the search text, the search result corresponding to the search text has the similarity between the comparison texts, and the search result also includes the comparison text information corresponding to the comparison text, that is, the label, so the similarity is directly used as the first score of the comparison text. However, if the comparison text does not match any of the search texts, then there is no similarity between the comparison text and any of the search texts, and the search results do not include the comparison text information corresponding to the comparison text. Therefore, the first score of the comparison text can be directly zero. The above-mentioned specified number is a preset value, such as 10. The above-mentioned specified number is the number of similar documents determined.
[0092] S3: Generate similar text pairs: match each search text of the search document with each comparison text of the similar document one by one to obtain similar text pairs.
[0093] In this specification, the server may generate similar text pairs by assigning a one-to-one correspondence between each search text of a search document and each comparison text of similar documents to obtain each similar text pair. Specifically, the server may assign a one-to-one correspondence between each search text of a search document and each comparison text of similar documents based on each search semantic vector of the search document and each comparison semantic vector of each similar document to obtain each similar text pair. Each similar text pair includes a search text and a comparison text, wherein the search text is derived from the search document and the comparison text is derived from the similar document.
[0094] In addition, the server can also determine the comparison text that matches each search text from the comparison texts of similar documents based on the search semantic vector of each search text and the comparison semantic vector of each comparison text of similar documents, and form each similar text pair with each search text and each determined comparison text.
[0095] Furthermore, the server may determine a first vector set consisting of the search semantic vectors of the search documents, and a second vector set consisting of the comparison semantic vectors of the similar documents. The Cartesian product between the first vector set and each second vector set is determined. The similarity between the vectors in each vector combination in the Cartesian product is determined. Based on the similarity between the vectors in each vector combination in the Cartesian product, comparison texts similar to each search text are determined from the comparison texts of the similar documents, and each search text and each determined comparison text are treated as similar text pairs. The Cartesian product includes all possible vector combinations between each second vector set and the first vector set. For example, if the first vector set includes search semantic vectors 1 to 3 and the second vector set includes comparison semantic vectors 4 to 5, then all possible vector combinations between the first vector set and the second vector set, i.e., the Cartesian product is search semantic vector 1 and comparison semantic vector 4, search semantic vector 1 and comparison semantic vector 5, search semantic vector 2 and comparison semantic vector 4, search semantic vector 2 and comparison semantic vector 5, search semantic vector 3 and comparison semantic vector 4, and search semantic vector 3 and comparison semantic vector 5. The similarity can be cosine similarity or Euclidean distance. The comparison text in each similar text pair is the comparison text in the similar documents that has the greatest similarity to the search text in each similar text pair, that is, the search text and the comparison text in each similar text pair in the similar documents have the greatest similarity, and the comparison text is also in the similar text.
[0096] In addition, when there are multiple similar documents, each similar document has a corresponding second vector set, and each second vector set has a corresponding Cartesian product with the first vector set, and the Cartesian product includes all possible vector combinations between each second vector set and the first vector set. Specifically, the server can determine a first vector set composed of the search semantic vectors of the search document, and for each similar document, determine a second vector set composed of the comparison semantic vectors of the similar document. Determine the Cartesian product between the first vector set and the second vector set of the similar document. Determine the similarity between the vectors in each vector combination in the Cartesian product corresponding to the similar document. Based on the similarity between the vectors in each vector combination in the Cartesian product corresponding to the similar document, determine the comparison text similar to each search text from each comparison text of the similar document, and use each search text and each determined comparison text as a similar text pair corresponding to the similar document.
[0097] In addition, in addition to constructing a first vector set based on each retrieval semantic vector, constructing a second vector set based on each comparison semantic vector, and then determining similar text pairs based on the first vector set and the second vector set, the server can also construct a first vector library based on each retrieval semantic vector, construct a second vector library based on each comparison semantic vector, and then determine similar text pairs based on the first vector library and the second vector library to improve efficiency.
[0098] S4: Fine-tune the large language model in the intelligent agent and construct a document comparison intelligent agent: Based on the retrieval scenario information, in the intelligent agent, a document comparison knowledge base is selected from several knowledge bases by calling the tool component in the tool component library, and based on the training data and the document comparison knowledge base, the general large language model in the intelligent agent is fine-tuned to obtain a document comparison intelligent agent, wherein the knowledge base includes a knowledge base constructed by local data and a knowledge base constructed by data obtained from the network.
[0099] In this specification, a server can fine-tune a large language model within an agent and construct a document comparison agent. Specifically, based on retrieval scenario information, the agent selects a document comparison knowledge base from a number of knowledge bases by invoking tool components from a tool component library. Based on training data and the document comparison knowledge base, the general large language model within the agent is fine-tuned to obtain the document comparison agent. The retrieval scenario information can be pre-set and may include the name, context, and description of the retrieval scenario. For example, the retrieval scenario information may be document relevance comparison in a science and technology novelty search scenario. The agent can be pre-built or acquired and include at least a tool component library, a general large language model, and various knowledge bases. Of course, other functions or databases may also be included, which are not specifically limited in this specification. The various knowledge bases can correspond to different scenarios. The server can select a knowledge base that meets the retrieval scenario information from the various knowledge bases by invoking tool components from the tool component library and use it as the document comparison knowledge base. The tool components invoked by the server can be used for knowledge base retrieval, i.e., retrieving knowledge bases that meet the retrieval scenario information. The training data includes the field values corresponding to the instruction field, input field, and label field. The training data can be in JSON format, as follows:
[0100] "instruction (i.e. instruction field)": instruction information (i.e. field value);
[0101] "input (i.e. input field)": training text pairs (i.e. field values).
[0102] "output (i.e. label field)": label (i.e. field value).
[0103] The above-mentioned instruction information may be pre-set information used to guide the large language model to complete the document matching task. The instruction information may include fixed text in the following prompt word template. The instruction information is information that must be input. The above-mentioned training text pair is used for input content related to the document matching task. The training text pair may include a first text and a second text, and the first text and the second text respectively represent the comparison text and the retrieval text. The training text pair may be a text pair that has been semantically compared in history. The training sample pair may be optional input information. The above-mentioned label is the correct result that the large language model should generate under a given input. The label is the label of the training text pair and is manually annotated in advance. The label includes the judgment result and the basis for judgment. The label may be information that must be input.
[0104] The general large language model is a large language model (LLM), which is a neural network model based on deep learning and has powerful semantic understanding and analysis capabilities. The general large language model can be any existing general large language model or a large language model pre-set according to actual needs, and this specification does not make specific restrictions. The general large language model can be an open source or non-open source large language model.
[0105] In this specification, in order to improve the judgment accuracy of the general large language model, a knowledge base of professional knowledge for use by the general large language model can be constructed. Since different industries and different fields have different specific requirements for document comparison, for example, there is a big difference between plagiarism annotation in paper duplication checking and novelty judgment in scientific and technological novelty checking, so when building a knowledge base, it is necessary to build a knowledge base based on the specific planning and professional knowledge of the specific field and industry, and then select a document comparison knowledge base from each knowledge base based on the retrieval scenario information. The data used to build the knowledge base includes but is not limited to relevant policies and regulations, domain knowledge, etc. The server can first obtain knowledge data and build a knowledge base based on the knowledge data. Specifically, a knowledge base composed of knowledge data can be directly constructed, or the knowledge data can be vectorized to obtain semantic vectors corresponding to the knowledge data, and then a knowledge base composed of semantic vectors can be constructed. This specification does not make specific restrictions. The above-mentioned knowledge data can be manually uploaded and stored locally in advance, that is, local data, or it can be obtained from the network (such as the Internet). That is, the above-mentioned knowledge base includes a knowledge base constructed by local data and a knowledge base constructed by data obtained from the network. This specification does not make specific restrictions. The knowledge data in the knowledge base may include relational data, semi-structured data, document data, online data and other data forms. The server can process the knowledge data according to the data form of the knowledge data to obtain the knowledge data in text form. The knowledge data in text form is then preprocessed to obtain each knowledge text. The semantic vectorization model is then used to determine the knowledge semantic vector corresponding to each knowledge text, and a knowledge base composed of knowledge semantic vectors is constructed. The above process of preprocessing the knowledge data in text form to obtain each knowledge text is similar to the above process of preprocessing the search documents to obtain each search text, and will not be repeated here.
[0106] When the general large language model in the agent is fine-tuned based on the training data and the document comparison knowledge base to obtain the document comparison agent, the server can use the retrieval enhancement generation technology to retrieve the knowledge data that matches the training data from the document comparison knowledge base. The field values corresponding to the instruction field and the input field in the knowledge data, training data, respectively, are input into the general large language model in the agent to obtain an output result. The general large language model is fine-tuned based on the output result and the field value corresponding to the label field in the training data. Among them, the above-mentioned retrieval enhancement generation (RGA) technology can be an existing technology, and the retrieval enhancement generation technology is used to retrieve the knowledge data that matches the training data from the document comparison knowledge base. Specifically, the server can determine the training semantic vector corresponding to the training data through the semantic vectorization model, and then determine the similarity between the training semantic vector and each knowledge semantic vector in the document comparison knowledge base, and screen each knowledge semantic vector based on each similarity to obtain the target knowledge semantic vector, and use the knowledge data corresponding to the target knowledge semantic vector as the knowledge data that matches the training data. When obtaining the target knowledge semantic vector, the knowledge semantic vector with a similarity greater than a preset threshold can be used as the target knowledge semantic vector. The knowledge semantic vectors can also be sorted from large to small according to the similarity to obtain a knowledge sequence, and a preset number of knowledge semantic vectors can be determined according to the order of the knowledge sequence and used as the target knowledge semantic vector.
[0107] When fine-tuning the general large language model in the agent based on the output results and the field values corresponding to the label fields in the training data, the server may use a pre-set loss function to determine the loss based on the output results and the field values corresponding to the label fields in the training data, and fine-tune the general large language model in the agent based on the loss. The fine-tuning of the general large language model may be terminated when a preset number of training cycles is reached or when the general large language model's performance on the training data meets expectations.
[0108] In addition, after fine-tuning training, the server can obtain verification data and evaluate the fine-tuned general large language model based on the verification data. After the evaluation is passed, the fine-tuned general large language model is used as a document comparison large language model to obtain a document comparison agent. Among them, the above-mentioned verification data is data independent of the training data, and the format of the verification data is consistent with the format of the above-mentioned training data. When evaluating the fine-tuned general large language model, the fine-tuned general large language model can be evaluated through evaluation indicators, and when the evaluation indicators reach a preset value, it is determined that the fine-tuned general large language model has passed the evaluation. The above-mentioned evaluation indicators may include accuracy, recall rate, F1 value, BLEU (Bilingual Evaluation Understudy), etc.
[0109] When the general large language model in the agent is fine-tuned based on the training data and the document comparison knowledge base to obtain the document comparison agent, the server can first configure the training environment, and based on the configured training environment, call the pre-written fine-tuning program to fine-tune the general large language model in the agent based on the training data and the document comparison knowledge base through the fine-tuning program to obtain the document comparison agent. Among them, the above-mentioned training environment includes the hardware resources (such as GPU) and software environment (such as deep learning framework) required for training, and may also include training parameters, which may include the number of training times, learning rate, batch size, etc. The document comparison agent may include a document comparison large language model, a tool component library and a document comparison knowledge base. Of course, it may also include other knowledge bases besides the document comparison knowledge base, which is not specifically limited in this description.
[0110] After obtaining the document comparison large language model, the model text corresponding to the document comparison large language model can be saved and published, and a service interface for calling the document comparison large language model can be generated.
[0111] The tool components in the tool component library of the agent can be selected based on tool requirements. That is, the server can determine the tool components that meet the tool requirements from the initial tool component library and use them as target tool components. A tool component library consisting of the target tool components is determined. Furthermore, the server can generate a processing flow for the document comparison agent based on the tool component library, the document comparison knowledge base (or knowledge bases), and the document comparison large language model. The processing flow of the document comparison agent is configured based on the generated processing flow. The tool requirements are pre-set. These tool requirements may include the types of tool components required. The initial tool component library may include several types of tool components, including tool components for data input, tool components for invoking the large language model, tool components for parsing files, tool components for vector retrieval, tool components for database queries, tool components for knowledge base queries, tool components for invoking interfaces for natural language processing technology, and tool components for network queries. The tool component library may include tool components for knowledge base queries, network queries, invoking interfaces for natural language processing technology, and invoking the large language model. It should be noted that the initial tool component library or the types of tool components included in the tool component library described above are only examples, and the types of tool components in the initial tool component library or the tool component library in this specification are not limited to the tool components described above.
[0112] The above-mentioned tool component for data input can support the input of various types of data such as text, audio, video, PDF, word, picture, http link, etc. The above-mentioned tool component for calling the large language model can be used to call the large language model in the intelligent agent. The above-mentioned tool component for parsing files can be used to parse various types of data files such as audio, video, PDF, word, picture, etc. The above-mentioned tool component for vector retrieval can be used to retrieve data from an already constructed vector library (such as the above-mentioned semantic vector library or document comparison knowledge base). The above-mentioned tool component for database query can be used to query data from an already constructed database through SQL statements. The above-mentioned tool component for knowledge base query can support querying knowledge data from various knowledge bases, the above-mentioned tool component for knowledge base retrieval can retrieve knowledge bases that meet the retrieval scenario information, the above-mentioned tool component for calling the interface of natural language processing technology is used to call the interface of natural language processing technology, and the above-mentioned tool component for network query can support searching for data from the Internet (i.e., the network).
[0113] The above-described processing flow represents the internal processing flow of the document comparison agent, i.e., a workflow. The server can use a pre-set generation algorithm to generate the processing flow of the document comparison agent based on the tool component library, the document comparison knowledge base (or knowledge bases), and the document comparison large language model. This processing flow can be in the form of a workflow. Based on the processing flow, the server can configure the internal processing sequence of the document comparison agent to obtain the document comparison agent.
[0114] S5: Construct prompt words and judge semantic relevance based on the document comparison agent: Based on the semantic similarity and text overlap of each similar text pair, splice the texts in each similar text pair and add them to the prompt words, input the prompt words into the document comparison agent, and use the document comparison agent to determine the knowledge data matching the prompt words from the document comparison knowledge base, and based on the knowledge data and the prompt words, determine the judgment result and judgment basis of whether each similar text pair is relevant through the document comparison large language model in the document comparison agent.
[0115] In this specification, the server can construct prompt words and judge the semantic relevance based on the document comparison agent, that is, based on the semantic similarity and text overlap of each similar text pair, the text in each similar text pair is spliced and supplemented into the prompt words, and the prompt words are input into the document comparison agent. The document comparison agent determines the knowledge data matching the prompt words from the document comparison knowledge base, and based on the knowledge data and the prompt words, the document comparison large language model in the document comparison agent is used to determine the judgment results and judgment basis of whether each similar text pair is relevant.
[0116] In this manual, there is a risk of misjudgment when relying solely on semantic vector similarity for semantic comparison. For example, there are a large number of fixed sentence paragraphs (such as experimental steps, background introduction) in the document, which makes it difficult to distinguish between structural similarity and relevance, or semantically similar texts may lack significant overlap in the literal sense, making it difficult to intuitively judge whether there is a correlation. Therefore, in order to overcome the above-mentioned technical defects, the server can splice the text in each similar text pair and add it to the prompt word based on the semantic similarity and text overlap between the retrieval text and the comparison text in each similar text pair, input the prompt word into the document comparison agent, determine the knowledge data that matches the prompt word from the document comparison knowledge base through the document comparison agent, and determine the judgment result and judgment basis of whether each similar text pair is relevant based on the knowledge data and the prompt word through the document comparison large language model in the document comparison agent. Among them, each similar text pair is a similar text pair between the retrieval document and each similar document. The judgment result is one of relevance and irrelevance, and the judgment basis is the reason corresponding to the judgment result.
[0117] Specifically, the method for constructing the prompt word in the above S5 is that the server can calculate the semantic similarity of each similar text pair and determine whether it exceeds the first threshold value. If it exceeds the first threshold value, the server can continue to calculate the text overlap in each similar text pair and determine whether it exceeds the second threshold value. If it does not exceed the second threshold value, the search text and the comparison text in each similar text pair are spliced and added to the prompt word. Among them, the semantic similarity is the cosine similarity or Euclidean distance between the search semantic vector corresponding to the search text in the similar text pair and the comparison semantic vector corresponding to the comparison text. The first threshold value and the second threshold value are respectively preset values. The above text overlap is the degree of literal overlap between the search text and the comparison text in the similar text pair. The text overlap can be determined using a preset algorithm or model, and this specification does not make specific limitations.
[0118] When the search text and comparison text in each similar text pair are spliced together and added to the prompt word, the server can splice the search text and comparison text in each similar text pair together through a preset prompt word template to obtain a prompt word. It should be noted that the prompt word at this time is a complete prompt word. In addition, the server can also splice the preset comparison rule, the search text and comparison text in each similar text pair together through a preset prompt word template to obtain a prompt word. Wherein, the above-mentioned prompt word template is pre-set, and the prompt word template includes fixed text and each card slot. Each card slot is used to fill in the search text and comparison text in the similar text pair and the comparison rule. The comparison rule is a principle for scientific and technological novelty search and document comparison, and the comparison rule can be pre-set. The fixed text can be a pre-set content, which is used to realize the comparison rule, the search text and comparison text in the similar text pair to be continuous. The above-mentioned prompt word template is used to realize the comparison rule, the splicing of the search text and comparison text in the similar text pair. The above prompt word template can be "Use [Card Slot 1] to determine whether [Card Slot 2] is related to [Card Slot 3], and generate a judgment result for representing whether it is related and a judgment basis corresponding to the generated judgment result:", "[Card Slot 1]" in the prompt word template is used to fill in the comparison rule, "[Card Slot 2]" in the prompt word template is used to fill in the search text, and "[Card Slot 3]" in the prompt word template is used to fill in the comparison text. Therefore, the above prompt word can be "Use [Comparison Rule] to determine whether [Retrieval Text] is related to [Comparison Text], and generate a judgment result for representing whether it is related and a judgment basis corresponding to the generated judgment result:". The fixed text in the above prompt word template can be pre-set, corresponding to the field value corresponding to the above instruction field, the search text and comparison text in the above similar text pair correspond to the field value corresponding to the above input field, and the above judgment result and judgment basis correspond to the field value corresponding to the above label field.
[0119] In addition, when there are multiple similar documents, the server can determine the semantic similarity between the search text and the comparison text in each similar document and each similar text pair between the similar document and the search document. When the semantic similarity of the similar text pair exceeds a first threshold, the text overlap between the search text and the comparison text in the similar text pair is determined. When the text overlap of the similar text pair does not exceed a second threshold, the preset comparison rules, the search text and the comparison text in the similar text pair are spliced together using a preset prompt word template to obtain a prompt word, and the prompt word is input into the document comparison agent, and the document comparison agent determines the knowledge data that matches the prompt word from the document comparison knowledge base through the document comparison agent, and based on the knowledge data and the prompt word, the document comparison large language model in the document comparison agent is used to obtain the judgment result and judgment basis of the similar text pair. For similar text pairs with relatively close semantics but low text overlap, it is necessary to combine the document comparison agent to make intelligent judgments. Through prompt engineering (Prompt Engineering), that is, prompt word templates, prompt words are constructed, and combined with knowledge data matching the prompt words, the document comparison agent is allowed to judge whether similar text pairs are related based on scientific and technological novelty search, document comparison principles and knowledge data, and further generate judgment basis to facilitate subsequent verification.
[0120] In this specification, after determining the semantic similarity between the search text and the comparison text in each similar text pair, when the semantic similarity of each similar text pair does not exceed the first threshold, that is, if the semantic similarity does not exceed the first threshold, the server can determine that the judgment result of each similar text pair is irrelevant, and determine that the judgment basis for each similar text pair is that the semantic similarity does not exceed the first threshold. Among them, since the semantic similarity does not exceed the first threshold, it means that there is a large difference in semantics between the search text and the comparison text, then the search text is irrelevant to the comparison text, and the corresponding judgment basis is that the semantic similarity does not exceed the first threshold. In addition, when there are multiple similar documents, when the semantic similarity of the similar text pair does not exceed the first threshold, the server can determine that the judgment result of the similar text pair is irrelevant, and determine that the judgment basis for the similar text pair is that the semantic similarity does not exceed the first threshold.
[0121] In addition, after the above determination that the semantic similarity exceeds the first threshold, the server determines that the judgment result of each similar text pair is related when the text overlap of each similar text pair exceeds the second threshold (that is, the text overlap exceeds the second threshold), and determines that the judgment basis for each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold. Among them, since the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold, it means that the semantics between the search text and the comparison text are similar, and the text overlap is high, then the search text is related to the comparison text, and the corresponding judgment basis is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold. In addition, when there are multiple similar documents, after the semantic similarity of the similar text pair exceeds the first threshold, the server can determine that the judgment result of the similar text pair is related when the text overlap of the similar text pair exceeds the second threshold, and determine that the judgment basis for the similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
[0122] S6: Generate a final search report: Generate a search report corresponding to the search document and the similar document based on the judgment result and judgment basis of each similar text pair, and determine a final search report based on the search report.
[0123] In this specification, the server may generate a final search report, that is, based on the judgment results and judgment basis of each similar text pair, generate a search report corresponding to the search document and similar documents, and determine a final search report based on the search report. The search report may be a summary of each judgment result and each judgment basis. In addition, the search report may also include the names or identifiers of similar text pairs and similar documents. The final search report may be a search report. The final search report may be displayed to a user, i.e., a scientific and technological novelty searcher or writer.
[0124] In addition, when there are multiple similar documents, the server can generate a search report corresponding to the search document and each similar document based on the judgment result and judgment basis of whether each similar text pair is related, and determine the final search report based on each search report. Among them, each similar text pair determined in the above S3 is determined based on each similar document and the search document, that is, one similar document and the search document can determine multiple similar text pairs. The above S5 is the judgment result and judgment basis corresponding to each similar text pair determined based on each similar document and the search document. Therefore, S6 can generate a search report corresponding to the search document and each similar document based on the judgment result and judgment basis of each similar text pair between the search document and each similar document. Specifically, the server can generate a search report corresponding to the similar text and the search document based on the judgment result and judgment basis corresponding to each similar text pair between the similar document and the search text for each similar document. The search report can be a summary of each judgment result and each judgment basis. In addition, the search report can also include the name or identification of the similar text pair and the similar document.
[0125] The above-mentioned final search report can be a summary of each search report, that is, each search report is directly used as the final search report. Of course, the above-mentioned final search report can also be one of each search report, that is, when determining the final search report based on each search report, the server can determine the relevance score corresponding to each similar document based on each search report, and then determine the final search report from each search report based on each relevance score. Among them, the relevance score can be the number of results that are judged to be relevant. Specifically, the server can determine, for each similar document, the number of similar documents and the search document corresponding to the search report that are judged to be relevant, and use it as the relevance score corresponding to the similar document. Based on the relevance score corresponding to each similar document, the search report between the similar document and the search document corresponding to the maximum relevance score is determined and used as the final search report.
[0126] In some embodiments of the present specification, during the vector search, a preset vector search strategy may be used to determine the comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library. The vector search strategy may be based on a preset relationship field and search from the comparison text that meets the relationship field. The relationship field may be at least one of the document title, document type, and chapter type, or other fields. This specification does not make specific limitations. The relationship field is pre-set by the user. Taking the relationship field as the chapter type as an example, the server may search from the comparison text of the same chapter type as the search text to perform semantic matching between the search text and the comparison text corresponding to the same type of chapter, thereby providing the accuracy of the semantic matching. Based on this, when determining the comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library in S2 above, the server may determine the search information corresponding to each search semantic vector based on the preset relationship field, and screen the comparison semantic vectors in the pre-built semantic vector library based on the search information. From the screened comparison semantic vectors, a comparison semantic vector that matches each search semantic vector is determined. Among them, the retrieval information is the field value of the preset relationship field of the retrieval text corresponding to the retrieval semantic vector. Taking the relationship field as the chapter type as an example, the retrieval information is the chapter type of the retrieval text corresponding to the retrieval semantics. Taking the retrieval document or the comparison document as a paper as an example, the chapter type may include title, abstract, introduction, literature review, method, conclusion, acknowledgments, references and appendix, etc. Each comparison text corresponding to each comparison semantic vector screened out above is a vector that meets the retrieval information of each retrieval semantic vector. When the comparison semantic vectors in the semantic vector library are screened according to the retrieval information, for each comparison semantic vector in the semantic vector library, when the field value corresponding to the relationship field of the comparison text of the comparison semantic vector is consistent with the retrieval information, the comparison semantic vector is used as the screened comparison semantic vector.
[0127] In addition, the server may also, for each search semantic vector, determine the search information of the search text corresponding to the search semantic vector based on the preset relationship field, and then determine, from the pre-built semantic vector library, the comparison semantic vectors that satisfy the search information corresponding to the search semantic vector, and determine, from the determined comparison semantic vectors, the comparison semantic vector that matches the search semantic vector. The process of determining, from the determined comparison semantic vectors, the comparison semantic vector that matches the search semantic vector is similar to the process of determining, from the semantic vector library, the comparison semantic vector that matches each search semantic vector in S3 above, and will not be repeated here.
[0128] In some embodiments of this specification, each search text may be matched with multiple comparison texts. Therefore, the search results may include multiple comparison semantic vectors that match the search semantic vector of the search document, labels corresponding to the multiple comparison semantic vectors, and similarities between the search semantic vector and the multiple comparison semantic vectors. The number of comparison texts matching each search text may be preset.
[0129] In some embodiments of the present specification, when generating a search report corresponding to the search document and each similar document based on the judgment results and judgment basis of each similar text pair between the search document and each similar document in the above S6, the server may generate a first report corresponding to the search document and each similar text based on the judgment results and judgment basis of each similar text pair between the search document and each similar document. Based on the judgment results and judgment basis of each similar text pair between the search document and each similar document, similar text pairs that meet the preset target conditions are determined, and based on the judgment results and judgment basis of the determined similar text pairs, a second report between the search document and each similar document is generated. The first report and the second report are used as search reports between the search text and each similar document. Among them, the above first report is a summary of all judgment results and judgment basis, and the second report is a summary of the judgment results and judgment basis that meet the preset target conditions. The target condition is pre-set, and the target condition includes semantic similarity detection items (i.e., the semantic similarity exceeds the first threshold, the text overlap does not exceed the second threshold, and the judgment result output by the document comparison agent is relevant), irrelevant items judged by the document comparison agent (i.e., the semantic similarity exceeds the first threshold, the text overlap does not exceed the second threshold, and the judgment result output by the document comparison agent is irrelevant), and highly similar semantic text items (i.e., the semantic similarity exceeds the first threshold, and the text overlap exceeds the second threshold). The above-mentioned semantic similarity detection items and irrelevant items judged by the document comparison agent are mainly for text pairs with high semantic similarity and low text overlap (i.e., the semantic similarity exceeds the first threshold and the text overlap does not exceed the threshold), and are obtained by the document comparison agent through the judgment of whether they are relevant. By additionally generating a second report, it is helpful for subsequent users to verify the judgment results and judgment basis. The specific implementation process of generating the retrieval report corresponding to the retrieval document and the similar document based on the judgment results and judgment basis of whether each similar text pair between the retrieval document and the similar document is similar to the above process and will not be repeated here.
[0130] In some embodiments of this specification, taking a text describing global environmental change as an example, assume that Text 1 reads: "With the intensification of global climate change, rising sea levels have become a problem that cannot be ignored. Scientists warn that if no measures are taken, many coastal cities may be submerged." If the order of the text is simply reversed, the original text (i.e., Text 1) can be transformed into Text 2, i.e., Text 2: "With the intensification of global climate change, scientists warn that rising sea levels have become a problem that cannot be ignored. If no measures are taken, many coastal cities may be submerged." Since the text itself has hardly changed, the literal overlap between Text 2 and Text 1 is close to 100%. This correlation can be easily detected using traditional literal overlap (i.e., text overlap) comparison methods.
[0131] However, if the key words of the original text are replaced with synonyms and transformed into text 3, that is, text 3: "In view of the worsening global warming, the rise in ocean levels has become a serious threat. Researchers have warned that if no response strategies are taken, many coastal cities may suffer a catastrophe." Then from the perspective of text overlap, text 3 and text 1 have less than 28.7% overlap, and the repeated words are relatively common and discontinuous non-academic words. It is difficult for the text overlap detection method to find this kind of related method. However, since the replacements are all semantically similar synonyms, from the perspective of semantic vector comparison, the similarity is 86.9%. The semantics between text 1 and text 3 are highly similar. This related text pair can be found through semantic similarity retrieval and comparison. The comparison results of the two cases of text overlap and semantic similarity are listed in the form of a table below, and the judgment results and judgment basis based on the document comparison agent are added, that is, using Figure 1 The judgment results and judgment basis obtained by comparing in the above manner are specifically shown in Table 1 below.
[0132] Table 1
[0133]
[0134] Table 1 shows a text comparison example. It includes the search text (i.e., Text 1) and the comparison texts (i.e., Text 2 and Text 3). The text similarity between Text 1 and Text 2 is 100%, and the semantic vector similarity (i.e., semantic similarity) is 99.4%. The document comparison agent's document comparison large language model outputs the judgment result (i.e., the large language model judgment result in Table 1) as related, based on "the two texts are basically consistent." The text similarity between Text 1 and Text 3 is 28.7%, and the semantic vector similarity (i.e., semantic similarity) is 86.9%. The large language model judgment result is related, based on "high consistency in key content and expression, and significant semantic repetition."
[0135] In some embodiments of the present specification, when the document comparison agent determines the knowledge data matching the prompt word from the document comparison knowledge base in step S5, the server may employ the document comparison agent, using retrieval enhancement generation technology, to retrieve the knowledge data matching the prompt word from the document comparison knowledge base. Specifically, for each similar text pair, the server may employ the document comparison agent, using retrieval enhancement generation technology, to retrieve the knowledge data matching the prompt word for that similar text pair from the document comparison knowledge base.
[0136] In some embodiments of this specification, Figure 1 The method is mainly divided into three parts. The first part is to build a semantic vector library, that is, to build a semantic vector library according to the semantic vectors corresponding to the compared texts of the compared documents. The specific process is the process of pre-building the semantic vector library as mentioned above, such as Figure 2 As shown, Figure 2 This is a flowchart of a semantic vector library construction process provided in this specification. Figure 2 Starting from the "Start" node, first obtain the uploaded comparison documents, and perform text extraction on each comparison document to obtain the comparison text data corresponding to each comparison document, that is, Figure 2 After that, the text data is split according to the chapter hierarchy of each comparison document to obtain the data of each comparison chapter. According to the pre-set filtering rules, the data of each comparison chapter is filtered to obtain the data of each comparison target chapter, that is, Figure 2 "Document Chapter Split and Filter" in the . According to the granularity requirements of each comparison document, the data of each comparison target chapter is split through natural language processing technology to obtain each comparison text, that is, Figure 2 The semantic vectorization model is used to determine the high-dimensional semantic vectors corresponding to each comparison text of each comparison document, that is, Figure 2 Generate labels corresponding to each semantic vector, and build a semantic vector library based on each semantic vector and the label corresponding to each semantic vector, that is, Figure 2 Finally, go to the "End" node.
[0137] The second part is to perform similar document retrieval based on the semantic vector library, that is, to pre-process the retrieval documents and search in the semantic vector library, integrate the retrieval results corresponding to each retrieval text, narrow the scope of similar documents through weighted sorting, and obtain similar documents. The specific process is the process of S2 above, such as Figure 3 As shown, Figure 3 This is a schematic diagram of a document semantic retrieval and comparison process provided in this specification. Figure 3Starting from the "start" node, first obtain the search document, and perform data preprocessing on the search document to obtain the search texts corresponding to the search document, that is, Figure 3 The semantic vectorization model is used to determine the search semantic vector corresponding to each search text, that is, Figure 3 "Semantic vectorization" in . From the pre-built semantic vector library, determine the comparison semantic vector that matches each search semantic vector to generate the search results corresponding to each search text, that is, Figure 3 "Vector Library Search" in . According to the similarities included in each search result and the comparison text corresponding to each label, the similarity score corresponding to each comparison document is determined. According to each similarity score, each comparison document is screened to determine similar documents, that is, Figure 3 The "Similar Document Filtering and Sorting" in the . Finally, go to the "End" node.
[0138] The third part is the document comparison assisted by the document comparison agent, which is based on semantic vector comparison, text comparison, document comparison agent judgment and other methods to compare the search document with similar documents, and combine the basis of whether the document comparison agent generates relevance to integrate all the comparison results to generate a search report. The specific process is the process of S3 to S6 above, as shown in the following example: Figure 4 As shown, Figure 4 This is a schematic diagram of a document comparison process assisted by a document comparison agent provided in this specification. Figure 4 Starting from the "start" node, first determine the similar text pairs between the retrieved document and the similar documents based on the retrieval semantic vectors of the retrieved document and the comparison semantic vectors of the similar documents, that is, Figure 4 According to the semantic similarity and text overlap between the search text and the comparison text in each similar text pair, the document comparison agent is used to determine the judgment result and judgment basis of each similar text pair, which is consistent with the process in S5 above, that is, Figure 4 "Determine whether it is relevant and generate judgment basis" in . Based on the judgment results and judgment basis of each similar text pair between the retrieved document and similar documents, generate retrieval reports corresponding to the retrieved document and similar documents, and determine the final retrieval report based on each retrieval report, that is, Figure 4 Then, go to the "Summarize Results and Generate Report" in the dialog box. Finally, go to the "End" node.
[0139] When "determining whether it is relevant and generating the basis for judgment" and "generating a report based on the summary results" are mentioned above, if Figure 5 As shown, Figure 5 This is a schematic diagram of a document comparison agent-assisted related determination process provided in this specification. Figure 5Starting from the "start" node, the semantic similarity between the search text and the comparison text in each similar text pair is first determined, and then the semantic similarity of each similar text pair is judged to be greater than the first threshold, that is, Figure 5 In the “determine whether the semantic similarity exceeds the first threshold”, if not (that is, the semantic similarity of each similar text pair does not exceed the first threshold), the server can determine that the judgment result of each similar text pair is irrelevant, and determine that the judgment basis for each similar text pair is that the semantic similarity does not exceed the first threshold, that is, Figure 5 “Judged to be irrelevant” in the
[0140] If yes (i.e. the semantic similarity of each similar text pair exceeds the first threshold), continue to determine the text overlap between the search text and the comparison text in each similar text pair, and judge whether the text overlap of each similar text pair exceeds the second threshold, i.e. Figure 5 In the “determine whether the text overlap exceeds the second threshold”, if so (i.e., the text overlap of each similar text pair exceeds the second threshold), the server can determine that the judgment result of each similar text pair is relevant, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold, i.e. Figure 5 If not (i.e., the text overlap degree of each similar text pair does not exceed the second threshold), the preset prompt word template is used to splice the preset comparison rules, the search text and the comparison text in each similar text pair to obtain the prompt word, and the prompt word is input into the document comparison agent, and the document comparison agent determines the knowledge data matching the prompt word from the document comparison knowledge base through the document comparison agent, and according to the knowledge data and the prompt word, the document comparison large language model in the document comparison agent is used to obtain the judgment result and judgment basis of each similar text pair, i.e. Figure 5 In the “Document Comparison Agent determines whether it is relevant and generates a basis for judgment”. If the document comparison agent determines that it is relevant (i.e., if it is), then the judgment result is relevant, i.e. Figure 5 If the document comparison agent determines that it is not relevant (i.e. if not), the result is irrelevant. Figure 5 After that, based on the judgment results and judgment basis of each similar text pair, a search report corresponding to the search document and similar documents is generated, and based on the search report, a final search report is generated, i.e. Figure 5 Then, go to the "Summarize Results and Generate Report" in the dialog box. Finally, go to the "End" node.
[0141] Based on the above Figure 1 The method shown in this specification provides a document semantic comparison system based on a large language model agent, such as Figure 6 As shown, Figure 6This is a schematic diagram of a document semantic comparison system based on a large language model agent provided in this specification. The system includes a data storage module 100, a data preprocessing module 101, a semantic vectorization module 102, a vector retrieval module 103, a document comparison module 104, a report generation module 105 and an agent construction module 106.
[0142] The above-mentioned data storage module 100 can store models, pre-built semantic vector libraries, search documents and each comparison document in a pre-uploaded comparison document collection. The model includes a semantic vectorization model and a document comparison agent. The semantic vector library includes at least the comparison semantic vectors corresponding to each comparison text of each comparison document.
[0143] The above-mentioned intelligent agent construction module 106 can be based on the retrieval scenario information. In the intelligent agent, by calling the tool component in the tool component library, a document comparison knowledge base is selected from several knowledge bases, and based on the training data and the document comparison knowledge base, the general large language model in the intelligent agent is fine-tuned to obtain a document comparison intelligent agent. Among them, the knowledge base includes a knowledge base constructed by local data and a knowledge base constructed by data obtained from the network. The training data includes the field values corresponding to the instruction field, the input field, and the label field. The specific execution process of the above-mentioned intelligent agent construction module 106 is consistent with the execution process of the above-mentioned S4, and will not be repeated here.
[0144] When the document comparison agent construction module 106 fine-tunes the general large language model in the agent based on the training data and the document comparison knowledge base to obtain the document comparison large language model, it can use retrieval enhancement generation technology to retrieve knowledge data that matches the training data from the document comparison knowledge base, and input the field values corresponding to the knowledge data, the instruction field in the training data, and the input field into the general large language model in the agent to obtain the output result. The general large language model is fine-tuned based on the output result and the field value corresponding to the label field in the training data. Verification data is obtained, and the fine-tuned general large language model is evaluated based on the verification data. After the evaluation passes, the fine-tuned general large language model is used as the document comparison large language model to obtain the document comparison agent.
[0145] The data preprocessing module 101 can perform data preprocessing on the search document to obtain the search texts corresponding to the search document. Figure 1The implementation process of performing data preprocessing on the search document in S1 to obtain the search texts corresponding to the search document is consistent. Specifically, the data preprocessing module 101 can perform text extraction on the search document to obtain the text data corresponding to the search document. According to the chapter hierarchy structure of the search document, the text data is split to obtain the data of each chapter. According to the pre-set filtering rules, the data of each chapter is filtered to obtain the data of each target chapter. According to the granularity requirements corresponding to the search document, the data of each target chapter is split using natural language processing technology to obtain the search texts.
[0146] The semantic vectorization module 102 can call the semantic vectorization model to determine the high-dimensional search semantic vector corresponding to each search text. Figure 1 In S1, the semantic vectorization model is used to determine the retrieval semantic vector corresponding to each retrieval text. The implementation process is consistent and will not be repeated here.
[0147] The vector retrieval module 103 can retrieve the comparison semantic vectors that match each search semantic vector from the semantic vector library, and determine similar documents based on the comparison text corresponding to the label of each searched comparison semantic vector. Figure 1 The implementation process of S2 is the same as that of , so I will not go into details here.
[0148] The above-mentioned semantic vector library also includes a label corresponding to each comparison semantic vector, and the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector. Based on this, when the vector retrieval module 103 determines the comparison semantic vector that matches each retrieval semantic vector in the semantic vector library, it can determine the retrieval information corresponding to each retrieval semantic vector based on the preset relationship field, and screen the comparison semantic vectors in the pre-built semantic vector library based on the retrieval information. From the screened comparison semantic vectors, the comparison semantic vector that matches each retrieval semantic vector is determined. Among them, the relationship field is at least one of the document title, document type and chapter type.
[0149] In addition, the vector retrieval module 103 can also determine the matching semantic vectors that match each search semantic vector from a pre-built semantic vector library to generate search results corresponding to each search text. Based on the matching text corresponding to each search result, the matching documents in the comparison document collection are sorted and screened to identify similar documents. The search results include the labels of the retrieved matching semantic vectors, and the semantic vector library includes at least the matching semantic vectors for each matching text corresponding to each matching document in the pre-uploaded comparison document collection.
[0150] The document comparison module 104 can match each search text with each comparison text of similar documents one by one to obtain each similar text pair. Based on the semantic similarity and text overlap of each similar text pair, the text in each similar text pair is spliced and added to the prompt word, and the prompt word is input into the document comparison intelligent agent. The document comparison intelligent agent determines the knowledge data that matches the prompt word from the document comparison knowledge base, and based on the knowledge data and the prompt word, the document comparison large language model in the document comparison intelligent agent determines the judgment result and judgment basis of whether each similar text pair is related. The specific implementation process of the document comparison module 104 is the same as the above Figure 1 The implementation process of S3 and S5 is the same, so I will not go into details here.
[0151] Specifically, when the document comparison module 104 matches each search text with each comparison text of a similar document one-to-one to obtain each similar text pair, it matches each search text of the search document with each comparison text of a similar document one-to-one based on each search semantic vector of the search document and each comparison semantic vector of each similar document to obtain each similar text pair.
[0152] When the text in each similar text pair is spliced and added to the prompt word, the document comparison module 104 can calculate the semantic similarity of each similar text pair and determine whether it exceeds a first threshold. If it exceeds the first threshold, the text overlap degree in each similar text pair is further calculated and determined whether it exceeds a second threshold. If it does not exceed the second threshold, the search text and the comparison text in each similar text pair are spliced and added to the prompt word. In addition, the document comparison module 104 can also splice the preset comparison rules, the search text in each similar text pair, and the comparison text using a preset prompt word template to obtain the prompt word.
[0153] The document comparison module 104 may further determine that each similar text pair is irrelevant if the first threshold is not exceeded, and determine that the basis for determining each similar text pair is that the semantic similarity does not exceed the first threshold. If the second threshold is exceeded, the document comparison module 104 may determine that each similar text pair is relevant, and determine that the basis for determining each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
[0154] The report generation module 105 can generate a search report corresponding to the search document and the similar document according to the judgment result and judgment basis of each similar text pair, and determine the final search report based on the search report. Figure 1 The implementation process of S6 is the same as that of , so I will not go into details here.
[0155] In some embodiments of the present specification, the data storage module may also add, delete, modify, and query the data stored in itself, such as the semantic vector library.
[0156] In some embodiments of the present specification, the data preprocessing module 101 may also obtain an uploaded comparison document set including each comparison document, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document. The semantic vectorization module 102 may also use a semantic vectorization model to determine the comparison semantic vectors corresponding to each comparison text of each comparison document. The data storage module 100 may generate a label corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector. The label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector.
[0157] In some embodiments of the present specification, when the report generation module 105 generates a search report corresponding to the search document and each similar document based on the judgment results and judgment basis of each similar text pair between the search document and each similar document, it can generate a first report corresponding to the search document and each similar text based on the judgment results and judgment basis of each similar text pair between the search document and each similar document. Based on the judgment results and judgment basis of each similar text pair between the search document and each similar document, similar text pairs that meet the preset target conditions are determined, and based on the judgment results and judgment basis of the determined similar text pairs, a second report between the search document and each similar document is generated. The first report and the second report are used as the search report between the search text and each similar document.
[0158] In some embodiments of this specification, the above system mainly provides services to the outside in the form of B / S architecture software services. The overall architecture of the system is divided into data layer, application layer, interaction layer and external data, such as Figure 7 As shown, Figure 7This is a schematic diagram of a document semantic comparison system architecture based on a large language model agent provided in this specification, wherein the data layer mainly includes a model and a semantic vector library, and the model includes a semantic vectorization model and a document comparison agent. The application layer mainly includes a data storage module 100, a data preprocessing module 101, a semantic vectorization module 102, a vector retrieval module 103, a document comparison module 104, a report generation module 105, and a document comparison agent construction module 106. The above-mentioned interaction layer mainly provides functions for interacting with users, mainly including providing a UI interface to interact with users, processing user input and output (such as form submission, page rendering), and displaying the results of application layer processing. The interaction layer provides document storage, document retrieval, and report generation functions, namely, a document storage page, a document retrieval page, and a report generation page. Users can upload each comparison document in batches or upload a single comparison document through the document storage page, and the document storage page sends each comparison document to the data storage module 100 for storage. Users can upload and retrieve documents through the document retrieval page, and the document retrieval page sends the retrieved documents to the data storage module 100 for storage. In addition, users can also set a vector search strategy through the document search page, which will then be sent to the vector search module 103 for subsequent use during vector searches. The report generation module 105 can display the generated final search report through the report generation page, which also provides a preview of the final search report and a report download function. In addition, users can also edit and modify the final search report on the report generation page. The external data primarily includes two parts: the comparison documents to be searched (i.e., the comparison document set) and the search document to be searched.
[0159] It should be noted that the specific embodiments described above can enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the present invention can still be modified or replaced with equivalents. In short, all technical solutions and improvements that do not depart from the spirit and scope of the present invention should be included in the scope of protection of the patent for the present invention.
Claims
1. A document semantic comparison method based on a large language model agent, characterized in that: include: S1: Split the search document and construct a search semantic vector: Obtain a search document and perform data preprocessing on the search document to obtain search texts corresponding to the search document; use a semantic vectorization model to determine the search semantic vectors corresponding to the search texts; S2: Retrieve similar documents: retrieve the comparison semantic vectors that match each retrieval semantic vector from the pre-built semantic vector library, and determine similar documents based on the comparison text corresponding to the label of each retrieved comparison semantic vector; S3: Generate similar text pairs: match each search text of the search document with each comparison text of the similar document one by one to obtain similar text pairs; S4: Fine-tuning the large language model in the agent and constructing a document comparison agent: Based on the retrieval scenario information, in the agent, a document comparison knowledge base is selected from a plurality of knowledge bases by calling a tool component in a tool component library, and the general large language model in the agent is fine-tuned based on the training data and the document comparison knowledge base to obtain a document comparison agent. The knowledge base includes a knowledge base constructed from local data and a knowledge base constructed from data acquired from the network. S5: Constructing prompt words and judging semantic relevance based on the document comparison agent: Based on the semantic similarity and text overlap of each similar text pair, the texts in each similar text pair are spliced and added to the prompt words, the prompt words are input into the document comparison agent, and the document comparison agent determines knowledge data matching the prompt words from the document comparison knowledge base, and determines the judgment result and judgment basis of whether each similar text pair is relevant based on the knowledge data and the prompt words and the document comparison large language model in the document comparison agent; S6: Generate a final search report: Generate a search report corresponding to the search document and the similar document based on the judgment result and judgment basis of each similar text pair, and determine a final search report based on the search report.
2. A document semantic comparison method based on a large language model agent as claimed in claim 1, characterized in that: The method for constructing the prompt word in S5 is: Calculate the semantic similarity of each similar text pair and determine whether it exceeds a first threshold; if it exceeds the first threshold, continue to calculate the text overlap in each similar text pair and determine whether it exceeds a second threshold; if it does not exceed the second threshold, splice the search text and comparison text in each similar text pair and add them to the prompt word.
3. The document semantic comparison method based on a large language model agent according to claim 2, characterized in that: The method further comprises: If the semantic similarity does not exceed the first threshold, determining that the judgment result of each similar text pair is irrelevant, and determining that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold; If the second threshold is exceeded, the judgment result of each similar text pair is determined to be relevant, and the judgment basis of each similar text pair is determined to be that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
4. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: The step S1 of performing data preprocessing on the search document to obtain the search texts corresponding to the search document specifically includes: Performing text extraction on the search document to obtain text data corresponding to the search document; Splitting the text data according to the chapter hierarchical structure of the search document to obtain data of each chapter; Filtering the chapter data according to pre-set filtering rules to obtain target chapter data; According to the granularity requirements corresponding to the search documents, each target chapter data is split using natural language processing technology to obtain various search texts.
5. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: The semantic vector library pre-built in S2 specifically includes: Obtaining an uploaded comparison document set including each comparison document, and performing data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; Using the semantic vectorization model, determining the comparison semantic vectors corresponding to the comparison texts of each comparison document; Generate a label corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector.
6. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: The S2 specifically includes: Determining, from a pre-built semantic vector library, a comparison semantic vector that matches each search semantic vector to generate a search result corresponding to each search text; wherein the search result includes a label of the retrieved comparison semantic vector; According to the comparison texts corresponding to the respective search results, the comparison documents in the comparison document set are sorted and screened to determine similar documents; wherein the semantic vector library at least includes the comparison semantic vectors of the comparison texts corresponding to the respective comparison documents in the comparison document set uploaded in advance.
7. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: The step S2 of retrieving a comparison semantic vector that matches each search semantic vector from the pre-built semantic vector library specifically includes: Determining the search information corresponding to each search semantic vector based on a preset relationship field, and screening the comparison semantic vectors in the pre-built semantic vector library based on the search information; wherein the relationship field is at least one of a document title, a document type, and a section type; From the screened comparison semantic vectors, a comparison semantic vector that matches each search semantic vector is determined.
8. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: In S4, based on the training data and the document comparison knowledge base, the general language model in the agent is fine-tuned to obtain the document comparison agent, which specifically includes: Retrieving knowledge data matching the training data from the document comparison knowledge base using retrieval enhancement generation technology; Inputting the field values corresponding to the instruction field and the input field in the knowledge data, the training data, and the general language model in the agent to obtain an output result; Fine-tuning the universal large language model according to the output result and the field value corresponding to the label field in the training data; Obtaining validation data and evaluating the fine-tuned general large language model based on the validation data; After the evaluation is passed, the fine-tuned universal large language model is used as the document comparison large language model to obtain a document comparison agent.
9. The document semantic comparison method based on a large language model agent according to claim 1, characterized in that: The tool component library includes tool components for knowledge base query, network query, calling interfaces of natural language processing technology and calling large language models.
10. A document semantic comparison system based on a large language model agent according to any one of claims 1 to 9, characterized in that: The system includes a data storage module, a data preprocessing module, a semantic vectorization module, a vector retrieval module, a document comparison module, a report generation module, and an agent construction module, wherein: The data storage module is used to store the model, the pre-built semantic vector library, the search documents and the comparison documents in the pre-uploaded comparison document set, wherein the model includes a semantic vectorization model and a document comparison agent; The agent construction module is configured to select a document comparison knowledge base from a plurality of knowledge bases by calling a tool component in a tool component library in the agent based on the retrieval scenario information, and fine-tune a general large language model in the agent based on the training data and the document comparison knowledge base to obtain a document comparison agent, wherein the knowledge base includes a knowledge base constructed from local data and a knowledge base constructed from data acquired from the network; The data preprocessing module is used to perform data preprocessing on the search document to obtain each search text corresponding to the search document; The semantic vectorization module is used to call the semantic vectorization model to determine the search semantic vectors corresponding to the respective search texts; The vector retrieval module is configured to retrieve a comparison semantic vector that matches each search semantic vector from the semantic vector library, and determine similar documents based on a comparison text corresponding to a label of each retrieved comparison semantic vector; The document comparison module is configured to match each of the search texts with each of the comparison texts of the similar documents one by one to obtain similar text pairs; based on the semantic similarity and text overlap of each similar text pair, the texts in each similar text pair are spliced and added to the prompt word; the prompt word is input into the document comparison agent; the document comparison agent determines knowledge data matching the prompt word from the document comparison knowledge base; and based on the knowledge data and the prompt word, the document comparison large language model in the document comparison agent determines the judgment result and judgment basis of whether each similar text pair is relevant; The report generating module is used to generate a search report corresponding to the search document and the similar document according to the judgment result and judgment basis of each similar text pair, and determine a final search report based on the search report.
11. A document semantic comparison system based on a large language model agent as claimed in claim 10, characterized in that: The document comparison module is specifically used to calculate the semantic similarity of each similar text pair and determine whether it exceeds a first threshold; if it exceeds the first threshold, continue to calculate the text overlap in each similar text pair and determine whether it exceeds a second threshold; if it does not exceed the second threshold, splice the search text and the comparison text in each similar text pair and add them to the prompt word.
12. A document semantic comparison system based on a large language model agent as claimed in claim 11, characterized in that: The document comparison module is also used to, if it does not exceed the first threshold, determine that the judgment result of each similar text pair is irrelevant, and determine that the judgment basis of each similar text pair is that the semantic similarity does not exceed the first threshold; if it exceeds the second threshold, determine that the judgment result of each similar text pair is relevant, and determine that the judgment basis of each similar text pair is that the semantic similarity exceeds the first threshold and the text overlap exceeds the second threshold.
13. The document semantic comparison system based on a large language model agent according to claim 10, characterized in that: The data preprocessing module is specifically used to perform text extraction on the search document to obtain text data corresponding to the search document; split the text data according to the chapter hierarchy structure of the search document to obtain data of each chapter; filter the data of each chapter according to pre-set filtering rules to obtain data of each target chapter; and split each target chapter data according to the granularity requirements corresponding to the search document through natural language processing technology to obtain each search text.
14. The document semantic comparison system based on a large language model agent according to claim 10, characterized in that: The data preprocessing module is further used to obtain an uploaded comparison document set including each comparison document, and perform data preprocessing on each comparison document to obtain each comparison text corresponding to each comparison document; The semantic vectorization module is further configured to determine, using the semantic vectorization model, the comparison semantic vectors corresponding to the comparison texts of each comparison document; The data storage module is specifically used to generate a label corresponding to each comparison semantic vector, and construct a semantic vector library based on each comparison semantic vector and the label corresponding to each comparison semantic vector; wherein the label includes the comparison document information, chapter information and comparison text information corresponding to the comparison semantic vector.
15. The document semantic comparison system based on a large language model agent according to claim 10, characterized in that: The vector retrieval module is specifically configured to determine, from a pre-built semantic vector library, a comparison semantic vector that matches each search semantic vector, to generate a search result corresponding to each search text; wherein the search result includes a label of the retrieved comparison semantic vector; and, based on the comparison text corresponding to each search result, sort and screen each comparison document in the comparison document collection to determine similar documents; wherein the semantic vector library includes at least the comparison semantic vectors of each comparison text corresponding to each comparison document in the comparison document collection uploaded in advance.
16. The document semantic comparison system based on a large language model agent according to claim 10, characterized in that: The vector retrieval module is specifically configured to determine, based on a preset relationship field, retrieval information corresponding to each retrieval semantic vector, and filter, based on the retrieval information, comparison semantic vectors in a pre-built semantic vector library; wherein the relationship field is at least one of a document title, a document type, and a section type; and determine, from the filtered comparison semantic vectors, a comparison semantic vector that matches each retrieval semantic vector.
17. The document semantic comparison system based on a large language model agent according to claim 10, characterized in that: The document comparison agent building module is specifically used to use retrieval enhancement generation technology to retrieve knowledge data that matches the training data from the document comparison knowledge base; input the field values corresponding to the knowledge data, the instruction field in the training data, and the input field into the universal large language model in the agent to obtain an output result; The general large language model is fine-tuned according to the output result and the field value corresponding to the label field in the training data; verification data is obtained, and the fine-tuned general large language model is evaluated based on the verification data; after the evaluation passes, the fine-tuned general large language model is used as the document comparison large language model to obtain a document comparison agent.
Citation Information
Cited By
Knowledge question-answering system based on intelligent agent
CN121029951A
Document retrieval and conflict detection system based on retrieval enhancement generation
CN121119168A