A method for accurately positioning clauses in contract Q&A
Through the big model technology of similarity calculation, combined with the information of multiple chapters of the contract, the problems of inaccurate and low efficiency in the existing technology are solved, accurate positioning and complete answers are achieved, and the reliability of the system is improved.
Patent Information
- Application Number
- CN202411039747.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-31
AI Technical Summary
The prior art is difficult to accurately understand the complex semantic and contextual information in the contract, resulting in insufficient correlation of information retrieval results, low information retrieval efficiency, and inaccurate and complete questions.
The big model technology that uses similarity calculation is used to divide the contract content into multiple chapters, and combine multiple chapter information using keyword semantic retrieval and multiple rounds of dialogue to calculate the similarity between small-grained blocks and the answers to the big model, similar terms and user questions, and accurately locate and answer.
It improves the accuracy and efficiency of information retrieval, enhances the completeness of question answers, reduces the risk of information omission, and ensures the reliability of the system.
Smart Images

Figure CN118916463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and specifically to a method for accurately positioning clauses in contract Q&A. Background Art
[0002] With the rapid development of information technology, the digitization of contracts and the massive growth of information have put forward higher requirements for the efficient management and accurate retrieval of clause content in contracts. The main challenges faced by existing technologies include: Limitations in semantic understanding: Traditional systems often have difficulty accurately understanding complex semantics and context information in documents, resulting in insufficient accuracy of the final answers and relevance of retrieval results. For example, contracts usually contain general terms, special terms, and attachments. Without combining all relevant chapters, it is difficult to draw accurate conclusions based on a single chapter or paragraph.
[0003] Efficiency issues in information retrieval: Facing a large collection of clauses, how to quickly locate relevant information is another challenge faced by existing technologies.
[0004] Accuracy of question answering: Existing systems may not be able to provide comprehensive or accurate answers when answering complex questions.
[0005] Therefore, in view of the above problems, a method for accurately positioning clauses in contract Q&A is needed. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for accurately positioning clauses in contract Q&A. The present invention utilizes large model technology for similarity calculation, focusing on information positioning and question answering in document chapters. Through efficient information retrieval and processing capabilities, it provides users with more accurate and in-depth answer services.
[0007] The present invention is implemented as follows:
[0008] The present invention provides a method for accurately positioning clauses in contract Q&A, which is specifically executed according to the following steps:
[0009] S1: Divide the content of the contract into multiple chapters according to the title. For the first-round questions raised, directly perform at least one round of screening through keyword semantic retrieval to answer the raised questions.
[0010] S2: Check whether other chapters not already given in the content of the answer in step S1 are needed. If so, find the cited chapter and add the cited chapter in a multi-round dialogue manner.
[0011] S3: Divide the sections obtained in steps S1 - S2 into smaller granularity blocks, calculate the similarity between these blocks and the final answer of the large model, similar clauses, and the question input by the user respectively through the vector similarity search model and summarize them. Retain the blocks with higher final similarity and number them one by one. The granularity blocks are paragraphs; as shown in Equation (1);
[0012] S final = α·S answer + β·S templates + γ·S question Equation (1)
[0013] Wherein,
[0014] S final - The finally obtained similarity;
[0015] S answer - The similarity between the small granularity block and the final answer generated by the large model;
[0016] S templates - The similarity between the small granularity block and the similar clauses input by the user;
[0017] S question - The similarity between the small granularity block and the question input by the user;
[0018] α, β, γ - are the weights of these similarities in the final similarity and are constants;
[0019] S4: Calculate the similarity between the unselected but numbered blocks in step S3 and the final answer of the large model using the similarity calculation model. When it is found that there are blocks with higher similarity among the unselected blocks, re - code these remaining blocks and repeat step S3; The similarity calculation model is specifically executed according to the following steps;
[0020] S 4.1 : First, perform data pre - processing on the chunked data of the contract content, and then perform data cleaning to remove noise data and outliers;
[0021] S 4.2 : Perform feature selection on the chunked data of the contract content, extract the features that have an important impact on similarity calculation from the original data; and perform feature standardization calculation: convert features with different dimensions to the same scale for comparison;
[0022] S 4.3 : Represent the selected features in vector form, with each feature corresponding to a dimension in the vector;
[0023] S 4.4: Select a similarity calculation method according to the specific application scenario and data characteristics, calculate the similarity, and apply the selected similarity calculation formula to calculate the similarity between two or more objects, such as formula (2);
[0024]
[0025] Among them, i·j represents the dot product of vectors i and j, and ││i││ and ││j││ represent the moduli of vectors i and j, respectively.
[0026] S5: The first round of questions raised in step S1 are accurately positioned again. Specifically, multiple rounds of dialogues are constructed by similarity calculation and renumbering of block contents. The final positioning is obtained by accurate similarity calculation based on the content of the dialogue, and the answer to the final positioning is output as information.
[0027] Furthermore, the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a main controller, the method described in any one of the above is implemented.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. Improve the accuracy and efficiency of information retrieval.
[0030] 2. By combining information from multiple chapters, the inaccuracy caused by a single information source is avoided and the accuracy of information retrieval is improved.
[0031] 3. Enhance the completeness of problem answers.
[0032] 4. The system can effectively quote and combine information from multiple chapters to provide more complete and accurate answers and reduce the risk of missing information.
[0033] 5. Multiple modules work together to ensure that important information is not missed, thus improving the reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It is understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 is a flow chart of the method of the present invention;
[0036] Figure 2 It is a multi-round dialogue reference chapter diagram of the present invention;
[0037] Figure 3 is the first-round dialogue call diagram of the present invention;
[0038] Figure 4 is the flowchart for finding the positioning block of the present invention. Specific Embodiments
[0039] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but is merely for the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0040] Please refer to Figures 1-4 , the present invention provides a method for accurately positioning clauses in contract questions and answers, which is specifically executed according to the following steps:
[0041] Specifically, it is executed according to the following steps:
[0042] S1: Divide the content of the contract into multiple chapters according to the title. For the first-round questions raised, directly perform at least one round of screening through keyword semantic retrieval and answer the raised questions;
[0043] S2: Check whether it is necessary to refer to other chapters that are not in the chapters already given in the content of the answer in step S1. If so, find the cited chapter and add the cited chapter in a multi-round dialogue manner;
[0044] S3: Divide the chapters obtained in steps S1 - S2 into smaller granularity blocks, and calculate the similarity between these blocks and the final answer of the large model, similar clauses, and the questions input by the user respectively through a vector similarity search model and summarize them. Retain the blocks with higher final similarity and number them one by one. The granularity blocks are paragraphs; as shown in formula (1);
[0045] S final = α·S answer + β·S templates + γ·S question Formula (1)
[0046] Wherein,
[0047] Sfinal - The similarity obtained finally;
[0048] S answer - The similarity between the small-grained blocks and the final answer generated by the large model;
[0049] S templates - The similarity between the small-grained blocks and the similar terms of the user input;
[0050] S question - The similarity between the small-grained blocks and the question of the user input;
[0051] α, β, γ - are the weights of these similarities in the final similarity and are constants;
[0052] S4: Calculate the similarity between the unselected numbered blocks in step S3 and the final answer of the large model using the similarity calculation model. When it is found that there are blocks with relatively high similarities among the unselected blocks, re-number these remaining blocks and repeat step S3; The similarity calculation model is specifically executed according to the following steps;
[0053] S 4.1 : First, perform data preprocessing on the chunked data of the contract content, and then perform data cleaning to remove noise data and outliers;
[0054] S 4.2 : Perform feature selection on the chunked data of the contract content, extract the features that have an important impact on similarity calculation from the original data; and perform calculation feature standardization: convert features with different dimensions to the same scale for comparison;
[0055] S 4.3 : Represent the selected features in vector form, with each feature corresponding to a dimension in the vector;
[0056] S 4.4 : Select a similarity calculation method according to the specific application scenario and data characteristics, calculate the similarity, and apply the selected similarity calculation formula to calculate the similarity between two or more objects, as shown in Equation (2);
[0057]
[0058] Among them, i·j represents the dot product of vectors i and j, and ││i││ and ││j││ represent the norms of vectors i and j respectively.
[0059] S5: Precise positioning is performed again on the first-round questions raised in step S1. Specifically, through similarity calculation, re-numbering the content of the blocks, constructing multi-round dialogues, and performing precise similarity calculation based on the content of the dialogues to obtain the final positioning, and output the answer of the final positioning as information.
[0060] In this embodiment, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a main controller, the method described in any one of the above is implemented.
[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for accurately positioning clauses in contract Q&A, characterized in that: Specifically, it is executed according to the following steps: S1: Divide the content of the contract into multiple chapters according to the title. For the first round of questions raised, directly perform at least one round of screening by combining keyword semantic retrieval to answer the raised questions. S2: Check whether other chapters not in the already given chapters need to be referred to in the content of the answer in step S1. If so, find the referred chapter and add the referred chapter in a multi-round dialogue manner. S3: Divide the chapters obtained in steps S1 - S2 into smaller granularity blocks, and calculate the similarity between these blocks and the final answer of the large model, similar clauses, and the questions input by the user respectively through a vector similarity search model and summarize them. Retain the blocks with higher final similarity and number them one by one, as shown in formula (1). S final = α·S answer + β·S templates + γ·S question Equation (1) Wherein, S final - The finally obtained similarity; S answer - The similarity between the small-grained chunks and the final answer generated by the large model; S templates - Similarity between small-grained blocks and similar terms of user input; S question - Similarity between small - grain blocks and the problems input by users; α, β, γ - are the weights of these similarities in the final similarity and are constants. S4: Calculate the similarity between the numbered but unselected blocks in step S3 and the final answer of the large model using a similarity calculation model. When it is found that there are blocks with higher similarity among the unselected blocks, re-number these remaining blocks and repeat step S3. S5: Precise positioning of the first round of questions raised in step S1 is performed again.
2. The method for accurately positioning clauses in contract Q&A according to claim 1, wherein In step S4, the similarity calculation model is specifically executed according to the following steps: S 4.1 : First, perform data preprocessing on the chunked data of the contract content, and then perform data cleaning to remove noise data and outliers; S 4.2 : Perform feature selection on the chunked data of the contract content, and extract features from the original data that have an important impact on similarity calculation; And perform calculation feature standardization: Convert features with different dimensions to the same scale for comparison. S 4.3 : Represent the selected features in vector form, where each feature corresponds to a dimension in the vector; S 4.4 : Select a similarity calculation method according to the specific application scenario and data characteristics, calculate the similarity, and apply the selected similarity calculation formula to calculate the similarity between two or more objects, as shown in Equation (2); Wherein, i·j represents the dot product of vectors i and j, and ││i││ and ││j││ respectively represent the norms of vectors i and j.
3. The method for accurately positioning clauses in contract Q&A according to claim 1, characterized in that: In step S3, the granularity block is a paragraph.
4. A method for accurately positioning clauses in contract Q&A according to claim 1, characterized in that: In step S5, specifically, through similarity calculation, re-numbering the content of the blocks, constructing a multi-round dialogue, and precisely calculating the similarity according to the content of the dialogue to obtain the final positioning, and output the answer of the final positioning as information.
5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the main controller, it implements the method described in any one of claims 1 - 4 above.
Citation Information
Patent Citations
Question and answer knowledge base using method and system based on large language model
CN118296104A
KR20220073432A