Evaluation review content generation method based on retrieval enhancement
Through the retrieval-enhanced evaluation review content generation method, the text embedding model and semantic similarity algorithm are used to filter relevant document fragments, and combined with the large language model to generate content, the problem of insufficient accuracy and traceability of the generated results in the existing technology is solved, and more efficient, accurate and verifiable evaluation review content generation is achieved.
Patent Information
- Application Number
- CN202510256198.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing method of directly generating the content to be evaluated using a large language model has the problems of low accuracy and low traceability of the generation results, which is difficult to meet the accuracy and verifiable needs of the evaluation and review task.
The evaluation and review content generation method based on search enhancement is adopted. By formatting and diced by the document to be evaluated, the semantic vector library is obtained by inputting it into the text embedding model, and the relevant document fragments are filtered out using the semantic similarity algorithm, and inputting them into the large language model to generate the evaluation and review content.
It improves the accuracy and relevance of the generated results, enhances the traceability of the generation process, reduces human subjectivity and cost, and improves efficiency and adaptability.
Smart Images

Figure CN120179783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of evaluation and review, and more specifically, to a method for generating evaluation and review content based on retrieval enhancement. Background Art
[0002] Currently, the generation of content to be evaluated and reviewed (i.e., obtaining the key content to be evaluated and reviewed in the document to be evaluated and reviewed) is a key link in the asset evaluation and review industry, aiming to ensure the quality and accuracy of the evaluation report. Traditional generation methods rely on manual methods, which have problems such as high costs, low generation efficiency, and low generation accuracy; with the development of technology, large language models have been introduced into this field, which can answer questions based on prompts and improve the generation efficiency.
[0003] However, there are significant problems with the existing method of directly generating content to be evaluated and reviewed using large language models. On the one hand, the accuracy of the generated results is relatively low. Since large language models often cannot accurately identify and extract key information when processing complex professional documents, the generated content may deviate from actual requirements; on the other hand, the traceability of the generated results is relatively low, and it is difficult for users to trace the processing process of the model, which is a major challenge for evaluation and review tasks that require high precision and verifiability.
[0004] Therefore, how to provide a method for generating evaluation and review content that can improve the efficiency, accuracy, and traceability of generating evaluation and review content is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for generating evaluation and review content based on retrieval enhancement.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] In the first aspect, a method for generating evaluation and review content based on retrieval enhancement is provided, including the following steps:
[0008] S1: Convert the format of the document to be evaluated and reviewed, perform chunking processing, and then input it into a text embedding model to obtain a vector library to be retrieved;
[0009] Input the question to be retrieved into the text embedding model to obtain a semantic vector to be retrieved;
[0010] S2: Use a semantic similarity algorithm and the vector library to be retrieved to screen out several document fragments with the highest semantic similarity to the semantic vector to be retrieved;
[0011] S3: Input several document fragments with the highest semantic similarity and the semantic vector to be retrieved into a large language model to obtain the content to be evaluated and reviewed of the document to be evaluated and reviewed.
[0012] Preferably, S1 specifically includes the following steps:
[0013] Convert the document to be evaluated and reviewed into a format;
[0014] Perform pagination and slicing on the document to be evaluated and reviewed after format conversion to obtain a number of paginated and sliced fragments; further perform context slicing on each paginated and sliced fragment to obtain a number of context sliced fragments;
[0015] Input the number of paginated and sliced fragments into the text embedding model to obtain a number of coarse-grained semantic vectors; input the number of context sliced fragments into the text embedding model to obtain a number of fine-grained semantic vectors; wherein, the number of coarse-grained semantic vectors and the number of fine-grained semantic vectors constitute the vector library to be retrieved;
[0016] Input the coarse-grained problem to be retrieved into the text embedding model to obtain a coarse-grained semantic vector to be retrieved; input the fine-grained problem to be retrieved into the text embedding model to obtain a fine-grained semantic vector to be retrieved; wherein, the coarse-grained problem to be retrieved and the fine-grained problem to be retrieved constitute the problem to be retrieved in S1.
[0017] Preferably, S2 specifically includes the following steps:
[0018] Select M coarse-grained semantic vectors with the highest semantic similarity to the coarse-grained semantic vector to be retrieved from the number of coarse-grained semantic vectors;
[0019] Select N fine-grained semantic vectors with the highest semantic similarity to the fine-grained semantic vector to be retrieved from the fine-grained semantic vectors corresponding to the M coarse-grained semantic vectors; wherein, the document content corresponding to the N fine-grained semantic vectors is the number of document fragments with the highest semantic similarity in S2; M and N are both positive integers.
[0020] Preferably, S3 specifically includes the following steps:
[0021] Input the document content corresponding to the N fine-grained semantic vectors, the coarse-grained problem to be retrieved, and the fine-grained problem to be retrieved into the large language model to obtain the content to be evaluated and reviewed of the document to be evaluated and reviewed.
[0022] Preferably, in S1, the document to be evaluated and reviewed is converted into a standardized Markdown format.
[0023] Preferably, the coarse-grained problem to be retrieved and the fine-grained problem to be retrieved are obtained based on the same retrieval entry; the coarse-grained problem to be retrieved and the fine-grained problem to be retrieved are retrieval problems with the same goal.
[0024] Preferably, the text embedding model is BERT; the large language model is GPT or llama.
[0025] In a second aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for generating evaluation and review content based on retrieval enhancement as described above is implemented.
[0026] In a third aspect, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for generating evaluation and review content based on retrieval enhancement as described above is implemented.
[0027] In a fourth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method for generating evaluation and review content based on retrieval enhancement as described above is implemented.
[0028] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for generating evaluation and review content based on retrieval enhancement. The following beneficial technical effects can be obtained:
[0029] 1) Efficiency improvement and cost reduction: By adopting a large language model, semantic embedding, and retrieval enhancement generation technology, the present invention can automatically output the evaluation and review content of the document to be evaluated and reviewed, significantly shortening the generation time. Especially when processing a large number of documents or peak tasks, parallel processing is supported, greatly reducing the manpower input and cost.
[0030] 2) Avoidance of human subjectivity: The evaluation and review content generation method of the present invention is implemented based on a large language model, semantic embedding, and retrieval enhancement technology, establishing a standardized process for the evaluation and review industry, unifying the quality benchmark, reducing the interference of human errors, and avoiding the subjectivity of human operations.
[0031] 3) Optimization of retrieval efficiency: The multi-scale text chunking strategy and progressive retrieval strategy provided by the present invention significantly improve the retrieval efficiency, and can quickly match the most relevant document fragments to the query item.
[0032] 4) Enhancement of accuracy and relevance: The present invention inputs the most relevant document fragments matched into the large language model, significantly improving the accuracy and relevance of the generation results.
[0033] 5) Improvement of traceability: The "retrieval-generation" combination method provided by the present invention avoids the black box problem, making the generation process more transparent and more traceable.
[0034] 6) Improved adaptability: The present invention can be applied to different review field scenarios, improving the applicability of large language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0036] Figure 1 It is a flowchart of a method for generating evaluation and review content based on retrieval enhancement provided by the present invention;
[0037] Figure 2 It is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0039] On the one hand, as Figure 1 shown, an embodiment of the present invention discloses a method for generating evaluation and review content based on retrieval enhancement, including the following steps:
[0040] Step 1: Convert the format of the document to be evaluated and reviewed;
[0041] In one embodiment, the present invention converts the document to be evaluated and reviewed into a standardized Markdown format.
[0042] In one embodiment, the document to be evaluated and reviewed according to the present invention includes multiple types of documents such as evaluation reports, evaluation descriptions, and evaluation result tables;
[0043] In one embodiment, the document to be evaluated and reviewed includes documents to be evaluated and reviewed in fields such as accounting and auditing.
[0044] Step 2: Perform paging and chunking processing on the document to be evaluated and reviewed after format conversion to obtain a number of paging and chunking fragments; further perform context chunking processing on each paging and chunking fragment to obtain a number of context chunking fragments;
[0045] Step 3: input the plurality of paged segments into the text embedding model to obtain a plurality of coarse-grained semantic vectors; input the plurality of context segments into the text embedding model to obtain a plurality of fine-grained semantic vectors; wherein the plurality of coarse-grained semantic vectors and the plurality of fine-grained semantic vectors constitute the vector library to be retrieved;
[0046] In one embodiment, the text embedding model is BERT or T5;
[0047] It can be understood that: the present invention designs a multi-scale text segmentation strategy, which segments the documents to be evaluated and reviewed according to different granularities, forming two fragment formats: page segmentation and context segmentation. Page segmentation is segmented according to the natural paging rules of the document, which is suitable for a larger range of content queries; context segmentation is based on semantic analysis, which subdivides the document into smaller context-related fragments, which is suitable for more accurate semantic matching. The segmented document fragments are all generated with their corresponding semantic vectors through the text embedding model, and are stored in a vector library in layers to obtain a vector library to be retrieved;
[0048] Assume that the document to be evaluated and reviewed is a 4-page evaluation report A. The paging and slicing process divides the evaluation report A into 4 paging and slicing segments A1, A2, A3, and A4 according to the number of pages. Next, the paging and slicing segment A1 is subjected to context slicing to obtain B1 context slicing segments. The paging and slicing segment A2 is subjected to context slicing to obtain B2 context slicing segments. The paging and slicing segment A3 is subjected to context slicing to obtain B3 context slicing segments. The paging and slicing segment A4 is subjected to context slicing to obtain B4 context slicing segments. Contextual slicing fragments; storing paging slicing fragment A1, paging slicing fragment A2, paging slicing fragment A3, paging slicing fragment A4, B1 contextual slicing fragments, B2 contextual slicing fragments, B3 contextual slicing fragments, and B4 contextual slicing fragments in layers (specifically: paging slicing fragment A1, paging slicing fragment A2, paging slicing fragment A3, and paging slicing fragment A4 are stored in one layer; B1 contextual slicing fragments, B2 contextual slicing fragments, B3 contextual slicing fragments, and B4 contextual slicing fragments are stored in another layer), and obtaining a vector library to be retrieved. The hierarchical vector library to be retrieved constructed by the present invention includes a large range of paging semantic vectors and a small range of contextual semantic vectors, providing flexible retrieval support for different retrieval tasks.
[0049] Step 4: Input the coarse-grained question to be retrieved into the text embedding model to obtain a coarse-grained semantic vector to be retrieved; input the fine-grained question to be retrieved into the text embedding model to obtain a fine-grained semantic vector to be retrieved; wherein the coarse-grained question to be retrieved and the fine-grained question to be retrieved constitute the question to be retrieved in S1.
[0050] It is understandable that the coarse-grained retrieval problem to be retrieved and the fine-grained retrieval problem to be retrieved are obtained based on the same entry to be retrieved; the coarse-grained retrieval problem to be retrieved and the fine-grained retrieval problem to be retrieved are retrieval problems with the same target.
[0051] It is understandable that the entry to be retrieved includes: risk-free interest rate Rf, risk premium, etc.;
[0052] It is understandable that the present invention converts a number of entries to be retrieved into retrieval problems in an artificial manner based on the review target; specifically: the present invention first converts a number of entries to be retrieved into coarse-grained retrieval problems with a larger semantic range, and then converts these entries to be retrieved into fine-grained retrieval problems with a smaller semantic range.
[0053] The coarse-grained retrieval problem to be retrieved is used for subsequent paged chunk retrieval; the fine-grained retrieval problem to be retrieved is used for subsequent context chunk retrieval.
[0054] Step 5: Screen out M coarse-grained semantic vectors with the highest semantic similarity to the coarse-grained semantic vector to be retrieved from the several coarse-grained semantic vectors;
[0055] Step 6: Screen out N fine-grained semantic vectors with the highest semantic similarity to the fine-grained semantic vector to be retrieved from the fine-grained semantic vectors corresponding to the M coarse-grained semantic vectors; wherein, the document content corresponding to the N fine-grained semantic vectors is the several document segments with the highest semantic similarity in S2; both M and N are positive integers.
[0056] Step 7: Input the document content corresponding to the N fine-grained semantic vectors, the coarse-grained retrieval problem, and the fine-grained retrieval problem into a large language model to obtain the content to be evaluated and reviewed for the document to be evaluated and reviewed.
[0057] The progressive retrieval strategy provided by the present invention can balance retrieval efficiency and generation accuracy, ensuring that the finally matched content is accurate and complete.
[0058] It is understandable that the content to be evaluated and reviewed for the document to be evaluated and reviewed finally obtained by the present invention is text that has been adjusted by the large language model and is highly relevant to the N fine-grained semantic vectors and highly corresponding to the coarse-grained retrieval problem and the fine-grained retrieval problem.
[0059] Furthermore, the large language model is GPT or llama.
[0060] On the other hand, the present invention also provides an electronic device, such as Figure 2As shown in the figure, the electronic device may include: a processor 201, a communications interface 202, a memory 203, and a communication bus 204. Among them, the processor 201, the communications interface 202, and the memory 203 communicate with each other through the communication bus 204. The processor 201 can call the logical instructions in the memory 203 to execute the method for generating an evaluation review content based on retrieval enhancement. In addition, when the logical instructions in the above-mentioned memory 203 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0061] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating an evaluation review content based on retrieval enhancement provided by the above-mentioned various methods.
[0062] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for generating an evaluation review content based on retrieval enhancement provided by the above-mentioned various methods.
[0063] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0064] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0065] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for the relevant parts.
[0066] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating evaluation and review content based on retrieval enhancement, characterized in that: The following steps are involved: S1: Convert the format of the document to be evaluated and reviewed, cut it into pieces, and then input it into the text embedding model to obtain a vector library to be retrieved; Inputting the question to be retrieved into the text embedding model to obtain the semantic vector to be retrieved; S2: using a semantic similarity algorithm and the to-be-searched vector library, screening out a number of document fragments having the highest semantic similarity to the to-be-searched semantic vector; S3: Inputting several document fragments with the highest semantic similarity and the semantic vectors to be retrieved into the large language model to obtain the content to be evaluated and reviewed of the document to be evaluated and reviewed.
2. A method for generating evaluation and review content based on retrieval enhancement according to claim 1, characterized in that: S1 specifically includes the following steps: Converting the format of the document to be evaluated and reviewed; The format-converted document to be evaluated and reviewed is processed by paging and slicing to obtain a number of paging and slicing segments; each paging and slicing segment is further processed by context slicing to obtain a number of context slicing segments; Input the plurality of paged segments into the text embedding model to obtain a plurality of coarse-grained semantic vectors; input the plurality of context segments into the text embedding model to obtain a plurality of fine-grained semantic vectors; wherein the plurality of coarse-grained semantic vectors and the plurality of fine-grained semantic vectors constitute the vector library to be retrieved; Input the coarse-grained question to be retrieved into the text embedding model to obtain a coarse-grained semantic vector to be retrieved; input the fine-grained question to be retrieved into the text embedding model to obtain a fine-grained semantic vector to be retrieved; wherein the coarse-grained question to be retrieved and the fine-grained question to be retrieved constitute the question to be retrieved in S1.
3. A method for generating evaluation and review content based on retrieval enhancement according to claim 2, characterized in that: S2 specifically includes the following steps: Selecting M coarse-grained semantic vectors having the highest semantic similarity with the coarse-grained semantic vector to be searched from the plurality of coarse-grained semantic vectors; N fine-grained semantic vectors with the highest semantic similarity to the fine-grained semantic vector to be retrieved are selected from the fine-grained semantic vectors corresponding to the M coarse-grained semantic vectors; wherein the document contents corresponding to the N fine-grained semantic vectors are several document fragments with the highest semantic similarity in S2; M and N are both positive integers.
4. The method for generating evaluation and review content based on retrieval enhancement according to claim 3, characterized in that: S3 specifically includes the following steps: The document contents corresponding to the N fine-grained semantic vectors, the coarse-grained questions to be searched, and the fine-grained questions to be searched are input into a large language model to obtain the contents to be evaluated and reviewed of the document to be evaluated and reviewed.
5. The method for generating evaluation and review content based on retrieval enhancement according to claim 1, characterized in that: In S1, the document to be evaluated and reviewed is converted into a standardized Markdown format.
6. The method for generating evaluation and review content based on retrieval enhancement according to claim 2, characterized in that: The coarse-grained question to be retrieved and the fine-grained question to be retrieved are obtained based on the same item to be retrieved; the coarse-grained question to be retrieved and the fine-grained question to be retrieved are questions to be retrieved with the same target.
7. The method for generating evaluation and review content based on retrieval enhancement according to claim 2, characterized in that: The text embedding model is BERT; the large language model is GPT or llama.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for generating evaluation and review content based on retrieval enhancement according to any one of claims 1 to 7 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating evaluation and review content based on retrieval enhancement according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating evaluation and review content based on retrieval enhancement according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Question and answer processing method and device
CN117972048A
Advanced retrieval enhanced blocking and vectoring method based on large model
CN118312579A
Text sorting method and device of retrieval system and electronic equipment
CN118885570A
Chain-based retrieval enhancement generation method and device and readable storage medium
CN118939776A
Multi-knowledge granularity text retrieval method and device for RAG
CN119415623A