A method and system for extracting scientific literature fields based on large models

By dividing the scientific literature into multiple fragments and integrating the answers using a differential comparison method, the problem of hallucination and noise in the field extraction of scientific literature is solved, and the completeness and accuracy of the answers are achieved.

CN119046444BActive Publication Date: 2025-05-06ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411533982.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-06
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

When extracting the existing large models in the scientific literature fields, it is difficult to effectively process long-form documents, making it difficult to correct the illusion and noise output from the model, and the answers to different fragments are scattered and difficult to integrate.

Method used

The completeness and accuracy of the answers are ensured by dividing the scientific literature into multiple fragments and using a differential comparison method to splice and integrate answers during the iteration process, identifying and correcting hallucinations.

Benefits of technology

The hallucination problem of a single chunk answer is effectively solved, and the effective integration of different chunk answers is realized, improving the accuracy and completeness of the big model in the field extraction of scientific literature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046444B_ABST
    Figure CN119046444B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for extracting scientific document fields based on a large model, comprising: dividing a scientific document into a plurality of fragments; iteratively generating answers to fields to be extracted, and excluding fragments in the chunks that do not contain the fields to be extracted; then, while generating candidate answers, excluding redundant answer information in different chunks by means of difference comparison, excluding erroneous answer information outputted due to hallucinations, and fusing chunks containing correct answers during the iterative process, thereby solving the problem that answers to different fragments are scattered and difficult to integrate; the method and system can maximize the performance of generating a large model for extracting scientific documents under the condition of limited computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and natural science, and specifically relates to a method and system for extracting scientific literature fields based on a large model. Background Art

[0002] After the emergence of large models such as ChatGPT and Tongyi Qianwen, people gradually tried to use large models to extract knowledge from scientific literature. At present, the main methods of using large models to extract scientific literature are: constructing a prompt input into the large model together with the entire scientific literature and the knowledge to be extracted. This method is simple, convenient, easy to implement and use. However, the context length of general open source models often cannot accept the length of the entire document. Some specialized long text models have a large number of parameters (extremely high computing resource consumption and high cost) or poor results (there is a lot of noise in the extracted results, requiring extremely high manual post-processing costs). Models that can handle the extraction of long scientific literature and have certain effects are often paid by token. Due to the long length of scientific literature, the cost of batch extraction is extremely high.

[0003] After slicing the document, each segment is asked to output the results separately, and finally the large model is used to integrate the output results. However, when hallucinations or noise appear in individual segments, or the extracted values ​​require segments at different positions to be inferred, due to the lack of the original text, the model cannot correct them at the integration stage, and generally directly integrates and outputs wrong hallucinations or noise.

[0004] After rag retrieves some fragments, it builds a large prompt input model together with the knowledge to be extracted to get the answer. If the retrieved fragments are small, there is often a problem of incomplete retrieval fragments, which leads to missing or wrong answers from the model. If the number of retrieved fragments is increased, the length of the text input to the model will increase, reducing the output effect of the model. Summary of the invention

[0005] In view of the above, the purpose of the present invention is to provide a scientific literature field extraction method and system based on a large model, which effectively solves the illusion of a single chunk answer by comparing the differences between different fragments, and at the same time realizes effective chunk fusion and splicing, solving the problem that answers from different chunks are scattered and difficult to integrate.

[0006] The present invention is implemented by the following technical solutions. In the first aspect of the present invention, the present invention provides a method for extracting scientific literature fields based on a large model, comprising the following steps:

[0007] (1) Divide scientific literature into several chunks = [chunk 1 , chunk 2 ,……,chunk N];

[0008] (2) Construct a prompt using the fields to be extracted, field explanations, and slice fragments;

[0009] (3) Traverse the several chunks of the segmentation and use step (2) to construct the first i Chunk i , the corresponding prompt input large model LLM gets the output out i , if out i If it is not empty, the traversal ends early;

[0010] (4) After the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. i Assign the content of to CHUNK, where CHUNK represents the chunk combination that is most suitable for extracting fields after iteration to the current chunk; if OUT is empty, jump to step (10); if OUT is not empty, jump to step (5);

[0011] (5) From several fragment chunks i +1 fragment to start traversal, use step (2) to build a chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early;

[0012] (6) After the traversal is completed, if out j is empty, jump to step (10); if out j If it is not empty, jump to step (7);

[0013] (7) CHUNK and chunk j Chunks generated by splicing +j And the prompt constructed by step (2) is input into the LLM model to obtain the output out +j ;

[0014] (8) Take OUT and out j The intersection out o ,out o = OUT ∩ out j , take out m = OUT - out o ,out n =outj -out o ;

[0015] (9) According to out +j ,out j Different division scenarios of OUT values;

[0016] (10) Output OUT as the value of the field to be extracted.

[0017] Furthermore, the step (1) specifically comprises the following sub-steps:

[0018] (1.1) Remove the references from the literature;

[0019] (1.2) Separate the table representation and text representation in the document, and each table represents a separate fragment chunk; the table has a title and annotations;

[0020] (1.3) Treat the part before the abstract of the document as a separate chunk and add "The following includes the title, author, unit, email address, publication year, journal name and other basic information of the document" at the beginning of the chunk; Chinese documents should be expressed in Chinese and English documents should be expressed in English;

[0021] (1.4) Divide the text representation of the document into chapters, including the abstract and appendix, and divide each chapter into segments of length no greater than L, where L is 768 or 1024;

[0022] (1.5) The basic information chunk, text representation chunk, and table chunk are sequentially combined to form the document chunks.

[0023] Specifically, in step (2), the prompt is constructed using the fields to be extracted, field explanations, and slice fragment chunks, including: the prompt requires the LLM to output in a specified format. If the extracted fields do not have hierarchical information, "extract the locations studied in this article" is returned in the form of a list; if the extracted fields have hierarchical information, "extract the locations studied in this article and the corresponding longitude and latitude" is returned in the format of json. When the fields to be extracted do not exist in the chunk, the return value is empty.

[0024] Specifically, the large model LLM in step (3) is qwen-70b-instruct.

[0025] Furthermore, the different scenarios in step (9) are specifically:

[0026] (a) If out +j ∩ outm != [] and out +j ∩ out n == [], will j The value of i , jump to step (5);

[0027] (b) If out +j ∩ out m == [] and out +j ∩ out n != [], out j , chunk j Assign the value of to OUT,CHUNK, j The value of i , jump to step (5);

[0028] (c) If out +j ∩ out m != [] and out +j ∩ out n != [], out +j , chunk +j Assign the value of to OUT, CHUNK, j The value of i , jump to step (5);

[0029] (d) If out +j ∩ out o != [], j The value of i , jump to step (5);

[0030] (e) If out +j ∩ out o == [] and Lev(OUT, out +j )>Lev(out j , out +j ),Will j The value of i , jump to step (5);

[0031] (f) If out +j ∩ out o == [] and Lev(OUT, out +j ) <Lev(out j , out +j ), out j ,chunk j Assign the value of to OUT, CHUNK, j The value ofi , jump to step (5).

[0032] Specifically, in step (9) scenario (e), out +j ∩ out o == [] and Lev(OUT, out +j )>Lev(out j , out +j ), where Lev represents Levenstein, and the specific process is:

[0033] (I) Set OUT and out +j Convert to a string, that is, a plain text string format;

[0034] (II) Calculate OUT and out +j Turn between Lewinstein than.

[0035] Specifically, in step (II), OUT and out are calculated. +j The Lewenstein ratio between is calculated as follows: ; sum For OUT and out +j The sum of the string lengths, For OUT and out +j The edit distance of the string, that is, OUT is converted to out +j The minimum number of operations required, where operations include insertion, deletion, and replacement.

[0036] The second aspect of the present invention is a scientific literature field extraction system based on a large model, comprising the following units:

[0037] Text slicing unit: divide scientific literature into several fragments chunks=[chunk 1 , chunk 2 ,……,chunkt N ];

[0038] Iterative answer generation unit: Use the fields to be extracted, field explanations, and slice fragments to build prompts; traverse the sliced ​​chunks, and use the first i Chunk i , the corresponding prompt input large model LLM gets the output out i , if out i If it is not empty, the traversal ends early; after the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. iAssign the content of to CHUNK, which represents the chunk combination that is most suitable for extracting fields after iteration to the current chunk; if OUT is empty, jump to the result output unit; if OUT is not empty, select the first chunk from several fragment chunks. i +1 fragment to start traversal, use the build chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early; after the traversal ends, if out j Is empty, jump to the step result output unit; if out j If it is not empty, go to the candidate answer fusion unit;

[0039] Candidate answer fusion unit: CHUNK and chunk j Chunks generated by splicing +j And input the LLM model through the constructed prompt to get the output out +j ; Take OUT and out j The intersection out o ,out o =OUT ∩ out j , take out m = OUT - out o ,out n = out j -out o ; According to out +j ,out j Different division scenarios of OUT values;

[0040] The result output unit: outputs OUT as the value of the field to be extracted.

[0041] The third aspect of the present invention: a terminal device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method when executing the computer program.

[0042] The fourth aspect of the present invention: a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the steps of the method when executed by a processor.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. The solution of the present invention divides the complete scientific literature into multiple fragments, and compares the answers of the current fragment, the next fragment, and the spliced ​​answers of the two fragments in continuous iteration. When a certain chunk has an illusion, it can effectively and timely determine the wrong answer and correct it in time. (There are two main situations of illusion when the large model is extracting: (1) There is no answer that the large model wants to extract in the corresponding fragment, and the large model answers based on its own knowledge. (2) The content that appears in the corresponding fragment is relatively close to the answer that needs to be extracted, but it is not the required answer, and the large model extracts it).

[0045] 2. In the process of iterative difference comparison, the solution of the present invention effectively splices different chunks containing correct answers in the order of document content, which solves the problem of scattered answers to different fragments that are difficult to integrate, and provides a strong guarantee for the integrity of the answers of the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0047] Figure 1 is a flow chart of a method for extracting scientific literature fields based on a large model provided in an embodiment;

[0048] Figure 2 It is a flowchart of iterative generation of answers provided by the embodiment;

[0049] Figure 3 is a candidate answer fusion flow chart provided in an embodiment;

[0050] Figure 4 It is a flow chart of the scientific literature field extraction system based on the big model of the present invention;

[0051] Figure 5 It is a schematic diagram of the electronic device of the present invention. DETAILED DESCRIPTION

[0052] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0053] The inventive concept of the present invention is as follows: in the face of the task of extracting scientific literature fields, the present invention divides the complete scientific literature into multiple fragments, and compares the difference between the current fragment, the next fragment, and the answer spliced ​​by the two fragments in continuous iteration. When a certain chunk has an illusion, it can effectively and timely determine the wrong answer and correct it in time. In the process of iterative difference comparison, different chunks containing correct answers are effectively spliced ​​according to the order of the literature content, solving the problem that the answers to different fragments are scattered and difficult to integrate.

[0054] Based on the above invention concept, Figure 1 As shown, the flowchart of the method for extracting scientific literature fields based on a large model provided in the embodiment includes the following steps:

[0055] Step 1: Divide the scientific document into chunks (chunks=[chunk 1 , chunk 2 , ……,chunk N ]), the specific segmentation process includes:

[0056] (1.1) Remove the references from the literature;

[0057] (1.2) Separate the table representation (table with title and notes) from the text representation in the document, and each table represents a separate chunk;

[0058] (1.3) Count the part before the abstract in the document as a separate chunk, and add "The following contains the basic information of the document, such as the title, author, unit, email address, publication year, journal name, etc." at the beginning of the chunk (expressed in Chinese for Chinese documents and in English for English documents);

[0059] (1.4) Divide the text representation of the document into chapters (including the abstract and appendix), and divide each chapter into segments of length not greater than L (L is 768 or 1024).

[0060] (1.5) The basic information chunk, text representation chunk, and table chunk are sequentially combined into the chunks of the document;

[0061] Step 2: Specific, such as Figure 2 The embodiment provides an iterative answer generation flow chart as shown in the figure, including:

[0062] (2.1) Use the fields to be extracted (Field), field explanations (description) and slice fragments (chunk) to construct the large model prompt text (prompt); prompt requires LLM to output in the specified format (if the extracted field has no hierarchical information (extract the location studied in this article), it is returned in the form of a list; if the extracted field has hierarchical information (extract the location studied in this article and the corresponding longitude and latitude), it is returned in json format). When the field to be extracted (Field) does not exist in the chunk, it is returned as empty;

[0063] (2.2): Traverse the chunks and build the chunk using step 2 i The corresponding prompt input LLM gets the output out i (chunk i Indicates the first i chunks), if out i If it is not empty, the traversal ends early;

[0064] Step 3: After the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. i Assign the content of OUT to CHUNK (CHUNK represents the best chunk combination for extracting fields after iteration to the current chunk (CHUNK is the best combination of all chunks before the current i-th chunk, named CHUNK)). If OUT is empty, jump to step 7; if OUT is not empty, jump to step 4;

[0065] Step 4: Specific, such as Figure 2 The answer provided in the embodiment is iteratively generated as shown in the flowchart, including: i +1 chunk to start traversal, use step 2.1 to build chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early;

[0066] Step 5: After the traversal is completed, if out j Is empty, jump to step 7; if out j If it is not empty, jump to step 6;

[0067] Step 6: Specific, such as Figure 3The candidate answer fusion flow chart provided in the embodiment includes: (6.1) combining CHUNK and chunk j Splice to generate chunks +j And construct the prompt input LLM through step 2, and get the output out +j ; The prompt is composed of chunk +j Constructed, input only chunk +j The corresponding prompt is OK;

[0068] (6.2) Take OUT and out j The intersection out o (out o = OUT ∩ out j ), take out m =OUT - out o ,out n =out j -out o .

[0069] (6.3) According to out +j ,out j Different division scenarios of OUT values;

[0070] Scenario 1: If out +j ∩ out m != [] and out +j ∩ out n == [], will j The value of i , jump to step 4;

[0071] Scenario 2: If out +j ∩ out m == [] and out +j ∩ out n != [], out j , chunk j Assign the value of to OUT, CHUNK, j The value of i , jump to step 4;

[0072] Scenario 3: If out +j ∩ out m != [] and out +j ∩ out n != [], out +j , chunk +j Assign the value of to OUT, CHUNK,j The value of i , jump to step 4;

[0073] Scenario 4: If out +j ∩ out o != [], j The value of i , jump to step 4;

[0074] Scenario 5: If out +j ∩ out o == [] and Lev(OUT, out +j )>Lev(out j , out +j ),Will j The value of i , jump to step 4; Lev in the formula represents Levenshtein ratio, Lev(OUT, out +j ) as an example, the specific process is:

[0075] (a) Connect OUT and out +j Convert to string (plain text string) format

[0076] (b) Calculate OUT and out +j Lewenstein ratio between turns: ; sum For OUT and out +j The sum of the string lengths, For OUT and out +j The edit distance of the string, that is, OUT is converted to out +j The minimum number of operations required, where operations include insertion, deletion, and replacement.

[0077] Scenario 6: If out +j ∩ out o == [] and Lev(OUT, out +j ) <Lev(out j , out +j ), out j ,chunk j Assign the value of to OUT, CHUNK, j The value of i , jump to step 4;

[0078] Step 7: Output OUT as the value of the field to be extracted.

[0079] In the present invention, the large model LLM used is qwen-70b-instruct (qwen represents the Qianwen series, 70b represents the model parameter size, and instruct represents that the model has been trained with supervised alignment).

[0080] Based on the same inventive concept, the embodiment also provides a scientific literature field extraction system based on a large model, such as Figure 4 As shown, it includes a text slicing unit, an answer iteration generation unit, a candidate answer fusion unit, and a result output unit: specifically:

[0081] Text slicing unit: divide scientific literature into several fragments chunks=[chunk 1 , chunk 2 ,……,chunk N ];

[0082] Iterative answer generation unit: Use the fields to be extracted, field explanations, and slice fragments to build prompts; traverse the sliced ​​chunks, and use the first i Chunk i , the corresponding prompt input large model LLM gets the output out i , if out i If it is not empty, the traversal ends early; after the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. i Assign the content of to CHUNK, which represents the chunk combination that is most suitable for extracting fields after iteration to the current chunk; if OUT is empty, jump to the result output unit; if OUT is not empty, select the first chunk from several fragment chunks. i +1 fragment to start traversal, use the build chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early; after the traversal ends, if out j Is empty, jump to the step result output unit; if out j If it is not empty, it will go to the candidate answer fusion unit; it is used to generate answers to the fields to be extracted urgently and exclude the fragments in the chunks that do not contain the fields to be extracted urgently;

[0083] Candidate answer fusion unit: CHUNK and chunk j Chunks generated by splicing+j And input the LLM model through the constructed prompt to get the output out +j ; Take OUT and out j The intersection out o ,out o =OUT ∩ out j , take out m = OUT - out o ,out n = out j -out o ; According to out +j ,out j Different division scenarios of OUT values; that is, for the scattered answers in different fragments and those that are difficult to integrate, by comparing the differences, redundant answer information in different chunks is eliminated, wrong answer information output due to hallucinations is eliminated, and chunks containing correct answers are merged in the iterative process;

[0084] The result output unit: outputs OUT as the value of the field to be extracted.

[0085] like Figure 5 As shown, the present invention further provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method when executing the computer program.

[0086] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.

[0087] It should be noted that the scientific literature field extraction system based on the big model provided in the above embodiment should be illustrated by the division of the above functional modules when performing field extraction. The above functional distribution can be completed by different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the scientific literature field extraction system based on the big model provided in the above embodiment and the scientific literature field extraction method embodiment based on the big model belong to the same concept. The specific implementation process is detailed in the scientific literature field extraction method embodiment based on the big model, which will not be repeated here.

[0088] Based on the same inventive concept, the embodiment also provides a computing device, which, at the hardware level, includes not only a processor and a memory, but also hardware required for other businesses such as an internal bus, a network interface, and a memory. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned method for generating samples of financial event extraction. Of course, in addition to the software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0089] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for extracting scientific literature fields based on a large model, characterized in that: The following steps are involved: (1) Divide scientific literature into several chunks = [chunk 1 , chunk 2 ,……,chunk N ]; Specifically, the complete scientific literature is divided into multiple segments, and the differences between the current segment, the next segment, and the answer spliced ​​by the two segments are compared in continuous iterations; (2) Construct a prompt using the fields to be extracted, field explanations, and slice fragments; (3) Traverse the several chunks of the segmentation and use step (2) to construct the first i Chunk i , the corresponding prompt input large model LLM gets the output out i , if out i If it is not empty, the traversal ends early; (4) After the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. i Assign the content of to CHUNK, where CHUNK represents the chunk combination that is most suitable for extracting fields after iteration to the current chunk; if OUT is empty, jump to step (10); if OUT is not empty, jump to step (5); (5) From several fragment chunks i +1 fragment to start traversal, use step (2) to build a chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early; (6) After the traversal is completed, if out j is empty, jump to step (10); if out j If it is not empty, jump to step (7); (7) CHUNK and chunk j Chunks generated by splicing +j And the prompt constructed by step (2) is input into the LLM model to obtain the output out +j ; (8) Take OUT and out j The intersection out o ,out o = OUT ∩ out j , take out m = OUT - out o ,out n = out j -out o ; (9) According to out +j ,out j Different division scenarios of OUT values; (10) Output OUT as the value of the field to be extracted.

2. According to claim 1, a method for extracting scientific literature fields based on a large model is characterized in that: The step (1) specifically comprises the following sub-steps: (1.1) Remove the references from the literature; (1.2) Separate the table representation and text representation in the document, and each table represents a separate fragment chunk; the table has a title and annotations; (1.3) Treat the part before the abstract of the document as a separate chunk and add "The following includes the title, author, unit, email address, publication year, journal name and other basic information of the document" at the beginning of the chunk; Chinese documents should be expressed in Chinese and English documents should be expressed in English; (1.4) Divide the text representation of the document into chapters, including the abstract and appendix, and divide each chapter into segments of length no greater than L, where L is 768 or 1024; (1.5) The basic information chunk, text representation chunk, and table chunk are sequentially combined to form the document chunks.

3. According to claim 1, a method for extracting scientific literature fields based on a large model is characterized in that: In the step (2), the prompt is constructed using the fields to be extracted, the field explanations, and the slice fragment chunk, including: the prompt requires the LLM to output in a specified format, if the extracted field has no hierarchical information, then "extract the location studied in this article" is returned in the form of a list; if the extracted field has hierarchical information, then "extract the location studied in this article and the corresponding longitude and latitude" is returned in the format of json, and when the field to be extracted does not exist in the chunk, the return is empty.

4. According to claim 1, a method for extracting scientific literature fields based on a large model is characterized in that: The large model LLM in step (3) is specifically qwen-70b-instruct.

5. According to claim 1, a method for extracting scientific literature fields based on a large model is characterized in that: The different scenarios in step (9) are specifically: (a) If out +j ∩ out m != [] and out +j ∩ out n == [], will j The value of i , jump to step (5); (b) If out +j ∩ out m == [] and out +j ∩ out n != [], out j , chunk j Assign the value of to OUT,CHUNK, j The value of i , jump to step (5); (c) If out +j ∩ out m != [] and out +j ∩ out n != [], out +j , chunk +j Assign the value of to OUT,CHUNK, j The value of i , jump to step (5); (d) If out +j ∩ out o != [], j The value of i , jump to step (5); (e) If out +j ∩ out o == [] and Lev(OUT, out +j ) > Lev(out j , out +j ),Will j The value of i , jump to step (5); (f) If out +j ∩ out o == [] and Lev(OUT, out +j ) < Lev(out j , out +j ), out j , chunk j Assign the value of to OUT, CHUNK, j The value of i , jump to step (5).

6. A method for extracting scientific literature fields based on a large model according to claim 5, characterized in that: Step (9) scenario (e) out +j ∩ out o == [] and Lev(OUT, out +j ) > Lev(out j , out +j ), where Lev represents Levenstein, and the specific process is: (I) Set OUT and out +j Convert to a string, that is, a plain text string format; (II) Calculate OUT and out +j Turn between Lewinstein than.

7. A method for extracting scientific literature fields based on a large model according to claim 6, characterized in that: In step (II), OUT and out are calculated. +j The Lewenstein ratio between is calculated as follows: ; sum For OUT and out +j The sum of the string lengths, For OUT and out +j The edit distance of the string, that is, OUT is converted to out +j The minimum number of operations required, where operations include insertion, deletion, and replacement.

8. A scientific literature field extraction system based on a large model, characterized in that: The following units are included: Text slicing unit: divide scientific literature into several fragments chunks=[chunk 1 , chunk 2 ,……,chunkt N ]; Specifically, the complete scientific literature is divided into multiple segments, and the differences between the current segment, the next segment, and the answer spliced ​​by the two segments are compared in continuous iterations; Iterative answer generation unit: Use the fields to be extracted, field explanations, and slice fragments to build prompts; traverse the sliced ​​chunks, and use the first i Chunk i , the corresponding prompt input large model LLM gets the output out i , if out i If it is not empty, the traversal ends early; after the traversal is completed, out i The value of is assigned to OUT, where OUT represents the optimal answer to the current chunk. i Assign the content of to CHUNK, where CHUNK represents the chunk combination that is most suitable for extracting fields after iteration to the current chunk; if OUT is empty, jump to the result output unit; If OUT is not empty, then select the first i +1 fragment to start traversal, use the build chunk j The corresponding prompt input LLM output is out j , the chunk j Indicates the jth chunk, j>=i+1; if out j If it is not empty, the traversal ends early; after the traversal ends, if out j Is empty, jump to the step result output unit; if out j If it is not empty, go to the candidate answer fusion unit; Candidate answer fusion unit: CHUNK and chunk j Chunks generated by splicing +j And input the LLM model through the constructed prompt to get the output out +j ; Take OUT and out j The intersection out o ,out o =OUT ∩ out j , take out m = OUT -out o ,out n = out j -out o ; According to out +j ,out j Different division scenarios of OUT values; The result output unit: outputs OUT as the value of the field to be extracted.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Question and answer method combining paragraph search and machine reading understanding

    CN115292469A

  • Scientific literature information extraction method and device, computer equipment and storage medium

    CN117668173A