LLM report interpretation method and system based on synthetic data

By applying RAG technology and synthetic data on large language models, combined with the Q&A caching mechanism, the problems of accuracy and coverage in the interpretation of genetic test reports and medical test reports are solved, and more efficient and accurate report interpretation is achieved.

CN120144732AActive Publication Date: 2025-06-13HARBIN INST OF TECH

Patent Information

Application Number
CN202510210795.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and incomplete information coverage when dealing with complex genetic test reports and medical test reports, especially when facing new fields or specific reports, the performance of large language models may be insufficient.

Method used

Using an intelligent report interpretation method based on RAG technology, combining synthetic data and question-answer cache mechanism, a professional question-and-answer pair is generated, and the system's knowledge base is enriched, and the answered question-and-answer pairs are stored through the cache mechanism to avoid repeated calculations and improve query response speed.

Benefits of technology

It significantly improves the accuracy and efficiency of report interpretation, can understand the professional content in the report more deeply, adapt to the needs of different fields, and continuously optimize the quality of answers through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144732A_ABST
    Figure CN120144732A_ABST
Patent Text Reader

Abstract

The invention discloses an LLM report interpretation method and system based on synthetic data, and relates to the technical field of natural language processing and large language model application, in particular to the LLM report interpretation method and system based on the synthetic data. The objective of the invention is to solve the problems of insufficient accuracy, universality and user adaptability in report interpretation in the prior art. Firstly, specialized question and answer pairs are generated through synthetic data, and a knowledge base of a system is enriched; secondly, answered question and answer pairs are stored through a cache mechanism, repeated calculation is avoided, and the query response speed is increased; and finally, by utilizing a context-enhanced generation model, when the user query is processed, an answer more conforming to the characteristics of the report field can be generated, so that the interpretation accuracy and specialty are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and large language model (LLM) applications, and particularly to an LLM report interpretation method and system based on synthetic data. Background Art

[0002] With the rapid development of fields such as genomics, medical testing, and clinical reports, the content of relevant test reports has become increasingly complex. The content of the reports often involves a large number of professional terms, complex structures, and multi-dimensional results, such as chart data, diagnostic results, and suggestions, etc., and the structures are diverse. When faced with these reports, traditional manual interpretation methods not only take a long time, but also are prone to incorrect interpretations due to human negligence or understanding deviations. In addition, some traditional report interpretation techniques usually rely on manually predefined rules or templates, and the interpretation of reports by these methods is often relatively simple and difficult to cover all potential complex information in the reports. Especially for the high-dimensional data, chart analysis, and their correlations in gene test reports and medical test reports, traditional methods cannot deeply understand and extract relevant information. Therefore, it is particularly important to develop an intelligent interpretation method that can automatically process such complex reports and adapt to the needs of different fields. How to accurately and quickly extract effective information from these reports has become a major challenge in the current technology.

[0003] Although the current large language model (LLM) performs well in text generation and semantic understanding, it still faces many challenges when processing highly specialized reports. These challenges mainly include: lack of domain-specific knowledge, insufficient accuracy in report interpretation, and limited accuracy of the model in generating answers when faced with ambiguous or complex queries. In recent years, in order to improve the accuracy and efficiency of report interpretation, retrieval-augmented generation (RAG) technology has been gradually widely applied to text generation tasks. The RAG technology combines information retrieval with a generation model, and by retrieving relevant context information and combining it with the generation model, it can significantly improve the accuracy of the generation results. However, in practical applications, the RAG technology still has certain limitations when faced with report interpretation, especially in terms of how to improve report interpretation efficiency and answer accuracy. Traditional RAG methods do not fully utilize domain-specific context information and caching mechanisms, resulting in the need to rely on a large amount of computing resources during queries and insufficient in-depth understanding of the professionalism of reports.

[0004] Existing report interpretation technologies have problems of insufficient accuracy and incomplete information coverage when processing professional reports in diverse fields. Especially when the data source is complex or the query is ambiguous, traditional methods are difficult to meet the precise needs of users. Although the large language model (LLM) has powerful generation capabilities, it may be insufficient in performance when faced with new fields or specific reports due to limitations in pre-trained data and context correlation capabilities.

[0005] To this end, the present invention proposes an intelligent report interpretation method based on the RAG technology and combined with synthetic data and a question-and-answer caching mechanism to better address this challenge. Summary of the Invention

[0006] The object of the present invention is to solve the problems of insufficient accuracy, extensiveness, and user adaptability existing in the prior art in report interpretation, and to propose an LLM report interpretation method and system based on synthetic data.

[0007] The specific process of an LLM report interpretation method based on synthetic data is as follows:

[0008] Step S0: Use a large language model to generate a synthetic database;

[0009] Step S1: Process each question-and-answer data in the synthetic database obtained in step S0 and store it in the question-and-answer cache library;

[0010] Step S2: Receive the PDF file uploaded by the user, use PyMuPDF to parse the PDF file to obtain text data;

[0011] Step S3: Segment the text data obtained in step S2 to obtain each segment of text data; process each segment of text data to generate vectors; establish an index for the vectors and store them in a temporary vector database;

[0012] Step S4: The user inputs a query text, and obtain the user's question Q through a large language model;

[0013] Step S5: Retrieve the most similar (Q q +A q ) in the question-and-answer cache library established in step S1 for the user's question Q;

[0014] Based on (Q q +A q ) to obtain the q-th question-and-answer pair data (Q q ,A q ) in the synthetic database, and let the q-th question-and-answer pair data C q =(Q q ,A q );

[0015] Wherein, (Q q +A q ) represents the q-th new text in the new text set; (Q q ,A q ) represents the q-th question-and-answer pair data in the synthetic database, and Q q represents the q-th question-and-answer pair data (Q q ,A q) The problem in A q represents the q-th question-and-answer pair data in the synthetic database (Q q , A q ) The answer in; 1 ≤ q ≤ M.

[0016] Step S6, write an instruction prompt word, and use a large language model to determine the most similar question-and-answer pair data C q = (Q q , A q ) Whether it can be directly answered. If so, proceed to step S7; otherwise, proceed to step S8;

[0017] Step S7, the most similar question-and-answer pair data C q = (Q q , A q ) The answer A in q is the answer A to the user's query text q , and end this Q&A;

[0018] Step S8, the user's question Q and the temporary vector database in step S3 obtain the corresponding text block data C t ;

[0019] Step S9, use the question-and-answer pair data C obtained in step S5 q , the text block data C obtained in step S8 t and the user's question Q in step S4, and the large language model obtains the response A;

[0020] Step S10,

[0021] The user judges whether the response A obtained by the large language model in step S9 is reasonable for the user's question Q;

[0022] If so, reply qualified;

[0023] If not, reply unqualified, the user provides a reasonable reply, and a new synthetic data is formed and added to the Q&A cache library in step S1.

[0024] A synthetic data-based LLM report interpretation system includes:

[0025] A data synthesis module, a document parsing module, a cache library retrieval module, a user query module, a document retrieval module, a Q&A determination module, a response generation module, and a response determination module.

[0026] The beneficial effects of the present invention are:

[0027] The objective of the present invention is to provide a method and system for enhancing the performance of LLM report interpretation based on synthetic data, so as to solve the problems of insufficient accuracy, comprehensiveness, and user adaptability existing in the prior art in report interpretation. The objective of the present invention is to provide an intelligent report interpretation method based on the RAG technology, especially in the fields of gene detection reports, medical detection reports, etc., aiming to improve the accuracy and efficiency of report interpretation. By introducing a cache library mechanism and combining Q&A pairs synthesized from professional structured data, the accuracy and response speed of report interpretation can be significantly improved. Specifically, the present invention aims to achieve this goal through the following means: First, generate professional Q&A pairs through synthetic data to enrich the knowledge base of the system; Second, store the answered Q&A pairs through the cache mechanism to avoid repeated calculations and improve the query response speed; Finally, utilize a context-enhanced generation model to generate answers that are more in line with the characteristics of the report field when processing user queries, thereby improving the accuracy and professionalism of the interpretation.

[0028] By combining the Q&A cache mechanism, synthetic data, and dynamic context learning ability, the report interpretation ability of the large language model can be effectively improved. The present invention proposes a method based on synthetic data, generates pseudo Q&A pairs through a structured database and text, and uses technologies such as sliding window chunking, vectorized indexing, and query rewriting to achieve the complete process from user queries to accurate report interpretation. This method aims to improve the accuracy, real-time performance, and field adaptability of report interpretation, and provide users with efficient and intelligent services.

[0029] With the help of context examples in the present invention, the answers generated by the system will be more in line with the professional field style of the report, improving the accuracy of the interpretation. This feature is especially applicable to highly professional fields such as gene detection reports and medical detection reports, and can help users quickly and accurately understand the key content in the report; In addition, the method of the present invention can continuously optimize the answer quality according to the user's feedback, enabling the system to have the ability of dynamic update and optimization, so as to continuously improve the service quality and adapt to the user's needs in the long-term application; Through this self-enhancing mechanism, the system can not only improve the response speed and accuracy of single queries, but also form a more accurate and efficient interpretation platform in the process of continuously accumulating Q&A pairs, adapting to the needs of more professional fields. Brief Description of the Drawings

[0030] Figure 1 It is a flowchart of a method for enhancing the performance of LLM report interpretation based on synthetic data according to the present invention;

[0031] Figure 2 It is a flowchart of the data synthesis of the Q&A cache library according to the present invention;

[0032] Figure 3 It is a vectorization and case indexing diagram of the Q&A data according to the present invention;

[0033] Figure 4 Build an index graph for the document vectorization of the present invention. Detailed implementation manners

[0034] Detailed implementation manner 1: Combine Figure 1 , Figure 2 to illustrate this implementation manner. The specific process of a method for interpreting LLM reports based on synthetic data in this implementation manner is as follows:

[0035] Step S0: Use a large language model to generate a synthetic database;

[0036] Step S1: As Figure 3 shown, process each Q&A data in the synthetic database obtained in step S0 and store it in the Q&A cache library;

[0037] Step S2: Receive a PDF file uploaded by the user, and use PyMuPDF to parse the PDF file to obtain text data;

[0038] Step S3: As Figure 4 shown, segment the text data obtained in step S2 to obtain each segment of text data; process each segment of text data to generate vectors; build an index for the vectors and store them in a temporary vector database;

[0039] Step S4: The user inputs a query text, and obtains the user's question Q through a large language model;

[0040] Step S5: Retrieve the most similar (Q q +A q ) (cosine similarity) of the user's question Q in the Q&A cache library established in step S1;

[0041] Based on (Q q +A q ), obtain the q-th Q&A pair data (Q q ,A q ) in the synthetic database, and let the q-th Q&A pair data C q =(Q q ,A q );

[0042] Among them, (Q q +A q ) represents the q-th new text in the new text set; (Q q ,A q ) represents the q-th Q&A pair data in the synthetic database, Q q represents the question in the q-th Q&A pair data (Q q ,A q ) in the synthetic database, and A qIndicates the answer in the q-th question-answer pair data (Q q , A q ) in the synthetic database; 1 ≤ q ≤ M.

[0043] Step S6, Compile an instruction prompt: "{Is the most similar question-answer pair data C q} directly answer the user's question {user's query}?" Use a large language model to determine whether the most similar question-answer pair data C q =(Q q , A q ) can directly answer. If so, proceed to step S7; otherwise, proceed to step S8;

[0044] Step S7, The answer A in the most similar question-answer pair data C q =(Q q , A q ) is the answer A q to the user's query text, and end this Q&A; q

[0045] Step S8, The user's question Q and the temporary vector database in step S3 obtain the corresponding text block data C t (C t contains K D j );

[0046] Step S9, Use the question-answer pair data C q obtained in step S5, the text block data C t obtained in step S8, and the user's question Q in step S4, and the large language model obtains a response A;

[0047] Step S10,

[0048] The user judges whether the response A obtained by the large language model in step S9 is reasonable for the user's question Q;

[0049] If so, reply qualified;

[0050] If not, reply unqualified, the user provides a reasonable reply, and form a new synthetic data to be added to the Q&A cache library in step S1.

[0051] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that in step S0, a large language model is used to generate a synthetic database.

[0052] The specific process is as follows:

[0053] Step S0-1,

[0054] From the existing medical structured database (including but not limited to open-source data extracted from, for example, the OncoKB tumor database,

[0055] Collect structured data from its structured storage (structured data stored in MySQL, excel, or JSON format). (In the OncoKB tumor database, the data includes structured data (tables) and unstructured data. The structured data is directly collected, and the unstructured data is converted into structured data. An example of the final data format, structured attributes and their attribute values, such as, {gene: ABL1, gene description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia.}, and another example, {gene: ABL1, variant: BCR-ABL1 Fusion, cancer type: B-Lymphoblastic Leukemia / Lymphoma, drug used: Dasatinib});

[0056] Step S0-2: Textualize the structured data in the structured database line by line. (For example, the structured data in Step S0-1 is directly texturized as: one line of the structured data (table) is one piece, and "gene: ABL1, gene description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia." is one piece);

[0057] Step S0-3: Based on each piece of data after textualization in S0-2, write instruction prompts and use a large language model to synthesize Q&A data to obtain the preliminarily synthesized Q&A data;

[0058] The specific process is as follows:

[0059] Design and write instruction prompts: "I hope you act as an expert in building Q&A data. Please build five Q&A data based only on the data I provide, and the generated statements should conform to the real conversation scenario; the following is the data to be built: gene: {data in the database}, gene description: {data in the database}";

[0060] Use the written instruction prompts to synthesize Q&A data to obtain the preliminary Q&A pairs;

[0061] {Data in the database} is the data after textualization in S0-2 (such as "gene: ABL1, gene description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia.");

[0062] Step S0-4: Compile an instruction prompt, such as "Please determine whether the synthesized Q&A data is reasonable", and use a large language model to perform quality screening on the preliminarily synthesized Q&A data. If it is unqualified, repeat Step S0-3 until it is qualified or the number of iteration rounds reaches the upper limit, to obtain a synthesized database {(Q 1 ,A 1 ),(Q 2 ,A 2 ),…,(Q i ,A i ),…,(Q M ,A M )};

[0063] Among them,

[0064] (Q 1 ,A 1 ) represents the first Q&A pair data in the synthesized database, Q 1 represents the question in the first Q&A pair data (Q 1 ,A 1 ) in the synthesized database, and A 1 represents the answer in the first Q&A pair data (Q 1 ,A 1 ) in the synthesized database;

[0065] (Q 2 ,A 2 ) represents the second Q&A pair data in the synthesized database, Q 2 represents the question in the second Q&A pair data (Q 2 ,A 2 ) in the synthesized database, and A 2 represents the answer in the second Q&A pair data (Q 2 ,A 2 ) in the synthesized database;

[0066] (Q i ,A i ) represents the i-th Q&A pair data in the synthesized database, Qi represents the question in the i-th Q&A pair data (Q i ,A i ) in the synthesized database, and A i represents the answer in the i-th Q&A pair data (Q i ,A i ) in the synthesized database;

[0067] (Q M ,A M ) represents the M-th Q&A pair data in the synthesized database, Q M represents the question in the M-th Q&A pair data (Q M ,A M) The problem in A M represents the M-th question-answer pair data (Q M , A M ) in the synthetic database.

[0068] Other steps and parameters are the same as those in the first specific implementation manner.

[0069] Specific implementation manner three: The difference between this implementation manner and the first or second specific implementation manner is that in step S1, as Figure 3 shown, each question-answer data in the synthetic database obtained in step S0 is processed and stored in the question-answer cache library;

[0070] The specific process is as follows:

[0071] Step S1-1: Concatenate each question-answer pair in the synthetic database {(Q 1 , A 1 ),…,(Q i , A i ),…,(Q M , A M )} (formally, {(Q 1 , A 1 ),…(Q n , A n ),...}) into a text to obtain a new text set {(Q 1 +A 1 ),…,(Q i +A i ),…,(Q M +A M )};

[0072] Among them,

[0073] (Q 1 +A 1 ) represents the first new text in the new text set, (Q i +A i ) represents the i-th new text in the new text set, (Q M +A M ) represents the M-th new text in the new text set;

[0074] Step S1-2: Encode each new text in the new text set obtained in step S1-1 into a vector using a vector encoder (using BGE1.5-large-zh);

[0075] Step S1-3: Build an index for the vectors obtained in step S1-2 and store them in a vector database as a question-answer cache library (using Chroma).

[0076] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0077] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that in step S3, as Figure 4 shown, the text data obtained in step S2 is segmented to obtain each segment of text data; each segment of text data is processed to generate vectors; an index is established for the vectors and stored in a temporary vector database; the specific process is as follows:

[0078] Step S3-1: Segment the text data obtained in step S2 to obtain each segment of text data; the specific process is as follows:

[0079] Step S3-2: Use a vector encoder (using bge1.5-large-zh) to encode each segment of text data obtained in step S3-1 into vectors;

[0080] Step S3-3: Establish an index for the vectors in step S3-2 and store them in a temporary vector database (using Chroma).

[0081] Other steps and parameters are the same as those in one of the first to third specific implementation manners.

[0082] Specific implementation manner five: The difference between this implementation manner and one of the first to fourth specific implementation manners is that in step S3-1, the text data obtained in step S2 is segmented to obtain each segment of text data; the specific process is as follows:

[0083] We use the method of sliding window, for example, set the window size to 512 characters and set each step to 256;

[0084] Use the sliding window method to divide the text data obtained in step S2 into N segments, each segment has a size of 512 characters, and the character overlap between adjacent segments is 256, to obtain the text data set D = {D 1 , D 2 , …, D j , …, D N};

[0085] Each segment of text data is represented as D j , j = 1, 2, …, N;

[0086] Among them, D 1 represents the first segment of text data, D 2 represents the second segment of text data, D j represents the jth segment of text data, D N represents the Nth segment of text data, j = 1, 2, …, N.

[0087] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.

[0088] Specific Embodiment Six: The difference between this embodiment and any one of the first to fifth specific embodiments is that in step S4, the user inputs a query text, and the user question Q is obtained through a large language model. The specific process is as follows:

[0089] Step S4-1: Perform re regular expression processing on the query text input by the user (including but not limited to eliminating chaotic punctuation marks) to obtain the user query text after re regular expression processing.

[0090] Step S4-2: Compile an instruction prompt, "Please standardize the user query text after re regular expression processing so that the user query text after re regular expression processing conforms to the problems in the field of gene detection", and use the large language model to perform standardization processing on the user query text after re regular expression processing to obtain the standardized user question Q. Standardization means eliminating its ambiguity so that it is more in line with the queries in the vertical field to which the report belongs.

[0091] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0092] Specific Embodiment Seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that in step S8, the user question Q and the temporary vector database in step S3 are used to obtain the corresponding text block data C t (C t contains K Ds j )

[0093] The specific process is as follows:

[0094] Use a vector encoder (using bge1.5-large-zh) to encode the user question Q into a vector.

[0095] Retrieve the K vectors that are most similar to the vector corresponding to the user question Q in the temporary vector database in step S3.

[0096] Based on the K most similar vectors, obtain the corresponding text block data C t (C t contains K Ds j )

[0097] Other steps and parameters are the same as those in any one of the first to sixth specific embodiments.

[0098] Specific Embodiment Eight: The difference between this embodiment and any one of the first to seventh specific embodiments is that in step S9, the Q&A pair data C obtained in step S5 q and the text block data C obtained in step S8 tObtain the user question Q in step S4, and the large language model obtains the response A; the specific process is as follows:

[0099] Use the Q&A pair data C obtained in step S5 q and the text block data C obtained in step S8 t Obtain the user question Q in step S4, write an instruction prompt, and use the large language model to obtain the response A;

[0100] Instruction prompt:

[0101] "The following are the retrieved relevant fragments,

[0102] {C t}

[0103] The following is an example of interpreting a Q&A. You can refer to the content, format, style, etc.:

[0104] {Q q}{A q}

[0105] Please answer the following question:

[0106] {Q}"

[0107] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.

[0108] Specific embodiment nine: A LLM report interpretation system based on synthetic data in this embodiment, the system includes:

[0109] A data synthesis module, a document parsing module, a cache library retrieval module, a user query module, a document retrieval module, a Q&A determination module, a response generation module, and a response determination module.

[0110] Specific embodiment ten: Different from specific embodiment nine, the data synthesis module is used to generate a synthetic database using a large language model, process each Q&A data in the synthetic database, and store it in the Q&A cache library;

[0111] The document parsing module is used to receive the PDF file uploaded by the user, parse the PDF file using PyMuPDF, and obtain the text data;

[0112] The cache library retrieval module is used to segment the text data obtained in step S2 to obtain each segment of text data; process each segment of text data to generate vectors; establish an index for the vectors and store them in a temporary vector database;

[0113] The user query module is used for the user to input query text and obtain the user question Q through the large language model;

[0114] The document retrieval module is used to retrieve the most similar (Q in the established Q&A cache library for the user's question q +A q )(cosine similarity); based on (Q q +A q ) to obtain the q-th Q&A pair data (Q q , A q ) in the synthesis database, and let the q-th Q&A pair data C q =(Q q , A q );

[0115] The Q&A determination module is used to write an instruction prompt word, "{Whether the retrieved most similar Q&A pair data C q} can directly answer the user's question {the user's query}?", and use the large language model to determine whether the most similar Q&A pair data C q =(Q q , A q ) can directly answer. If so, the answer A q in the most similar Q&A pair data C q =Q q , A q is the answer A q to the user's query text, and this Q&A ends; otherwise, the user's question Q and the temporary vector database in step S3 obtain the corresponding text block data C t (C t contains K Ds j );

[0116] The response generation module is used to use the Q&A pair data C q , the text block data C t and the user's question Q to obtain the response A from the large language model;

[0117] The response determination module is used for the user to determine whether the response A obtained by the large language model is reasonable for the user's question Q; if it is, the reply is qualified; if not, the reply is unqualified, and the user provides a reasonable reply, which is composed into new synthesis data and added to the synthesis database in step S0.

[0118] Other steps and parameters are the same as those in the specific implementation method eight.

[0119] Example:

[0120] The following combines the accompanying drawings and specific implementation methods to further clarify the present invention. It should be understood that the following schematic embodiments of the present invention are only used to explain the present invention and not to limit the scope of the present invention.

[0121] The present invention involves two different types of models. The first type is the dense vector retrieval model, which can vectorize and encode text. Currently, the open-source model BGE1.5-large-zh is selected. The second type is the large language model (LLM), which can receive instructions and questions and generate responses. In the present invention, the open-source model Qwen2-72B-Instruct model is selected.

[0122] The prompts used in the present invention are only examples for achieving effects using large models and can be generally understood as a series of prompts, rather than being construed as a limitation to the present invention.

[0123] Example 1

[0124] This example provides a method for enhancing the performance of LLM report interpretation based on synthetic data, which can be applied to the interpretation of gene detection reports, and can also be applied to the interpretation of detection reports or body reports during disease diagnosis and treatment, such as Figure 1 shown, including the following steps:

[0125] Step S0 / As Figure 2 shown, use the large language model to generate a synthetic database. The specific process is as follows:

[0126] Step S0-1,

[0127] Collect structured data from existing medical structured databases (including but not limited to open-source data extracted from databases such as the OncoKB tumor database,

[0128] the structured storage of which is structured data stored in MySQL, excel, or JSON format). The data in the OncoKB tumor database includes structured data (tables) and unstructured data. The structured data is directly collected, and the unstructured data is converted into structured data. Examples of the final data format, structured attributes and their attribute values are as follows: {gene: ABL1, gene description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia.}, and another example: {gene: ABL1, mutation: BCR-ABL1 Fusion, cancer type: B-Lymphoblastic Leukemia / Lymphoma, drug used: Dasatinib}.

[0129] Step S0-2: Textualize the structured data in the structured database line by line (for example, directly textualize the structured data in Step S0-1 as: one line of the structured data (table) is one piece, and "Gene: ABL1, Gene Description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia." is one piece);

[0130] Step S0-3: Based on each piece of data after textualization in S0-2, write instruction prompts, and use the large language model to synthesize Q&A data to obtain the preliminarily synthesized Q&A data;

[0131] The specific process is as follows:

[0132] Design and write instruction prompts: "I hope you act as an expert in constructing Q&A data. Please construct five Q&A data only based on the data I provide, and the generated statements should conform to the real conversation scenario; the following is the data to be constructed: Gene: {data in the database}, Gene Description: {data in the database}";

[0133] Use the written instruction prompts to synthesize Q&A data to obtain the preliminary Q&A pairs;

[0134] {Data in the database} is the data after textualization in S0-2 (such as "Gene: ABL1, Gene Description: a tyrosine kinase, is frequently altered by chromosomal translocations in leukemia.");

[0135] Step S0-4: Write instruction prompts, such as "Please judge whether the synthesized Q&A data is reasonable", and use the large language model to perform quality screening on the preliminarily synthesized Q&A data. If it is unqualified, repeat Step S0-3 until it is qualified or the iteration round reaches the upper limit to obtain the synthesized database {(Q 1 ,A 1 ),(Q 2 ,A 2 ),…,(Q i ,A i ),…,(Q M ,A M )};

[0136] Among them,

[0137] (Q 1 ,A 1 ) represents the first Q&A pair data in the synthesized database, and Q 1 represents the first Q&A pair data in the synthesized database (Q 1,A 1 ) the problem in, A 1 represents the first question - answer pair data (Q in the synthesized database 1 ,A 1 ) the answer in;

[0138] (Q 2 ,A 2 ) represents the second question - answer pair data in the synthesized database, Q 2 represents the second question - answer pair data (Q in the synthesized database 2 ,A 2 ) the question in, A 2 represents the second question - answer pair data (Q in the synthesized database 2 ,A 2 ) the answer in;

[0139] (Q i ,A i ) represents the i - th question - answer pair data in the synthesized database, Q i represents the i - th question - answer pair data (Q in the synthesized database i ,A i ) the question in, A i represents the i - th question - answer pair data (Q in the synthesized database i ,A i ) the answer in;

[0140] (Q M ,A M ) represents the M - th question - answer pair data in the synthesized database, Q M represents the M - th question - answer pair data (Q in the synthesized database M ,A M ) the question in, A M represents the M - th question - answer pair data (Q in the synthesized database M ,A M ) the answer in.

[0141] Step S1, as Figure 2 shown, process each question - answer data in the synthesized database obtained in Step S0 and store it in the question - answer cache library;

[0142] The specific process is:

[0143] Step S1 - 1, Take the synthesized database {(Q 1 ,A 1 ),…,(Q i ,A i ),…,(Q M ,A M )} (Formally, {(Q 1 ,A 1 ),…(Qn , A n ),...}) into a single text for each Q&A pair to obtain a new text set {(Q 1 + A 1 ), …, (Q i + A i ), …, (Q M + A M )};

[0144] Among them,

[0145] (Q 1 + A 1 ) represents the first new text in the new text set, (Q i + A i ) represents the i-th new text in the new text set, (Q M + A M ) represents the M-th new text in the new text set;

[0146] Step S1-2: Encode each new text in the new text set obtained in Step S1-1 into a vector using a vector encoder (using BGE1.5-large-zh);

[0147] Step S1-3: Build an index for the vectors obtained in Step S1-2 and store them in a vector database as a Q&A cache library (using Chroma).

[0148] Step S2: Receive the PDF file uploaded by the user and parse the PDF file using PyMuPDF to obtain text data;

[0149] Step S3: As Figure 4 shown, segment the text data obtained in Step S2 to obtain each segment of text data; process each segment of text data to generate vectors; build an index for the vectors and store them in a temporary vector database;

[0150] The specific process is as follows:

[0151] Step S3-1: Segment the text data obtained in Step S2 to obtain each segment of text data; the specific process is as follows:

[0152] We use the method of sliding window, such as setting the window size to 512 characters and setting each step to 256;

[0153] Use the sliding window method to divide the text data obtained in Step S2 into N segments, each segment with a size of 512 characters, and the character overlap between adjacent segments is 256, to obtain a text data set D = {D 1 , D 2 , …, D j,…,D N};

[0154] Each piece of text data is represented as D j , j = 1, 2, …, N;

[0155] Among them, D 1 represents the first piece of text data, D 2 represents the second piece of text data, D j represents the j-th piece of text data, D N represents the N-th piece of text data, j = 1, 2, …, N;

[0156] Step S3-2: Use a vector encoder (using bge1.5-large-zh) to encode each piece of text data obtained in Step S3-1 into a vector;

[0157] Step S3-3: Index the vectors in Step S3-2 and store them in a temporary vector database (using Chroma).

[0158] Step S4: The user inputs a query text, and obtains the user's question Q through a large language model;

[0159] The specific process is as follows:

[0160] Step S4-1: Perform re regular expression processing on the query text input by the user (including but not limited to eliminating chaotic punctuation marks) to obtain the user's query text after re regular expression processing;

[0161] Step S4-2: Compile an instruction prompt, "Please standardize the user's query text after re regular expression processing so that the user's query text after re regular expression processing conforms to the questions in the field of gene detection", and use a large language model to perform standardization processing on the user's query text after re regular expression processing to obtain the standardized user question Q; Standardization means eliminating its ambiguity so that it is more in line with the queries in the vertical field to which the report belongs.

[0162] Step S5: Retrieve the most similar (Q q +A q )(cosine similarity) of the user question Q in the Q&A cache library established in Step S1;

[0163] Based on (Q q +A q ) to obtain the q-th Q&A pair data (Q q ,A q ) in the synthesis database, and let the q-th Q&A pair data C q =(Q q ,A q );

[0164] Among them, (Q q +A q ) represents the q-th new text in the new text set; (Q q ,A q ) represents the q-th question-answer pair data in the synthetic database, Q q represents the question in the q-th question-answer pair data (Q q ,A q ) in the synthetic database, and A q represents the answer in the q-th question-answer pair data (Q q ,A q ) in the synthetic database; 1 ≤ q ≤ M.

[0165] Step S6, write an instruction prompt: "{Can the most similar question-answer pair data C q} directly answer the user's question {user's query}?" Use a large language model to determine whether the most similar question-answer pair data C q =(Q q ,A q ) can directly answer. If so, proceed to step S7; otherwise, proceed to step S8;

[0166] Step S7, the answer A q in the most similar question-answer pair data C q ,A q ) is the answer A q to the user's query text, and end this Q&A; q

[0167] Step S8, obtain the corresponding text block data C t (C t contains K D j ) from the user's question Q and the temporary vector database in step S3; the specific process is as follows:

[0168] Use a vector encoder (using bge1.5-large-zh) to encode the user's question Q into a vector;

[0169] Retrieve the K most similar vectors to the vector corresponding to the user's question Q in the temporary vector database in step S3;

[0170] Obtain the corresponding text block data C t (C t contains K D j ) based on the K most similar vectors;

[0171] Step S9, utilize the question-answer pair data C q obtained in step S5, the text block data C tObtain the user question Q in step S4, and the large language model obtains the response A; the specific process is as follows:

[0172] Utilize the Q&A pair data C obtained in step S5 q and the text block data C obtained in step S8 t Obtain the user question Q in step S4, compile an instruction prompt, and use the large language model to obtain the response A;

[0173] Instruction prompt:

[0174] "The following are the retrieved relevant fragments,

[0175] {C t}

[0176] The following is an example of interpreting a Q&A, and you can refer to the content, format, style, etc.:

[0177] {Q q}{A q}

[0178] Please answer the following question:

[0179] {Q}"

[0180] Step S10,

[0181] The user determines whether the response A obtained by the large language model in step S9 is reasonable for the user question Q;

[0182] If it is, then reply qualified;

[0183] If not, then reply unqualified, and the user provides a reasonable reply, which is combined into new synthetic data and added to the synthetic database in step S0.

[0184] This embodiment proposes a method for enhancing the report interpretation performance of a large language model (LLM) based on synthetic data, which is applicable to the interpretation of gene detection reports and other medical detection reports. By using a structured database in a professional field to generate pseudo Q&A data and combining a dense vector retrieval model to vectorize and encode the text, a Q&A cache library is constructed. When the user queries, the system first optimizes the query through standardization processing, then finds the most relevant Q&A data through vector retrieval, and combines the report content to generate a response, improving the interpretation accuracy and efficiency. Finally, the system will optimize the Q&A data according to the user feedback and dynamically update the cache library to continuously improve the performance.

[0185] Embodiment 2

[0186] This embodiment provides a system for enhancing the report interpretation performance of an LLM (Large Language Model) based on synthetic data, which can be applied to the interpretation of gene detection reports, and can also be applied to the interpretation of detection reports or physical examination reports during the disease diagnosis and treatment process. It includes the following modules:

[0187] (1) The data synthesis module is used to generate a synthetic database using a large language model, process each Q&A data in the synthetic database, and store it in the Q&A cache library;

[0188] The data processing unit collects information from (such as a MySQL database) and unstructured texts (such as medical reports, academic articles, etc.). By using the SQLAlchemy library to connect to the database, the pandas library is used to organize tabular data and convert it into a format suitable for Q&A.

[0189] The Q&A synthesis unit: According to the information extracted from the data, through the Prompt technology, use a large language model (such as Qwen2-72B-Instruct or GPT series) to synthesize pseudo Q&A pairs.

[0190] The quality screening unit: Conduct quality assessment on the synthesized Q&A through a large language model, and screen out Q&A pairs that meet the standards. If the Q&A does not meet the expectations, the system will synthesize again. Example instructions used: "Please judge whether the following Q&A is reasonable."

[0191] The vectorization module: Vectorize the screened Q&A pairs and store them in the vector database. Use the sentence-transformers library to encode the Q&A pairs into vectors (we use BGE1.5-large-zh), and use vector databases such as Chroma or FAISS for storage and management for subsequent rapid retrieval.

[0192] (2) The document parsing module is used to receive the PDF file uploaded by the user, and use PyMuPDF to parse the PDF file to obtain text data;

[0193] The PDF parsing unit, when the user uploads a PDF report, the system uses the PyMuPDF (also known as fitz) library to extract the text in the report to ensure accurate parsing.

[0194] The text cleaning unit uses regular expressions (re) for cleaning, removes these redundant parts, and retains the parts related to the report content.

[0195] Chunk Processing Unit: To better retrieve reports, the text is divided into multiple smaller segments. The size of each segment is usually 512 characters, and the sliding window method cuts the text with a step size of 256 characters. Each piece of text is vectorized separately for subsequent retrieval and interpretation.

[0196] (3) Cache Library Retrieval Module, which uses the LangChain framework and retrieves the user's query through a vector database to find the most relevant Q&A data. This module ensures that the system can respond to user queries efficiently and quickly.

[0197] The cache library retrieval module is used to segment the text data obtained in step S2 to obtain each segment of text data; process each segment of text data to generate vectors; establish an index for the vectors and store them in a temporary vector database;

[0198] The user query module is used for the user to input query text and obtain the user's question Q through a large language model;

[0199] The document retrieval module is used to retrieve the most similar (Q q +A q )(cosine similarity) of the user's question in the established Q&A cache library; based on (Q q +A q ) to obtain the q-th Q&A pair data (Q q ,A q ) in the synthesis database, and let the q-th Q&A pair data C q =(Q q ,A q );

[0200] Vectorization Unit: With the support of the data synthesis module, all Q&A pairs have been converted into vector form and stored in the vector database. Each Q&A pair is encoded as a fixed-length vector representing its semantic features.

[0201] Similarity Unit: When the user submits a query, the system first vectorizes the query text and retrieves it through a vector database (such as FAISS or Chroma). The retrieval process is based on cosine similarity or other similarity algorithms to find the most similar Q&A pair. We use LangChain to implement the database connection and encoding query process. (4)

[0203] The Q&A determination module is used to write an instruction prompt word, "{Whether the retrieved most similar Q&A pair data C q} can directly answer the user's question {the user's query}?", and use a large language model to judge the most similar Q&A pair data C q =(Q q ,Aq ) Whether it can be directly answered. If so, the most similar Q&A pair data C q = (Q q , A q ) The answer A in q is the answer A to the user's query text q , and this Q&A session ends; otherwise, the user's question Q and the temporary vector database in step S3 obtain the corresponding text block data C t (C t contains K Ds j );

[0204] Text Vectorization Unit: Each block of text in the report (such as the fragments segmented by the PDF parsing module) will be vectorized. Using the sentence-transformers library, the BGE1.5 large zh model is used to encode each text block into a vector and store it in the temporary vector database.

[0205] Relevant Text Retrieval Unit: When the user queries, the system searches the temporary vector database for the text fragment most relevant to the query. This retrieval process also relies on vector similarity calculation to ensure that the returned fragment is highly relevant to the query.

[0206] (5) The Response Generation Module is used to utilize the Q&A pair data C q , the text block data C t and the user's question Q, and the large language model obtains the response A;

[0207] Answer Generation Unit, the integrated context is input into the large language model (such as Qwen2-72B-Instruct), and through the generation ability of the model, the final answer is generated. The generated answer will consider all the provided context information to ensure accuracy.

[0208] Answer Output Unit: The finally generated answer will be returned to the user. If the query is relatively simple, the answer may only come from the Q&A pairs retrieved from the cache library; if the query is complex, the answer will incorporate more report fragments and Q&A examples.

[0209] Query Standardization Unit: The query submitted by the user may be ambiguous or non-standard. The system will standardize the query through the large language model. This includes correcting incorrect punctuation marks and eliminating semantic ambiguities to make the query more explicit.

[0210] Context integration: In addition to standard queries, the system also integrates the queries with other relevant information (such as report fragments, Q&A examples) so that the generator module can generate more accurate answers. Instructions used are like "The following is an example of an interpreted Q&A. You can refer to its content, format, style, etc.: {example Q&A pair}. The following is a relevant report fragment: {retrieved relevant report content}"

[0211] (8) The response determination module is used for the user to determine whether the response A obtained by the large language model is reasonable for the user question Q; if it is, it replies qualified; if not, it replies unqualified, and the user provides a reasonable reply, which is composed into new synthetic data and added to the synthetic database in step S0.

[0212] Through the collaboration of multiple modules, the system realizes a closed loop from data synthesis, document parsing, query processing to user feedback. By continuous learning and optimization, it can provide users with high-quality report interpretation services.

[0213] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for interpreting LLM reports based on synthetic data, characterized by: The specific process of the method is: Step S0: Generate a synthetic database using a large language model; Step S1, processing each question and answer data in the synthetic database obtained in step S0, and storing it in the question and answer cache library; Step S2, receiving the PDF file uploaded by the user, and using PyMuPDF to parse the PDF file to obtain text data; Step S3, segmenting the text data obtained in step S2 to obtain each segment of text data; processing each segment of text data to generate a vector; indexing the vector and storing it in a temporary vector database; Step S4: The user inputs a query text, and the user question Q is obtained through the large language model; Step S5: Search the most similar (Q q +A q ); Based on (Q q +A q ) Get the qth question-answer pair data (Q q ,A q ), let the qth question-answer pair data C q =(Q q ,A q ); Among them, (Q q +A q ) represents the qth new text in the new text set; (Q q ,A q ) represents the qth question-answer pair data in the synthetic database, Q q represents the qth question-answer pair data (Q q ,A q ) in the problem, A q represents the qth question-answer pair data (Q q ,A q )The answer is; 1≤q≤M. Step S6: Write the instruction prompt words and use the large language model to determine the most similar question-answer pair data C q =(Q q ,A q ) whether it can be answered directly, if yes, proceed to step S7, otherwise proceed to step S8; Step S7: most similar question-answer pair data C q =(Q q ,A q ) q That is the answer A to the user's query text q , and end this Q&A session; Step S8: The user question Q and the temporary vector database of step S3 obtain the corresponding text block data C t ; Step S9, using the question-answer pair data C obtained in step S5 q , the text block data C obtained in step S8 t The user question Q in step S4 is obtained, and the large language model obtains the response A; Step S10: The user judges whether the response A obtained by the large language model in step S9 is reasonable for the user question Q; If yes, then reply qualified; If not, the reply is unqualified. The user provides a reasonable response to form new synthetic data and add it to the question and answer cache library in step S1.

2. The method for interpreting an LLM report based on synthetic data according to claim 1, characterized in that: In step S0, a synthetic database is generated using a large language model; the specific process is: Step S0-1, collecting structured data from an existing structured database; Step S0-2, formatting the structured data in the structured database by item; Step S0-3: based on each piece of data converted into text in S0-2, write a command prompt word, use a large language model to synthesize question and answer data, and obtain preliminary synthesized question and answer data; Step S0-4: Write instruction prompt words and use the large language model to screen the quality of the initially synthesized question and answer data. If it fails, repeat step S0-3 until it passes or the number of iterations reaches the upper limit, and obtain the synthetic database {(Q1, A1), (Q2, A2), …, (Q i ,A i ),…,(Q M ,A M )}; in, (Q1, A1) represents the first question-answer pair data in the synthetic database, Q1 represents the question in the first question-answer pair data (Q1, A1) in the synthetic database, and A1 represents the answer in the first question-answer pair data (Q1, A1) in the synthetic database; (Q2, A2) represents the second question-answer pair data in the synthetic database, Q2 represents the question in the second question-answer pair data (Q2, A2) in the synthetic database, and A2 represents the answer in the second question-answer pair data (Q2, A2) in the synthetic database; (Q i ,A i ) represents the i-th question-answer pair data in the synthetic database, Qi represents the i-th question-answer pair data in the synthetic database (Q i ,A i ), Ai represents the i-th question-answer pair data in the synthetic database (Q i ,A i ) in the answer; (Q M ,A M ) represents the Mth question-answer pair data in the synthetic database, Q M represents the Mth question-answer pair data (Q M ,A M ) in the problem, A M represents the Mth question-answer pair data (Q M ,A M ) in the answer.

3. The method for interpreting an LLM report based on synthetic data according to claim 2, characterized in that: In step S1, each question and answer data in the synthetic database obtained in step S0 is processed and stored in the question and answer cache library; the specific process is: Step S1-1: Synthesize the database {(Q1,A1),…,(Q i ,A i ),…,(Q M ,A M )} are connected into a text, and the new text set {(Q1+A1),…,(Q i +A i ),…,(Q M +A M )}; in, (Q1+A1) represents the first new text in the new text set, (Q i +A i ) represents the i-th new text in the new text set, (Q M +A M ) represents the Mth new text in the new text set; Step S1-2, encoding each new text in the new text set obtained in step S1-1 into a vector using a vector encoder; Step S1-3: Create an index for the vector obtained in step S1-2 and store it in a vector database as a question and answer cache.

4. The method for interpreting an LLM report based on synthetic data according to claim 3, characterized in that: In the step S3, the text data obtained in step S2 is segmented to obtain each segment of text data; each segment of text data is processed to generate a vector; the vector is indexed and stored in a temporary vector database; The specific process is: Step S3-1, segmenting the text data obtained in step S2 to obtain each segment of text data; Step S3-2, using a vector encoder to encode each piece of text data obtained in step S3-1 into a vector; Step S3-3: Create an index for the vector in step S3-2 and store it in a temporary vector database.

5. The method for interpreting an LLM report based on synthetic data according to claim 4, characterized in that: In step S3-1, the text data obtained in step S2 is segmented to obtain each segment of text data; the specific process is: The text data obtained in step S2 is divided into N segments using the sliding window method. The size of each segment is 512 characters, and the character overlap between adjacent segments is 256. The text data set D = {D1, D2, ..., D j ,…,D N }; Each piece of text data is represented by D j , j = 1, 2, ..., N; Among them, D1 represents the first paragraph of text data, D2 represents the second paragraph of text data, and D j represents the jth paragraph of text data, D N Represents the Nth paragraph of text data, j=1,2,…,N.

6. The method for interpreting an LLM report based on synthetic data according to claim 5, characterized in that: In step S4, the user inputs a query text, and the user question Q is obtained through the large language model; the specific process is: Step S4-1, performing re regular expression processing on the query text input by the user to obtain the user query text processed by re regular expression; Step S4-2: Write instruction prompt words, and use the large language model to standardize the user query text processed by the re regular expression to obtain the standardized user question Q.

7. The method for interpreting an LLM report based on synthetic data according to claim 6, characterized in that: The user question Q in step S8 and the temporary vector database in step S3 obtain the corresponding text block data C t ; The specific process is: Use a vector encoder to encode the user question Q into a vector; Retrieve K vectors that are most similar to the vector corresponding to the user question Q from the temporary vector database in step S3; Based on the most similar K vectors, the corresponding text block data C is obtained t .

8. The method for interpreting an LLM report based on synthetic data according to claim 7, characterized in that: In step S9, the question-answer pair data C obtained in step S5 is used q , the text block data C obtained in step S8 t The user question Q in step S4 is obtained, and the large language model obtains the response A; The specific process is: Using the question-answer pair data C obtained in step S5 q , the text block data C obtained in step S8 t Get the user question Q in step S4, write the instruction prompt words, and use the large language model to get the response A.

9. A LLM report interpretation system based on synthetic data, characterized by: The system comprises: a data synthesis module, a document parsing module, a cache retrieval module, a user query module, a document retrieval module, a question and answer determination module, a response generation module and a response determination module.

10. The LLM report interpretation system based on synthetic data according to claim 9, characterized in that: The data synthesis module is used to generate a synthetic database using a large language model, process each question and answer data in the synthetic database, and store it in a question and answer cache library; The document parsing module is used to receive PDF files uploaded by users and use PyMuPDF to parse the PDF files to obtain text data; The cache library retrieval module is used to segment the text data obtained in step S2 to obtain each segment of text data; process each segment of text data to generate a vector; index the vector and store it in a temporary vector database; The user query module is used for users to input query text and obtain user questions Q through the large language model; The document retrieval module is used to retrieve the most similar (Q q +A q ); Based on (Q q +A q ) Get the qth question-answer pair data (Q q ,A q ), let the qth question-answer pair data C q =(Q q ,A q ); The question-answer determination module is used to write instruction prompts and use the large language model to determine the most similar question-answer pair data C q =(Q q ,A q ) can be answered directly, if so, the most similar question-answer pair data C q =(Q q ,A q ) q That is the answer A to the user's query text q , and end this question and answer session; otherwise, the user question Q and the temporary vector database of step S3 obtain the corresponding text block data C t ; The response generation module is used to utilize the question-answer pair data C q , text block data C t The large language model gets the response A based on the user question Q. The response judgment module is used for the user to judge whether the response A obtained by the large language model is reasonable to the user question Q; if it is, the reply is qualified; if not, the reply is unqualified. The user provides a reasonable reply to form new synthetic data and add it to the synthetic database of step S0.

Citation Information

Patent Citations

  • Method and device for realizing online question and answer processing by automatically analyzing financial report and converting financial report into structured data, processor and medium

    CN117493518A

  • Intelligent document question and answer method based on large model

    CN117932018A

  • Enterprise annual report analysis method based on LLM and RAG

    CN118520867A

  • Oil and gas field report generation method and system based on plug-in knowledge base and large model

    CN119358529A

  • Rag-based legal information question-and-answer system and method to improve search ability and increase generative ai accuracy

    KR102765364B1

Cited By

  • LLM-RAG sample construction method and device, and storage medium

    CN120705581A