Report interpretation method and system based on intention recognition and intelligent query guidance

By applying intention recognition and intelligent query guidance technology in report interpretation, combining multiple searchers and large language models, the problem of low accuracy of report interpretation in the existing technology is solved, and more efficient and intelligent report interpretation is achieved.

CN120146186APending Publication Date: 2025-06-13HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210799.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing report interpretation methods have low accuracy and are unable to effectively understand and adapt to the complex needs of users, resulting in insufficient information interpretation efficiency and accuracy.

Method used

Using a method based on intent identification and intelligent query guidance, the user input problems are rewrite and vectorized, and combined with the BM25 searcher and dense searcher, the report content and external information are dynamically integrated to identify the user's next query intent.

Benefits of technology

It significantly improves the accuracy and intelligence of report interpretation, can better understand and adapt to user needs, provide accurate search results and answers, and lower the threshold for users to learn.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146186A_ABST
    Figure CN120146186A_ABST
Patent Text Reader

Abstract

The invention discloses a report interpretation method and system based on intention recognition and intelligent query guidance, belongs to the field of natural language processing in artificial intelligence in the field of computers, and particularly relates to the report interpretation method and system. The objective of the invention is to solve the problem that an existing report interpretation method is low in accuracy. The method comprises the following steps of: converting a report in a PDF format into a report in a text format to obtain each segmented text; creating a hybrid retriever; rewriting the problem through the large model; the rewritten questions are retrieved in all the text segments, and related information is obtained; inputting the rewritten question and the obtained related information into a large model, judging whether retrieval is needed or not, and if yes, extracting webpage content; if not, the obtained related information, the extracted webpage content and the rewritten question are input to an open source large language model together, and the large language model generates an answer; and inputting the generated answers and historical dialogue records to an open source large language model, and identifying the next query intention of the user by the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing in artificial intelligence in the computer field, and specifically relates to a report interpretation method and system based on intent recognition and intelligent query guidance. Background Art

[0002] With the rapid development of information technology and the wide application of big data, report interpretation has become an important part of the daily work in various industries. Whether in the fields of medicine, finance or market research, complex data reports and analysis documents often contain a large amount of information, and users need to spend a lot of time and energy to sort out and understand them. However, the traditional report interpretation methods have many limitations:

[0003] (1) The amount of information is huge and scattered: The report content usually has a complex structure, involving various data and analysis conclusions, and it is difficult for users to quickly extract key information.

[0004] (2) The interpretation process depends on manual experience: Currently, most interpretations rely on the user's own professional knowledge and analysis ability, and important information is easily overlooked due to lack of experience or understanding deviation.

[0005] (3) Lack of intelligent assistance: Traditional interpretation tools mainly focus on keyword retrieval or fixed-format abstract generation, and cannot dynamically adjust the interpretation strategy according to the actual needs of users, nor can they predict the possible next needs of users.

[0006] In recent years, intent recognition technology, as an important research direction in the field of natural language processing (NLP), has been widely applied in intelligent assistants, question-and-answer systems and search engines. Through intent recognition technology, the real needs behind the user's language can be understood and accurate solutions can be provided for them. However, in the scenario of report interpretation, the application of intent recognition is still in the initial exploration stage, and existing methods are difficult to achieve a comprehensive understanding of users' complex needs and targeted feedback. Summary of the Invention

[0007] The purpose of the present invention is to solve the problem of low accuracy of existing report interpretation methods, and to propose a report interpretation method and system based on intent recognition and intelligent query guidance.

[0008] The specific process of a report interpretation method based on intent recognition and intelligent query guidance is as follows:

[0009] Step S1: Preprocess the report file in PDF format uploaded by the user, convert it into text format, segment the text format, and obtain each segmented text;

[0010] Step S2: Perform vectorization processing on each text obtained in step S1 to obtain the text vector of each text;

[0011] Store the text vector in the Chroma database and convert the Chroma database after storing the text vector into a dense retriever;

[0012] Build an inverted index based on the N segments of text obtained in step S1; Create a BM25 retriever according to the inverted index;

[0013] Create a hybrid retriever of BM25 and the dense retriever through the EnsembleRetriever class in the langchain.retrievers library;

[0014] Step S3: Rewrite the question input by the user through the large model and store it in the historical conversation record;

[0015] Step S4: Retrieve the rewritten question in all the text segments obtained in step S1 based on step S2, and use the most relevant k segments of text as relevant information;

[0016] Step S5: Input the question rewritten in S3 and the relevant information obtained in S4 into the large model. The large model determines whether retrieval is needed. If so, perform a web search, and use a crawler tool to crawl the returned search results and extract the web page content; If not, execute S6;

[0017] Step S6: Input the relevant information obtained in S4, the web page content extracted in S5, and the question rewritten in S3 into the open-source large language model. The large language model generates an answer and stores the answer in the historical conversation record;

[0018] Step S7: Input the answer generated in S6 and the historical conversation record into the open-source large language model. The large language model identifies the user's next query intention, and repeats to execute S3 to S7 until the conversation ends.

[0019] A report interpretation system based on intention recognition and intelligent query guidance, including:

[0020] A document parsing module, an intention recognition module, a report retrieval module, a web retrieval module, an answer generation module, and a question recommendation module.

[0021] The beneficial effects of the present invention are:

[0022] The purpose of the present invention is to provide a report interpretation method and system based on intention recognition and intelligent query guidance, to solve the problem that the products on the existing market cannot well understand PDF report files, can better understand and adapt to various questions raised by users, so as to provide more accurate search results and more accurate answers, and can also identify the user's next query intention according to the historical conversation and the professional field where the report is located, improving the accuracy of report interpretation.

[0023] The present invention proposes a report interpretation method and system based on intent recognition and intelligent query guidance. By combining the three technologies of question intent recognition, network search intent recognition and next query intent recognition, it can fully capture user needs, dynamically integrate report content and external information, and significantly improve the efficiency and intelligence level of report interpretation. The introduction of this technology fills the gap in the current intelligent application of report interpretation and has broad technical development prospects and practical application value.

[0024] The report interpretation method based on intention recognition and intelligent query guidance described in the present invention has the following beneficial effects by integrating three intention recognitions (question intention recognition, network search intention recognition, and next query intention recognition):

[0025] Improve the accuracy and intelligence of report interpretation: Through problem intent recognition, it can quickly locate the core issues that users are concerned about, and extract key information from the report in a targeted manner, avoiding the interference of redundant data, and improving the efficiency and accuracy of information interpretation;

[0026] Realize rapid retrieval and expansion of information: By using network search intention recognition, the present invention can intelligently determine the supplementary information required by users, and obtain relevant background knowledge or cutting-edge data from the network in real time based on report interpretation, thereby helping users form a more comprehensive understanding;

[0027] Enhance the convenience and predictability of user interaction: By identifying the next query intention, the present invention can proactively predict the content that the user may further focus on, and provide guidance or recommendations in advance, thereby reducing the burden of repeated operations on the user and improving the fluency of the user experience;

[0028] Adapt to the needs of multiple scenarios and fields: The present invention is applicable to the interpretation needs of different types of reports (such as medical reports, financial reports, market research reports, etc.), and has a wide range of application value. At the same time, the organic combination of the three intent recognitions can dynamically adjust the analysis strategy to adapt to the usage habits and industry backgrounds of different users;

[0029] Lower the user learning threshold: Compared with traditional report interpretation methods, the present invention provides users with intuitive and intelligent interpretation results through intent recognition technology. Users can quickly understand the report content without professional knowledge, thereby greatly lowering the usage threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flowchart of a report interpretation method based on intention recognition and intelligent query guidance according to the present invention;

[0031] Figure 2 A flowchart of a report interpretation system based on intention recognition and intelligent query guidance according to the present invention;

[0032] Figure 3 This is the flowchart of the intent recognition module in a report interpretation system based on intent recognition and intelligent query guidance according to the present invention;

[0033] Figure 4 This is the flowchart of the report retrieval module in a report interpretation system based on intent recognition and intelligent query guidance according to the present invention. Detailed implementation manners

[0034] Detailed implementation manner 1: The specific process of a report interpretation method based on intent recognition and intelligent query guidance in this implementation manner is as follows:

[0035] It should be noted that two different types of models are involved in the present invention. The first type is the dense vector retrieval model, which can vectorize and encode text. The open-source model BGE1.5-large-zh is selected in the invention. The second type is the large language model (LLM), which can receive instructions and questions and generate responses. The open-source model Qwen2-72B-Instruct model is selected in the present invention.

[0036] It should be noted that the prompts used in the present invention are only examples for achieving effects using large models, and can be generally understood as a series of prompts, rather than being understood as a limitation to the present invention.

[0037] Step S1: Preprocess the PDF format report file uploaded by the user, convert it into text format, and segment the text format to obtain each segmented text;

[0038] Step S2: Perform vectorization processing on each text obtained in step S1 to support dense retrieval, and obtain the text vectors of each text;

[0039] Store the text vectors in the Chroma database, and convert the Chroma database after storing the text vectors into a dense retriever;

[0040] Construct an inverted index based on the N texts obtained in step S1 to support keyword matching; create a BM25 retriever according to the inverted index;

[0041] Create a hybrid retriever of BM25 and the dense retriever through the EnsembleRetriever class in the langchain.retrievers library;

[0042] Step S3: Rewrite the question input by the user through the large model and store it in the historical conversation record;

[0043] Step S4: Retrieve the rewritten question in all the text segments obtained in Step S1 based on Step S2, and take the most relevant k text segments as relevant information;

[0044] Step S5: Input the question rewritten in S3 and the relevant information obtained in S4 into the large model. The large model determines whether retrieval is needed. If so, conduct a web search, and use a crawler tool to crawl the returned search results and extract the web content; if not, execute S6;

[0045] Step S6: Input the relevant information obtained in S4, the web content extracted in S5, and the question rewritten in S3 together into an open-source large language model. The large language model generates an answer and stores the answer in the historical conversation record;

[0046] Step S7: Input the answer generated in S6 and the historical conversation record into an open-source large language model. The large language model identifies the user's next query intention (question), and repeat the execution of S3 to S7 until the conversation ends.

[0047] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that in Step S1, the PDF format report file uploaded by the user is preprocessed, converted into a text format, and the text format is segmented to obtain the segmented text; specifically including:

[0048] Step S11: Convert the Python script into a dynamic and interactive web application through the Streamlit library;

[0049] Step S12: Optimize the web application through CSS and JavaScript to obtain an optimized web application. The web application accepts the PDF document uploaded by the user;

[0050] The style and interaction behavior of the tooltips are customized and enhanced through CSS and JavaScript, enabling users to obtain a better visual and interaction experience when using;

[0051] Step S13: Use the PyMuPDF library to parse the PDF document uploaded by the user into a text format;

[0052] Clean the text (clean it using regular expressions (re) to remove redundant parts such as headers and footers), and obtain the cleaned text;

[0053] Step S14: To better retrieve the report, divide the text cleaned in S13 into N smaller segments, where N is a positive integer, and each text segment represents D j , j = 1, 2, …, N; the specific process is as follows:

[0054] Use the sliding window method to divide the cleaned text into N smaller segments. The size of each segment is usually 1024 characters, and the character overlap between adjacent segments is 512 (the first 512 characters of the latter segment are the last 512 characters of the former segment). Each segment of text represents D j 。

[0055] Other steps and parameters are the same as those in the first specific implementation manner.

[0056] Specific implementation manner three: The difference between this implementation manner and the first or second specific implementation manner is that in step S2, each segment of text obtained in step S1 is vectorized to support dense retrieval, and the text vector of each segment of text is obtained;

[0057] Store the text vectors in the Chroma database, and convert the Chroma database after storing the text vectors into a dense retriever;

[0058] Construct an inverted index based on the N segments of text obtained in step S1 to support keyword matching; Create a BM25 retriever according to the inverted index;

[0059] Create a hybrid retriever of BM25 and the dense retriever through the EnsembleRetriever class in the langchain.retrievers library;

[0060] The specific process is as follows:

[0061] Step S21: Use the HuggingFaceBgeEmbeddings class in the langchain_community.embeddings library to load the BGE (Bayesian Generative Embedding) embedding model;

[0062] Step S22: Use the Chroma.from_documents method in the langchain_chroma library to create a Chroma database;

[0063] Input each segment of text obtained in step S1 into the BGE (Bayesian Generative Embedding) embedding model, and the BGE (Bayesian Generative Embedding) embedding model outputs the text vector of each segment of text (1 segment of text corresponds to 1 text vector);

[0064] Store the text vectors in the Chroma database;

[0065] Convert the Chroma database after storing the text vectors into a dense retriever;

[0066] Step S23: Tokenize each piece of text using the BM25Retriever.from_documents method in the langchain_community.retrievers library;

[0067] Use the BM25Retriever.from_documents method in the langchain_community.retrievers library to build an inverted index for the N pieces of tokenized text (build 1 inverted index for N pieces of text), and create a BM25 retriever based on the inverted index;

[0068] Step S24: Create a hybrid retriever of BM25 and dense retriever through the EnsembleRetriever class in the langchain.retrievers library (score Hybrid (Q,D j ) = α × score BM25 (Q,D j ) + β × score Dense (Q,D j ))。

[0069] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0070] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that in step S3, the question input by the user is rewritten through a large language model, and the rewritten question is stored in the historical conversation record; specifically including:

[0071] Input the prompt1 into the large language model, and the large language model outputs the rewritten question. The large language model stores the rewritten question in the historical conversation record; specifically including:

[0072] Step S31: Decompose the ambiguity of the question input by the user through the large language model to obtain the question after ambiguity decomposition (write out all possible ambiguities);

[0073] Step S32: The large language model performs anaphora resolution on the question after ambiguity decomposition (replace pronouns with specific objects) to obtain the question after anaphora resolution;

[0074] Step S33: The large language model identifies the most reasonable question based on the question after anaphora resolution, the large language model rewrites the question based on the most reasonable question, and the large language model stores the rewritten question in the historical conversation record;

[0075] The prompt1 includes the question input by the user;

[0076] The prompt1 used is:

[0077] You are a linguist; complete the task according to the following steps;

[0078] Task steps:

[0079] 1. Please break down the following potentially ambiguous query into 2 - 5 questions, each question representing a possible meaning. Please only output the questions;

[0080] 2. Combine the historical chat records, identify the referential relationships in the query and resolve the references;

[0081] 3. Identify the most reasonable intention of the query;

[0082] 4. Reformulate the most reasonable intention of the query into a new query in a clearer and more understandable way. The new query should not contain pronouns;

[0083] The historical chat records are as follows:

[0084] {history}

[0085] The query content is as follows:

[0086] {query};

[0087] Each line in the above format is each line input to the large language model.

[0088] Prompt1 example:

[0089] Input:

[0090] You are a linguist; complete the task according to the following steps;

[0091] Task steps:

[0092] 1. Please break down the following potentially ambiguous query into 2 - 5 questions, each question representing a possible meaning. Please only output the questions;

[0093] 2. Combine the historical chat records, identify the referential relationships in the query and resolve the references;

[0094] 3. Identify the most reasonable intention of the query;

[0095] 4. Reformulate the most reasonable intention of the query into a new query in a clearer and more understandable way. The new query should not contain pronouns;

[0096] The historical chat records are as follows:

[0097] {[{"role":"user","content":"What is PD - L1?"},

[0098] {"role":"assistant","content":"PD-L1 is a protein located on the cell surface that inhibits T cell activity by binding to PD-1, helping cancer cells evade the immune system. Blocking the interaction between PD-1 and PD-L1 can reactivate the immune system and be used in cancer immunotherapy, such as PD-1 or PD-L1 inhibitors."},

[0099] {"role":"user","content":"What is an anti - detachment section?"},

[0100] {"role":"assistant","content":"An anti - detachment section is a technique in pathological or histological experiments, mainly used to prevent tissue sections from falling off the glass slide during staining or processing. Such sections usually involve coating the glass slide with a special adhesive (such as polylysine or gelatin), or by pre - treating the slide surface to ensure that the tissue sample can firmly adhere to the glass slide, avoiding detachment or damage during the operation. The anti - detachment section technique is very important in pathological diagnosis and research, especially in experiments that require repeated processing and staining."}

[0101] The query content is as follows:

[0102] {Can this detection be done without using this section?};

[0103] Output:

[0104] Decomposed questions:

[0105] 1. Can the detection of this protein be carried out without using an anti - detachment section?

[0106] 2. Without using an anti - detachment section, can the detection of this protein be effectively carried out?

[0107] 3. Are there any other types of sections that can replace the anti - detachment section for the detection of this protein?

[0108] 4. If the anti - detachment section is not used, will the detection effect of this protein be affected?

[0109] Anaphora resolution:

[0110] "This protein": refers to "PD - L1" mentioned previously, a protein located on the cell surface.

[0111] "This detection": refers to the detection method related to PD - L1.

[0112] "This section" or "anti - detachment section": refers to the "anti - detachment section" mentioned previously, a technique used to prevent tissue samples from falling off the glass slide.

[0113] The most reasonable intention:

[0114] The user wants to know whether the detection of PD-L1 protein can be carried out without using anti - detachment sections and is worried about whether this will affect the detection effect.

[0115] New query statement:

[0116] "Can the detection of PD-L1 protein be carried out without using anti - detachment sections without affecting its detection effect?"

[0117] Other steps and parameters are the same as one of the specific embodiments one to three.

[0118] Specific embodiment five: The difference between this embodiment and one of the specific embodiments one to four is that in step S4, based on step S2, the rewritten question is retrieved in all text segments obtained in step S1, and the most relevant k text segments are used as relevant information; the specific process is as follows:

[0119] Use a hybrid retriever to obtain the k pieces of relevant information with the highest scores. Among them, for each piece of information in the report, the scores of BM25 and dense retrieval are calculated respectively, and the scores of BM25 and dense retrieval are weighted and fused to obtain the final score;

[0120] Step S41: Based on the rewritten question Q = {q 1 , q 2 , …, q n} and the j - th text segment D j , calculate the BM25 retriever score score BM25 (Q, D j );

[0121] Step S42: Based on the rewritten question Q = {q 1 , q 2 , …, q n} and the j - th text segment D j , calculate the dense retrieval score score Dense (Q, D j );

[0122] Step S43: Based on the BM25 retriever score score BM25 (Q, D j ) obtained in step S41 and the dense retrieval score score Dense (Q, D j ) obtained in step S42, calculate the hybrid retriever score;

[0123] Step S44: Select the k text segments corresponding to the k highest hybrid retriever scores as relevant information.

[0124] 1 ≤ k ≤ 100, usually k = 5 or k = 10.

[0125] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.

[0126] Specific Embodiment Six: Different from any one of the first to fifth specific embodiments, in step S41, based on the rewritten question Q = {q 1 , q 2 , …, q n} and the j-th text segment Dj, calculate the BM25 retriever score score BM25 (Q, D j ); The formula is as follows:

[0127]

[0128] Where:

[0129] Q = {q 1 , q 2 , …, q n} represents the rewritten question, q 1 represents the first word in the rewritten question, q 2 represents the second word in the rewritten question, q n represents the n-th word in the rewritten question, including n word items;

[0130] Words in the rewritten question can be repeated. It's just that when calculating the BM25 retriever score score BM25 (Q, D j ), q 1 , q 2 , …, q n are non-repeated. For example, the sentence "I love you. Do you love me?" is represented as {I, love, you, do}.

[0131] D j represents the j-th text segment obtained in step S1, j = 1, 2, …, N;

[0132] f(q i , D j ) represents the frequency of the word in the j-th text segment D j ;

[0133] |D j | represents the length (total number of words) of the j-th text segment D j ;

[0134] avgdl represents the average length of all text segments;

[0135] k 1a and b respectively represent the hyperparameters of the BM25 retriever, usually k 1 = 1.2 and b = 0.75;

[0136] represents the term q i 's inverse document frequency; expressed as:

[0137]

[0138] where:

[0139] N represents the total number of text segments obtained in step S1;

[0140] n(q i ) represents the number of text segments containing the term q i , i = 1, 2,..., n.

[0141] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0142] Specific embodiment seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that in step S42, based on the rewritten question Q = {q 1 , q 2 ,..., q n} and the j-th text segment D j , the dense retrieval score score Dense (Q, D j ) is calculated; the formula is as follows:

[0143]

[0144] where,

[0145] v Q represents the text vector of the rewritten question Q = {q 1 , q 2 ,..., q n};

[0146] represents the text vector of the j-th text segment D j ;

[0147] · represents the dot product operation of vectors;

[0148] ||v Q || represents the norm (vector length) of the text vector v Q ;

[0149] represents the norm of the text vector of the j-th text segment D j ;

[0150] The rewritten question \(Q = \{q 1 , q 2 , \ldots, q n \}\)'s text vector \(v Q is obtained as follows:

[0151] The rewritten question \(Q\) is input into the BGE embedding model, and the BGE embedding model outputs the text vector of the rewritten question \(Q\);

[0152] The text vector of the \(j\)-th paragraph of text \(D j is obtained as follows:

[0153] The \(j\)-th paragraph of text \(D j is input into the BGE embedding model, and the BGE embedding model outputs the text vector of the \(j\)-th paragraph of text \(D j BM25 .

[0154] Other steps and parameters are the same as those in any one of the first to sixth specific embodiments.

[0155] Specific Embodiment 8: The difference between this embodiment and any one of the first to seventh specific embodiments is that in step S43, based on the BM25 retriever score \(score BM25 (Q, D j ) obtained in step S41 and the dense retriever score \(score Dense (Q, D j ) obtained in step S42, the hybrid retriever score is calculated; the formula is:

[0156] score Hybrid (Q, D j ) = \(\alpha\times score BM25 (Q, D j ) + \(\beta\times score Dense (Q, D j )

[0157] where \(\alpha, \beta\) are the weight parameters of the hybrid retriever, and \(\alpha+\beta = 1\).

[0158] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.

[0159] In step S5, the rewritten question in S3 and the relevant information obtained in S4 are input into the large model, and the large model determines whether retrieval is required. If so, network search is performed, and the search results returned by the network search are crawled using a crawler tool to extract the web page content; if not, S6 is executed;

[0160] The specific process is as follows:

[0161] Input the prompt2 into the large language model. The large language model outputs whether retrieval is needed. If so, conduct a web search, crawl the search results returned by the web search using a crawler tool, and extract the web content; if not, execute S6;

[0162] prompt2 includes the question rewritten in S3 and the relevant information obtained in S4;

[0163] Prompt2:

[0164] Background information:

[0165] You are an xx scientist. Please complete the task according to my requirements.

[0166] Task:

[0167] Please determine whether the knowledge I provided is sufficient for you to answer the query based on the query and knowledge I gave you (if you don't need this knowledge, also answer "yes").

[0168] Output template:

[0169] Reason: {xxxxxxx}

[0170] Answer: {yes or no}

[0171] The query content is as follows:

[0172] {query}

[0173] The knowledge content is as follows:

[0174] {knowledge};

[0175] Each line in the above format is each line input into the large language model;

[0176] {knowledge} is the relevant information obtained in S4, and {query} is the relevant information obtained in S4.

[0177] Prompt2 example:

[0178] Input:

[0179] Background information:

[0180] You are a biologist. Please complete the task according to my requirements.

[0181] Task:

[0182] Please determine whether the knowledge I provided is sufficient for you to answer the query based on the query and knowledge I gave you (if you don't need this knowledge, also answer "yes").

[0183] Output template:

[0184] Reason: {xxxxxxx}

[0185] Answer: {Yes or No}

[0186] The query content is as follows:

[0187] {Can the detection of PD-L1 protein be performed without using anti - detachment sections without affecting its detection effect?}

[0188] The knowledge content is as follows:

[0189] {Currently, 5 antibodies targeting PD - 1 / PD - L1 have been approved by the US FDA for marketing. They are pembrolizumab and nivolumab targeting PD - 1, and atezolizumab, avelumab, and durvalumab targeting PD - L1. Besides achieving breakthrough efficacy in their respective targeted cancer fields, they have become the first class of therapies that do not target the cancer site but rather specific genetic characteristics, and are one of the best examples in the era of precision medicine. With the release of a large amount of clinical data, researchers have been able to analyze the cancer types that can most benefit from PD - 1 inhibitor treatment. Currently, Hodgkin lymphoma, Merkel cell carcinoma, microsatellite - unstable / mismatch - repair - deficient cancers, and adult melanoma are the main battlefields where PD - 1 inhibitors shine, with remission rates reaching 50% - 90%. Melanoma induced by carcinogens is also an area where PD - 1 inhibitors can have a significant effect, with a remission rate of 35% - 40%. For cancers such as non - small cell lung cancer, head and neck cancer, gastric cancer, and bladder cancer, the remission rates are in the range of 15% - 25% (Ribas et al., Science 359, 1350–1355 (2018)).}

[0190] Output:

[0191] Reason:

[0192] The knowledge provided mainly focuses on the clinical applications, efficacy, and applicable cancer types of PD - 1 / PD - L1 inhibitors. However, the query content is concerned with whether the detection of PD - L1 protein can be performed without using anti - detachment sections without affecting its detection effect. This part of the content has no direct relevance to the provided knowledge because the former involves issues in pathology or laboratory techniques, while the latter is about clinical treatment and drug efficacy data.

[0193] The knowledge provided is insufficient to answer questions regarding the technical details of PD-L1 protein detection, particularly whether it can be performed without using anti - detachment sections without affecting the detection results. More specific information about PD-L1 detection methods and related experimental techniques is needed to accurately answer this question.

[0194] Answer:

[0195] No.

[0196] In S6, the relevant information obtained in S5, the web content extracted in S5, and the rewritten question in S3 are input to the open - source large - language model. The large - language model generates an answer and stores the answer in the historical conversation record;

[0197] The specific process is as follows:

[0198] Input the prompt3 into the large - language model. The large - language model generates an answer and stores the answer in the historical conversation record;

[0199] The prompt3 includes the relevant information obtained in S5, the web content extracted in S5, and the rewritten question in S3;

[0200] The used prompt3:

[0201] Please answer the user's question based on the report context and Internet information.

[0202] The report context information includes: {report_context}`.

[0203] The Internet information includes: {internet_context}`.

[0204] Please answer the user's question by combining the above information: {norm_query}.

[0205] Please output strictly according to the format of the following output template:

[0206] [Answer]: “xxxx”;

[0207] Each line of the above format is each line input to the large - language model.

[0208] Prompt3 example:

[0209] Input:

[0210] Please answer the user's question based on the report context and Internet information.

[0211] The reported context information includes: {Currently, five antibodies targeting PD-1 / PD-L1 have been approved by the US FDA for marketing. They are pembrolizumab and nivolumab targeting PD-1, and atezolizumab, avelumab, and durvalumab targeting PD-L1. Besides achieving breakthrough efficacy in their respective targeted cancer fields, they have also become the first class of therapies that do not target the cancer site but specific genetic characteristics, and are one of the best examples in the era of precision medicine. With the release of a large amount of clinical data, researchers have been able to analyze the cancer types that can most benefit from PD-1 inhibitor treatment. Currently, Hodgkin lymphoma, Merkel cell carcinoma, microsatellite instability / mismatch repair defect cancers, and adult melanoma are the main battlefields where PD-1 inhibitors show great prowess, with remission rates reaching 50%-90%. Melanoma induced by carcinogens is also an area where PD-1 inhibitors can have obvious effects, with a remission rate of 35%-40%. For cancers such as non-small cell lung cancer, head and neck cancer, gastric cancer, and bladder cancer, the remission rates are in the range of 15%-25% (Ribas et al., Science 359, 1350–1355 (2018)).}`。

[0212] The Internet information includes: {Frequently Asked Questions about PD-L1 Immunohistochemistry Detection

[0213] In recent years, immunotherapy mainly based on immune checkpoint inhibitors has brought survival benefit opportunities to patients with solid tumors, and more and more immune checkpoint inhibitors have been approved by the FDA or NMPA for the treatment of various cancers. Immune checkpoints mainly based on programmed cell death 1 (PD-1) / programmed death ligand 1 (PD-L1) have become effective biomarkers for screening suitable populations and predicting the efficacy of current immunotherapy. Therefore, it is particularly important to understand immune checkpoints, immune checkpoint inhibitors, detection methods, and judgment criteria.

[0214] Regarding some common problems in the current PD-L1 immunohistochemistry detection, the editor has summarized and answered them. The editor will continue to organize and update them later, so stay tuned!

[0215] Question

[0216] Answer

[0217] I. What is the significance of immune checkpoints and immune checkpoint inhibitors in tumor treatment?

[0218] PD-1 belongs to a transmembrane protein on the cell membrane and is expressed on immune cells such as T cells and B cells, as well as tumor cells. PD-L1 is programmed death ligand 1, the ligand of PD-1, and is usually highly expressed on tumor cells. Studies have shown that when PD-L1 on the tumor cell membrane binds to PD-1 on immune cells such as T cells, the tumor cells send inhibitory signals, inhibiting the recognition and killing of tumor cells by T cells. The tumor cells inhibit the body's immune function and then grow and invade wantonly. Immune checkpoint inhibitors can bind to PD-1 or PD-L1, blocking the inhibition of immune function by tumor cells and ensuring the clearance of tumors by T cells. Based on the relationship between PD-1 / L1 and immune checkpoint inhibitors, it is recommended to perform PD-L1 immunohistochemical detection on tumor cells before immunotherapy, which is beneficial for screening patients with high PD-L1 expression to guide the treatment of immune checkpoint inhibitors.

[0219] II. What specimen types can be submitted for PD-L1 detection?

[0220] Currently, the specimens for PD-L1 immunohistochemical detection are generally histological samples and some cytological samples (such as pleural effusion, ascites, etc.). Generally, it is required that there are no less than 100 tumor cells in the sample for quality control. The "Expert Consensus on PD-L1 Immunohistochemical Detection of Solid Tumors (2021 Edition)" and the "Clinical Pathology Expert Consensus on PD-L1 Expression Detection in Chinese Non-Small Cell Lung Cancer" summarized the sample types that can be submitted for PD-L1 detection:

[0221] (1) It is recommended to give priority to performing PD-L1 immunohistochemical detection on the anti-droplet sections of tumor tissues in paraffin-embedded specimens, and at the same time, it should be ensured that there are enough tumor cells for evaluation. Some studies have shown that the consistency of PD-L1 expression rate of tumor cells among multiple wax blocks of the same tissue tumor in lung cancer patients is high (94%).

[0222] (2) Both surgically resected specimens and biopsy specimens can be used for PD-L1 detection. When tissue specimens cannot be obtained, cytological samples such as pleural effusion and ascites can be embedded in wax blocks for detection.

[0223] (3) Since PD-L1 has temporal heterogeneity and is affected by treatment, it is recommended to perform PD-L1 detection before the initial diagnosis of patients and before timely changing the treatment plan.

[0224] (4) Both primary lesions and metastatic lesions can be used for PD-L1 detection. Since there may be differences in PD-L1 expression between metastatic lesions and primary lesions, it is recommended to perform PD-L1 detection on the primary lesion and metastatic lesion respectively when necessary to clarify the PD-L1 expression status.

[0225] (5) It is not recommended to perform PD-L1 immunohistochemical detection on decalcified specimens of bone metastases.

[0226] III. What are the antibody numbers and evaluation indicators for PD-L1 immunohistochemical detection?

[0227] "Expert Consensus on Immunohistochemical Detection of PD-L1 in Solid Tumors (2021 Edition)" points out that there are mainly 5 kinds of immunohistochemical antibodies (platforms) in current clinical research: 22C3 (DaKo), 28-8 (DaKo), SP142 (Ventana), SP263 (Ventana), and 73-10 (DaKo). For the determination of PD-L1 positive rate, different companies use different antibodies for detection, and the reference indicators or threshold standards for each antibody are also different. In tumor tissues, in addition to cancer cells, there are also surrounding stromal cells (such as immune cells, endothelial cells, fibroblasts, etc.). Therefore, after staining a tumor tissue specimen, the stained cells may be either cancer cells or other cells. For the determination methods of PD-L1 detection antibodies, there are TPS, CPS, and separately calculating TC and IC.

[0228] IV. How to select PD-L1 detection antibodies before detection?

[0229] The PD-L1 immunohistochemical antibody numbers are diverse, and the PD-L1 protein scoring indicators and thresholds are different for different cancer types. Even for some cancer types, there are no relevant standards. Moreover, the PD-L1 detection antibodies, detection and evaluation indicators, and efficacy interpretation thresholds corresponding to different types of immunotherapy drugs are all different.

[0230] Exactly how should cancer patients choose to better complete the detection and more accurately match relevant immune drugs? Here's the answer. The key to achieving "precision" tumor immunotherapy is to select antibodies according to drug indications. Therefore, before a patient undergoes PD-L1 immunohistochemical detection, mainly 2 questions need to be considered: 1. Is there an approved or recommended PD-L1 antibody number for this cancer type in the FDA / NMPA / guidelines? (For example, for advanced non-small cell lung cancer, 22C3 antibody detection is recommended for the treatment of Keytruda), if so, it should be given priority. 2. Is there a corresponding clinical study in this cancer type that refers to the detection antibody number? If so, it can be referred to (for example, the KEYNOTE-028 study showed that in the treatment of advanced recurrent ovarian cancer patients with PD-L1 positive (22C3, CPS≥1) using Keytruda, the objective response rate was 11.5% (3 / 26) and the continuous remission exceeded 30 months, 23.1% (6 / 26) of the patients had tumor shrinkage, and the median overall survival was 13.8 months)

[0231] V. What is the consistency of detection results among different antibodies in lung cancer?

[0232] In the PD-L1 detection of lung cancer, the Blueprint Program indicated that the detection results of the three approved companion diagnostic antibodies, 28-8, 22C3, and SP263, had high consistency; however, the detection sensitivity of the SP142 antibody was relatively low, while the 73-10 antibody had higher detection sensitivity compared to the above-mentioned antibodies. Based on this study, a review proposed that the detection results of the 28-8, 22C3, and SP263 monoclonal antibodies in non-small cell lung cancer were interchangeable. Given that the FDA has approved multiple detection methods for evaluating PD-L1 expression, which are usually used in combination with specific therapeutic drugs; moreover, the sensitivity and repeatability of different detection methods vary. It is recommended that before PD-L1 detection, the antibody number should be selected in combination with the corresponding indication and the drug associated with the companion diagnosis. It is worth noting that the SP263 antibody has been approved to guide the treatment of nivolumab, pembrolizumab, and durvalumab simultaneously, with more associated immune checkpoint inhibitors.

[0233] VI. What risks may exist in the interchange between different antibodies?

[0234] The "Expert Consensus on PD-L1 Immunohistochemical Detection in Solid Tumors (2021 Edition)" pointed out that the consistency studies on different PD-L1 antibodies mainly focused on the field of lung cancer (the consistency of 22C3, 28-8, and SP263 in detecting PD-L1 expression in tumor cells was good), and there were also reports in the fields of breast cancer and urothelial cancer. However, the consistency reports in other cancer types were relatively rare. In fact, using different antibodies to analyze PD-L1 expression in the same sample may lead to different results and affect the decision on whether a patient is suitable for immunotherapy with immune checkpoint inhibitors.

[0235] For example, Keynote-028 was a Phase Ib study. In this study, for advanced solid tumor patients receiving pembrolizumab treatment, the PD-L1 positivity was defined as tumor or stromal cell positive staining ≥ 1% (i.e., CPS threshold ≥ 1) using the 22C3 antibody, suggesting the possibility of clinical efficacy in tumor patients. Similarly, in a Phase I / II study, advanced urothelial cancer patients received durvalumab monotherapy, and it was found using the SP263 antibody that patients with CPS ≥ 25% had a higher response rate. It can be seen that the thresholds for the same indicator also vary in different cancer types or antibodies.

[0236] Similarly, the analysis of the expression of PD-L1 in different cancer types using the same antibody can also lead to different results. For example, when 22C3 is used to match the indications of pembrolizumab, the positive thresholds for different cancer types are different: for non-small cell lung cancer, the TPS is evaluated as ≥1%; for gastric cancer, the CPS is evaluated as ≥1; and for urothelial cancer, the CPS is ≥10. If the detection results of urothelial cancer patients are evaluated using the gastric cancer standard (CPS ≥ 1), false positive results (CPS = 1 - 9) may occur. Based on the different positive thresholds of PD-L1 expression in the same antibody for different cancer types, the detection result thresholds cannot be interchanged among different cancer types.

[0237] It can be seen from this that the thresholds of the same detection index for different antibody numbers may not be consistent; the judgment criteria for different antibody numbers in the same cancer type are also different, and the interchange of antibody numbers and thresholds among cancer types may cause false negative and false positive results.

[0238] VII. Why are the recommendations for some indications inconsistent in each guideline?

[0239] Currently, for the detection of PD-L1, there are many indicators included in domestic and foreign guidelines. Usually, the corresponding cancer types and indicators are included based on the indications approved by the FDA. Therefore, if the diagnostic information approved by the FDA, the NCCN guidelines and CSCO guidelines usually update accordingly. However, in fact, there are some antibodies that, after being approved by the FDA on an accelerated basis, subsequent clinical trial results show that the efficacy of this indication does not reach the clinical trial endpoint or other reasons lead to the withdrawal of the approved indication. For example, the indication of pembrolizumab in gastric cancer:

[0240] The FDA accelerated the approval of pembrolizumab for the treatment of advanced gastric cancer patients with PD-L1 expression (CPS ≥ 1) as early as 2017. In view of this, the gastric cancer NCCN guidelines and CSCO guidelines subsequently recommended pembrolizumab for the treatment of gastric cancer patients with PD-L1 CPS ≥ 1. However, in July this year, MSD's official website issued a news stating that it voluntarily withdrew this indication. Subsequently, the updated gastric cancer NCCN guidelines (version V4) revoked this indication. However, for the domestic CSCO guidelines, due to the long update interval (once a year), the 2021 CSCO guidelines also recommended the treatment of pembrolizumab in the relevant treatment of gastric cancer.

[0241] Therefore, due to the different update frequencies and time differences of the guidelines, it will lead to inconsistencies in the indications of immune checkpoint inhibitors among different guideline recommendations, FDA / NMPA approvals, or before and after.

[0242] NCCN guidelines (version v5) recommend the treatment plan for gastric cancer

[0243] VIII. What is the impact of the age of the specimen on PD-L1 detection?

[0244] As is well known, whether it is gene detection or PD-L1 immunohistochemical detection, generally speaking, the closer the sampling date of the specimen is, the better. However, in actual clinical tests, the sampling of many patients has been separated by several months or even years. Especially for advanced metastatic cancer patients who have undergone multiple lines of treatment after surgery, the previous surgical specimens are submitted for PD-L1 immunohistochemical detection to seek immunotherapy options. Then, what is the difference between fresh specimens and old specimens? How should it be standardized when submitting samples clinically?

[0245] The "Chinese Expert Consensus on the Standardization of PD-L1 Immunohistochemical Detection in Non-Small Cell Lung Cancer (2020 Edition)" clearly states that to ensure the staining quality and the accuracy of test results, it is recommended to use specimens as recent as possible before immunotherapy for PD-L1 detection in clinical practice. If recent specimens are not available, tissue / cell wax block specimens within 3 years can be considered. Research shows that there is a relatively high consistency in PD-L1 expression rates (76.2%) between tissue specimens obtained within ≤3 years and recently obtained specimens, while the proportion of high PD-L1 expression in specimens stored for >3 years decreases significantly, and the detection effect is poor.

[0246] IX. Is there heterogeneity in PD-L1 expression in different specimens / lesions?

[0247] For the PD-L1 immunohistochemical detection project, most of the specimens submitted are surgical specimens, biopsy specimens, and cytological specimens such as pleural effusion and ascites. PD-L1 also shows heterogeneous expression in intratumoral and intertumoral tissues. The "Expert Consensus on PD-L1 Immunohistochemical Detection in Solid Tumors (2021 Edition)" points out that the heterogeneous expression of PD-L1 in tumor tissues can lead to biopsy specimens not being able to represent the actual expression status of the entire tumor. Research results of multiple cancers such as lung cancer, gastric cancer, breast cancer, and bladder cancer suggest that the number of biopsy specimens taken, the diameter of the specimen, and the number of tumor cells in the specimen are important factors affecting the consistency of PD-L1 expression between biopsy and surgical specimens.

[0248] The "Chinese Expert Consensus on the Standardization of PD-L1 Immunohistochemical Detection in Non-Small Cell Lung Cancer" points out the intertumoral heterogeneity of PD-L1 expression. For example, in lung cancer patients, there are differences in PD-L1 expression between the primary focus and metastatic foci, and the inconsistency rate of this intertumoral heterogeneity is 11.4%-39%. Research shows that the difference between the lesions with natural progression in the tumor (recurrence or metastatic foci without treatment) and the primary focus is actually not large (<20%). However, before and after patients receive anti-tumor treatment (chemotherapy, chemoradiotherapy, targeted therapy, immunotherapy, etc.), the difference in PD-L1 expression levels between the primary focus and metastatic foci will be more significant.

[0249] X. What is the relationship between PD-L1 expression and the prognosis of cancer?

[0250] For tumor patients with high PD-1 / PD-L1 expression, the treatment of immune checkpoint inhibitors has played a very important role clinically and can achieve good clinical effects in many different cancers. However, multiple studies have shown that the increased expression of PD-L1 in many cancer types is also a predictor and prognostic marker for tumor treatment.

[0251] In lung cancer, positive PD-L1 expression is more common in men, smokers, and patients with poorly differentiated and metastatic tumors. The results of a Meta-analysis showed that in patients with non-small cell lung cancer (HR = 1.35), adenocarcinoma (HR = 1.79), squamous cell carcinoma (HR = 1.79), and pulmonary lymphoepithelioma-like carcinoma (HR = 3.04), high PD-L1 expression was significantly associated with reduced overall survival, while this phenomenon was not observed in small cell lung cancer (HR = 1.05). Therefore, PD-L1 expression can be an important predictor of poor prognosis in non-small cell lung cancer. Similarly, PD-L1 is an important factor for poor prognosis in breast cancer, indicates a shorter survival time in gastric cancer, also means a shorter overall survival and a higher recurrence risk in papillary thyroid cancer. In colorectal cancer, studies have shown that PD-L1 expression may provide a means of immune escape for normally immunogenic colorectal cancer. Among patients with microsatellite instability (MSI-H), those with high PD-L1 expression have a poor prognosis.

[0252] In summary, the detection of PD-L1 expression is not only an indicator for cancer patients to select immune checkpoint inhibitor treatment strategies, but has also been confirmed as a prognostic biomarker for certain cancers, playing multiple roles in the driving and treatment of tumors as an immunosuppressive checkpoint.

[0253] Huizhen Medicine's PD-L1 immunohistochemistry detection project includes four antibody detections with high consistency and excellent sensitivity, namely 22C3, 28-8, SP263, and 73-10. It can accept different types of samples such as histology and cytology. No matter which antibody you choose, the report will provide the detection values of four different indicators, namely CPS, TPS, TC, and IC for the sample, and comprehensively evaluate the PD-L1 immunohistochemistry expression detection results. Clinically, treatment decisions can be made based on the pre-selected immune drugs with reference to different antibody numbers and related thresholds. Moreover, a 5-natural-day report cycle and a complete after-sales interpretation service will provide you with fast and accurate detection services.

[0254] Please answer the user's question based on the above information: {Can the detection of PD-L1 protein be carried out without using anti-drop sections without affecting its detection effect?}

[0255] Please strictly output according to the format of the following output template:

[0256] [Answer]: "xxxx";

[0257] Output:

[0258] Answer: "For the detection of PD-L1 protein, it is recommended to give priority to performing it on the non - detachment sections of tumor tissues in paraffin - embedded specimens to ensure that there are sufficient tumor cells for evaluation and to guarantee the detection effect. Although in some cases, such as when fresh tissue specimens cannot be obtained, other types of samples (such as cytological samples like pleural effusion and ascites) can be considered and detected by embedding them into wax blocks, these alternative methods should be carefully selected and it should be ensured that the number of tumor cells in the sample is not less than 100 to maintain the accuracy of the detection. Therefore, performing the detection not on the non - detachment sections may affect the detection effect."

[0259] In S7, the answer generated by S6 and the historical conversation records are input into an open - source large - language model, and the large - language model identifies the user's next query intention;

[0260] The specific process is as follows:

[0261] The prompt4 is input into the large - language model, and the large - language model identifies the user's next query intention;

[0262] The prompt4 includes the answer generated by S6 and the historical conversation records;

[0263] The used prompt such as prompt4:

[0264] The conversation history is as follows:

[0265] {history}

[0266] Based on the conversation history, generate 3 recommended questions related to the user's interests and context. The recommended questions should help the user explore more content related to your answer.

[0267] Please output strictly in the format of the following output template:

[0268] [Recommended Question 1]: "xxxx"

[0269] [Recommended Question 2]: "xxxx"

[0270] [Recommended Question 3]: "xxxx".

[0271] Prompt4 example:

[0272] Input:

[0273] The conversation history is as follows:

[0274] {[{"role":"user","content":"What is PD - L1?"},

[0275] {"role":"assistant","content":"PD-L1 is a protein located on the cell surface that inhibits T cell activity by binding to PD-1, helping cancer cells evade the immune system. Blocking the interaction between PD-1 / PD-L1 can reactivate the immune system and is used in cancer immunotherapy, such as PD-1 or PD-L1 inhibitors."},

[0276] {"role":"user","content":"What is an anti - detachment section?"},

[0277] {"role":"assistant","content":"An anti - detachment section is a technique in pathological or histological experiments, mainly used to prevent tissue sections from detaching from the glass slide during staining or processing. Such sections are usually prepared by coating the glass slide with a special adhesive (such as polylysine or gelatin), or by pretreating the surface of the slide to ensure that the tissue sample can firmly adhere to the glass slide, avoiding detachment or damage during the operation. The anti - detachment section technique is very important in pathological diagnosis and research, especially in experiments that require repeated processing and staining."}

[0278] {"role":"user","content":"Can the detection of PD - L1 protein be carried out without using anti - detachment sections without affecting the detection effect?"},

[0279] {"role":"assistant","content":"The detection of PD - L1 protein is preferably carried out on anti - detachment sections of tumor tissue in paraffin - embedded specimens to ensure that there are sufficient tumor cells for evaluation and to guarantee the detection effect. Although in some cases, such as when fresh tissue specimens are unavailable, other types of samples (such as cytological samples from pleural effusion or ascites) can be considered and detected by embedding them into wax blocks, these alternative methods should be carefully selected and it should be ensured that the number of tumor cells in the sample is not less than 100 to maintain the accuracy of the detection. Therefore, detection not performed on anti - detachment sections may affect the detection effect."}

[0280] Based on the conversation history, generate 3 recommended questions related to the user's interests and context. The recommended questions should help the user explore more content related to your answer.

[0281] Please strictly follow the format of the following output template:

[0282] [Recommended question 1]: "xxxx"

[0283] [Recommended question 2]: "xxxx"

[0284] [Recommended Question 3]: "xxxx".

[0285] Output:

[0286] [Recommended Question 1]: "When performing PD-L1 detection on different types of samples, what are the specific precautions?"

[0287] [Recommended Question 2]: "Besides the anti-detachment sectioning technique, what other methods can be used to improve the stability of tissue samples during the detection process?"

[0288] [Recommended Question 3]: "In cancer immunotherapy, how to select the most appropriate treatment plan based on the expression level of PD-L1?"

[0289] Each line in the above format is each line input to the large language prediction model.

[0290] Specific Embodiment Nine: A report interpretation system based on intention recognition and intelligent query guidance according to this embodiment includes:

[0291] A document parsing module, an intention recognition module, a report retrieval module, a network retrieval module, an answer generation module, and a question recommendation module.

[0292] Specific Embodiment Ten: Different from Specific Embodiment Nine, the document parsing module is used to preprocess the PDF format report file uploaded by the user, convert it into text format, segment the text format, and obtain each segmented text;

[0293] The intention recognition module is used to rewrite the question input by the user through a large model and store it in the historical conversation record;

[0294] The report retrieval module is used to retrieve the rewritten question in all the obtained text segments based on the BM25 retriever, the dense retriever, and the hybrid retriever, and use the most relevant k segments of text as relevant information;

[0295] The network retrieval module is used to input the rewritten question and the obtained relevant information into the large model. The large model determines whether network search is needed. If so, it performs network search and uses a crawler tool to crawl the returned search results and extract the web page content; if not, it executes the answer generation module;

[0296] The answer generation module is used to input the obtained relevant information, the extracted web page content, and the rewritten question into an open-source large language model. The large language model generates an answer and stores the answer in the historical conversation record;

[0297] The question recommendation module is used to input the generated answer and the historical conversation record into an open-source large language model. The large language model recognizes the user's next query intention.

[0298] Other steps and parameters are the same as those in the ninth specific implementation.

[0299] Example 1:

[0300] This example provides a report interpretation method based on intent recognition and intelligent query guidance, which can be applied to the interpretation of various field reports, such as Figure 1 As shown, it includes the following steps:

[0301] Step S1: Preprocess the PDF format report file uploaded by the user, convert it into text format, and segment the text format to obtain each segmented text.

[0302] Step S2: Perform vectorization processing on each text obtained in step S1 to obtain the text vector of each text.

[0303] Store the text vectors in the Chroma database, and convert the Chroma database after storing the text vectors into a dense retriever.

[0304] Build an inverted index based on the N texts obtained in step S1; create a BM25 retriever according to the inverted index.

[0305] Create a hybrid retriever of BM25 and the dense retriever through the EnsembleRetriever class in the langchain.retrievers library.

[0306] Step S3: Rewrite the question input by the user through a large model and store it in the historical conversation record.

[0307] Step S4: Retrieve the rewritten question in all the text segments obtained in step S1 based on step S2, and use the most relevant k texts as relevant information.

[0308] Step S5: Input the rewritten question in step S3 and the relevant information obtained in S4 into the large model. The large model determines whether retrieval is required. If so, perform a network search, and use a crawler tool to crawl the returned search results and extract the web page content; if not, execute S6.

[0309] Step S6: Input the relevant information obtained in step S4, the web page content extracted in S5, and the rewritten question in S3 into an open-source large language model. The large language model generates an answer and stores the answer in the historical conversation record.

[0310] Step S7: Input the answer generated in step S6 and the historical conversation record into the open-source large language model. The large language model recognizes the user's next query intent, and repeats the execution of S3 to S7 until the conversation ends.

[0311] Example 2:

[0312] This example provides a report interpretation system based on intent recognition and intelligent query guidance, which can be applied to the interpretation of reports in various fields, such as Figure 2 shown, including the following modules:

[0313] (1) Document parsing module, used to process PDF files into a text format that can be understood by open-source large models, including:

[0314] Front-end interaction unit, which converts Python scripts into dynamic and interactive Web applications through the Streamlit library and accepts PDF documents uploaded by users;

[0315] PDF parsing unit, which uses the PyMuPDF library to parse PDF into text format;

[0316] Text cleaning unit, which uses regular expressions (re) for cleaning to remove redundant parts such as headers and footers;

[0317] Chunk processing unit, in order to better retrieve the report, the text will be divided into multiple smaller segments. The size of each segment is usually 1024 characters, and the sliding window method is used to cut the text with a step size of 512 characters.

[0318] (2) Intent recognition module, used to convert the unclear questions input by users into clear, independently retrievable questions without pronouns, and store them in the historical conversation record, such as Figure 3 shown, including:

[0319] Ambiguity decomposition unit, which decomposes ambiguity according to the historical record and the questions input by users;

[0320] Anaphora resolution unit, which resolves anaphora in the questions after ambiguity decomposition;

[0321] Intent recognition unit, which recognizes the most reasonable intent and rewrites the questions, and stores the rewritten questions in the historical conversation record.

[0322] (3) Report retrieval module, which performs vector retrieval and sparse retrieval on the rewritten questions and PDF files respectively, and fuses the results with weights, such as Figure 4 shown, including:

[0323] Sparse retrieval unit, which uses the BM25Retriever.from_documents method in the langchain_community.retrievers library to tokenize the documents and build an inverted index, create a BM25 retriever, and calculate BM25 scores;

[0324] For the query Q = {q 1 , q 2 , …, q n} and the document D, the BM25 score calculation formula is as follows:

[0325]

[0326] Where:

[0327] Q = {q 1 , q 2 , …, q n}: The tokenized set of the query, containing n terms;

[0328] D: A certain document;

[0329] f(q i , D): The frequency of the term in document D;

[0330] |D|: The length (total number of terms) of document D;

[0331] avgdl: The average length of all documents;

[0332] k 1 and b: Hyperparameters of BM25, commonly k 1 = 1.2 and b = 0.75;

[0333] i represents the inverse document frequency of the term q

[0334]

[0335] N: The total number of documents in the document collection;

[0336] n(q i ): The number of documents containing the term q i ;

[0337] Dense retrieval unit, use the Chroma.from_documents method in the langchain_chroma library to create a Chroma database object from the documents, generate the embedding vectors of the reports through the BGE model and store them in this database object, and then use the as_retriever method to convert the database object into a dense retriever to calculate the dense retrieval score;

[0338] The dense retrieval score calculation formula is as follows:

[0339]

[0340] Where, v Q represents the embedding vector of the query Q; vD Represents the embedding vector of document D;

[0341] · Represents the dot product operation of vectors;

[0342] ||v Q || Represents the norm (vector length) of the query vector v; ||v Q ; ||v D || Represents the norm of the document vector v; D ;

[0343] Weighted fusion unit, creates a hybrid retriever of BM25 and dense retrieval through the EnsembleRetriever class in the langchain.retrievers library, combines the BM25 score and the dense retrieval score, calculates the final score, and selects the top k document fragments with the highest scores. The formula is:

[0344] score Hybrid (Q, D j ) = α · score BM25 (Q, D j ) + β · score Dense (Q, D j )

[0345] where α and β are the weight parameters for hybrid retrieval, satisfying α + β = 1.

[0346] (IV) Network retrieval module, determines the user's network retrieval intention. If needed, uses a search engine to perform a network search on the question, and uses a crawler tool to crawl the returned search results and extract the web page content, including:

[0347] Retrieval intention recognition unit, used to analyze the query input by the user, determine whether a network search is needed, and identify the retrieval topic and keywords;

[0348] Network retrieval unit, responsible for interacting with the search engine, performing a search based on the keywords, and returning relevant search results;

[0349] Content crawling unit, extracts the specified web page content from the search results, performs structured processing on it, and extracts the information that the user cares about, such as text, pictures, or data tables.

[0350] (V) Answer generation module, inputs the retrieved report, network knowledge, and the rewritten question into an open-source large language model to generate an answer, and stores the answer in the historical conversation record.

[0351] (VI) Question recommendation module, based on the historical conversation, identifies the user's next query intention and gives recommended questions.

[0352] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should fall within the protection scope of the appended claims of the present invention.

Claims

1. A report interpretation method based on intention recognition and intelligent query guidance, characterized in that: The specific process of the method is: Step S1, pre-processing the report file in PDF format uploaded by the user, converting it into text format, segmenting the text format, and obtaining each segmented text; Step S2, respectively performing vectorization processing on each paragraph of text obtained in step S1 to obtain a text vector for each paragraph of text; The text vector is stored in a Chroma database, and the Chroma database after the text vector is stored is converted into a dense search engine; Construct an inverted index based on the N text segments obtained in step S1; create a BM25 searcher based on the inverted index; Create a hybrid retriever of BM25 and dense retriever through the EnsembleRetriever class in the langchain.retrievers library; Step S3: rewrite the question input by the user through the large model and store it in the historical conversation record; Step S4, based on step S2, searching the rewritten question in all the text segments obtained in step S1, and taking the most relevant k text segments as relevant information; Step S5, input the rewritten question in step S3 and the relevant information obtained in step S4 into the big model, and the big model determines whether retrieval is required. If so, a network search is performed, and the returned search results are crawled with a crawler tool to extract the web page content; if not, execute S6; Step S6: input the relevant information obtained in step S4, the webpage content extracted in step S5, and the question rewritten in step S3 into the open source large language model, and the large language model generates an answer, and stores the answer in the historical conversation record; Step S7: Input the answer generated in step S6 and the historical conversation record into the open source large language model. The large language model identifies the user's next query intention and repeats steps S3 to S7 until the conversation ends.

2. A report interpretation method based on intention recognition and intelligent query guidance according to claim 1, characterized in that: The step S1 pre-processes the report file in PDF format uploaded by the user, converts it into text format, and segments the text format to obtain segmented text; specifically includes: Step S11, converting the Python script into a Web application through the Streamlit library; Step S12: Optimize the Web application through CSS and JavaScript to obtain an optimized Web application, and the Web application accepts the PDF document uploaded by the user; Step S13: Use the PyMuPDF library to parse the PDF document uploaded by the user into text format; Perform data cleaning on the text to obtain the cleaned text; Step S14: Divide the text cleaned in S13 into N segments, where N is a positive integer and each segment represents D j , j = 1, 2, ..., N; the specific process is: The cleaned text is divided into N segments using the sliding window method. The size of each segment is 1024 characters, and the character overlap between adjacent segments is 512. Each segment of text represents D j .

3. A report interpretation method based on intention recognition and intelligent query guidance according to claim 2, characterized in that: In the step S2, each paragraph of text obtained in the step S1 is vectorized to obtain a text vector for each paragraph of text; The text vector is stored in a Chroma database, and the Chroma database after the text vector is stored is converted into a dense search engine; Construct an inverted index based on the N text segments obtained in step S1; create a BM25 searcher based on the inverted index; Create a hybrid retriever of BM25 and dense retriever through the EnsembleRetriever class in the langchain.retrievers library; The specific process is: Step S21, use the HuggingFaceBgeEmbeddings class in the langchain_community.embeddings library to load the BGE embedding model; Step S22, using the Chroma.from_documents method in the langchain_chroma library to create a Chroma database; Each paragraph of text obtained in step S1 is input into the BGE embedding model, and the BGE embedding model outputs the text vector of each paragraph of text; Store the text vector into the Chroma database; Convert the Chroma database stored in text vectors into a dense search engine; Step S23, use the BM25Retriever.from_documents method in the langchain_community.retrievers library to segment each text segment; Use the BM25Retriever.from_documents method in the langchain_community.retrievers library to build an inverted index for the N segments of text after word segmentation, and create a BM25 retriever based on the inverted index; Step S24: Create a hybrid retriever of BM25 and dense retriever through the EnsembleRetriever class in the langchain.retrievers library.

4. The report interpretation method based on intention recognition and intelligent query guidance according to claim 3 is characterized by: In step S3, the question input by the user is rewritten by using the large language model, and the rewritten question is stored in the historical conversation record; specifically, it includes: Step S31: Decompose the question input by the user through the large language model to obtain the question after the decomposition; Step S32: The large language model performs reference resolution on the question after the ambiguous decomposition to obtain the question after reference resolution; Step S33: The large language model identifies the most reasonable question based on the question after the coreference resolution, rewrites the question based on the most reasonable question, and stores the rewritten question in the historical conversation record.

5. The report interpretation method based on intention recognition and intelligent query guidance according to claim 4 is characterized by: In step S4, the rewritten question is searched in all the text segments obtained in step S1 based on step S2, and the most relevant k segments of text are used as relevant information; The specific process is: Step S41, based on the rewritten question Q = {q1, q2, ..., q n } and the jth paragraph of text D j , calculate the BM25 retriever score BM25 (Q,D j ); Step S42: Based on the rewritten question Q = {q1, q2, ..., q n } and the jth paragraph of text D j , calculate the dense retrieval score score Dense (Q,D j ); Step S43: BM25 searcher score obtained in step S41 BM25 (Q,D j ) and the dense search score score obtained in step S42 Dense (Q,D j ), calculate the hybrid retriever score; Step S44, select k text segments corresponding to the k highest hybrid retriever scores as relevant information.

6. A report interpretation method based on intention recognition and intelligent query guidance according to claim 5, characterized in that: The step S41 is based on the rewritten question Q={q1,q2,…,q n } and the jth paragraph of text D j , calculate the BM25 retriever score BM25 (Q,D j ); the formula is as follows: in: Q={q1,q2,…,q n } represents the rewritten question, q1 represents the first word in the rewritten question, q2 represents the second word in the rewritten question, n It represents the nth word in the rewritten question, which contains n terms; D j represents the jth segment of text obtained in step S1, j = 1, 2, ..., N; f(q i ,D j ) indicates that the word is in the jth paragraph of text D j The frequency of occurrence in |D j | indicates the jth paragraph of text D j Length; avgdl represents the average length of the text of all segments; k1 and b represent the hyper parameters of the BM25 retriever, respectively; Representation word q i The inverse document frequency of ; expressed as: in: N represents the total number of text segments obtained in step S1; n(q i ) means it contains word q i The number of text segments, i = 1, 2,…, n.

7. A report interpretation method based on intention recognition and intelligent query guidance according to claim 6, characterized in that: The step S42 is based on the rewritten question Q={q1,q2,…,q n } and the jth paragraph of text D j , calculate the dense retrieval score score Dense (Q,D j ); the formula is as follows: in, v Q The rewritten problem Q = {q1,q2,…,q n }'s text vector; Denotes the jth paragraph of text D j The text vector of Represents the dot product operation of vectors; ||v Q || represents the text vector v Q The norm of the text vector; Denotes the jth paragraph of text D j The norm of the text vector; The rewritten problem Q = {q1, q2, ..., q n }'s text vector v Q The acquisition process is: The rewritten question Q is input into the BGE embedding model, and the BGE embedding model outputs the text vector of the rewritten question Q; The jth paragraph of text D j Text vector The acquisition process is: The jth paragraph of text D j Input BGE embedding model, BGE embedding model outputs the jth text D j Text vector.

8. The report interpretation method based on intention recognition and intelligent query guidance according to claim 7 is characterized by: The step S43 is based on the BM25 searcher score obtained in step S41. BM25 (Q,D j ) and the dense search score score obtained in step S42 Dense (Q,D j ), calculate the hybrid retriever score; the formula is: score Hybrid (Q,D j )=α×score BM25 (Q,D j )+β×score Dense (Q,D j ) Among them, α and β are weight parameters of the hybrid retriever, satisfying α+β=1.

9. A report interpretation system based on intention recognition and intelligent query guidance, characterized in that: The system comprises: Document parsing module, intention recognition module, report retrieval module, network retrieval module, answer generation module, question recommendation module.

10. The report interpretation system based on intention recognition and intelligent query guidance according to claim 9, characterized in that: The document parsing module is used to pre-process the report file in PDF format uploaded by the user, convert it into text format, segment the text format, and obtain each segmented text; The intent recognition module is used to rewrite the questions input by the user through the large model and store them in the historical conversation records; The report retrieval module is used to retrieve the rewritten question in all the obtained text segments based on the BM25 retriever, dense retriever and hybrid retriever, and take the most relevant k segments of text as relevant information; The network retrieval module is used to input the rewritten questions and the obtained related information into the big model. The big model determines whether retrieval is needed. If so, it will conduct a network search and crawl the returned search results with a crawler tool to extract the web page content. If not, the answer generation module will be executed. The answer generation module is used to input the acquired relevant information, extracted webpage content and rewritten questions into the open source large language model. The large language model generates answers and stores the answers in the historical conversation records. The question recommendation module is used to input the generated answers and historical conversation records into the open source large language model, which identifies the user's next query intention.