System for answering medical questions
The medical question answering system addresses inaccuracies and resource inefficiencies by using a two-stage search and hallucination detection, ensuring timely and accurate responses through a user interface that distinguishes between clinical and non-clinical sources.
Patent Information
- Application Number
- DE202025100184
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2035-01-31
AI Technical Summary
Existing question answering systems, particularly in the medical domain, often generate inaccurate or outdated answers due to reliance on outdated information and lack of scientific reliability, and they consume excessive computational resources.
A medical question answering system that employs a two-stage search process involving coarse filtering based on embeddings and fine-grained classification using neural networks, combined with hallucination detection and a user interface that distinguishes between responses from clinical practice guidelines and other documents, to provide timely and accurate answers.
The system generates scientifically reliable and computationally efficient answers by leveraging up-to-date medical databases, reducing computational resources and minimizing hallucinations, suitable for real-world clinical decision-making.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No. 63 / 695,309, filed September 16, 2024, U.S. Application No. 18 / 975,621, filed December 10, 2024, and U.S. Application No. 18 / 975,915, filed December 10, 2024. The disclosure of each of the prior applications is considered part of the disclosure of the present application and is incorporated herein by reference. BACKGROUND
[0002] This specification concerns question answering (QA). Question answering is an area of computer technology that attempts to automatically provide answers to questions entered by humans, often in natural language format.
[0003] For example, in response to the input question "What is POTS?" a medical question answering system may use one or more machine learning models to process the input question and possibly other text data and return the answer beginning with "Postural orthostatic tachycardia syndrome (POTS) is a chronic disorder of the autonomic nervous system characterized by orthostatic intolerance and an excessive increase in heart rate upon standing."
[0004] Machine learning models receive an input and generate an output, such as a predicted output, based on the received input. Some machine learning models are parametric models and generate output based on the received input and the values of the model's parameters.
[0005] Some machine learning models are deep models, which use multiple layers of models to produce an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a nonlinear transformation to a received input to produce an output. SUMMARY
[0006] This specification describes a medical question answering system, implemented as computer programs on one or more computers at one or more locations, that receives a medical question and uses neural networks and other components of the system to generate an answer to the medical question.
[0007] To generate the answer, the medical question answering system searches a medical database containing medical documents to obtain a set of document snippets from the medical database that are relevant to the medical question. The medical question answering system then integrates the set of relevant document snippets into a prompt before processing the prompt using a generative neural network to generate the answer to the medical question.
[0008] According to one aspect, a method is provided performed by one or more computers, the method comprising: obtaining question data representing a medical question; obtaining a plurality of document snippets from a medical database storing medical documents, comprising performing a search in the medical database for document snippets relevant to the medical question using: (i) a respective embedding of each of the plurality of document snippets, and (ii) an embedding of the medical question; for each document snippet in the plurality of document snippets, determining a relevance score for the document snippet by using a ranking neural network based on the document snippet and the medical question;selecting a subset of the plurality of document excerpts based at least in part on the relevance scores; generating a prompt comprising (i) the medical question and (ii) the subset of the plurality of document excerpts; and generating a response to the medical question based on processing the prompt using a generative neural network.
[0009] Determining, for each document snippet in the plurality of document snippets, the relevance score for the document snippet may comprise: processing, by the ranking neural network, an input comprising: (i) the medical question and (ii) the document snippet to generate a respective score for each of a plurality of different relevance levels, wherein for each relevance level, the respective score indicates a probability that a relevance between the medical question and the document snippet has the relevance level; and determining the relevance score for the document snippet based on the respective score for each of the plurality of relevance levels.
[0010] Determining the relevance score for the document excerpt may include determining a linear combination of the respective scores.
[0011] Selecting the subset of the plurality of document snippets based at least in part on the relevance scores may comprise selecting, from the plurality of document snippets and as an initial subset of the plurality of document snippets, highest-scoring document snippets having the highest relevance scores; for each of a plurality of aspects: assigning a respective weight to each document snippet in the initial subset of the plurality of document snippets with respect to the aspect; and selecting one or more document snippets from the document snippets in the initial subset of the plurality of document snippets based on the respective weight assigned to each document snippet.
[0012] The plurality of aspects may include one or more of the following: a timeliness of a medical document that includes the document excerpt, a quality of a provider of the medical document, or a relevance between a user who asked the medical question and an author of the medical document.
[0013] Selecting the subset of the plurality of document snippets based at least in part on the relevance scores may include: combining the one or more document snippets selected for each of the plurality of aspects to generate a combined set of document snippets; and applying semantic filtering, deduplication, or both to the combined set of document snippets to generate the subset of the plurality of document snippets.
[0014] The request may further comprise one or both of the following: (iii) metadata associated with each document snippet in the subset of the plurality of document snippets, or (iv) predetermined instructions.
[0015] The answer to the medical question may include (i) an answer generated by the generative neural network from processing the prompt, (ii) a bibliography of medical documents comprising the subset of the plurality of document excerpts, and (iii) data providing a rationale for the said medical documents.
[0016] The reasons may include, for each document snippet in the subset of the plurality of document snippets, one or more of the following: a similarity statement, an impact statement, or a timeliness statement.
[0017] Generating the answer to the medical question may comprise: processing at least the answer using a hallucination detection neural network to generate one or more hallucination detection outputs indicating (i) whether there is a contradiction between (a) the answer generated by the generative neural network from processing the prompt and (b) the subset of the plurality of document snippets, (ii) whether at least one document snippet in the subset of the plurality of document snippets supports the answer generated by the generative neural network from processing the prompt, or both (i) and (ii).
[0018] The method may further comprise: determining that a contradiction exists; and in response, modifying the response generated by the generative neural network to generate a modified response.
[0019] Obtaining question data representing a medical question may comprise: receiving an initial user input; and preprocessing the initial user input to generate the medical question, wherein the medical question (i) is in a predetermined natural language, (ii) has a question format, (iii) resolves any acronyms or abbreviations in the initial input, and (iv) replaces any brand names in the initial input with generic names.
[0020] The classification neural network may be trained based on an optimization of a supervised learning objective function on a classification training dataset comprising a plurality of classification training pairs, each classification training pair comprising (i) a medical question and a document snippet, and (ii) being associated with a ground truth score for each of the plurality of different relevance levels.
[0021] Obtaining the plurality of document snippets may comprise: for each document snippet, determining a distance between (i) a respective embedding of the document snippet and (ii) the embedding of the medical question; and selecting document snippets as the plurality of document snippets based on the distances.
[0022] For each of the multitude of document excerpts, the respective embedding can be generated by an embedding model from the processing of the document excerpt.
[0023] The embedding model can be trained based on an optimization of an objective function for contrastive learning on an embedding training dataset comprising a plurality of embedding training pairs, each embedding training pair comprising a medical question and a document snippet.
[0024] For each embedding training pair, the document snippet can be extracted from a medical source document, and the medical question can be generated by a neural language model network based on processing the document snippet.
[0025] Two or more of the plurality of embedding training pairs may comprise a same document fragment, and the objective function for contrastive learning may comprise a term that penalizes the embedding model for generating respective embeddings of medical questions contained in the two or more of the plurality of embedding training pairs that differ from each other.
[0026] According to another aspect, a method is provided performed by one or more computers, the method comprising: receiving a query for medical information from a user and by means of a user interface presented to the user on a display of a user device; generating a response to the query from the user by automatically retrieving and parsing data from a document corpus, comprising: determining, based on an automated search of the document corpus, that one or more clinical practice guideline documents from the document corpus include information answering the query; in response to determining that one or more clinical practice guideline documents include information answering the query, generating a first response to the query based only on clinical practice guideline documents;and generating a second response to the query based at least in part on one or more other documents that are not clinical practice guideline documents; and presenting, by means of the user interface and on the display of the user device: a first user interface element presenting the first response generated based only on clinical practice guideline documents, the first user interface element visually highlighting that the first response is derived only from clinical practice guideline documents and identifying one or more clinical practice guideline documents that were processed to generate the first response; and a second user interface element presenting the second response generated at least in part on documents that are not clinical practice guideline documents.
[0027] The first user interface element may be displayed above and before the second user interface element within the user interface.
[0028] The first user interface element may comprise a walled garden environment presented within the user interface, and the first generated response, which may be based only on the clinical practice guideline documents, is presented within the walled garden environment.
[0029] The first user interface element may include a header indicating that the first response was created based only on clinical practice guideline documents.
[0030] Generating the first response to the query based only on clinical practice guideline documents may comprise making a first call to a generative neural network, wherein generating the second response to the query based at least in part on one or more other documents that are not clinical practice guideline documents may comprise making a second call to the generative neural network after the generative neural network has generated the first response in response to the first call.
[0031] Generating the first response to the query based only on clinical practice guideline documents may comprise: processing, by the generative neural network, a first request comprising (i) the medical information query and (ii) the clinical practice guideline documents to generate the first response.
[0032] Generating the second response to the query based at least in part on one or more other documents that are not clinical practice guideline documents may comprise: processing, by the generative neural network, a second prompt comprising (i) the medical information query, (ii) the one or more other documents, and (iii) the first response generated by the generative neural network to generate the second response.
[0033] According to a further aspect, one or more computer-readable storage media are provided encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the steps of the above method aspects.
[0034] According to a further aspect, a system is provided comprising one or more computers and one or more storage devices storing instructions that, when executed by one or more computers, cause the one or more computers to perform the respective steps of the above method aspects.
[0035] The subject matter described in this specification may be implemented in certain embodiments to realize one or more of the following advantages.
[0036] Vast amounts of expert knowledge are stored in databases, e.g., in the form of medical documents describing clinical trials, academic journal articles describing scientific and medical research, and the like. By performing the two-step process described in this specification to retrieve relevant and timely information from a database in a computationally effective and accurate manner, and then causing a generative neural network to generate answers in response to questions by leveraging the retrieved information (rather than merely outdated information available at the time the generative neural network was trained), the medical question-answering system can generate answers that are more accurate and timely.
[0037] Compared to existing question answering systems, the medical question answering system described in the specification is capable of generating answers that are much more likely to be scientifically reliable and factually correct, e.g. answers that are more likely to be based on external, vetted medical databases or other bodies of knowledge, such as published treatises, medical guidelines or established reference materials that are continually kept up to date, and at the same time are much less likely to include hallucinations as part of the answers.
[0038] The medical question answering system is particularly beneficial in a variety of real-world scenarios that require scientifically sound clinical decision-making. For example, the medical question answering system can generate answers that help a physician in an active environment (e.g., a treatment room) quickly and efficiently answer free-form questions relevant to a patient's current and future treatment, such as dosages, contraindications, side effects, drug interactions, disease symptoms, possible second- and third-line treatments, safety aspects, etc.
[0039] Viewed from another perspective, the medical question answering system acts as an interface between a user and a large corpus of medical documents that would otherwise be too large to be searched in response to any question the user might pose regarding the medical documents. This interface enables the user to practically utilize the information contained in the unmanageably large corpus of medical documents to retrieve a set of reliable and, most importantly, provably accurate information.
[0040] Various methods described in this specification enable the medical question answering system to achieve these benefits with a reduced use of computational resources. First, through the two-stage search process, which includes a relatively less computationally intensive first stage followed by a relatively more computationally intensive second stage, as described in this specification, the medical question answering system can generate the answers with a reduced use of computational resources such as memory and processing power. In other words, the two-stage search process described in this specification makes it computationally feasible to search a huge corpus of documents.
[0041] Specifically, the first stage of the two-stage search process involves coarse filtering based on embeddings and is less computationally intensive, as it involves computing a medical question embedding and then measuring similarities between the medical question embedding and several previously computed medical snippet embeddings that can be reused for other medical question embeddings. The second stage of the two-stage search process, on the other hand, involves performing fine-grained classification using a classification neural network and is more computationally intensive, as it requires processing each of several document snippets using the classification neural network.
[0042] Second, many existing question-answering systems frequently generate answers that contain hallucinations. In contrast, the described methods related to hallucination detection improve the possibility that answers generated by the medical question-answering system are scientifically reliable, factually correct, and free of hallucinations by performing hallucination detection at various levels, e.g., sentence level, paragraph level, etc., of the content contained in the generated answer based on the content of the medical documents stored in the medical database. The medical question-answering system described in this specification is therefore suitable for use in production environments, e.g., in an educational or medical institution, where false or misleading information can have serious consequences.
[0043] Second, the described methods, which relate to contrastive training, improve the efficiency of computational resources when training a neural embedding model network used to enable the two-stage search process. In contrast to many conventional contrastive training approaches, which require an equal or approximately equal number of medical question embeddings and document snippet embeddings to be generated during training, the specification describes a training approach in which the neural embedding model network is designed to generate more medical question embeddings than document snippet embeddings.Since a medical question may be shorter than a document snippet, generating an embedding of the medical question can save computational costs compared to generating an embedding of the document snippet because a smaller amount of data needs to be processed by the neural embedding model network.
[0044] The savings in computational costs can be substantial when embeddings of only a relatively small number of document snippets relative to the number of medical questions need to be generated, e.g., generating one document snippet embedding for every 3, 5, or 10 medical question embeddings, rather than generating one document snippet embedding for each individual medical question embedding.
[0045] Second, the described methods related to the improved user interface can help users more clearly distinguish between responses generated solely based on clinical practice guideline documents and those not, and provide insights to improve review efficiency. The described methods also represent an improvement in user interface technology.By generating an answer based on the clinical practice guidelines and presenting it to the user with a corresponding visual display, and then, after presenting the answer to the user, generating another answer based on the other medical documents for presentation to the user at a later time with a different visual display, the described methods reduce latency in delivering and presenting content that satisfies the user's information needs and minimize potential confusion regarding different answers to the same medical question generated based on different source documents.
[0046] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a diagram of an example medical question answering system. Fig. Figure 2 is a representation of an example medical database. Fig. Figure 3 is a representation of example operations performed by a planner. Fig. Figure 4 is a flowchart of an example process for generating an answer to a medical question. Fig. 5 is a flowchart of substeps of one of the steps of the process of Fig. 4 to determine relevance ratings for a document excerpt. Fig. 6 is a flowchart of substeps of another of the steps of the process of Fig. 4 to select a suitable subset from a variety of document excerpts. Fig. Figure 7 is a flowchart of an example process for generating a response to a user query. Fig. Figure 8 is an illustration of an example of a user interface.
[0055] Like reference numerals and labels in the various drawings indicate like elements. SUMMARY
[0047] Fig. 1 shows an example medical question answering system 100. The medical question answering system 100 is an example of a system implemented as computer programs on one or more computers at one or more locations that receives a medical question 102 and uses neural networks and other components of the system to generate an answer 126 to the medical question.
[0048] The medical question answering system 100 may obtain data representing the medical question 102 in various ways. In some cases, the data representing the medical question 102 is also referred to as "question data." In some cases, the medical question 102 is also referred to as a "query," "prompt," or "question."
[0049] In some cases, the question data comprises text data, and the medical question answering system 100 receives the question data as a text query from a user submitted via a user interface of a user device. For example, the medical question 102 may be entered by the user by typing using a data input device, such as a keyboard, a touchscreen, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart speaker), or another portable or desktop input device.
[0050] In some other cases, the question data includes audio data, and the medical question answering system 100 receives the question data as a natural language voice query from the user and converts the speech into the medical question 102 by applying a speech recognition engine to the speech. For example, the medical question 102 may be received in the form of a sound (speech) signal captured by a microphone of the user's computer and converted into the medical question 102 by a speech recognition engine, i.e., a speech-to-text converter.
[0051] The medical question answering system 100 has access to a medical database 150. The medical database 150 may be any database that stores a corpus of medical documents in the form of document clippings. That is, the medical database 150 stores a plurality of document clippings included in the corpus of medical documents.
[0052] A medical document is an electronic document that contains medical content. Examples of medical documents include web pages, word processing documents, spreadsheet documents, presentation documents, PDF (Portable Document Format) documents, and so on.
[0053] In practice, the medical database 150 may include a large number of medical documents, e.g., at least 100,000 medical documents, or at least 1,000,000 medical documents, or at least 10,000,000 medical documents. Furthermore, the medical database 150 may change dynamically over time, e.g., when new medical documents are added or medical documents are removed (e.g., when new clinical practice guidelines replace old ones, or when published papers are retracted, etc.).
[0054] The corpus of medical documents may, for example, include web pages and other electronic documents accessible via the internet. Additionally or alternatively, the corpus of medical documents may, for example, be part of a proprietary medical database, e.g., of a scientific content publisher (e.g., a physical science content publisher, a life science content publisher, a health science content publisher, or a social science content publisher), a technical content publisher, a medical content publisher, or another organization. Optionally, the corpus of medical documents may also include clinical practice guideline documents.
[0055] A clinical practice guideline document is a document that includes information about clinical practice guidelines. Clinical practice guidelines are statements with recommendations for optimizing patient care based on a systematic review of evidence and an assessment of the benefits and harms of alternative treatment options. Clinical practice guidelines can be used, for example, to support individual clinical decisions, to provide best practice recommendations in the treatment and care of people by healthcare professionals, to develop standards to guide and assess the clinical practice of individual healthcare professionals and healthcare organizations, to contribute to the education and training of healthcare professionals, and to help patients make informed decisions.
[0056] A "document snippet" includes at least a portion of the content of an entire electronic document. In some implementations, the length of the plurality of document snippets may vary. For example, each document snippet may correspond to a paragraph, page, or chapter of an electronic document.
[0057] In some implementations, the plurality of document snippets may each be approximately the same length. For example, a document snippet may include 100 words, 300 words, 500 words, or the like. As another example, each document snippet may include approximately 100 partial words, 300 partial words, or 500 partial words, or the like. A partial word is generally an incomplete word, although there may also be partial words that correspond to complete words in a vocabulary. For example, the word "certainly" may include a partial word "surely" and a partial word "likely."
[0058] In some implementations, the medical question answering system (100) or other database system may generate the plurality of document snippets by applying a sliding window method to each medical document in the corpus of medical documents to extract every possible sequence of a predetermined fixed length of words or subwords from the medical document. The extracted sequences of words or subwords may then be stored as document snippets in the medical database 150.
[0059] Fig. Figure 2 is a representation of an example of the medical database 150.
[0060] As illustrated, the medical database 150 may store document snippets from a variety of guideline documents, e.g., clinical practice guideline documents or regulatory guideline documents containing information about policies, standards, regulations, etc. The medical database 150 may store document snippets contained in a variety of clinical trial documents containing clinical information about various entities, including disease entity, drug name, line of therapy, etc. The medical database 150 may store document snippets contained in medical labeling documents, e.g., drug labels, which include labels for drugs approved by the Food and Drug Administration (FDA) or other government agency.The medical database 150 may store document excerpts contained in a variety of government communications, such as copies of notices issued by the Centers for Disease Control and Prevention (CDC) or other government agencies. The medical database 150 may also store document excerpts contained in a variety of clinical research documents, such as scientific or clinical research publications.
[0061] After receiving the medical question 102, the medical question answering system 100 uses a scheduler 110 to orchestrate the operations performed by the neural networks and the other components of the system when generating the answer 126 in response to the medical question 102.
[0062] Fig. 3 is an illustration of example operations performed by the scheduler 110.
[0063] In some implementations, the medical question answering system 100 includes or has access to multiple candidate generative neural networks, and the planner 110 may act as a module selector. The module selector may select one of the multiple candidate generative neural networks as the generative neural network 120 to generate the answer 126 in response to the medical question 102. The module selector may use any suitable approach to select between the multiple candidate generative neural networks.
[0064] For example, the module selector may make this selection based on the medical question, e.g., based on a length, a reception time, or another aspect of the medical question. As another example, the module selector may make this selection based on the availability or utilization of the multiple candidate generative neural networks.
[0065] As another example, the module selector may make this selection based on the availability of the computational resources (e.g., processing power, memory and other storage, bandwidth, etc.) of the medical question answering system 100 - for example, by selecting a candidate generative neural network that has a larger memory footprint and a higher processing power requirement when a larger amount of computational resources is available, while selecting a candidate generative neural network that has a smaller memory footprint and a lower processing power requirement when only a smaller amount of computational resources is available.
[0066] Typically, a candidate generative neural network with more parameters or a more complex architecture, e.g., with more layers, has a larger memory footprint and higher processing power requirements than another candidate generative neural network with fewer model parameters or a less complex architecture, e.g., with fewer layers.
[0067] The availability of computing resources may vary depending on the total number of medical questions received simultaneously, the geographic regions from which the medical questions were received, or other factors. In some particular examples, the module selector may select a lightweight generative neural network candidate if more than a threshold number of medical questions were received during a given time window, or select a lightweight generative neural network for medical questions received from users in geographic regions more than a threshold distance from the medical question answering system 100.
[0068] In some implementations, the scheduler 110 may act as a gatekeeper. The gatekeeper ensures that the medical question answering system 100 does not waste computing resources answering irrelevant queries that the system may receive from various users.
[0069] To this end, the gatekeeper may parse question data submitted to system 100 to determine whether the question data contains a medical question, i.e., whether a question submitted by a user relates to a medical topic. For example, the gatekeeper may use generative neural network 120 or another machine learning model for text classification to process the question data and generate a classification output indicating whether a question represented by the question data is a medical question or not.
[0070] In cases where the gatekeeper determines that the question data does not contain a medical question, it could reject the question—for example, it could provide the user with a predetermined answer without using other components of system 100 to further process the question data. For example, the gatekeeper can reject the question "What is 2+2?" with the predetermined answer "The question is outside the scope of the medical question-answering system."
[0071] The gatekeeper may also convert the medical question submitted by the user, which may be in an initial form (e.g., as free text), into a refined form that is more suitable for further processing by the other components of the system 100.
[0072] In some implementations, the gatekeeper may remove erroneous content such as typos or grammatical errors from the medical question in its initial form. In some implementations, the gatekeeper may resolve acronyms, abbreviations, and other shorthand contained in the medical question in its initial form. In some implementations, the gatekeeper may replace certain words or phrases contained in the medical question in its initial form with more precise medical terminology. For example, the gatekeeper may convert a medical question in the initial form, "How does the keto diet affect diabetes?" into a medical question in the refined form, "How does the ketogenic diet affect type 2 diabetes?"
[0073] In some implementations, in response to receiving the medical question in the initial form, the gatekeeper may preprocess the medical question in the initial form to generate the medical question in a predetermined refined form. Specifically, the medical question in the predetermined refined form (i) is in a predetermined natural language, (ii) has a question format, (iii) resolves any acronyms or abbreviations that may be included in the medical question in the initial form, and (iv) replaces any brand names that may be included in the medical question in the initial form with generic names.
[0074] The gatekeeper may do this, for example, by transforming the medical question from the initial form to the predetermined refined form based on a predetermined transformation template. As another example, the gatekeeper may do this by using the generative neural network 120 or another neural network—for example, by using the generative neural network 120 to process an input comprising the medical question in the initial form and data defining the predetermined refined form to produce an output comprising the medical question in the predetermined refined form.
[0075] In some implementations, the planner 110 may function as a translator, e.g., an omni-translator that translates text in various natural source languages into text in a common natural target language. For example, the translator may translate a medical question in the form of a phrase in a natural source language into a medical question in the form of a phrase in a natural target language. In this example, the answer 126 may be in the form of another phrase, also in the natural source language.
[0076] In some implementations, the medical question answering system 100 includes or has access to multiple external tools, and the planner 110 may act as a tool selector. The multiple external tools are separate from the generative neural network 120 and, in some implementations, separate from the system 100, e.g., remote.
[0077] An external tool can, in principle, be any software function that can be queried by the medical question answering system 100, e.g., through API (application programming interface) calls, to provide data in response to a query. Examples of these external tools include a calculator tool (e.g., a medication dosage calculator tool) and a calendar tool, to name a few.
[0078] The tool selector may select one or more of the multiple external tools to generate responses that may be incorporated into the response 126 to be generated by the generative neural network 120 in response to the medical question 102. Like the module selector, the tool selector may use any suitable approach to select between the multiple external tools, e.g., based on the medical question, the availability or use of the multiple external tools, etc.
[0079] In some implementations, the planner 110 may act as an ontology incorporator. The ontology incorporator maps specific words or phrases in the medical question to nodes in an ontology graph, which is used to augment a retrieval process to be performed by a retrieval engine 130 by providing the retrieval engine 130 with access to synonymous, hyponymous, or other labels. The ontology graph includes a set of nodes connected by edges. The nodes of the ontology graph may represent medical concepts, such as the names of medications or medical conditions. The edges of the ontology graph may represent relationships between medications, such as "is a precursor to," "is equivalent to," "is a version of," or other relationships.
[0080] In some implementations, the planner 110 may function as a question intent analyzer, which may augment the retrieval process to be performed by the retrieval engine 130. The question intent analyzer may use the generative neural network 120 or another text classification model to determine an intent of the medical question and then, based on the intent, determine which set(s) of document snippets should be searched in the medical database 150.
[0081] For example, in response to determining that an intent of the medical question relates to clinical trial purposes, the question intent analyzer may determine that document excerpts contained in clinical trial documents should be searched. As another example, in response to determining that an intent of the medical question relates to academic research purposes, the question intent analyzer may determine that document excerpts contained in academic journal articles should be searched.
[0082] The question intent analyzer can thus reduce the computational resource consumption of the medical question answering system 100 by avoiding an exhaustive search of the entire medical database 150. In some implementations, the question intent analyzer can be implemented as a classification model, e.g., a classification neural network configured to process the medical question to generate a classification output comprising a score distribution across a set of possible document categories. Accordingly, only documents in some of the categories of the set whose score meets a threshold need to be searched.For example, the set(s) of document excerpts that need to be searched in response to the medical question based on the classification results may be less than 50%, less than 10%, or less than 5% of the entire medical database.
[0083] To generate the answer 126 in response to the medical question 102, the medical question answering system 100 uses the retrieval engine 130 to search the medical database 150 to obtain, based on the medical question 102, a smaller subset of document excerpts from the medical database 150 that are relevant to the medical question 102.
[0084] The medical question answering system 100 then incorporates the smaller subset of document snippets into a prompt 118 before processing the prompt 118 using the generative neural network 120 to generate the answer 126 to the medical question 102.
[0085] For example, the medical question answering system 110 may generate a prompt 118 that includes the medical question and the smaller subset of document snippets, and then cause the generative neural network 120 to generate the answer 126 to the medical question 102 based on the processing of the prompt 118.
[0086] Optionally, the prompt 118 also includes a predetermined set of system instructions. For example, the predetermined set of system instructions may be represented by natural language text, e.g., text describing what can and cannot be included in the response 126. Optionally, the prompt 118 also includes the results obtained from one or more external tools, e.g., a dosage amount calculated by a medication dosage calculator tool. Optionally, the prompt 118 further includes one or more historical medical questions received by the system 100 and one or more responses generated by the system 100 in response to the historical medical questions. Optionally, the prompt 118 further includes metadata associated with each document snippet in the subset of the plurality of document snippets.The metadata may, for example, include the publication dates of the medical documents from which the document excerpts in the subset are obtained.
[0087] In this way, Answer 126 will include information contained in the smaller subset of document snippets, including actual information that was not available during training of the generative neural network and / or proprietary information that was not publicly available and therefore may have been excluded from the training data used to train the generative neural network, thereby improving the quality, e.g., usefulness, factual accuracy, completeness, timeliness, or some combination thereof, of Answer 126 with respect to Medical Question 102.
[0088] The generative neural network 120 may be or include a (large) language model that has been trained to receive an input sequence of tokens selected from a vocabulary and auto-regressively generate an output sequence of tokens from the vocabulary. For example, the input sequence may represent the prompt 118, while the output sequence may represent the answer 126 to the medical question 102.
[0089] The token vocabulary may contain any of a variety of tokens representing text symbols or other symbols. For example, the token vocabulary may include one or more of the characters, subwords, words, punctuation marks, numbers, or other symbols found in a corpus of natural language text.
[0090] For example, the language model may comprise any of a variety of different transformer-based neural network architectures, e.g., pure encoder-transformer architectures, encoder-decoder-transformer architectures, pure decoder-transformer architectures, other attention-based architectures, etc. As another example, the language model may comprise any of a variety of recurrent neural network architectures.
[0091] Example implementations of such a language model are described in more detail in Anil, Rohan, et al., "Palm 2 technical report." arXiv preprint arXiv:2305.10403; and Touvron, Hugo, et al., "Llama 2: Open foundation and fine-tuned chat models." arXiv preprint arXiv:2307.09288 (2023), Jiang, Albert Q, et al., "Mistral 7B." arXiv preprint arXiv:2310.06825 (2023), but others may also be used.
[0092] More specifically, the auto-regressively generated output sequence is generated by generating each particular token in the output sequence under the condition of a current input sequence that includes any tokens that precede the particular text token in the output sequence, i.e., the tokens already generated for all previous positions in the output sequence that precede the particular position of the particular token, and the tokens included in the prompt 118.
[0093] To generate a specific token at a specific position within an output sequence, the generative neural network 120 may process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the token vocabulary. The generative neural network 120 may then select a token from the vocabulary as the specific token based on the score distribution. For example, the neural network of the generative neural network 120 may select the token with the highest score using a greedy approach or pick a token from the distribution, e.g., using nucleus sampling or another sampling method.
[0094] In many scenarios, the medical database 150 may be a large database storing a very large number, e.g., ten million, one hundred million, one billion, or more, of medical documents and thus an even larger number of document snippets. This very large number of document snippets that must be searched in these scenarios presents a challenge for a system designed to computationally effectively search the medical database and accurately select from it a smaller number of document snippets relevant to the medical question.
[0095] To address these challenges, as described below with reference to Fig. 4 and Fig. 5, some implementations of the medical question answering system 150 use the retrieval engine 130 to retrieve the smaller subset of document snippets by performing the search using a two-step process: an initial retrieval step that retrieves a plurality of document snippets from the document snippets stored in the medical database 150, followed by a reranking step that selects an appropriate subset of document snippets from the plurality of document snippets retrieved from the medical database 150 in the initial retrieval step.
[0096] A “suitable subset” of the plurality of document clippings includes at least one document clipping, but fewer than all of the plurality of document clippings retrieved in the initial retrieval step from the document clippings stored in the medical database 150.
[0097] The retrieval engine 130 performs the two-step process using multiple neural networks. The initial retrieval step is performed using an embedding model neural network 134. Subsequently, the reclassification step is performed using a classification neural network 136.
[0098] The neural embedding model network 134 (or "embedding model 134" for short) may comprise any suitable neural network architecture that enables the embedding model 134 to process a medical question to generate an embedding of the medical question (a "medical question embedding") or to process a document snippet to generate an embedding of the document (a "document snippet embedding").
[0099] An embedding is an ordered collection of numeric values in an embedding space. For example, an embedding can comprise one or more vectors of floating-point or other numeric values with a fixed dimensionality. The medical question embedding and the document snippet embeddings generally have the same dimensionality, meaning the medical question embedding and the document snippet embedding have the same number of numeric values.
[0100] For example, the embedding model 134 may comprise any type of neural network layers (e.g., fully connected layers, embedding layers, attention layers, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and in any suitable configuration (e.g., as a linear sequence of layers).
[0101] In some implementations, the embedding model 134 may be initialized with a base language model pre-trained using unsupervised learning. That is, the embedding model 134 may start with the same architecture and weights as (at least part of) the base language model trained on a text corpus to perform one or more language modeling tasks that do not require labeled training examples. For example, the embedding model 134 may be initialized with the base language model described in Devlin, Jacob, "Bert: Pre-training of deep bidirectional transformers for language understanding." arXiv preprint arXiv:1810.04805 (2018).
[0102] After initialization, the embedding model 134 can be fine-tuned using a user-defined embedding training dataset to learn fine-tuned weights based on an optimization of a fine-tuning objective function. The user-defined embedding training dataset includes a plurality of embedding training pairs. Each embedding training pair includes a medical question and a document snippet.
[0103] By design, each embedding training pair can be a positive embedding training pair, such that the document snippet includes information relevant to answering the medical question, or alternatively can be a negative embedding training pair, such that the document snippet includes information irrelevant to answering the medical question.
[0104] Such a custom embedding training dataset may be automatically generated by system 100 or another training system based on available medical documents, e.g., the medical documents stored in medical database 150. In some implementations, the document snippet for each embedded training pair may be one of the document snippets extracted from a medical source document, and the medical question may be generated by a neural language model network based on processing the document snippet. The neural language model network may, for example, be used to process a prompt comprising: (i) the document snippet and (ii) an instruction containing text in a natural language to generate a medical question that may (or may not) be answered, at least in part, by the document snippet.
[0105] In some implementations, multiple medical questions may be generated using the neural language model network based on processing the same medical document or even the same document snippet. Thus, two or more of the plurality of embedding training pairs may include the same document snippet.
[0106] Since a medical question may be shorter than a document snippet, generating an embedding of the medical question may save computational costs compared to generating an embedding of the document snippet because a smaller amount of data needs to be processed by the embedding model 134. The savings in computational costs during training may be substantial when embeddings of only a relatively small number of document snippets need to be generated relative to the number of medical questions, e.g., in a training setup where the embedding model 134 is designed to generate embeddings of multiple medical questions for each document snippet.
[0107] In implementations where multiple medical questions are generated for a document snippet, these questions can be generated such that one of the multiple medical questions can be answered by the document snippet, while others of the multiple medical questions cannot be answered. After generating the multiple medical questions, multiple embedding training pairs—including a positive embedding training pair and one or more negative embedding training pairs—can then be generated.
[0108] That is, a positive embedding training pair can be generated, comprising the document snippet and the medical question that can be answered by the document snippet. Additionally, one or more negative embedding training pairs can be generated, each comprising the document snippet and a medical question that cannot be answered by the document snippet.
[0109] In these implementations, the fine-tuning objective function may be a contrastive learning objective function. The contrastive learning objective function includes a term that encourages the embedding model 134 to generate similar embeddings (according to a similarity measure) for the medical question and document snippet included in each positive embedding training pair.
[0110] For example, the similarity measure may be a distance measure in an embedding space determined based on one of the following: a Euclidean distance, a Manhattan distance, or another distance measure, and the term (when used to compute gradient-based updates to the parameters of the embedding model 134) may bring the embeddings produced by the embedding model 134 for the medical question and the document snippet contained in each positive training pair closer together in the embedding space, i.e., decrease the distance between the medical question embedding and the document snippet embedding produced by the embedding model 134.
[0111] The objective function for contrastive learning also includes another term that encourages the embedding model 134 to generate different embeddings (according to a similarity measure) for the medical question and the document snippet included in each negative embedding training pair. In other words, for each negative embedding training pair, the other term penalizes the embedding model 134 for generating similar embeddings (according to a similarity measure) for the medical question and the document snippet included in the negative embedding training pair.
[0112] For example, the similarity measure may also be a distance measure in an embedding space, and the other term (when used to compute gradient-based updates to the parameters of the embedding model 134) may push the embeddings generated by the embedding model 134 for the medical question and the document snippet contained in each negative embedding training pair apart in the embedding space, i.e., increase the distance between the medical question embedding and the document snippet embedding generated by the embedding model 134.
[0113] Examples of contrastive loss functions that the system can use to train the embedding model are described in: Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). “A Simple Framework for Contrastive Learning of Visual Representations”. In Proceedings of the 37th International Conference on Machine Learning (ICML); and Hadsell, R., Chopra, S., & LeCun, Y. (2006). “Dimensionality Reduction by Learning an Invariant Mapping”. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR).
[0114] In general, for some or all training pairs (i.e., each comprising a medical question and a document snippet), the medical question has a shorter length than the document snippet. For example, the medical question may be <50%, <10%, or <5% of the length of the document snippet. In particular, medical questions may contain a sentence, a paragraph, or a series of paragraphs, whereas document snippets may comprise entire documents with many paragraphs of text, or part of them. Thus, processing a medical question using the embedding model to generate an embedding of the medical question may require significantly fewer computational resources (e.g., memory and processing power) than processing a document snippet using the embedding model to generate an embedding of the document snippet.
[0115] The system can exploit the asymmetry in the length of medical questions compared to document snippets to increase the efficiency of training the embedding model. Specifically, at each training iteration, the system can identify, for each of one or more document snippets, one "positive" medical question for which the document snippet answers the medical question, and N "negative" medical questions for which the document snippet does not answer the medical question, where N is any positive integer value, e.g., N=3 or N=5 or N=10 or N=100. The system then processes the document snippet, the positive medical question, and the N medical questions using the embedding model (and according to current values of the set of embedding model parameters) to generate corresponding embeddings.The system then measures distances between the embeddings and propagates gradients of a contrastive loss objective function, which depends on these embeddings, back through the embedding model (e.g., the neural embedding network). In this way, the system can drastically reduce the consumption of computational resources compared to an implementation where, for each of one or more medical questions, the system identifies one "positive" document snippet that answers the medical question and N "negative" document snippets that do not answer the medical question.
[0116] The neural ranking network 136 is configured to process, for a document snippet, an input comprising: (i) the medical question and (ii) the document snippet to produce an output comprising a score for each of a plurality of different relevance levels. For each relevance level, the score may be a probability value (e.g., between 0 and 1, inclusive of both ends) indicating a probability that a relevance between the medical question and the document snippet has the relevance level.
[0117] The classification neural network 136 may comprise any suitable neural network architecture that enables it to perform its described functions. For example, the classification neural network 136 may include any suitable types of neural network layers (e.g., fully connected layers, activation layers, attention layers, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and in any suitable configuration (e.g., as a linear sequence of layers).
[0118] In some implementations, the classification neural network 136 may be initialized with a base language model pre-trained with unsupervised learning and then fine-tuned by fine-tuning on a user-defined dataset comprising question-document snippet pairs with relevance annotations. For example, the neural network 136 may be initialized with the base language model described in Touvron, Hugo, et al., "Llama 2: Open foundation and fine-tuned chat models." arXiv preprint arXiv:2307.09288 (2023).
[0119] That is, the classification neural network 136 may start with the same architecture and weights as the base language model (at least a portion thereof), trained on a text corpus to perform one or more language modeling tasks that do not require labeled training examples, and then trained on user-defined data to learn fine-tuned weights based on an optimization of a fine-tuning objective function, e.g., a supervised learning objective function. The user-defined data may, for example, comprise a plurality of classification training pairs. Each classification training pair (i) comprises a medical question and a document snippet, and (ii) is associated with a ground truth score for each of a plurality of different relevance levels.
[0120] After performing the two-step process of retrieving the subset of document snippets, the medical question answering system 100 incorporates the subset of document snippets into the prompt 118 and then causes the generative neural network 120 to generate the answer 126 to the medical question 102 based on the processing of the prompt 118.
[0121] The answer 126 to the medical question 102 thus includes an output sequence of tokens generated by the generative neural network 120 based on the processing of the prompt 118, which is represented as an input sequence of tokens.
[0122] In some implementations, the response 126 includes a bibliography of the medical documents from which the subset of the plurality of document excerpts is obtained. In some implementations, the response 126 includes data indicating a rationale for the listed medical documents. For example, the rationale for each document excerpt in the subset of document excerpts may include one or more of the following: a similarity statement, an impact statement, or a recency statement.
[0123] For example, in response to a medical question “What health risks are associated with GLP-1 receptor agonists such as Ozempic and Mounjaro?”, the medical question answering system 100 may generate an answer that includes: Glucagon-like peptide-1 receptor agonists (GLP-1 RAs) such as semaglutide (Ozempic) and Tirzepatide (Mounjaro) is associated with various health risks. The most common adverse effects are gastrointestinal (GI) in nature, including nausea, vomiting, diarrhea, and abdominal pain. [1-2] There is also a risk of pancreatitis and gallbladder disease. [1] ... Cited literature 1. GLP-1 Agonists: A Review for Emergency Clinicians. Lang B, Pelletier J, Koyfman A, Bridwell RE. The American Journal of Emergency Medicine. 2024;78:89-94. doi:10.1016 / j.ajem.2024. 01.010. New research 2. Glucagon-Like Peptide-1 Receptor Agonists Associated Gastrointestinal Adverse Events: A Cross-Sectional Analysis of the National Institutes of Health All of Us Cohort. Aldhaleei WA, Abegaz TM, Bhagavathula AS. Pharmaceuticals (Basel, Switzerland). 2024;17(2):199. doi:10.3390 / ph17020199. New research
[0124] In this example, the group of tokens beginning with "glucagon-like" and ending with "biliary disease" represents the answer to the question, where "[1]" and "[1-2]" are tokens representing in-context citation markers, and the group of tokens under "Cited Literature" represents a bibliography of medical documents comprising the subset of the plurality of document excerpts, where "New Research" are tokens representing a justification for the listed medical documents (in this example, because they were published within the predetermined time period, e.g., within the last week, month, or year from which the question originates). If available, the publication dates can be stored as metadata associated with the medical documents in the medical database. In other examples, the justification for the medical documents could be different, e.g.,a declaration of similarity or a declaration of effects.
[0125] The similarity statement is based on the medical question embedding and the document snippet embedding generated using the neural embedding model network. For example, if the similarity measure between the document snippet embedding for a given document snippet and the medical question embedding meets a predetermined threshold (e.g., a distance threshold in an embedding space), the medical question answer may include "Highly Relevant" tokens as part of the reference list to highlight the similarity between a medical source document containing the given document snippet and the medical question.
[0126] The impact statement is based on citation metrics of a medical source document to quantify the medical source document's impact in the research community. If the medical source document was published in a journal that ranks above the 90th or 95th percentile of impact factor, the medical question answer can include the tokens "Leading Journal" or "Top Journal" as part of the reference list to highlight the medical source document's impact. If available, the journal's impact factors can be stored as metadata associated with the medical documents in the medical database.
[0127] In some implementations, the medical question answering system 100 may use the same generative neural network to generate multiple different candidate answers in response to the medical question. For example, if the generative neural network 120 is designed as an auto-regressive language model, some implementations of the medical question answering system 100 may do this by using beam search decoding from score distributions generated by the generative neural network 120, a sample-and-rank decoding strategy, or another decoding strategy that leverages the auto-regressive nature of the generative neural network 120.
[0128] The medical question answering system 100 may then select one or more selected candidate answers from the plurality of different candidate answers as a final answer 126 in response to the medical question 102 for output on a display of a user device.
[0129] For example, the selection may be based on an evaluation model implementing an evaluation function. As another example, the selection may be made by using the generative neural network 120 to "discriminate" between the generated candidate answers to determine which, if any, candidate answers should be provided in response to the medical question 102.
[0130] In some implementations, after receiving question data representing a medical question 102, the medical question answering system 100 may generate various prompts 118 based on the same medical question 102 and then generate a candidate answer for each various prompt by using the generative neural network 120 to process the prompt 118. The various prompts 118 may, for example, include document snippets selected from various types of medical documents stored in the medical database 150.
[0131] For example, as described below with respect to Fig. 6 and Fig. 7, some implementations of the medical question answering system 100 may generate a first prompt that includes document snippets from only one or more clinical practice guideline documents stored in the medical database 150 (and not from other medical documents stored in the medical database) and a second prompt that includes document snippets from the other medical documents stored in the medical database 150.
[0132] In this example, the first response generated by the generative neural network 120 in response to the first prompt may be presented using a first user interface element that is different from a second user interface element that presents a second response generated by the generative neural network 120 in response to the second prompt.
[0133] After generating an answer 126 and before providing the generated answer for presentation on the display, the medical question answering system 100 may use a quality assurance engine 140 to check the quality of the answer 126 generated using the generative neural network 120 to ensure the quality of the answer.
[0134] In some implementations, the quality assurance engine 140 may use a hallucination detection neural network configured to process a hallucination detection input comprising the response 126 and the subset of document snippets to generate one or more hallucination detection outputs that may be used by the quality assurance engine 140 to determine whether the response 126 contains any hallucination content, i.e., to determine whether the generative neural network 120 is synthesizing any nonexistent, distorted, or inaccurate information.
[0135] The hallucination detection neural network may comprise any suitable neural network architecture that enables it to perform its described functions. For example, the hallucination detection neural network may comprise any type of neural network layer (e.g., fully connected layers, embedding layers, attention layers, etc.) in any suitable number (e.g., 5 layers, 10 layers, or 20 layers) and in any suitable configuration (e.g., as a linear sequence of layers).
[0136] In some implementations, the hallucination detection neural network can be initialized with a base language model pre-trained using unsupervised learning and then fine-tuned, e.g., by supervised fine-tuning on a user-defined dataset comprising a variety of hallucination detection training examples.
[0137] For example, each hallucination detection training example may include a hallucination detection training input and a hallucination detection target output. The hallucination detection training input may include a first text string and a second text string. The hallucination detection target output may indicate whether the first text string includes content that contradicts the content contained in the second text string, whether the first text string includes content that is relevant to, similar to, or supports the content contained in the second text string, and so on.
[0138] In practice, hallucination detection may occur at different levels, e.g., at the sentence level, the paragraph level, and so on. For example, the one or more hallucination detection outputs may include a hallucination detection output indicating whether there is a contradiction between (i) each sentence or paragraph included in the response generated by the generative neural network 120 from processing the prompt 118 and (ii) the subset of document snippets.
[0139] As another example, the one or more hallucination detection outputs may include a hallucination detection output indicating whether at least one document snippet in the subset of document snippets supports each sentence or paragraph included in the response generated by the generative neural network from processing the prompt.
[0140] In these implementations, the medical question answering system 100 may withhold providing the answer to the user if it determines that there is a contradiction or lack of support from the document snippets in an answer. Instead, the system 100 may augment the answer generated by the generative neural network 120, e.g., apply a modification or correction, to generate an augmented answer. As another example, the system 100 may repeat the process to generate another answer.
[0141] Once the answer 126 generated by the generative neural network 120 has been quality-checked by the quality assurance engine 140, the medical question answering system 100 may provide the answer 126 to the user. Additionally or alternatively, the system 100 may provide the answer 126 to another system for further processing or store the answer 126 in a storage device for later use.
[0142] For example, the medical question answering system 100 may provide the answer 126 for display on a user interface of a user device, e.g., the user device through which the user asked the medical question 102.
[0143] As another example, the medical question answering system 100 may be implemented as part of or in communication with a digital assistance device, e.g., a mobile device, a smartwatch or other wearable device, or a smart speaker, and the digital assistance device may provide the answer 126 to the user, e.g., by generating speech representing the answer 126 and playing the speech to the user via a speaker.
[0144] Fig. 4 is a flowchart of an example process 400 for generating an answer to a medical question. For simplicity, the process 400 is described as being performed by a system of one or more computers located at one or more locations. For example, a medical question answering system, e.g., the medical question answering system 100 of Fig. 1, suitably programmed in accordance with this specification, perform process 400.
[0145] The system receives question data representing a medical question (step 402), e.g., from a user via a user interface.
[0146] The system obtains a plurality of document clippings from a medical database that stores medical documents (step 404). Step 404 may correspond to the initial retrieval step in the two-step search process.
[0147] To obtain the plurality of document snippets, the system obtains an embedding of the medical question (a "medical question embedding") generated by the embedding model based on the medical question, and for each of the document snippets stored in the medical database, an embedding of the document snippet (a "document snippet embedding") generated by the embedding model based on the document snippet. In some implementations, the embeddings of the document snippets may be pre-computed and stored in a data store for reuse during each iteration of process 400 to reduce runtime latency.
[0148] The system then searches for the k document snippet embeddings that are most similar to the medical question embedding according to a given similarity measure. K can in principle be any positive integer, i.e., an integer greater than or equal to one, but is generally much smaller than the total number N of document snippets stored in the medical database. The system performs a search of the medical database to retrieve the plurality of document snippets, each corresponding to the k document snippet embeddings.
[0149] For some similarity measures, e.g., Manhattan distance, Euclidean distance, or other distance measures, the most similar document snippet embeddings are those that are closest to the medical question embedding (have the smallest similarity measure with the medical question embedding). For some other similarity measures, e.g., inner product, the most similar document snippet embeddings are those that have the greatest similarity measure with the medical question embedding.
[0150] For each document snippet in the plurality of document snippets, the system determines a relevance score for the document snippet using a ranking neural network based on the document snippet and the medical question (step 406). Then, based at least in part on the relevance scores determined for the plurality of document snippets, the system selects an appropriate subset of the plurality of document snippets (step 408). Steps 406 through 408 may correspond to the reranking step in the two-step search process. By performing the reranking step, the system further reduces the plurality of document snippets to the appropriate subset of the plurality of document snippets.
[0151] The determination of the relevance ratings for the document excerpts is described below in Fig. 5, which is a flowchart of substeps 502-504 of step 406 of process 400 of Fig. 4.
[0152] For each document snippet in the plurality of document snippets, the system processes an input comprising the medical question and the document snippet using the ranking neural network to generate a respective score for each of a plurality of different relevance levels (step 502). For each relevance level, the respective score indicates a probability that a relevance between the medical question and the document snippet has the relevance level.
[0153] For example, the output of the classification neural network for each document snippet in the plurality of document snippets may include: P(A | Question, Excerpt), P(B | Question, Excerpt), ... P(F | Question, Excerpt) where A, B, ... F are tokens representing the different relevance levels. For example, A represents the highest relevance level (i.e., the document question is very relevant to the question), B represents the second highest relevance level (i.e., the document question is somewhat relevant to the question, etc.). In other examples, there may be more or fewer different relevance levels. The different relevance levels can also be represented by different tokens.
[0154] For each document snippet in the plurality of document snippets, the system determines the relevance score for the document snippet based on the output of the neural ranking network, ie, based on the respective scores generated by the neural ranking network for the plurality of relevance levels (step 504).
[0155] For example, the relevance score can be determined as a linear combination of the scores: X_0 * P(A | Question, Excerpt) + X_1 * P(B | Question, Excerpt) + ... X_6 * P(F | Question, Excerpt), where X_0, X_1, ... X_6 are predetermined weights, which can be, for example, tunable hyperparameters of the system.
[0156] The selection of the appropriate subset from the multitude of document excerpts based at least in part on the relevance ratings is described below in Fig. 6, which is a flowchart of substeps 602-608 of step 408 of process 400 of Fig. 4.
[0157] The system generates a ranking of the plurality of document snippets based on the relevance score determined for each document snippet using the neural ranking network (step 602). For example, the plurality of document snippets are arranged in the ranking such that the document snippet with the highest score is at the top of the ranking and the document snippet with the lowest score is at the bottom of the ranking.
[0158] The system selects, as an initial subset of the plurality of document snippets, the highest-ranked document snippets that have the highest relevance scores according to the ranking (step 604). The initial subset of the plurality of document snippets includes document snippets that have the highest relevance scores among the plurality of document snippets. For example, the system may select the document snippets that are at the top of the ranking.
[0159] For each of a plurality of aspects, the system assigns a weight to each document snippet in the initial subset of the plurality of document snippets with respect to the aspect (step 606) and then selects one or more document snippets from the document snippets in the initial subset of the plurality of document snippets based on the respective weight assigned to each document snippet (step 608), e.g., one or more document snippets having the highest weights. The weights assigned to the same document snippet may be different for different aspects.
[0160] For example, the plurality of aspects may include one or more of the following: a recency of a medical source document containing the document excerpt (i.e., a difference between the publication date and the date the medical question is received), a quality of a provider of the medical document (e.g., an impact factor of a journal in which the medical document was published), or a relevance between a user who asked the medical question and an author of the medical document (e.g., an affiliation or other relationship). For example, the system may generate a higher relevance weight for a particular medical document if the user who asked the medical question is also the author of the medical document in question.
[0161] In some implementations, the system selects an equal fixed number of document snippets based on the respective weights for each of the plurality of aspects, while in other implementations the system selects a different number of document snippets based on the respective weights for each of the plurality of aspects, i.e., the system may select more document snippets for one aspect than document snippets for another aspect.
[0162] The system combines, e.g., concatenates, the one or more document snippets selected for each of the plurality of aspects to generate a combined set of document snippets (step 610). The combined set of document snippets may then be used as the appropriate subset of the plurality of document snippets. In some implementations, the system applies further processing, e.g., semantic filtering, deduplication, or both, to the combined set of document snippets to generate the subset of the plurality of document snippets.
[0163] The system generates a prompt that includes at least the medical question and the appropriate subset of the plurality of document excerpts (step 410). Optionally, the prompt may also include additional data, e.g., one or more of the data types (or metadata) described above with reference to Fig. 1 were mentioned.
[0164] The system causes a generative neural network to process the prompt as input to generate an answer to the medical question as output (step 412). For example, the answer may be represented as an output sequence comprising tokens selected from a vocabulary, and the generative neural network may generate the answer auto-regressively by sequentially generating the tokens comprising the output sequence, taking into account any tokens already generated in the output sequence.
[0165] Fig. 7 is a flowchart of an example process 700 for generating a response to a user query. For simplicity, the process 700 is described as being performed by a system of one or more computers located at one or more locations. For example, a medical question answering system, e.g., the medical question answering system 100 of Fig. 1, suitably programmed in accordance with this specification, perform process 700.
[0166] The system receives a medical information query from a user and via a user interface displayed to the user on a display of a user device (step 702).
[0167] The system generates multiple responses to the query received from the user by automatically retrieving and parsing data from a corpus of medical documents stored in a medical database using a generative neural network (step 704). The multiple responses may include a first response and a second response.
[0168] The generative neural network is designed to generate a response based on processing a prompt generated by the system based on the query and data retrieved from the corpus of medical documents, e.g., document snippets. Specifically, the system generates various prompts based on the query and then processes each individual prompt using the same generative neural network to generate a respective response to the prompt.
[0169] As part of generating the first response, the system determines, based on an automated search of the medical document corpus, that one or more clinical practice guideline documents from the medical document corpus include information that answers or is relevant to the query (step 706).
[0170] In some implementations, the automated search may be a two-step process performed using an embedding model and a classification neural network, as described above. For example, the document snippets selected in the reclassification step each have a similarity measure with respect to the medical question embedding that satisfies a similarity threshold (e.g., a distance threshold in an embedding space).
[0171] In some other implementations, the automated search may be a one-step process performed using the embedding model. The one-step process may include the initial retrieval step but not the reclassification step. For example, the system may select a document snippet from the document snippets stored in the medical database according to the similarity measure. The selected document snippet has a similarity measure with respect to the medical question embedding that meets the similarity threshold.
[0172] In response to determining that one or more clinical practice guideline documents include information that answers or is relevant to the query, the system generates a first prompt that includes (i) the medical information query and (ii) data extracted from the one or more clinical practice guideline documents, e.g., one or more document snippets included in the one or more clinical practice guideline documents, and then processes the first prompt using the generative neural network to generate the first response to the query (step 708).
[0173] The first request specifically includes data extracted from the one or more clinical practice guideline documents, but excludes or does not include data extracted from other medical documents that may be stored in the medical database and that are not clinical practice guideline documents.
[0174] For example, the first request does not include data from clinical trial documents. Furthermore, the first request does not include data from medical labeling documents. Furthermore, the first request does not include data from regulatory notification documents. Furthermore, the first request does not include data from clinical research documents.
[0175] The system generates a second prompt that includes at least (i) the query for medical information and (ii) data extracted from one or more other medical documents that are not clinical practice guideline documents, e.g., one or more document excerpts included in the one or more other medical documents, and then processes the second prompt using the generative neural network to generate the second response to the query (step 710). In some implementations, the second prompt may also include the first prompt, the first response, or both.
[0176] For example, the one or more other medical documents may have been retrieved through the same automated search as the one or more clinical practice guideline documents or through a separate automated search of the medical database. Generally, the one or more other medical documents are distinct from the one or more clinical practice guideline documents; for example, the one or more other medical documents may include one or more clinical trial documents, one or more medical labeling documents, or both.
[0177] In some implementations, the system may generate the first response to the query and the second response to the query through separate calls to the generative neural network. In other words, the system may make a first call to the generative neural network to cause the generative neural network to generate the first response to the query based on the first prompt. Then, after the generative neural network generates the first response in response to the first call, the system may make a second call to the generative neural network to cause the generative neural network to generate the second response to the query based on the second prompt.
[0178] The system presents a first user interface element via the user interface on the display of the user device (step 712). The first user interface element displays the first response generated based on the clinical practice guideline documents, but not based on other medical documents.
[0179] The first user interface element visually highlights that the first response is derived only from clinical practice guideline documents (and not from other medical documents that are not clinical practice guideline documents) and identifies the one or more clinical practice guideline documents that were processed to generate the first response.
[0180] The system presents a second user interface element via the user interface on the display of the user device (step 714). The second user interface element presents the second response generated at least in part based on other medical documents that are not clinical practice guideline documents. That is, the second user interface element is presented within the same user interface, but lacks the visual cue that the second response originates only from clinical practice guideline documents.
[0181] Fig. 8 is an illustration of an example user interface 800.
[0182] The user interface 800 shows an input window 810 into which a user can enter a query. In the example of Fig. 8, the user has entered the question "How is psoriasis treated?" and in response, the user interface 800 presents a first user interface element 820 in which a first answer 825 "Treatment of psoriasis involves a multifaceted approach depending on the severity and extent of the disease..." may be displayed, and further presents a second user interface element 830 in which a second answer 835 "In addition to the guidelines of the American Academy of Dermatology and the National Psoriasis Foundation..." may be displayed.
[0183] In the example of Fig. 8, the first user interface element 820 is displayed above and before the second user interface element 830 within the user interface 800. In the example of Fig. 8, the first user interface element 820 is a bounded environment presented within the user interface 800. The first response 825, generated based on the clinical practice guideline documents but not based on other medical documents, is presented within the bounded environment.
[0184] Within the bounded environment, a bibliography of one or more clinical practice guideline documents is also displayed (which in the example of Fig. 8 includes three clinical practice guideline documents published by the American Academy of Dermatology), on the basis of which the first answer 825 is generated.
[0185] Optionally, the first user interface element 820 may include a “Practice Guideline” header 826 indicating that the first response 825 was generated based on the clinical practice guideline documents, but not on other medical documents.
[0186] Optionally, the first user interface element 820 may be displayed as a brief summary that may be expanded upon selection (e.g., double-click, hover, etc.) to show a more detailed representation.
[0187] This specification uses the term "designed" in connection with systems and computer program components. For a system comprising one or more computers to be designed to perform specific operations or acts means that the system has software, firmware, hardware, or a combination of them installed thereon which, when operated, causes the system to perform the operations or acts. For one or more computer programs to be designed to perform specific operations or acts means that the one or more programs include instructions which, when executed by a data processing device, cause the device to perform the operations or acts.
[0188] Embodiments of the subject matter and the functional operations described in this specification may be implemented in digital electronic circuitry, in tangible computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, that is, one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by a data processing device or for controlling its operation.The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. Alternatively or additionally, the program instructions may be encoded in a synthetically generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal used to encode information for transmission to suitable receiving means for execution by a data processing device.
[0189] The term "data processing device" refers to data processing hardware and includes all types of devices, apparatus, and machines for processing data, including, for example, a programmable processor, one or more processors or computers. The device may also be or further comprise special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The device may optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that establishes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof.
[0190] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not have to, correspond to a file in a file system. A program may be embodied in part of a file containing other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in several coordinated files, e.g.,Files that store one or more modules, subprograms, or pieces of code. A computer program can be designed to run on one or more computers located at one or more locations and connected by a data communications network.
[0191] In this specification, the term "database" is used broadly to refer to any collection of data: The data need not be structured in any particular way, or at all, and it may be stored on storage devices in one or more locations. For example, the index database may comprise multiple data collections, each of which may be organized and accessed differently.
[0192] Likewise, in this specification, the term "engine" is used broadly to mean a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers at one or more locations. In some cases, one or more computers are dedicated to a specific engine; in other cases, multiple engines may be installed and running on the same computer or computers.
[0193] The processes and logic flows described in this specification may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output data. The processes and logic flows may also be performed by specialized logic circuitry, such as an FPGA or ASIC, or by a combination of specialized logic circuitry and one or more programmed computers.
[0194] Computers suitable for executing a computer program may be based on general-purpose or special-purpose microprocessors, or both, or any other type of central processing unit. Generally, a central processing unit receives instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by, or incorporated into, special-purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, e.g.magnetic, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to them, or both. However, a computer need not include such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GPS (Global Positioning System) receiver, or a portable storage device, such as a USB (Universal Serial Bus) flash drive, to name a few.
[0195] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0196] For interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor, for displaying information to the user, and a keyboard and a pointing device, e.g., a mouse or trackball, through which the user can provide input to the computer. Other types of devices may also be used to enable interaction with a user; for example, feedback provided to the user may be any form of sensory feedback, e.g.,visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including auditory, speech, or tactile input. Furthermore, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser. Furthermore, a computer may interact with a user by sending text messages or other types of messages to a personal device, e.g., a smartphone running a text messaging application, and receiving response messages from the user.
[0197] Data processing devices for the implementation of machine learning models may also include, for example, special hardware acceleration units for processing general and computationally intensive parts of the training or production of machine learning, i.e., for inference.
[0198] Machine learning models can be implemented and deployed using a machine learning framework, such as a TensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.
[0199] Embodiments of the subject matter described in this specification may be implemented in a computing system that comprises a backend component, e.g., such as a data server, or that comprises a middleware component, e.g., an application server, or that comprises a frontend component, e.g., a client computer having a user interface, web browser, or app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0200] The computing system may include clients and servers. A client and a server are generally remote from each other and typically interact through a communications network. The client and server relationship is established by means of computer programs executing on the respective computers, which have a client-server relationship with each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for displaying data and receiving user input from a user interacting with the device acting as a client. Data generated on the user device, e.g., a result of the user interaction, may be received at the server from the device.
[0201] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.Although features may be described above as operating in certain combinations and may even initially be claimed as such, in some cases one or more features from a claimed combination may be removed from the combination and the claimed combination may relate to a sub-combination or a variation of a sub-combination.
[0202] Likewise, although operations are illustrated in a particular order in the drawings and in the claims, this should not be understood to require that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve desirable results. Multitasking and parallel processing may be advantageous under certain circumstances. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood to require such separation in all embodiments, and it should be understood that the described program components and systems may, in principle, be integrated together into a single software product or packaged into multiple software products.
[0203] Specific embodiments of the subject matter have been described. Other embodiments are also within the scope of the following claims. For example, the acts recited in the claims may be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or a sequential order to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Zitierte Patentliteratur
[0000] US 63 / 695,309
[0001] US 18 / 975,621
[0001] US 18 / 975,915
[0001] Zitierte Nicht-Patentliteratur
[0000] Anil, Rohan, et al. „Palm 2 technical report.“ arXiv preprint arXiv:2305.10403
[0091] Touvron, Hugo, et al. „Llama 2: Open foundation and fine-tuned chat models.“ arXiv preprint arXiv:2307.09288 (2023 [0091, 0118] Jiang, Albert Q, et al. „Mistral 7B.“ arXiv preprint arXiv:2310.06825 (2023
[0091] Devlin, Jacob „Bert: Pre-training of deep bidirectional transformers for language understanding.“ arXiv preprint arXiv:1810.04805 (2018
[0101] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). „A Simple Framework for Contrastive Learning of Visual Representations“. In Proceedings of the 37th International Conference on Machine Learning (ICML)
[0113] Hadsell, R., Chopra, S., & LeCun, Y. (2006). „Dimensionality Reduction by Learning an Invariant Mapping“. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)
[0113] GLP-1 Agonists: A Review for Emergency Clinicians. Lang B, Pelletier J, Koyfman A, Bridwell RE. The American Journal of Emergency Medicine. 2024;78:89-94. doi:10.1016 / j.ajem.2024. 01,010
[0123] Glucagon-Like Peptide-1 Receptor Agonists Associated Gastrointestinal Adverse Events
[0123] A Cross-Sectional Analysis of the National Institutes of Health All of Us Cohort
[0123] Aldhaleei WA, Abegaz TM, Bhagavathula AS
[0123] Pharmaceuticals (Basel, Schweiz). 2024;17(2):199. doi:10.3390 / ph17020199
[0123]
Claims
[1] System comprising: one or more computers; and one or more memory devices communicatively coupled to the one or more computers, the one or more memory devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising: Receiving a medical information query from a user and via a user interface presented to the user on a display of a user device; Generating multiple answers to the user's query by automatically retrieving and parsing data from a document corpus, comprising: Determining, based on an automated search of the document corpus, that one or more clinical practice guideline documents from the document corpus include information that answers the query; in response to determining that one or more clinical practice guideline documents include information that answers the query, generating a first response to the query based only on clinical practice guideline documents; and Generating a second response to the query based at least in part on one or more other documents that are not clinical practice guideline documents; and Present, through the user interface and on the display of the user device: a first user interface element presenting the first response generated based only on clinical practice guideline documents, the first user interface element visually highlighting that the first response is derived only from clinical practice guideline documents and identifying one or more clinical practice guideline documents processed to generate the first response; and a second user interface element that presents the second response generated at least in part based on documents that are not clinical practice guideline documents. [2] The system of claim 1, wherein the first user interface element is presented above and prior to the second user interface element within the user interface. [3] The system of any of claims 1 and 2, wherein the first user interface element comprises a bounded environment presented within the user interface, and wherein the first response generated based only on the clinical practice guideline documents is presented within the bounded environment. [4] The system of any one of claims 1 to 3, wherein the first user interface element comprises a header indicating that the first response is generated based only on clinical practice guideline documents. [5] The system of any one of claims 1 to 4, wherein generating the first response to the query based only on clinical practice guideline documents comprises making a first call to a generative neural network, and wherein generating the second response to the query based at least in part on one or more other medical documents that are not clinical practice guideline documents comprises making a second call to the generative neural network after the generative neural network has generated the first response in response to the first call. [6] The system of claim 5, wherein generating the first response to the query based only on clinical practice guideline documents comprises: Processing a first request by the generative neural network comprising (i) the query for medical information and (ii) document snippets contained in the clinical practice guideline documents to generate the first response. [7] The system of claim 6, wherein the second response to the query, based at least in part on one or more other documents other than clinical practice guideline documents, comprises: Processing, by the generative neural network, a second request comprising (i) the query for medical information, (ii) document snippets contained in the one or more other medical documents, and (iii) the first response generated by the generative neural network to generate the second response. [8] One or more computer storage media having stored thereon instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of the respective system according to any one of claims 1 to 7.
Citation Information
Patent Citations
18/975,621
US-ANMELDUNGNR.63/695,309
US-ANMELDUNGNR.18/975,915