Language model neural networks for domain-specific conversational agents
Fine-tuning a general-purpose language model through self-play techniques enhances its performance in domain-specific communication, addressing the challenge of adapting to complex interactions like medical diagnostics, improving diagnostic accuracy and user engagement.
Patent Information
- Application Number
- PCT/US2025/011204
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
General-purpose language model neural networks struggle to effectively engage in domain-specific communication sessions due to the need for domain-specific expertise, particularly in complex interactions like medical diagnostics, where context and skillful history-taking are crucial.
The use of 'self-play' techniques to fine-tune a general-purpose language model neural network by generating simulated communication sessions and incorporating context-based feedback to enhance its performance in specific domains, such as medical diagnostics.
The fine-tuned model improves diagnostic accuracy and user engagement by effectively modeling expert agents in the domain, providing accurate diagnoses and treatment plans, outperforming primary care physicians in simulated clinical examinations.
Smart Images

Figure US2025011204_17072025_PF_FP_ABST
Abstract
Description
[0001] LANGUAGE MODEL NEURAL NETWORKS FOR DOMAIN-SPECIFIC CONVERSATIONAL AGENTS
[0002] CROSS REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to U.S. Provisional Application No. 63 / 620.130, filed on January 11, 2024. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.
[0004] BACKGROUND
[0005] This specification relates to processing data using machine learning models.
[0006] As one example, neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to another layer in the network, e.g., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of weights.
[0007] SUMMARY
[0008] This specification describes a system implemented as computer programs on one or more computers in one or more locations that trains and uses a language model neural network as a domain-specific conversational agent for a particular domain.
[0009] More specifically, this disclosure describes techniques for “fine-tuning;’ i.e., further training, the language model neural network using “self-play” techniques.
[0010] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.
[0011] Many domains require domain-specific expertise in order for agents to engage in effective communication sessions with users, making it difficult for general-purpose language model neural networks, e.g., LLMs, to be effectively adapted for domain-specific dialogue. For example, engaging in a medical diagnostic interaction with a patient requires skillful history -taking in order to use the interaction to make an accurate diagnosis or effectively manage providing healthcare to the patient. Moreover, domain-specific interactions are a complex skill whose optimal conduct is highly dependent on context.
[0012] To account for this, this specification describes “self-play” techniques for effectively fine-tuning a general-purpose language model neural network, i.e., one that has been pre- trained on a large unsupervised data set, to serve as a domain-specific dialogue agent for a particular domain.
[0013] This specification also describes techniques to, at inference, use a language model neural network to iteratively generate a dialogue turn during a communication session with a user. In particular, by first causing the language model neural network to analyze the user information that is available in the current communication session and then use the analyzed information to generate a response to the most recent communication, the language model neural network can more effectively model an expert agent in the domain.
[0014] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and description below.
[0015] According to a first aspect there is provided a method performed by one or more computers for fine-tuning a language model neural network for use in participating in user communication sessions in a particular domain. The method includes, at each fine-tuning iteration in a sequence of fine-tuning iterations, generating one or more simulated communication sessions, where each simulated communication session is a communication session between one or more agents that include an expert agent for the particular domain and one or more user agents. Generating one or more simulated communication sessions includes, for each simulated communication session, determining a topic for the simulated communication session, determining, based on the topic, a context for the simulated communication session and then generating, using the context, a simulated dialogue that includes a sequence of one or more dialogue rounds, where each dialogue round includes a sequence of dialogue turns that each correspond to one of the one or more agents. To generate the simulated dialogue, for each dialogue turn within each dialogue round, the method includes generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a respective prompt for the agent corresponding to the dialogue turn. The method further includes processing the dialogue input using the language model neural network to generate the dialogue turn. After generating the one or more simulated communication sessions, the method includes generating, from each simulated communication session, one or more training examples for training the language model neural network. Then, the language model neural network is trained on training data comprising the one or more training examples. In some implementations, generating the simulated dialogue includes, for each dialogue turn within each dialogue round, generating a moderator input for the dialogue turn from at least (i) the dialogue round after the dialogue turn and (ii) a moderator prompt that specifies criteria for terminating the dialogue round. The moderator input is then processed using the language model neural network to generate a moderator output. Then, it can be determined, from the moderator output, whether to terminate the dialogue round after the dialogue turn and before performing any additional dialogue turns.
[0016] In some implementations, where the sequence of one or more dialogue rounds includes a plurality of dialogue rounds, generating the simulated dialogue includes, after each dialogue round other than a last dialogue round in the sequence, generating a critic input for the dialogue round from at least (i) the dialogue turns in the dialogue round and (ii) a critic prompt that specifies criteria for evaluating the expert agent. Then, the critic input can be processed using the language model neural network to generate a critic output that provides natural language feedback to the expert agent. Further, for each dialogue round after a first dialogue round in the sequence of dialogue rounds, generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a prompt for the agent corresponding to the dialogue turn includes, when the dialogue turn corresponds to the expert agent, generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) a respective prompt for the expert agent, and (iii) the critic output generated after the preceding dialogue round.
[0017] Further in some implementations, for each dialogue round after the first dialogue round in the sequence of dialogue rounds, generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a prompt for the agent corresponding to the dialogue turn includes, when the dialogue input corresponds to the expert agent, generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) the respective prompt for the expert agent, (iii) the critic output generated after the preceding dialogue round, and (iv) dialogue turns corresponding to the expert agent in the preceding dialogue round.
[0018] Further in some implementations, generating, from each simulated communication session, one or more training examples for training the language model neural network includes generating, from only the last dialogue round within each simulated communication session, one or more training examples for training the language model neural network. In some cases, determining a topic for the simulated communication session includes selecting the topic from a set of topics for the particular domain.
[0019] In some cases, determining, based on the topic, a context for the simulated communication session includes generating a respective set of values for each of a plurality of properties associated with the topic. Then, a context input can be generated from at least (i) the respective sets of values of the properties and (ii) a context prompt that characterizes a specified format for the context. Then, the context input is processed using the language model neural network to generate the context.
[0020] In some implementations, generating a respective set of values for each of a plurality of properties associated with the topic includes obtaining passages specify ing ranges of values for each of the plurality of properties and then filtering the obtained passes using a second language model neural network.
[0021] In some cases, the sequence of fine-tuning iterations includes one or more fine- tuning iterations.
[0022] In some implementation, generating, from each simulated communication session, one or more training examples for training the language model neural network includes generating, from only one or more of the dialogue rounds within each simulated communication session and not from the context for the simulated communication session, one or more training examples for training the language model neural network.
[0023] In some cases, training data comprises the training examples and a static set of training examples relating to the particular domain.
[0024] In some implementations, the particular domain is a medical diagnostic domain, the expert agent for the particular domain is a clinician, and the topic for the simulated communication is a medical condition to which the simulated communication session relates.
[0025] In some cases, the context is a patient vignette of a patient having the medical condition.
[0026] According to a second aspect there is provided a method performed by one or more computers and for carrying out a communication session in a particular domain with a user, where the communication session includes a respective dialogue turn at each of a sequence of time steps, each time step corresponding either to the user or to an expert agent. The method includes, at each time step corresponding to the expert agent, generating, from the dialogue turns at preceding time steps in the sequence and a user analysis prompt that specifies one or more properties of the user, a first input. The method further includes processing the first input using a language model neural network to generate a user analysis output that characterizes the one or more properties of the user. Then, from at least the user analysis output and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, a second input is generated. The method then includes processing the second input using the language model neural network to generate an initial response to the dialogue turn at the immediately preceding time step and generating, from the initial response, the dialogue turn at the time step.
[0027] In some implementations, dialogue turns at time steps corresponding to a user are received through a user interface of a user device, and where the method further includes, at each time step corresponding to the expert agent, providing the generated dialogue turn to the user through a user interface of the user device.
[0028] In some cases, the generated dialogue turn includes providing the dialogue turn as text displayed in the user interface.
[0029] Further in some cases, providing the generated dialogue turn includes providing the dialogue turn as speech through a speaker of the user device.
[0030] In some implementations, generating, from at least the user analysis output and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, a second input includes generating, from at least the user analysis output, the dialogue turns at preceding time steps in the sequence, and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, the second input.
[0031] Further in some cases, generating, from the initial response, the dialogue turn at the time step includes generating, from at least the initial response, the dialogue turns at preceding time steps in the sequence, and a refinement prompt that specifies requirements for dialogue turns generated by the expert agent, a third input. Then, the third input is processed using the language model neural network to generate the dialogue turn at the time step.
[0032] Further in some cases, the particular domain is a medical diagnostic domain and the expert agent for the particular domain is a clinician dialogue agent.
[0033] According to a third aspect there is provided the methods of the first and second aspect performed by a system that includes one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the respective operations of the respective method. According to a fourth aspect there is provided the method of the first and second aspect performed by one or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the respective method.
[0034] Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and the claims.
[0035] BRIEF DESCRIPTION OF DRAWINGS
[0036] FIG. 1 shows an example training system.
[0037] FIG. 2 is a diagram that illustrates an example fine-tuning iteration of the example training system.
[0038] FIG. 3 demonstrates an example medical diagnostic language model training system.
[0039] FIG. 4A illustrates the improvement of diagnostic dialogue using the example training system for a medical diagnostic language model.
[0040] FIG. 4B illustrates the improved accuracy using the example training system for a medical diagnostic language model.
[0041] FIG. 5 is a flow diagram of an example process for training the language model neural network.
[0042] FIG. 6 is a flow-diagram of sub-steps of one of the steps of the process of FIG. 5.
[0043] FIG. 7 is a flow diagram of sub-steps of one of the steps of the process of FIG. 5.
[0044] FIG. 8 is a flow diagram of an example process of using the fine-tuned language model neural network as a domain specific conversational agent.
[0045] DETAILED DESCRIPTION
[0046] FIG. 1 shows an example training system 100 that includes a domain specific language model neural network 110.
[0047] The training system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components and techniques described below can be implemented.
[0048] The system 100 trains a domain specific language model neural network 110 for use as a domain-specific conversational agent.
[0049] A domain-specific conversational agent, as used in this specification, is software that carries out communication sessions with users that are specific to a particular domain, i.e., that serv es as an expert agent for the particular domain during communication sessions with users.
[0050] For example, the particular domain can be a medical diagnostic domain. In this example, the expert agent for the particular domain is a clinician and the users can be patients seeking medical advice or treatment. In this particular example, an input 102 could be data representative of a medical condition. For example, the input 102 may comprise physiological data. The input data 102 may comprise data describing a patient’s symptoms or condition, for example textual and / or audio data describing a physiological condition. For example, the input may describe physical characteristics such as the presence, size, placement, and morphology' of a rash or wound, the presence or absence of pain in a particular body part, etc. Output data 112 could be a diagnosis or treatment plan.
[0051] As another example, the particular domain can be physical, electronic, or electrical system diagnostic domain. In this example, the expert agent for the particular domain may be an engineer and the users can be users’ seeking to diagnose or fix a fault in a real-world system. An input 102 could be data representative of a fault in the physical, electronic, or electrical system and output data 112 could be a diagnosis or repair instructions. For example, the input may be a description of sensor readings, visual or audio characteristics of a system, or a description of mechanical properties, for example, that a component will not move or is loose.
[0052] As another example, the particular domain can be an agricultural domain. In this domain, the expert agent may be an agricultural consultant and the user can be a farmer seeking advice regarding planting time, irrigation, or pest control. An input 102 could be data collected from soil sensors, weather forecasts and crop health monitoring systems. For example, the input may be a description of pH levels, moisture content and nutrient composition of soil or a description of the state of pest affected plants, e.g., their color, physical feel, etc. The output data 112 could be an optimal planting time, an irrigation schedule or pest control measures.
[0053] As another example, the particular domain can be a financial domain. In this example, the expert agent may be a financial consultant and the users can be users seeking investment advice or assistance in financial planning. An input 102 could be data representing spending patterns, risk metrics, market trends or financial statements and output data 112 could be an investment plan, a risk assessment or other relevant financial advice. For example, the input may be a description of stock prices, values of market indices, e.g., S&P 500 and NASDAQ, a value at risk (VaR) metric, etc. As another example, the particular domain can be a legal domain. In this domain, the expert agent for the particular domain is a legal expert and the users can be asking questions relating to laws or regulations. In this particular example, an input 102 could be text data describing a legal question and the output data 112 could be a legal analysis.
[0054] As yet another example, the particular domain can be a customer support domain for a particular product or company.
[0055] As yet another example, the particular domain can be an education domain, where the expert agent is an educator, and the users are students.
[0056] The domain specific language model neural network 110 can be any appropriate language model neural network with any appropriate neural network architecture. As a particular example, the language model neural network can be a large language model (LLM).
[0057] The large language model neural network 110 can have any of a variety of Transformer-based neural network architectures, e.g., encoder-only Transformer architectures, encoder-decoder Transformer architectures, decoder-only Transformer architectures, diffusion Transformer architectures, other attention-based architectures, and so on.
[0058] Examples of such Transformer-based neural network architectures include those described in Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910. 10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR. abs / 2001.09977, 2020: Tom B Brow n. Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla DhariwaL Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005. 14165, 2020; Aakanksha Chowdhery. et al. PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv:2204.02311; Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv:2305. 10403, 2023; and Gemini Team, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023).
[0059] The language model neural network 110 can receive an input sequence made up of text tokens selected from a vocabulary and auto-regressively generates an output sequence made up of text tokens from the vocabulary. The language model neural network 110 operates auto-regressively; it executes an auto-regressive token generation process to auto- regressively generate an output sequence of tokens by generating each particular token in the output sequence conditioned on a current input sequence that includes any tokens that precede the particular token in the output sequence, i.e., the tokens that have already been generated for any previous positions in the output sequence that precede the particular position of the particular token, and in some cases, a context input that provides context for the output sequence.
[0060] To generate a particular token at a particular position within an output sequence, the language model neural network 110 can process the current input sequence to generate a score distribution, e.g., a probability distribution, that assigns a respective score, e.g., a respective probability, to each token in the vocabulary of tokens. The output sequence generation system can then select, as the particular token, a token from the vocabulary using the score distribution. For example, the output sequence generation system can greedily select the highest-scoring token or can sample, e.g., using top-k sampling, nucleus sampling or another sampling technique, a token from the distribution.
[0061] As a particular example, the language model neural network 110 can include a sequence of attention blocks followed by an output subnetwork, and, during the processing of the current input sequence, each attention block in the sequence receives a respective input hidden state for each token in the current input sequence. The attention block then updates at least the hidden state for the last token in the current input sequence at least in part by applying self-attention to generate an output hidden state for the last token.
[0062] In some implementations, the attention block can make use of key-value (KV) caching or other caching techniques to avoid recomputing, at each time step, the respective output hidden states for tokens other than the last one in the current input sequence by storing the key and value vectors that have been generated by the attention block in a cache, such that they are later be reused.
[0063] The input hidden states for the first attention block are embeddings of the input tokens in the input sequence and the input hidden states for each subsequent attention block are the output hidden states generated by the preceding attention block. The output subnetwork processes the output hidden state generated by the last attention block in the sequence for the last token in the current input sequence to generate the score distribution.
[0064] The vocabulary of tokens can include any of a variety of tokens that represent text symbols or other symbols. For example, the vocabulary of tokens can include one or more of: characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code.
[0065] Once the language model neural network 110 has generated the output sequence of tokens, the tokens are then decoded to generate the output data 112 that is human-readable, e.g., an output sentence. The model uses the vocabulary of tokens to map the output tokens to a corresponding word using decoder architecture.
[0066] In particular, the system 100 can fine-tune the domain specific language model neural network 110 for use in participating in user communication sessions in a particular domain.
[0067] In some implementations, the system 100 can fine-tune the domain specific language model neural network, initialized 120 from a ‘"base,’7pre-trained language model neural network 130, e.g., a foundation LLM model.
[0068] More specifically, the system 100 can fine-tune the domain specific language model neural network 110 by using a “self-play” based simulated learning environment. That is, the system 100 can train the domain specific language model neural network 110 by using the neural network 110 to generate simulated communication sessions for the particular domain and then generate training examples based on those simulated communication sessions. The system 100 can perform the fine-tuning by training the neural network 110 at each fine-tuning iteration in a sequence of one or more fine-tuning iterations.
[0069] In some implementations, the domain specific language model neural network 110 can be trained on real-world datasets and dialogues as well as the simulated communication sessions. For example, in a medical diagnostic domain, the domain specific language model neural network 110 can be trained on multiple choice medical question-answering data, expert-curated long-form medical reasoning data, electronic health record note summaries, and large-scale transcribed medical conversation interactions. In this particular example, the training tasks can include dialogue generation tasks as well as medical question answering, reasoning and summarization tasks.
[0070] In the particular example of medical diagnosis, both real-world dialogues and simulated communication sessions can be used to train the domain specific language model neural network 110 as (i) existing real -world data often fails to capture the vast range of medical conditions and scenarios, hindering the scalability and comprehensiveness of the model 110, and (ii) the data derived from real-world dialogue transcripts tend to be noisy, containing ambiguous language, e.g., slang, jargon, sarcasm, etc., interruptions, ungrammatical utterances and implicit references. By incorporating simulated communication sessions and other “self-play” techniques, described in further detail below, the domain specific language model neural network 1 10’s knowledge and capabilities are able to be scaled across a multitude of medical conditions and contexts.
[0071] The training process will be described in more detail below in reference to FIG. 2 and FIG. 3.
[0072] FIG. 2 is a diagram that illustrates an example fine-tuning iteration of the example training system.
[0073] The training system 200 trains the domain specific language model 110 on training data that includes one or more training examples 260 generated at each of one or more finetuning iterations 220 in a sequence of fine-tuning iterations.
[0074] At each fine-tuning iteration 220. the system generates one or more simulated communication sessions, e.g., simulated communication session 240, using the language model neural network and then generates, from each simulated communication session 240, one or more training examples 260 for training the domain specific language model 110. Each fine-tuning iteration 220 represents an outer “self-play” loop as the domain specific language model 110 is fine-tuned on training examples that the model 110 itself generates based on simulated communication sessions.
[0075] For each simulated communication session 240, the system 200 can determine a topic for the simulated communication session 240 and, based on the topic, the system 200 can determine a context 232 for the session 240. The system 200 can select the topic from a set of topics for the particular domain. For example, if the desired domain is a legal domain, the topic can be selected from a set of topics related to the legal domain, e.g., criminal charges, civil case, contracts, wills, etc.
[0076] To determine the context 232 for simulated communication session 240 based on the topic, the system 200 can generate a respective set of values for each of one or more properties associated with the topic. The system 200 can then generate a context input from at least (i) the respective sets of values of the properties and (ii) a context that characterizes a specified format for the context. The system 200 can process the context input using the domain specific language model neural network 110 to generate the context 232.
[0077] The context 232 is described in further detail with reference to FIG. 7.
[0078] Each of the simulated communication sessions 240 of the one or more generated simulated communication sessions includes a simulated dialogue 250 generated, using the context 232, between an expert agent 242 and one or more user agents, e.g.. user agent 244. As an example, the system 200 can use the domain specific language model 1 10 and the context 232 to iteratively generate the simulated dialogue 250 between the expert agent 242 and the user agent 244.
[0079] Any given simulated dialogue 250 includes a sequence of one or more dialogue rounds where each dialogue round includes a respective dialogue turn at each of a sequence of time steps, each time step corresponding either to the user agent 244 or to the expert agent 242. For example, a dialogue round in the simulated dialogue 250 includes dialogue turn A 252. dialogue turn B 254 and dialogue turn C 256.
[0080] In particular, the system 200 can first generate a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a respective prompt for the agent corresponding to the dialogue turn. That is, the system 200 can generate a dialogue input based on the current status of the dialogue round and a respective prompt for the agent.
[0081] For example, in a medical diagnostic domain, the respective prompt for the user agent 244, e.g., patient agent, can be to reply to the expert agent 242’s, e.g., doctor agent’s, questions about their symptoms drawing upon information provided in the context 232, e.g., symptoms, demographics, and management plan. As a particular example, the respective prompt for the patient agent can be the following:
[0082] “Y ou are a patient chatting with a doctor over an online chat interface. The doctor has never met you before. <Context> Respond to the doctor's questions honestly as they interview you, asking any questions that may come up.”
[0083] In the same example, the respective prompt for the expert agent 242, e.g., doctor, can be to act as an empathetic clinician, interviewing patients about their medical history and symptoms to ultimately arrive at an accurate diagnosis. As a particular example, the respective prompt for the doctor can be the following:
[0084] “You are an empathetic clinician asking a patient about their medical history' over an online chat interface. Y ou know nothing about the patient in advance. Respond to the patient with a single-tum response to better understand their history and symptoms. Do not ask more than two questions. If the patient asks a question, be sure to answer appropriately.”
[0085] The system 200 can then process the dialogue input using the domain specific language model neural network 110 to generate the dialogue turn. In some implementations, after each dialogue round other than the last dialogue round in the sequence, the system 200 can generate a critic input for the dialogue round from at least (i) the dialogue turns in the dialogue round and (ii) a critic prompt that specifies criteria for evaluating the expert agent. The system can then process the critic input using the domain specific language model neural network 110 to generate a critic output that provides natural language feedback to the expert agent. This “inner” self-play loop allows the model 110 to leverage in context critic feedback to refine its behavior during a simulated conversation session.
[0086] As a particular example, the criteria specified by the critic prompt can include general setup criteria, response criteria, and diagnosis and treatment criteria. For example, the general setup criteria can include the following:
[0087] • The doctor agent engages in an online chat with a patient.
[0088] • It cannot perform a physical exam due to the virtual setting.
[0089] • The doctor agent needs to gather information about the patient's medical history to determine their condition.
[0090] • The doctor agent needs to "dig deep" by asking follow-up questions to understand the patient's symptoms fully.
[0091] • The doctor agent must provide at least two possible diagnoses before narrowing down to the correct one.
[0092] In this particular example, the general setup criteria can give the critic input context to the doctor agent’s setup, e.g., online chat, role, goals, and limitations. The setup criteria focuses on giving the critic agent context to better understand what the doctor agent should be doing in the dialogue and how it should be doing it. As another example, the response criteria specified in the critic input can include the following:
[0093] • Empathy and Conciseness: Responses should be empathetic and professional, addressing the patient's concerns without being overly long or asking too many questions at once.
[0094] • Human-like Interaction: The doctor agent must convincingly portray a human clinician and avoid revealing its Al identity .
[0095] • Natural Flow: Responses should be logical, realistic, and encourage the patient to share more information.
[0096] • Response Variety: Avoid repetitive language or response patterns (e.g., starting every7sentence with "Okay").
[0097] • Information Gathering: Ask relevant questions without repeating itself or requesting information it already has.
[0098] • Demographics: It's okay to ask for basic demographic information.
[0099] In this particular example, the response criteria can help the critic agent evaluate the dialogue of the doctor agent by7giving specific evaluation guidelines, e.g., empathy. response variety', etc. The response criteria can help the critic agent give tailored feedback that will actually improve the doctor agent's responses and enhance the experience and satisfaction of the patient agent. As another example, the diagnosis and treatment criteria can include the following:
[0100] • Differential Diagnoses: The doctor agent should explore multiple possible diagnoses based on the patient's information.
[0101] • Ground Truth: The doctor agent should ultimately guide the conversation towards the correct diagnosis and treatment provided in the prompt.
[0102] • Confidence Level: Express increasing confidence in the correct diagnosis as information is gathered, but avoid absolute certainty.
[0103] • Limitations: The doctor agent cannot schedule appointments or take real-world actions.
[0104] In this particular example, the diagnosis and treatment criteria can help the critic agent evaluate the final diagnosis and treatment plan given to the patient agent by the doctor agent. Similar to the response criteria, the diagnosis and treatment criteria give the critic agent specific evaluation guidelines to use when evaluating the performance of the doctor agent to improve the diagnosis and treatment plan in a helpful way, e.g., the doctor agent cannot schedule appointments versus the doctor agent scheduled the wrong appointment. In addition, the critic prompt can include the patient condition and background information, written as the following:
[0105] “The patient has the following condition: <condition>, with the following background information: <context>.'’
[0106] In this particular implementation, generating a dialogue input for the dialogue turn corresponding to the expert agent includes generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) a respective prompt for the expert agent, and (iii) the critic output generated after the preceding dialogue round.
[0107] The system 200 can then begin a subsequent dialogue round with the same user agent 244, the “fine-tuned” expert agent 244, that has learned from the critic input, and the same context 232. The system 200 can then take the previous dialogue round into account when generating a dialogue turn in the subsequent dialogue round. More specifically, for each dialogue round after the first dialogue round in the sequence of dialogue rounds, generating a dialogue input for the dialogue turn corresponding to the expert agent can include generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) the respective prompt for the expert agent, (iii) the critic output generated after the preceding dialogue round, and (iv) dialogue turns corresponding to the expert agent in the preceding dialogue round.
[0108] Each training example 260 can be generated from a dialogue turn within a respective simulated communication session 240 and can include (i) a corresponding prompt, (ii) one or more preceding dialogue turns within the simulated communication session 240. and, as the target output within the training example, (iii) the dialogue turn.
[0109] In some implementations, the one or more training examples 260 can be generated from one or more dialogue turns from only the last dialogue round within each simulated communication session 240.
[0110] In some implementations, the one or more training examples 260 can be generated from one or more dialogue turns from only one or more of the dialogue rounds within each simulated communication session and not from the context for the simulated communication session.
[0111] The system then trains 270 the domain specific language model 110 on training data that includes the one or more training examples 260 generate from the fine-tuning iteration 232. The training system 200 can perform multiple fine-tuning iterations in a training sequence, and at each iteration 232, the system 200 can train the neural network 110 on an evolving set of simulated dialogues 250, i.e.. dialogues that improve in quality because the language model neural network is used to generate the simulated dialogues and improves in quality as fine-tuning progresses.
[0112] In some implementations, the training data can include the training examples and a static set of training examples relating to the particular domain.
[0113] FIG. 3 demonstrates an example medical diagnostic language model training system.
[0114] The medical diagnostic language model training system 300 is an example of a training system, e.g., the training system 200 of FIG. 2, that trains a neural network for use in a medical diagnostic domain. For example, the medical diagnostic language model 310 can act as medical diagnostic conversation agent to help diagnose and recommend treatment plans to users. The medical diagnostic language model 310 can be trained on data 380 including medical reasoning data 382, long form medical question answering (Q / A) data 384. medical summary data 386, real world dialogues 388 and simulated dialogues 350.
[0115] The medical reasoning data 382, long form medical question answering (Q / A) data 384, medical summary data 386, and real-world dialogues 382 can be used to train the model on a variety of real-world datasets.
[0116] For example, the medical reasoning data 382 can be the MedQA (multiple-choice) dataset consisting of US Medical Licensing Examination (USMLE) multiple-choice style open domain questions with four or five possible answers. The training set consists of 11,450 questions and the test set has 1,273 questions. The training set can also include 191 curated MedQA questions from the training set where clinical experts crafted step by step reasoning leading to the correct answer.
[0117] As another example, the long-form medical question answering data 384 can be a dataset that consist of expert-crafted long form responses to 64 questions from HealthSearchQA. LiveQA, and Medication QA in MultiMedBench.
[0118] For example, the medical summarization data 386 can be a dataset consisting of 65 clinician-written summaries of medical notes from MIMIC-III, a large publicly available database containing medical records of intensive care unit patients. MIMIC-III contains approximately 2 million notes spanning 13 types including cardiology, respiratory, radiology, physician, general, discharge, case management, consult, nursing, pharmacy, nutrition, rehabilitation, and social work. 5 notes from each category were selected, with a minimum total length of 400 tokens and at least one nursing note per patient. Clinicians were instructed to write abstractive summaries of individual medical notes, capturing key information while also permitting the inclusion of new information and clarifying phrases and sentences not present in the original note.
[0119] As another example, the real-world dialogues data 388 can be a de-identified dataset licensed from a dialogue research organization comprising 98,919 audio transcripts of medical conversations during in-person clinical visits from over 1,000 clinicians over a 10- year period in the United States. It covers 51 medical specialties (primary care, rheumatology, hematology, oncology, internal medicine, and psy chiatty' among others) and 168 medical conditions and visit reasons (type II diabetes, rheumatoid arthritis, asthma, depression among common conditions). Audio transcripts contained utterances from different speaker roles such as doctors, patients, and nurses. On average a conversation had 149.8 dialogue turns. For each conversation, the metadata contained information about patient demographics, reason for visit (follow-up for pre-existing condition, acute needs, annual exam and more) and diagnosis type (new, existing, or other unrelated). In this particular example, selected dialogues involved only doctors and patients, not other roles such as nurses.
[0120] The training system 300 can also train the medical diagnostic language model 310 on simulated dialogues 350 generated by the domain specific language model 310 as described above.
[0121] The system 300 can use a simulated dialogue generator 355 including one or more agents and a context 332 to generate a simulated dialogue 350 and provide in context feedback to the expert (e.g., doctor) agent for self-improvement. The operations of each agent, e.g., the expert agent and one or more user (e.g. patient) agents, are performed by processing different input using the domain specific language model neural network. That is, the system 300 can use the same domain specific language model neural network to act as multiple agents by providing different prompts, e.g., the expert agent prompt and user agent prompts described above, to the neural network.
[0122] In this particular implementation, the topic of the simulated communication is a medical condition to which the simulated communication sessions relates.
[0123] In this particular implementation, the context 332 can be a patient vignette, or description, that can include essential background information such as patient demographics, symptoms, past medical history, past surgical history, past social history, and patient questions, as well as an associated diagnosis and management plan.
[0124] In some implementations, the system 300 can include a vignette generator aimed to create varied and realistic patient scenarios at scale, which can be subsequently used as the context for generating simulated dialogues.
[0125] Given the context 332 detailing a specific medical condition, the simulated dialogue generator 355 can generate a realistic simulated dialogue 350 among a patient agent 342 and a doctor agent 344 by instructing the medical diagnostic language model to act as the patient agent 342 in one dialogue turn and as the doctor agent 344 in the subsequent dialogue turn.
[0126] The patient agent 342 can embody the individual experiencing the medical condition outlined in the context 332. The patient agent 342’s role can involve truthfully responding to the doctor agent 344’ s inquiries as well as raising any additional questions or concerns they may have had.
[0127] The doctor agent 344 can play the role of the empathetic clinician seeking to comprehend the patient agent 342’s medical history within the simulated dialogue. The doctor agent 344's objective can be to formulate questions that could effectively reveal the patient agent 342’s symptoms and background, leading to an accurate diagnosis and treatment plan.
[0128] In some implementations, a moderator agent 346 can be used to determine when the simulated dialogue has reached an end. The moderator agent 346 can continually assess the ongoing dialogue between the patient agent 342 and the doctor agent 344 to determine when the conversation has reached a natural conclusion. The moderator agent 346 will be described in further detail with reference to FIG. 6. By providing a moderator agent 346, implementations can prevent unnecessary or wasted use of resources such as energy and compute, which would be used by continuing a conversation that is not developing further and may be ended.
[0129] The simulated dialogue generator 355 can further include an inner feedback loop 340 that includes the moderator agent 346, the doctor agent 344, a critic agent 348 and the simulated dialogue 350. Essentially, the inner feedback loop 340 can represent the inner self-play loop of the medical diagnostic language model 310, allowing the model 310 to receive in context feedback when generating simulated dialogues. For example, the moderator 346 can give feedback on when to end the simulated dialogue 350.
[0130] As another example, the critic agent 348 can be implemented to provide in context feedback to the doctor agent 344 and enhance the doctor agent 344's performance in subsequent simulated dialogues 350. The critic agent 348 is aware of the ground truth diagnosis and can evaluate the doctor agent 344’s responses based on if (i) the doctor agent 344 exhibits empathy and professionalism while addressing the patient agent 342’s questions or comments in a concise manner, (ii) the doctor agent 344 avoids asking too many or repetitive questions, focusing on a maximum of one or two per response, (iii) the responses reveal that the doctor agent 344 is a chatbot, e.g., the responses should flow naturally, maintain factual accuracy and facilitate further engagement from the patient, and (iv) the doctor agent 344 asks sufficient questions to identify at least two of the most likely- differential diagnoses and further refines their understanding through targeted questions towards the ground truth diagnosis and offers the corresponding treatment.
[0131] That is, the critic agent 348 can receive the first round of simulated dialogue between the doctor agent 344 and the patient agent 342 and give feedback based on the above criteria. As a specific example, the critic feedback for a simulated dialogue between a patient agent 342 with carpal tunnel syndrome and a doctor agent 344 can look like the following: "‘Overall, the doctor did a good job of gathering information and explaining the patient’s condition in a clear concise manner. The questions were targeted to differentiate between carpal tunnel syndrome and other potential causes, leading to a more confident diagnosis.
[0132] Here are a few specific suggestions for improvement:
[0133] 1. Early Reassurance: After the initial symptom description, a brief reassurance like ‘These are concerning symptoms, but we’ll work together to figure this out.’ can build rapport early on.
[0134] 2. Symptom Specificity: Instead of asking general weakness, ask ‘Which fingers are weak? Is it gripping, pinching or fine movements?’ This helps pinpoint nerve involvement.
[0135] 3. Neck Pain: Instead of just asking about presence, ask ‘Does neck movement make hand symptoms better / worse? Any tingling down the BACK of your arm?’ This helps rule out cervical issues more definitively.
[0136] 4. Differential: Mentioning other possibilities, like cubital tunnel syndrome or even arthritis, shows broader thinking, even if less likely.
[0137] 5. Treatment Nuance: Instead of just listing options, tailor them: ‘Splinting helps MOST at night, NSAIDS are for WHEN pain flares, ergonomics is KEY to PREVENTING worsening.”
[0138] Following the critic agent 348 ’s feedback, the doctor agent 344 incorporates the suggestions to improve its responses in subsequent rounds of dialogue with the same patient agent 342 (where each dialogue round starts from scratch).
[0139] This inner feedback loop 340, or inner self-play loop, is repeated one or more times to generate the simulated dialogue 350 that is then used for each iteration of fine-tuning. That is. within every fine-tuning training iteration, or outer self-play loop in which the model learns from generated simulated dialogues 350, the model 310 includes an inner selfplay loop during simulated dialogue generation in which the model receives in-context feedback to improve generation.
[0140] FIG. 4A illustrates the improvement of diagnostic dialogue using the example finetuned medical diagnostic language model. FIG. 4A visually compares the performance of primary care physicians (PCPs) and the fine-tuned medical diagnostic language model described in this disclosure, denoted as the articulate medical intelligence explorer (AMIE). The two approaches, e.g., PCP versus AMIE, were assessed in a blinded remote objective structured clinical examination (OSCE) with validated patient actors interacting via a text interface. That is, 20 board-certified PCPs and 20 validated patient actors participated in online-text based consultations, with each patient actor completing one consultation with a PCP and one with AMIE for a specific context, e.g., medical condition. The contexts covered conditions from cardiovascular, respiratory, gastroenterology, neurology, urology', obstetric, and gynecology domains, and internal medicine.
[0141] Upon conclusion of the consultations, the patient actor and the OSCE agent, i.e., the PCPs or the medical diagnostic language model, each filled in a post-questionnaire. A pool of specialist physicians also evaluated PCPs and AMIE with respect to the quality of their consultation and their responses to the post-questionnaire.
[0142] The PCPs and AMIE were compared across multiple axes including management plan, escalation recommendation, empathy, perceived openness and honesty, patient’s confidence in care and top-3 diagnostic accuracy. Across multiple axes corresponding to both specialist physicians (28 out of 32) and actor (24 out of 26) perspective, AMIE was rated as superior to PCPs while being non-inferior on the rest.
[0143] FIG. 4B illustrates the improved diagnostic accuracy using the example fine-tuned medical diagnostic language model.
[0144] FIG. 4B visually compares the diagnostic accuracy of primary7care physicians (PCPs) and the fine-tuned medical diagnostic language model described in this disclosure, denoted as the articulate medical intelligence explorer (AMIE). The top-k diagnosis accuracy was compared with respect to the ground truth diagnosis in graph 410 and with respect to diagnosis matches with any item on the accepted differential in graph 420.
[0145] Both graphs show a significantly higher top-k accuracy for AMIE, the medical diagnostic language model described in this disclosure, than that of PCPs across all values of k.
[0146] The medical diagnostic language model has been trained on large sets of real-world datasets as well as its own generated simulated dialogues, enabling the model to have wide knowledge and understanding of a variety of medical conditions and diagnoses.
[0147] FIG. 5 is a flow diagram of an example process for training the language model neural network. The system 500 can generate one or more simulated communication sessions (step 502). The below steps 504-512 further describe the process of generating a simulated communication session.
[0148] First, the system 500 can determine a topic for the simulated communication session (step 504). The system 500 can select the topic from a set of topics for the particular domain. For example, if the desired domain is a legal domain, the topic can be selected from a set of topics related to the legal domain, e.g.. criminal charges, civil case, contracts, wills, etc.
[0149] The system 500 can then determine, based on the topic, a context for the communication session (step 506). To determine the context for simulated communication session based on the topic, the system 500 can generate a respective set of values for each of one or more properties associated with the topic. The system 500 can then generate a context input from at least (i) the respective sets of values of the properties and (ii) a context that characterizes a specified format for the context. The system 500 can process the context input using the domain specific language model neural network to generate the context.
[0150] The topic and context are described in further detail with reference to FIG. 7.
[0151] The system 500 can generate, using the context, a simulated dialogue that includes a sequence of one or more dialogue rounds where each dialogue round includes a sequence of dialogue turns (step 508). Any given simulated dialogue includes a sequence of one or more dialogue rounds where each dialogue round includes a respective dialogue turn at each of a sequence of time steps, each time step corresponding either to the user agent or to the expert agent. For example, a dialogue round in a simulated dialogue can include dialogue turn A by the expert agent, dialogue turn B by the user agent, and dialogue turn C by the expert agent.
[0152] The below steps, 510 and 512, are substeps of step 508, describing the process of generating each dialogue turn in a dialogue round within a simulated dialogue.
[0153] For each dialogue turn, the system 500 can generate a dialogue input for the turn (step 510). The system 500 can generate a dialogue input based on the current status of the dialogue round and a respective prompt for the agent.
[0154] The system 500 can then process the dialogue input using the language model neural network to generate the dialogue turn (step 512).
[0155] The simulated dialogue is described in further detail with reference to FIG. 2.
[0156] The system 500 can then generate, from each simulated communication session, one or more training examples (step 514). Each training example can be generated from a dialogue turn within a respective simulated communication session and can include (i) a corresponding instruction, (ii) one or more preceding dialogue turns within the simulated communication session, and, as the target output within the training example, (iii) the dialogue turn.
[0157] The system 500 can then train the language model neural network on the one or more training examples (step 516). The language model neural network can be trained on task-specific instructions to fine-tune the model in playing either the patient or the doctor role in the simulated dialogues. The language model neural network can be trained to predict the next dialogue turn based on all previous interactions, assuming the doctor or patient role, and then compared to a target output. From each training example, one or more dialogue turns can be randomly sampled for each the doctor and patient roles as the target output to predict based on the conversation leading up to that target output. For example, 3 turns could be sampled for each the doctor and patient role as the target output.
[0158] Training the language model neural network is described in further detail with reference to FIG. 2.
[0159] FIG. 6 is a flow-diagram of step 508 of the process depicted in FIG. 5.
[0160] In some implementations, a moderator can be used to determine when the simulated dialogue has reached an end. The moderator can continually assess the ongoing dialogue between the user agent and the expert agent to determine when the dialogue round has reached a natural conclusion.
[0161] For each dialogue turn within a dialogue round, the system 600 can generate a moderator input for the dialogue turn (step 602). The moderator input for the dialogue turn can be generated from at least (i) the dialogue round after the last dialogue turn and (ii) a moderator prompt that specifies criteria for terminating the dialogue round.
[0162] For example, in a medical diagnostic domain, the moderator prompt can include criteria to terminate the dialogue round after the doctor agent provides a diagnosis, a treatment plan, and adequately addressed any remaining patient questions. In some implementations, the moderator prompt can include criteria to terminate the dialogue round if either agent, e.g.. expert or user, initiated a farewell, e.g., says goodbye. As a particular example, the moderator prompt can be the following:
[0163] ‘‘The following is a conversation between a doctor and a patient. <simulated dialogue> The conversation should only come to an end if the doctor has finished giving the patient a diagnosis and treatment plan and the patient has no questions left. A conversation also comes to an end if the doctor or patient says goodbye.’7
[0164] The system 600 can then process the moderator input using the language model neural network to generate a moderator output (step 604). As an example, the moderator output can be a single value denoting whether the conversation between the agents has come to an end. e.g., a ‘"yes” or "‘no”.
[0165] The system 600 can then determine, from the moderator output, whether to terminate the dialogue round after the dialogue turn and before performing any additional dialogue turns (step 606). For example, if the moderator output is a “yes’", then the system 600 can break out of the dialogue round between the expert agent and the user agent, and the dialogue history, e.g., the dialogue turns between the expert and the user agent up to that time step, is the final simulated dialogue for that dialogue round.
[0166] FIG. 7 is a flow diagram of step 506 of the process depicted in FIG. 5.
[0167] When determining, based on the topic, a context for the simulated communication session, the system 700 can first generate a respective set of values for each of one or more properties associated with the topic (step 702). In a particular example, the one or more properties associated with a medical condition topic can include patient demographics, symptoms, and a management plan.
[0168] When generating the respective set of values for each of one or more of the properties associated with the topic, the system 700 can obtain passages specifying ranges of values for each of the one or more properties and then filter the obtained passages using a second language model neural network.
[0169] For example, a topic of the simulated communication session can be a medical condition. In this particular example, the one or more properties associated with the topic can include patient demographics, symptoms, and management plans. The system 700 can then obtain passages specify ing ranges for each of these properties.
[0170] In a particular implementation, the system 700 can obtain passages for each of the properties using an internet search engine. The system can then use a general purpose LLM to filter the obtained passages to ensure that the passages are relevant to the given topic, e.g., condition.
[0171] A general purpose LLM is a large language model trained on diverse datasets so that the model can perform a variety of tasks such as text generation, summarization, and filtering. An example of a general purpose LLM are PaLM and PaLM 2. To use the general purpose LLM to filter the obtained passages, the LLM can receive the topic and the obtained passages and then can encode both into embeddings for comparison. The model can then calculate the similarity between the topic embedding and each passage embedding. As an example, the LLM can use cosine similarity to calculate the similarity between the two embeddings, e.g., the topic embedding and the passage embedding.
[0172] The LLM can then filter out the passages that do not reach a threshold score of similarity.
[0173] The general purpose LLM is used as the fine-tuned domain specific language model described in this specification has been specifically tailored for generation tasks, e.g., to create natural and contextually relevant conversations between two user agents in a particular domain as well as summarization. Thus, in some cases, the general purpose LLM that is trained for a variety of tasks is better equipped for performing a filtering task than the domain specific language model neural network described in this specification.
[0174] The system 700 can then generate a context input using the domain specific language model neural network (step 704). The context input can be generated from at least (i) the respective sets of values of the properties and (ii) a context prompt that characterizes a specified format for the context. For example, the context prompt can be a one-shot exemplar that demonstrates the specified format. As a particular example, the context input in a medical diagnostic domain, e.g.. a patient vignette, can be generated from the following information:
[0175] Condition: [Condition]
[0176] Demographic Passages: [Retrieved Demographic Passages] Symptoms Passages: [Retrieved Symptoms Passages]
[0177] Management Plan Passages: [Retrieved Management Plan Passages] Example Format: [Oneshot example]
[0178] The system 700 can then process the context input using the language model neural network to generate the context (step 706). For example, the system 700 can use the context input in a medical diagnostic domain to generate a context, e.g., patient vignette, for the given topic, e g., condition, by prompting the language model neural network to generate a context based on the above context input. Essentially, the general purpose LLM can simply filter the relevant passages and it is then the domain specific language model neural network that can use those relevant passages in the context input to generate a context for the language model neural network to use when generating the simulated dialogues.
[0179] FIG. 8 is a flow diagram of an example process of using the fine-tuned language model neural network as a domain specific conversational agent.
[0180] The system 800 can use the fine-tuned domain specific language model neural network as a domain-specific conversational agent, i.e., to serve as an expert agent when cartying out communication sessions with users. Any given communication session includes a respective dialogue turn at each of a sequence of time steps, each time step corresponding either to the user or to an expert agent and the system 800 uses the domain specific language model neural network to generate the dialogue turns at time steps corresponding to the expert agent (while receiving the dialogue turns at time steps corresponding to the user from the user).
[0181] The below steps describe the process of generating a dialogue turn as the expert agent using the domain specific language model neural network.
[0182] In particular, the system 800 can first generate, from the dialogue turns at preceding time steps in the sequence and a user analysis prompt that specifies one or more properties of the user, a first input (step 802). As an example, in a medical diagnostic domain, a user analysis prompt can specify summarizing the positive and negative symptoms provided by the user as well as any relevant medical / family / social history and demographic information, producing a current differential diagnosis, noting missing information needed for a more accurate diagnosis, and / or assessing confidence in the current differential diagnosis and highlighting its urgency. The system 800 can generate a first input based on a user analysis prompt that includes one or more of the above properties and the current conversation history, i.e., the dialogue turns at preceding time steps.
[0183] The system 800 can then process the first dialogue input using a language model neural network to generate a user analysis output that characterizes the one or more properties of the user (step 804).
[0184] The system 800 can then generate, from at least the user analysis output and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, a second input (step 806). As an example, in a medical diagnostic domain, a response prompt can include a response to the user’s, e.g.. patient’s, last message, further questions to acquire missing information to refine the differential diagnosis, and if necessary, a recommendation for immediate action, e.g., an emergency room visit. In some cases, the response prompt can include presenting the differential diagnosis if confident based on the available information.
[0185] In some implementations, the second dialogue input can be generated from at least the dialogue turns at preceding time steps in addition to the user analysis output and the response prompt.
[0186] The system 800 can then process the second dialogue input using the domain specific language model neural network to generate an initial response to the dialogue turn at the immediately preceding time step (step 808).
[0187] The system 800 can generate, from the initial response, the dialogue turn at the time step, e.g., by providing the initial response as the dialogue turn or by further refining the initial response, e.g., using the language model neural network again (step 810).
[0188] In some implementations, further refining the initial response includes generating a third input from at least the initial response, the dialogue turns at preceding time steps in the sequence, and a refinement prompt that specifies requirements for dialogue turns generated by the expert agent. For example, in a medical diagnostic domain, the requirements can be related to factuality and formatting of the response, e.g., avoid factual inaccuracies on patient facts and unnecessary' repetition, show empathy, and display in a clear format. The system 800 can then process the third input using the domain specific language model neural network to generate the dialogue turn at the time step.
[0189] In some implementations, the dialogue turns at time steps corresponding to the user are received through a user interface of a user device.
[0190] In some implementations, the generated dialogue turns at time steps corresponding to the expert agent are provided to the user through a user interface of the user device. For example, the generated dialogue turn can be provided as text displayed in the user interface. As another example, the generated dialogue turn can be provided as speech through a speaker of the user device.
[0191] This specification uses the term ‘"configured'’ in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly- embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0192] The term "‘data processing apparatus"’ refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0193] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0194] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0195] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0196] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0197] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory’ devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory' devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0198] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory' feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0199] Data processing apparatus for implementing machine learning models can also include, for example, special -purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0200] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Jax framework.
[0201] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g.. a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0202] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g.. an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0203] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0204] Similarly, while operations are correspond toed in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0205] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes correspond toed in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0206] What is claimed is:
Claims
CLAIMS1. A method performed by one or more computers and for fine-tuning a language model neural network for use in participating in user communication sessions in a particular domain, the method comprising, at each fine-tuning iteration in a sequence of fine-tuning iterations: generating a plurality of simulated communication sessions, wherein each simulated communication session is a communication session between a plurality' of agents that comprise an expert agent for the particular domain and one or more user agents, and wherein the generating comprises, for each simulated communication session: determining atopic for the simulated communication session; determining, based on the topic, a context for the simulated communication session; and generating, using the context, a simulated dialogue that comprises a sequence of one or more dialogue rounds, wherein each dialogue round comprises a sequence of dialogue turns that each correspond to one of the plurality of agents, and wherein generating the simulated dialogue comprises, for each dialogue turn within each dialogue round: generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a respective prompt for the agent corresponding to the dialogue turn; and processing the dialogue input using the language model neural network to generate the dialogue turn; generating, from each simulated communication session, one or more training examples for training the language model neural network; and training the language model neural network on training data comprising the one or more training examples.
2. The method of any preceding claim, wherein generating the simulated dialogue comprises, for each dialogue turn within each dialogue round: generating a moderator input for the dialogue turn from at least (i) the dialogue round after the dialogue turn and (ii) a moderator prompt that specifies criteria for terminating the dialogue round; processing the moderator input using the language model neural network to generate a moderator output: and determining, from the moderator output, whether to terminate the dialogue round after the dialogue turn and before performing any additional dialogue turns.
3. The method of any preceding claim, wherein: the sequence of one or more dialogue rounds comprises a plurality of dialogue rounds, generating the simulated dialogue comprises, after each dialogue round other than a last dialogue round in the sequence: generating a critic input for the dialogue round from at least (i) the dialogue turns in the dialogue round and (ii) a critic prompt that specifies criteria for evaluating the expert agent; and processing the critic input using the language model neural network to generate a critic output that provides natural language feedback to the expert agent, and for each dialogue round after a first dialogue round in the sequence of dialogue rounds, generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a prompt for the agent corresponding to the dialogue turn comprises, when the dialogue turn corresponds to the expert agent, generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) a respective prompt for the expert agent, and (iii) the critic output generated after the preceding dialogue round.
4. The method of claim 3, wherein, for each dialogue round after the first dialogue round in the sequence of dialogue rounds, generating a dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn and (ii) a prompt for the agent corresponding to the dialogue turn comprises, when the dialogue input corresponds to the expert agent, generating the dialogue input for the dialogue turn from (i) the dialogue round as of the dialogue turn, (ii) the respective prompt for the expert agent, (iii) the critic output generated after the preceding dialogue round, and (iv) dialogue turns corresponding to the expert agent in the preceding dialogue round.
5. The method of claim 3 or claim 4, wherein generating, from each simulated communication session, one or more training examples for training the language model neural network comprises: generating, from only the last dialogue round within each simulated communication session, one or more training examples for training the language model neural network.
6. The method of any preceding claim, wherein determining a topic for the simulated communication session comprises: selecting the topic from a set of topics for the particular domain.
7. The method of any preceding claim, wherein determining, based on the topic, a context for the simulated communication session comprises: generating a respective set of values for each of a plurality of properties associated with the topic; generating a context input from at least (i) the respective sets of values of the properties and (ii) a context prompt that characterizes a specified format for the context; and processing the context input using the language model neural network to generate the context.
8. The method of claim 7, wherein generating a respective set of values for each of a plurality of properties associated with the topic comprises: obtaining passages specifying ranges of values for each of the plurality of properties; and filtering the obtained passes using a second language model neural network.
9. The method of any preceding claim, wherein the sequence of fine-tuning iterations comprises a plurality of fine-tuning iterations.
10. The method of any preceding claim, wherein generating, from each simulated communication session, one or more training examples for training the language model neural network comprises: generating, from only one or more of the dialogue rounds within each simulated communication session and not from the context for the simulated communication session, one or more training examples for training the language model neural network.
11. The method of any preceding claim, wherein the training data comprises the training examples and a static set of training examples relating to the particular domain.
12. The method of any preceding claim, wherein the particular domain is a medical diagnostic domain, the expert agent for the particular domain is a clinician, and the topic for the simulated communication is a medical condition to which the simulated communication session relates.
13. The method of claim 12, wherein the context is a patient vignette of a patient having the medical condition.
14. The method of claim 12 or 13, wherein language model neural network for use in participating in user communication sessions in a particular domain is configured to output, for a particular user communication session, a diagnosis of a medical condition.
15. The method of any of claims 1 to 11. wherein the particular domain is a physical, electronic or electrical system diagnostic domain, the expert agent for the particular domain is an engineer, and the topic for the simulated communication is a system fault condition to which the simulated communication relates.
16. The method of claim 15, wherein the context is a description of the physical, electronic or electrical system diagnostic domain having the system fault condition.
17. The method of claim 15 or 16, wherein language model neural network for use in participating in user communication sessions in a particular domain is configured to output, for a particular user communication session, a diagnosis of a system fault condition.
18. A method performed by one or more computers and for carrying out a communication session in a particular domain with a user, wherein the communication session comprises a respective dialogue turn at each of a sequence of time steps, each timestep corresponding either to the user or to an expert agent, and the method comprising, at each time step corresponding to the expert agent: generating, from the dialogue turns at preceding time steps in the sequence and a user analysis prompt that specifies one or more properties of the user, a first input; processing the first input using a language model neural network to generate a user analysis output that characterizes the one or more properties of the user; generating, from at least the user analysis output and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, a second input; processing the second input using the language model neural network to generate an initial response to the dialogue turn at the immediately preceding time step; and generating, from the initial response, the dialogue turn at the time step.
19. The method of claim 18, wherein dialogue turns at time steps corresponding to a user are received through a user interface of a user device, and wherein the method further comprises, at each time step corresponding to the expert agent, providing the generated dialogue turn to the user through a user interface of the user device.
20. The method of claim 19, wherein providing the generated dialogue turn comprises providing the dialogue turn as text displayed in the user interface.
21. The method of claim 19, wherein providing the generated dialogue turn comprises providing the dialogue turn as speech through a speaker of the user device.
22. The method of any one of claims 19-21, wherein generating, from at least the user analysis output and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, a second input comprises: generating, from at least the user analysis output, the dialogue turns at preceding time steps in the sequence, and a response prompt that specifies requirements for a response to the dialogue turn at an immediately preceding time step, the second input.
23. The method of any one of claims 19-22, wherein generating, from the initial response, the dialogue turn at the time step comprises: generating, from at least the initial response, the dialogue turns at preceding time steps in the sequence, and a refinement prompt that specifies requirements for dialogue turns generated by the expert agent, a third input; and processing the third input using the language model neural network to generate the dialogue turn at the time step.
24. The method of any one of claims 19-23, wherein the particular domain is a medical diagnostic domain and the expert agent for the particular domain is a clinician dialogue agent.
25. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the operations of the respective method of any preceding claim.
26. A computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1 to 24.
Citation Information
Patent Citations
Method and system for generating conversation summary
US11709989B1
Cognitive natural language processing software framework optimization
US20230072003A1
Systems and methods for generating training data for sequential conversational responses
US20230108855A1
Multimode Conversational Agent using a Pattern-Completion Engine
US20230336504A1
Cited By
Tool calling method in task-based dialogue, medium, equipment and product
CN120523922A
Text generation method and device based on large model, training method and device, equipment and medium
CN121599118A