Systems and methods for facilitating conversation based on text data and speech data
The system addresses the limitations of existing machine learning models by using enhanced token generation and state determination to retrieve relevant data, resulting in contextually accurate and user-centric responses.
Patent Information
- Application Number
- PCT/EP2024/087991
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Existing machine learning models for generating responses are prone to errors, limited by context windows, and susceptible to hallucinations, failing to consider user knowledge, interests, and evolving states during conversations.
A system that receives user inputs, generates enhanced tokens using a first machine learning model, determines a state based on previous interactions, retrieves predetermined data from a database, and uses a second machine learning model to generate responses, ensuring contextual and topical relevance.
The system effectively generates responses that are contextually relevant, reduce hallucinations by relying on predetermined data, and adapt to the user's evolving state, providing more accurate and user-centric interactions.
Smart Images

Figure EP2024087991_26062025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR FACILITATING CONVERSATION BASED ON TEXTDATA AND SPEECH DATACROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 613,839, filed December 22, 2023, and titled “SYSTEMS AND METHODS FOR GENERATING RESPONSES BASED ON PREDETERMINED DATA AND USING MACHINE LEARNING MODELS,” which is incorporated herein by reference in its entirety.FIELD
[0002] One or more embodiments described herein relate to systems and computerized machine learning methods for generating responses based on predetermined data.BACKGROUND
[0003] Some known machine learning models can generate responses that indicate erroneous information. As such, it can be desirable to have systems configured to retrieve predetermined data to be used as context by machine learning models to generate responses.SUMMARY
[0004] In an embodiment, a method includes receiving a user input from a user device and generating a set of user input tokens based on the user input. The method also includes generating a set of enhanced input tokens by providing the set of user input tokens as input to a first machine learning model. A state is determined based on a previous state and at least one of the set of user input tokens or the set of enhanced input tokens. Predetermined data is retrieved from a database based on the state and at least one of the set of user input tokens or the set of enhanced input tokens. The method also includes generating a set of response tokens by providing the set of user input tokens and the predetermined data as input to a second machine learning model. Based on the set of response tokens, a response is sent to the user device.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] FIG. 1 is a schematic diagram of a system, according to an embodiment.
[0006] FIG. 2 is a schematic diagram of a compute device included in a system, according to an embodiment.
[0007] FIG. 3 is a schematic diagram of logic components included in a system, according to an embodiment.
[0008] FIG. 4 is a signal diagram showing a plurality of interactions implemented by a system, according to an embodiment.
[0009] FIG. 5 is a flowchart showing a method of using a system to generate a response based on predetermined data, according to an embodiment.DETAILED DESCRIPTION
[0010] Some known machine learning models are configured to generate text. For example, some known large language models (LLMs) can perform as chat agents (e.g., “chat hots”), generating responses based on inputs received from a user. Such known LLMs can, in some instances, use previous conversations (e.g., dialogue) to generate a next response. However, given that context windows (e.g., the tokens and / or text segments used as input to an LLM to generate text) can be limited by known scaling laws, memory constraints, processing resources, and / or the like, the amount of previous conversation data used as context for some known LLMs is limited. Moreover, in some instances, some known LLMs over-emphasize previous conversation data as compared to instruction prompts (e.g., inputs, queries, etc.) received more recently from a user. As a result, previous dialogue can outweigh and / or dominate prompts to these known LLMs, disproportionately influencing responses generated by these known LLMs.
[0011] Additionally, some known LLMs can be prone to hallucination, generating fictional and / or misleading content. Moreover, hallucinations generated by such known LLMs can fail to be mutually consistent, given that context to these known LLMs is limited to the instruction prompt and the conversation history. Some known LLMs also do not determine or update a state based on a progression of a conversation. More specifically, such known LLMs are not affected by a state that can indicate, for example, a perceived and / or predicted level of comprehension, understanding, appreciation, etc., associated with the user and with respect to atopic, subject, knowledge domain, etc. As a result, such known LLMs do not consider a user’s knowledge, interests, mental state, etc., while generating a response. Thus, a need exists for a machine-learning based system that can implement an agent (e.g., a teacher agent and / or the like, described herein) having extended conversation context from which to generate aresponse. A need also exists for a machine-learning based system that can generate a response having contextual and / or topical content based on data retrieved from a memory. Additionally, a need exists for a machine-learning based system that can generate a response based on an evolving state (e.g., a state indicating a user’s comprehension) that is updated based on, for example, instruction prompts received from the user.
[0012] FIG. 1 is a schematic diagram of a system 100 for generating responses based on received inputs, according to an embodiment. The system 100 includes a user device 110, respondent compute device 120, database 130, and network Nl. The system 100 can include alternative configurations, and various steps and / or functions of the processes described below can be shared among the various devices of the system 100 or can be assigned to specific devices (e.g., the user device 110, the respondent compute device 120, the database 130, and / or the like). For example, in some configurations, a user can provide inputs directly to the respondent compute device 120 rather than via the user device 110, as described herein.
[0013] In some embodiments, the user device 110 and / or the respondent compute device 120 can include any suitable hardware-based computing devices and / or multimedia devices, such as, for example, a server, a desktop compute device, a smartphone, a tablet, a wearable device, a laptop and / or the like. In some implementations, the user device 110 and / or the respondent compute device 120 can be implemented at an edge node or other remote computing facility. In some implementations, each of the user device 110 and / or respondent compute device 120 can be a data center or other control facility configured to run a distributed computing system and can communicate with other compute devices.
[0014] In some implementations, the user device 110 can include a peripheral device(s) (not shown in FIG. 1) configured to receive an input from a user. For example, the user device 110 can include a keyboard that the user can use to generate text data to be used as an instruction prompt by the respondent compute device 120. In another implementation, the user device 110 can include a graphical user interface (GUI) that a user can use to define and / or articulate the instruction response (e.g., by choosing from a list of predefined instruction prompts). In yet another implementation, the user device 110 can include a microphone, camera, and / or video camera, and the user device 110 and / or the respondent compute device 120 can be configured to convert audio, image, and / or video data to text data and / or tokens (described herein) to be used to generate a response. For example, the user device 110 and / or the respondent compute device 120 can use facial recognition techniques to infer a user’s mental state (e.g., apathy,boredom, concern, shock, confusion, anxiety, and / or the like) from a facial expression of the user depicted in video data. This inferred mental state can be communicated to the respondent compute device 120 (e.g., as a set of tokens) to be considered while generating a response.
[0015] The respondent compute device 120 can be configured to generate a response to a user based on input data received from the user (e.g., via the user device 110). For example, the respondent compute device 120 can be configured to execute (e.g., via a processor) a conversation management application 112, which can be functionally and / or structurally equivalent to the conversation management application 212 of FIG. 2, described herein. The conversation management application 112 can be implemented via software and / or hardware and can, for example, act as a participant (e.g., an agent) in a dialogue with a user. Similarly stated, the conversation management application 112 can assume the role of one person (e.g., an educator, instructor, counsellor, medical professional, therapist, etc.) in a two-person dialogue. In some instances, this dialogue can occur in the context of educating the user. For example, the user can be a student receiving instruction on a subject, a patient receiving counselling for a symptom and / or treatment, and / or the like. To provide instructive content to the user that the user would find helpful, the conversation management application 112 can generate responses based on predetermined data. This predetermined data can include, for example, stored content that is (1) factual, educational, instructive, and / or the like, (2) relevant to the user’s input (e.g., query), and / or (3) appropriate based on the user’s perceived and / or predicted knowledge level. The conversation management application 112 can be further configured to evaluate the performance and / or progression of a user (e.g., a student, patient, etc.) based on the inputs received from the user and / or the predicted state associated with the user.
[0016] As described herein (e.g., in relation to FIGS. 2-4), the conversation management application 112 can also determine and / or estimate a state (e.g., a state associated with a user’s comprehension and / or knowledge of a subject) and generate responses based on this state. For example, the conversation management application 112 can retrieve predetermined data (e.g., backstory data) from a memory (e.g., the database 130, described herein) based on the state and cause that predetermined data to be considered within the context window. The conversation management application 112 can also be configured to summarize previous conversations with the user and generate subsequent responses based on those summaries.
[0017] In some implementations, the respondent compute device 120 can generate text data that can cause a response to be communicated to the user. For example, in some implementations, the user device 110 and / or the respondent compute device 120 can display the text data to the user (e.g., via a display operably coupled to the user device 110 and / or the respondent compute device 120). Alternatively and / or in addition, the user device 110 and / or the respondent compute device 120 can cause an audio signal (e.g., speech data) to be generated based on the generated text data, such that the response can be vocalized to the user (e.g., via a speaker operably coupled to the user device 110 and / or the respondent compute device 120). In some implementations, the audio signal can indicate a tone of voice (e.g., enthusiastic, pessimistic, empathetic, concerned, etc.) using, for example, a synthesized speech markup language (SSML) and based on a determined state (described herein in relation to, for example, the state determinator 310 of FIG. 3) of the agent and / or user. Alternatively and / or in addition, the user device 110 and / or the respondent compute device 120 can cause video data to be generated based on the text data, such that, for example, the user can view (e.g., via the operably coupled display) an animation of a simulated entity (e.g., instructor, educator, counsellor, professional, and / or the like) providing the generated response.
[0018] The database 130 can include at least one memory, repository and / or other form of data storage. The database 130 can be in communication with the user device 110 and / or the respondent compute device 120 (e.g., via the network Nl). In some implementations, the database 130 can be housed and / or included in one or more of the user device 110, the respondent compute device 120, or a separate compute device(s). The database 130 can be configured to store, for example, predetermined data and / or previous response data that can be retrieved or otherwise accessed by one or more compute devices, such as, for example, the respondent compute device 120, to perform at least some of the features (e.g., in relation to the conversation management application 112) described herein.
[0019] The database 130 can include a computer storage, such as, for example, a hard drive, memory card, solid-state memory, ROM, RAM, DVD, CD-ROM, write-capable memory, and / or read-only memory. In addition, the database 130 may include a distributed storage system where data is stored on a plurality of different storage devices, which may be physically located at a same or different geographic location (e.g., in a distributed computing system). In some implementations, the database 130 can be associated with cloud-based / remote storage.
[0020] Database 130 can be networked, via the network Nl, to the user device 110 and / or the respondent compute device 120 directly using wired connections and / or wireless connections. The network Nl can include various configurations and protocols, including, for example, short range communication protocols, Bluetooth®, Bluetooth® LE, the Internet, World Wide Web, intranets, virtual private networks, wide area networks, local networks, private networks using communication protocols proprietary to one or more companies, Ethernet, WiFi and HTTP, cellular data networks, satellite networks, free space optical networks and / or various combinations of the foregoing. Such communication can be facilitated by any device capable of transmitting data to and from other compute devices, such as a modem(s) and / or a wireless interface(s).
[0021] In some implementations, although not shown in FIG. 1, the system 100 can include a plurality of user devices 110, respondent compute devices 120, and / or databases 130. For example, in some implementations, the system 100 can include a plurality of user devices 110, where each user device 110 can be associated with a different user from a plurality of users. Each user can be associated with a state (e.g., a comprehension level, etc., as described herein), and the respondent compute device 120 can be configured to track each state associated with each user and / or generate responses for respective users based on the states of those respective users. In some implementations, a plurality of user devices 110 can be associated with a single user, where each user device 110 can be associated with a different input modality (e.g., text input, audio input, video input, etc.).
[0022] FIG. 2 is a schematic diagram of a respondent compute device 201 of a system, according to an embodiment. The respondent compute device 201 can be structurally and / or functionally similar to, for example, the respondent compute device 120 of the system 100 shown in FIG. 1. The respondent compute device 201 can be a hardware-based computing device, a multimedia device, or a cloud-based device such as, for example, a computer device, a server, a desktop compute device, a laptop, a smartphone, a tablet, awearable device, a remote computing infrastructure, and / or the like. The respondent compute device 201 includes a memory 210, a processor 220, and a network interface 230 operably coupled to a network N2.
[0023] The processor 220 can be, for example, a hardware based integrated circuit (IC), or any other suitable processing device configured to run and / or execute a set of instructions or code (e.g., stored in memory 210). For example, the processor 220 can be a general-purpose processor, a central processing unit (CPU), an accelerated processing unit (APU), anapplication specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic array (PLA), a complex programmable logic device (CPLD), a programmable logic controller (PLC), a remote cluster of one or more processors associated with a cloud-based computing infrastructure and / or the like. The processor 220 is operatively coupled to the memory 210 (described herein). In some embodiments, for example, the processor 220 can be coupled to the memory 210 through a system bus (for example, address bus, data bus and / or control bus).
[0024] The memory 210 can be, for example, a random-access memory (RAM), a memory buffer, a hard drive, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and / or the like. The memory 210 can store, for example, one or more software modules and / or code that can include instructions to cause the processor 220 to perform one or more processes, functions, and / or the like. In some implementations, the memory 210 can be a portable memory (e.g., a flash drive, a portable hard disk, and / or the like) that can be operatively coupled to the processor 220. In some instances, the memory can be remotely operatively coupled with the respondent compute device 201, for example, via the one or more network interface controllers 240. For example, a remote database server can be operatively coupled to the respondent compute device 201.
[0025] The memory 210 can store various instructions associated with processes, algorithms and / or data, including machine learning models (e.g., LLMs, classifiers, tokenizers, embedders, etc., as described herein). Memory 210 can further include any non-transitory computer-readable storage medium for storing data and / or software that is executable by processor 220, and / or any other medium which may be used to store information that may be accessed by processor 220 to control the operation of the respondent compute device 201. For example, the memory 210 can store data associated with the conversation management application 212. The conversation management application 212 can be functionally and / or structurally similar to the conversation management application 112 of FIG. 1 and / or the conversation management application 312 of FIG. 3, described herein.
[0026] The conversation management application 212 can include a summary generator 214 and a data retriever 216. The summary generator 214 can be functionally and / or structurally similar to the summary generator 314 of FIG. 3 and / or the summary generator 414 of FIG. 4, described in further detail herein. The data retriever 216 can be functionally and / or structurally similar to the data retriever 316 of FIG. 3 and / or the data retriever 416 of FIG. 4, described infurther detail herein. As described herein, the summary generator 214 can be configured generate a summary of a conversation history (e.g., previous dialogue) included in a context window. The conversation management application 212 can then replace, within the context window, the set of tokens representing the conversation history with a set of summary tokens representing the generated summary. As described herein, the data retriever 216 can be configured to retrieve predetermined data (e.g., data representing instructive, educational, and / or factual content) and inject that predetermined data into the context window. As a result, the conversation management application 212 can generate responses based on the predetermined data. The predetermined data can be retrieved based on a state, a topic of conversation, and / or an importance metric associated with the topic of conversation, as described herein.
[0027] The network interface 230 can be configured to connect to the network N2, which can be functionally and / or structurally similar to the network N1 of FIG. 1. For example, network N2 can use any of the wired and wireless short range communication protocols described above with respect to network N1 of FIG. 1.
[0028] In some instances, the respondent compute device 201 can further include a display, an input device, and / or an output interface (not shown in FIG. 2). The display can be any display device by which the respondent compute device 201 can output and / or display data. The input device can include a mouse, keyboard, touch screen, voice interface, and / or any other handheld controller or device or interface via which a user may interact with the respondent compute device 201. The output module can include a bus, port, and / or other interfaces by which the respondent compute device 201 may connect to and / or output data to other devices and / or peripherals, such as a speaker.
[0029] FIG. 3 is a schematic diagram of logic components 300 for generating a response, according to an embodiment. The logic components 300 can be associated with a compute device (e.g., a compute device that is structurally and / or functionally similar to the respondent compute device 201 of FIG. 2 and / or the respondent compute device 120 of FIG. 1). In some instances, for example, the logic components 300 can be implemented as software stored in memory 210 and configured to be executed via the processor 220 of FIG. 2. In some instances, for example, at least a portion of the logic components 300 can be implemented in hardware. The logic components 300 include an input 302, a response 304, and a conversation management application 312 (which can be functionally and / or structurally similar to, forexample, the conversation management application 112 of FIG. 1 and / or the conversation management application 212 of FIG. 2). The conversation management application 312 includes a tokenizer 306, a context window 308, an input enhancer 309, a state determinator 310, a summary generator 314, a data retriever 316, and a machine learning model 320.
[0030] The input 302 can include text data (e.g., string data, phrases, sentences, questions, natural language, etc.) generated (e.g., written, dictated, acted out through gestures and / or sign language, etc.) by a user (e.g., a patient, student, learner, etc.). In some implementations, the input 302 can include or be based on image data and / or video data. For example, a user can submit an image of a symptom and / or the like. The user can provide the text data as an instruction prompt to the conversation management application 312 to solicit a response and / or cause a state transition within the conversation management application 312 (e.g., to an improved state, such as a state indicating improved comprehension, as described in relation to the state determinator 310 herein). The input 302 can include, for example, a question associated with a subject and / or topic. Such queries can include, for example: “What do I need to do before my session?”, “How long does treatment take?”, “What will the drug do to my brain?”, “Will psilocybin make me feel like I am dying?”, and / or the like. In some implementations, the input 302 (e.g., the text data) can be stored in a memory (e.g., the database 130 of FIG. 1), such that an entity (e.g., an educator, medical professional, content moderator, reviewer, etc.) can review the input 302 (e.g., in substantially real-time or at a later point in time).
[0031] The tokenizer 306 can segment a string of text (e.g., natural language), and the resulting segments (e.g., words, sub-words, characters, punctuation, and / or the like) can be referred to as “tokens.” In some instances, a token can be a numerical representation of a discrete element (e.g., a word, sub-word, etc.) of natural language text. For example, a token can be assigned an integer value that can range from zero to the size of a vocabulary to be considered. This process (referred to herein as tokenization) can permit a machine learning model (e.g., a large language model and / or a transformer) to infer semantic meaning from natural language text based on a decomposed representation of that text. Tokenization can be performed by, for example, a byte pair encoding algorithm and / or any other suitable tokenization method.
[0032] The tokenizer 306 can take as input the text data included in the input 302 and generate a set of input tokens. In some implementations, the tokenizer 306 can be included in the machine learning model 320, as described herein. In some implementations, as described inrelation to the data retriever 316, the tokenizer 306 (and / or a tokenizer structurally and / or functionally similar to the tokenizer 306) can generate a set of predetermined tokens. In some implementations, as described in relation to the input enhancer 309, the tokenizer 306 (and / or a tokenizer structurally and / or functionally similar to the tokenizer 306) can generate a set of enhanced input tokens. Such sets of tokens can be included in the context window 308, as can a set of summary tokens generated by the summary generator 314, as described herein.
[0033] The context window 308 can be a collection of tokens (e.g., a set(s) of input tokens, a set(s) of response tokens, a set(s) of predetermined tokens, a set(s) of previous response tokens, a set(s) of summary tokens, etc.) stored in a memory (e.g., a memory structurally and / or functionally similar to the memory 210 of FIG. 2) and to be used as context by the machine learning model 320 while interpreting the input 302 and / or generating the response 304. Tokens can be added to and / or removed from such a collection of tokens (e.g., the context window 308) as a conversation between the user and the agent progresses. For example, tokens associated with older dialogue can be removed from the collection of tokens (e.g., the context window 308) in response and / or contemporaneous to tokens associated with more recent dialogue being included in the collection of tokens (e.g., the context window 308). As a result, the conversation management application 312 can infer the meaning of an instruction prompt and / or an appropriate response based on the more recent dialogue (as opposed to the older dialogue).
[0034] Illustrating the context window 308 in use, the machine learning model 320 can interpret (e.g., determine a semantic meaning of) a token (e.g., a target token) included in a sequence of input tokens within the context window 308 based on the remaining tokens within the context window 308. For example, if the input 302 indicates the question “What’s best?”, the machine learning model 320 can interpret “best” to refer to, for example, “the best treatment” if a token representing “treatment” was included in the context window 308 (e.g., as a result of a previous dialogue discussing treatment options and / or symptoms). The machine learning model 320 can then generate a response (e.g., “I recommend rest, ice, and elevation”) based on the context inferred from the tokens in the context window 308. In this example, the machine learning model 320 would not respond with “Dogs are best,” for example, even though older dialogue (e.g., dialogue associated with tokens since removed from the context window 308) between the user and the agent may have been related to, for example, pets. The context window 308 (e.g., the collection of tokens) can have a fixed size (e.g., a fixed number of tokens)or a variable size (e.g., a number of tokens based on a position of the target token in the sequence).
[0035] The input enhancer 309 can include a machine learning model (e.g., a first machine learning model, such as an LLM, in addition to the machine learning model 320) configured to generate a set of enhanced input tokens. The set of enhanced input tokens can represent a derived input based on the input 302 (e.g., a paraphrase of the input 302, a restatement of a meaning of the input 302, and / or the like). In some instances, the set of enhanced input tokens can have semantic meaning that is less apparent in the input 302. For example, the input 302 provided by the user can be ill-posed, ambiguous, vague, etc. To illustrate the input enhancer 309 in use, the input 302 can include the question, “Will psilocybin make me feel like I am dying?” Taking as input the input 302 (e.g., text data) and / or a set of input tokens generated based on the input 302, the input enhancer 309 can generate one or more sets of enhanced input tokens representing derived inputs (e.g., a paraphrase and / or restatement). Such derived inputs can include, for example, “Can the consumption of psilocybin induce sensations that mimic the feeling of dying?”, “Is it possible for psilocybin to create an experience that resembles the sensation of dying?”, and / or “Does psilocybin have the potential to make me feel as if I am on the verge of death?” The one or more sets of enhanced input tokens and / or the set of input tokens can be provided as input to at least one of the state determinator 310 (described herein) or the context window 308.
[0036] The state determinator 310 can determine a state based on the set of input tokens generated by the tokenizer 306 in response to receiving the input 302 (e.g., the user’s questions, prompts, etc.). Alternatively and / or in addition, the state determinator 310 can determine the state based on the set of enhanced input tokens generated by the input enhancer 309. Alternatively and / or in addition, although not shown in FIG. 3, in some implementations, the state determinator 310 can receive text data (e.g., rather than or in addition to tokens generated by the tokenizer 306) as input to determine a state. In such implementations, the state determinator can receive text data directly from the input 302. The state can indicate a predicted and / or perceived level (e.g., elementary, intermediate, advanced, etc.) of comprehension and / or knowledge that a user has with respect to a topic and / or subject. In some implementations, the state can indicate a percentage (e.g., a percentage of content that a user has knowledge of and / or mastery of). The state can also indicate, for example, a predicted emotional state, psychological state, mood, and / or the like, to be considered by the conversation management application 312in generating the response 304. Specific examples of a state can include, for example, a degree of interest in and / or attention paid to a topic, a degree of enjoyment of a topic, a degree of comprehension of a topic, an evolving mood state associated with overall contentment, depression, etc., and / or the like. The state can be influenced by, for example, the content, tone, etc., of the user’s prompt included in the input 302. As described below, the state can evolve over time based on the user’s knowledge, mood, etc. By determining the state, the state determinator 310 can cause the data retriever 316 (described herein) to retrieve predetermined data (e.g., educational and / or instructive content) that is appropriate for and / or can be understood by the user given the user’s level of comprehension, interest, etc., as indicated by the state.
[0037] In some implementations, the conversation management application 312 can determine a state independent of the input 302 provided by the user. For example, the conversation management application 312 can cause a transition to a state based on a passing of a predefined period of time (e.g., a predefined education timeline, described herein in relation to the data retriever 316) and / or a number of exchanges in a conversation. As a result, the agent, implemented by the conversation management application 312, can simulate, for example, a teacher that can challenge the user (e.g., a student) rather than base the instruction rate on the user’s learning and / or comprehension rate.
[0038] In some implementations, the state determinator 310 can classify the instruction prompt included in the input 302 by the level of comprehension indicated by the instruction prompt and / or by topic. For example, if the input 302 includes the statement “What is depression,” the state determinator 310 can perform a classification of the statement to predict that the user has a novice understanding of depression. Alternative, if the input 302 includes, for example, the statement “What does a score of 9 on the Montgomery-Asberg Depression Rating Scale [MADRS] mean?”, the state determinator 310 can perform a classification of the statement to predict that the user has an advanced understanding of depression. While the agent is simulating a therapist and the user is a patient, for example, a topic indicated by the state can include a treatment, symptom, event (e.g., a death in the family), diagnosis, and / or the like.
[0039] The state determinator 310 can implement zero-shot topic classification using, for example, a machine learning model, such as BART-MNLI, TARS, GPT, and / or the like. As described herein in relation to the data retriever 316, the topic classification can be used by the conversation management application 312 to locate and retrieve relevant data from a memoryto add to the context window 308. In some implementations, the state determinator 310 can use a machine learning model to generate an embedded vector based on the input 302. This machine learning model can include, for example, a transformer model and / or a similar model configured to identify relationships in sequential data and / or natural language, such as the machine learning model 320 or a machine learning model separate from the machine learning model 320. This embedded vector can indicate a semantic meaning of an instruction prompt of the input 302 and can be used by the data retriever 316 to locate and retrieve additional data from a memory to add to the context window 308, as described herein.
[0040] In some implementations, the state determinator 310 can generate a psychological metric based on the input 302 and / or the topic. The psychological metric can include, for example, an importance metric (e.g., an importance score) and / or a valence metric (e.g., a valence score). The importance metric can be a subjective (as to the user) measure of how important and / or significant a topic is (e.g., a believed importance). The valence metric can indicate a degree to which a topic is associated with a positive sentiment, neutral sentiment, negative sentiment, a sentiment therebetween, and / or the like. For example, a topic related to a death of a parent can have a high importance metric and a low valence metric, whereas a topic related to a recently viewed movie can have a low importance metric and a high valence metric.
[0041] In some instances, the state (e.g., the psychological state, the importance metric, the valence metric, etc.) can be represented by a numerical metric (e.g., a continuous variable) and / or a discrete state. The state can have an initial state (e.g., based on configuration data, described herein), and, as a conversation and / or relationship with a user progresses, the state can change and / or evolve. A user receiving counselling, for example, can deem a topic to be important or unimportant (as indicated by an importance metric), but can change opinion on that topic’s importance over time. In some instances, the agent implemented by the conversation management application 312 can be tasked with causing the user to transition to a goal state (e.g., a symptom-free state, a relaxation state, an improved state, etc.).
[0042] In some implementations, a state can be represented as a numerical value (e.g., a predicted percentage of content grasped by the user), which can be rounded to a nearest discrete value. This discrete value can be used as a key to search for and / or look up (e.g., in a lookup table, described herein in relation to the data retriever 316) and retrieve data (e.g., predetermined natural language text) indexed according to the key. Alternatively and / or inaddition, a state can be represented as a discrete node within a graph (e.g., a knowledge graph 318, described herein), a Markov decision process (MDP), and / or the like. The MDP can include, for example, a plurality of nodes (e.g., a state space), where each node can be associated with a different state (e.g. a level of comprehension, an interest, etc.). A transition from one state (e.g., node) to another state can be caused by, for example, atopic and / or subject being referenced in the input 302 by the user (e.g., as determined using zero-shot topic classification, described above), an importance and / or valence metric for the topic being above or below a threshold value, a specific word and / or phrase being included in the instruction prompt of the input 302, etc.
[0043] In some instances, the state determinator 310 can perform zero-shot classification on an input 302 to generate a high-risk topic classification if, for example, the input 302 suggests that a user might be discussing and / or contemplating self-harm and / or harm to others (e.g., based on keywords suggestive of harm being included in the input 302 and / or a tone of voice reflected in an audio signal associated with the input 302). For example, a high-risk topic classification can indicate, based on inferences generated from the input 302, that the user is expressing suicidal ideation and / or suicidal contemplation. In response to generating a high- risk topic classification based on the input 302, the conversation management application 312 can be configured to cause transmission of a signal indicating the high-risk classification to, for example, a crisis professional and / or the like.
[0044] In some implementations, the conversation management application 312 can have an initial state (e.g., an initial prediction of a user’s knowledge level and / or comprehension level) that can be defined based on an indication of preliminary content conveyed to the user prior to the conversation management application 312 receiving an input 302 from the user. For example, for a user that is a patient, the user can be provided with content (e.g., online content, printed materials, video content, audio content, etc.) that provides the user with an introduction to a topic (e.g., a treatment, symptom, diagnosis, etc.) to be discussed further with the agent. Based on a signal generated (1) automatically based on the user interacting with (e.g., clicking on, viewing, etc.) the content and / or (2) based on a user filling out a questionnaire to indicate the content that the user has already viewed, the conversation management application 312 can generate an initial state. Based on this initial state, the conversation management application 312 can retrieve predetermined data that is not redundant to the preliminary content already consumed by the user, so as to prevent annoying and / or boring the user.
[0045] In some instances, the initial state can represent, for example, an initial psychological state (e.g., a state of depression), level of interest, etc., that the agent is tasked with improving through conversation. In some implementations, the initial state can be a classification generated based on a psychometric instrument and / or a questionnaire (e.g., a questionnaire associated with: the Montgomery-Asberg Depression Rating Scale (MADRS), the Hamilton Depression Rating Scale (HAM-D), the Hamilton Anxiety Rating Scale (HAM-A), etc.). Alternatively or in addition, the psychometric instrument can be associated with a structured interview. For example, rather than a questionnaire that is completed by a user who is selfreporting symptoms, a structured interview can include a human expert (e.g., a therapist, mental health expert, etc.) who can administer the questioning, interpret responses from the user, and / or record the responses.
[0046] The instrument and / or questionnaire (which can refer to, for example, 50 symptoms, less than 50 symptoms, or more than 50 symptoms) can be completed by the user and / or human expert prior to engaging with the agent. Based on the instrument and / or questionnaire, initial symptoms that the user presents can be identified and / or predicted. For example, a covariance structure of the symptoms presented by the user (as indicated by the instrument and / or questionnaire) can be analyzed to determine relationships between symptoms. Based on these relationships, representations of symptoms (e.g., embedded symptom vectors) can be clustered within an embedding space, such that the symptoms can be collectively classified to determine an initial state (e.g., apparent sadness, pessimistic thoughts, severe depression, etc.). In some instances, the psychometric instrument can be used to determine a state subsequent to the initial state and / or in the midst of the conversation.
[0047] The data retriever 316 can be configured to retrieve data from a storage memory (e.g., the database 130 of FIG. 1, the memory 210 of FIG. 2, and / or the like) based on the state determined by the state determinator 310. The retrieved data can include, for example, predetermined data representing educational, instructive, and / or factual content (e.g., content associated with a scientific study, medical journal, textbook, expert author, and / or the like). The retrieval of the predetermined data can be caused and / or triggered by a transition to a particular state and / or by a topic being brought up by the user as communicated in the input 302. In this way, the agent can mimic a professional (e.g., an educator, counsellor, and / or the like) who can impart knowledge at a pace and / or level appropriate for the user (e.g., a learner). The agent can also retrieve the predetermined data based on the state indicating a user’s degreeof interest towards a topic. As a result of retrieving predetermined data associated with trusted sources, the data retriever 316 can cause the conversation management application 312 to exclude from the response 304 a fact not indicated by the predetermined data. Similarly stated, the data retriever 316 can prevent and / or reduce hallucinations caused by the machine learning model 320 (e.g., an LLM, etc., as described herein) by deriving responses from a closed corpus of trusted information (and not using other information not in the trusted information).
[0048] In some implementations, the data retriever 316 can include a knowledge graph 318, which can include a graph data structure that can represent and / or track the state and / or map the state to the predetermined data. For example, the knowledge graph 318 can specify an order in which predetermined data is to be delivered to the user. To illustrate the knowledge graph 318 in use, an example knowledge graph 318 can map to first predetermined data that can include, for example, first instructive content. The example knowledge graph 318 can also map to second predetermined data that can include, for example, second instructive content that is associated with a more basic and / or elementary concept than that of the first instructive content. For example, the first instructive content can include a more advanced discussion of a topic, building off of the second instructive content that includes a less advanced discussion of the topic. As a result, the example knowledge graph 318 can specify that the first instructive content is to be retrieved and / or conveyed to the user after retrieval and / or conveyance of the second instructive content. Similarly stated, the example knowledge graph 318 can specify a dependency of the first instructive content on the second instructive content. The example knowledge graph 318 can further specify a threshold state (e.g., a threshold level of understanding, as determined by the state determinator 310) to be achieved before the first instructive content is to be retrieved. As a result, the conversation management application 312 can graduate the complexity of the responses 304 (e.g., from less complex to more complex) based on the progression of the user’s understanding.
[0049] In some instances, the order of predetermined data defined by the knowledge graph 318 can be overridden based on, for example, the state. For example, if the state indicates that a user has a level of comprehension associated with more advanced content that has a dependency on more elementary content that has yet to be retrieved and / or conveyed to the user, the conversation management application 312 can be configured to override the dependency. Thus, if, for example, the user has an initial state associated with the more advanced content as a result of consuming preliminary content before engaging with the agentor as determined based on interactions with the agent, the conversation management application 312 can be configured to retrieve the more advanced content without first retrieving more elementary content so as to avoid boring and / or frustrating the user.
[0050] In some instances, based on a state generated by the state determinator 310 and representing an inferred mood of the user, the conversation management application 312 can be configured to adjust, for example, the rate at which predetermined data is retrieved and / or conveyed to the user. For example, the state can indicate that the user is depressed, tired, and / or having a mood not conducive to learning. In response, the data retriever 316 can be configured to retrieve predetermined data at a slower rate (e.g., over a longer period time, over a greater number of inputs and / or responses, etc.), so as to avoid making the user feel overwhelmed. Alternatively or in addition, such a state can cause the data retriever 316 to prevent more advanced content (as specified by the knowledge graph 318) from being retrieved until the state determinator 310 determines a state indicating that the user is more receptive to learning.
[0051] In some instances, the knowledge graph 318 can define a decision point (e.g., a fork) that can determine predefined data to be retrieved and conveyed to the user based on a predicted state (e.g., as determined by the state determinator 310) indicating the user’s interest in and / or believed importance of a topic. For example, the knowledge graph 318 can map to first content having a topic of physics and second content having a topic of organic chemistry. If, for example, the predicted state determined based on the input 302 indicates that the user enjoys physics more than organic chemistry, the knowledge graph 318 can cause the data retriever 316 to retrieve the first content instead of the second content. In some instances, the knowledge graph 318 can also cause predefined data to be retrieved based on the state indicating the user’s predicted mood. For example, if the state indicates that the user is perceived to be sad, the knowledge graph 318 can prevent distressing content from being retrieved until the user’s perceived mood improves, causing content having lighter subject matter to be retrieved instead.
[0052] In some implementations, the knowledge graph 318 can specify a goal (e.g., a learning goal) to be achieved by a predetermined time (e.g., date). This goal can include, for example, conveyance of time-sensitive predetermined data to the user within a predefined range of the current date. For example, the conversation management application 312 can be configured to receive as input a first date (e.g., the current date) and a second date (e.g., a target date, such as a scheduled surgery date if the user is a patient scheduled for surgery, or an exam date if the user is a student). If the user is to be presented with time-sensitive predetermined data (e.g.,information about the surgery, exam, etc.) at a time in relation to the target date (e.g., within a predefined range of the target date, such as within 5 days before the target date), the knowledge graph 318 can be configured to map to the time-sensitive predetermined data based on the current date being within the predefined range. As a result, the conversation management application 312 can generate a response that includes the time-sensitive predetermined data and can be conveyed to the user within the predetermined range of the target date. In some instances, the conversation management application 312 can be configured to convey the timesensitive predetermined data via the response 304 independent of any input 302 received from the user. As a result, the user can be informed of time-sensitive information without having to provide an input 302 to trigger a response from the agent.
[0053] The storage memory from which the predetermined data is retrieved can include, for example, a database (e.g., a SQL database and / or a searchable vector store) configured for keyvalue search and / or embedded vector (e.g., text embedding) search. To construct the database searchable by key-value pair, the predetermined data can be labelled (e.g., via manual labelling and / or automatic classification using a machine learning model) as being associated with a state, topic, importance metric, etc. The predetermined data can be stored within the database in a location (e.g., at an index) associated with the label. The predetermined data can then be located if, for example, a classification of the input 302 (e.g., a classification of the state, topic, importance metric, etc., associated with the input 302) is associated with, equivalent to, and / or similar to the label. The state determinator 310 can generate the classification using, for example, a machine learning model, as described herein.
[0054] Alternatively and / or in addition, to construct the database searchable by embedded vector, the predetermined data can be associated with an embedded vector that can indicate a semantic meaning (e.g., a triggering state) within an embedding space. To search for the predetermined data, the data retriever 316 can receive an indication of a topic of the input 302 from the state determinator 310. This indication can include, for example, an embedded vector that encodes a semantic meaning of the input 302 and that is associated with the embedding space. Within the embedding space, the position of the embedded vector generated based on the input 302 and / or received from the state determinator 310 can be compared to the distance of the embedded vector associated with the predetermined data. If this distance is sufficiently small (e.g., below a threshold distance), the associated predetermined data can be located and retrieved.
[0055] Having retrieved the data, the conversation management application 312 can be configured to tokenize the predetermined data (e.g., using the tokenizer 306 or a functionally and / or structurally similar tokenizer) and include the resulting set of predetermined tokens in the context window 308. For example, the predetermined data can include at least one document (e.g., at least one article) having text data. This text data can be concatenated and included in a prompt with the input 302 prior to tokenization. Alternatively, in some instances, the predetermined data can be stored in token form prior to retrieval of the data (e.g., during an initialization and prior to use), such that the set of predetermined tokens can be retrieved and included in the context window 308. As a result of the set of predetermined tokens being included in the context window 308, the machine learning model 320 (described herein) can consider the set of predetermined tokens as context while interpreting the input 302 and / or generating the response 304. For example, the machine learning model 320 can select at least one predetermined token from the set of predetermined tokens for inclusion in a set of response tokens representing the response 304. Such selection of the at least one predetermined token can be based on, for example, a confidence value and / or a probability that the predetermined token(s) could be included in a possible (e.g., relevant, sensible, coherent, etc.) response. In some implementations, the conversation management application 312 can be configured to include in the response 304 a citation associated with predetermined data used while generating the response 304. For example, if the predetermined data included content from a specific article in a medical journal, the citation can indicate that article to the user.
[0056] In some implementations, the conversation management application 312 can retrieve the predetermined data asynchronously to the receiving of the input 302 that triggers the retrieval. For example, using a first machine learning model, the input 302 can be classified and / or embedded into a vector by the state determinator 310 a period of time after a response 304 is generated in response to the input 302. Specifically, the conversation management application 312 can include the set of input tokens associated with the input 302 in the context window 308, generate the response 304 using the machine learning model 320 and based on that context window 308, and then use the input 302 to determine that predetermined data should be retrieved. After retrieving the predetermined data and generating the set of predetermined tokens, the conversation management application 312 can generate a response 304 based on the predetermined data in response to a later input, even if that input (unlike the previous input that caused the retrieval) would not itself be sufficient to cause retrieval of the predetermined data.
[0057] In some implementations, the predetermined data can be retrieved independent of a determined state and / or the input 302. For example, the conversation management application 312 can be configured to retrieve and / or receive (e.g., from a news feed) predetermined data based on a real-world current event (e.g., a new conflict, an election, breaking news, etc.), where the predetermined data can include data (e.g., a text description) associated with the current event. In some instances, the category of current event (e.g., politics, finance, etc.) can be predetermined. For example, for an agent having a role of teacher, the knowledge graph 318 can specify a curriculum to be conveyed to the user (e.g., a student). If the curriculum includes a discussion of, for example, recent scientific breakthroughs, the conversation management application 312 can be configured to relevant scrape news feeds. As a result of the retrieving and / or receiving, the conversation management application 312 can cause tokens representing the current event data to be injected into the context window 308, such that the conversation management application 312 can generate a response 304 that addresses the current event.
[0058] The summary generator 314 can generate a summary of a conversation history (e.g., previous dialogue) included in the context window 308. The conversation management application 312 can then replace, within the context window 308, the set of tokens representing the conversation history with a set of summary tokens representing the generated summary. In some instances, the set of summary tokens can be smaller (e.g., in data size) than the set of tokens previously in the context window 308 and that were replaced by the set of summary tokens.
[0059] The summary generator 314 can use a machine learning model (e.g., the machine learning model 320, a machine learning model separate from the machine learning model 320 and configured for natural language processing, a transformer model, and / or the like) to generate the set of summary tokens based on the tokens in the context window. In some implementations, the summary generator 314 can emphasize and / or include in the generated summary more details about previously discussed topics that have a higher importance as compared to topics that have a lower importance (e.g., as indicated by an importance metric). For example, a conversation history represented in the context window 308 can include a discussion of a significant topic (e.g., a symptom, illness, issue of concern, etc.) having a higher importance and a discussion of an insignificant topic (e.g., small talk, an exchange of pleasantries, etc.) having a lower importance. The summary generator 314 can generate a summary of this conversation that excludes reference to the insignificant topic or summarizesthe insignificant topic more briefly (e.g., with less text and / or data) as compared to the significant topic. In some implementations, a set of summary tokens can be generated for and replace tokens within the context window 308 that are associated with a less important topic (as determined by the importance metric), while tokens within the context window 308 associated with a more important topic can remain in the context window 308 without being summarized and / or replaced (e.g., for a period of time after the set of summary tokens are generated for the less important topic).
[0060] The summary can be generated when the conversation history has exceeded a threshold length and / or time. For example, the threshold can be a number of tokens in the context window 308, a number of inputs 302 and / or responses 304 represented in the context window 308 by (respectively) a set of input tokens and / or a set of response tokens, a length of time elapsed without a summary being generated, etc. The threshold length can be selected based on, for example, memory availability, processing resource availability, and / or the like, such that memory and / or processor usage can be improved. Alternatively and / or in addition, the threshold length can be selected to limit and / or prevent previous dialogue remaining in the context window 308 from outweighing and / or dominating a recently received input 302 (e.g., an instruction prompt). The threshold length can also be selected to limit and / or prevent the machine learning model 320 from hallucinating, such as generating an incorrect, untruthful, fictitious, misleading, and / or irrelevant response. By reducing the context window using the summary, resource usage, prompt domination, and / or hallucinations can be mitigated while context of a previous conversation(s) is still maintained when generating a response.
[0061] In some instances, a set of summary tokens can be further summarized after a period of time and / or after a number of exchanges within a conversation. Similarly stated, a set of summary tokens can be generated based on one or more previous summaries of dialogue. For example, an existing set of summary tokens within the context window 308 can represent a summary of multiple topics having a range of importance metrics and / or valence metrics. After a period of time and / or number of exchanges (e.g., based on predetermined thresholds of time and / or exchanges), the summary can be further compacted and / or condensed by generating a new (e.g., smaller) set of summary tokens that excludes, for example, a topic(s) having lower importance relative to a remaining topic(s). As a result of gradually reducing a summary (e.g., by gradually excluding topics based on importance metrics and / or valence metrics), the conversation management application 112 can emulate a human that gradually remembersfewer details over time and / or that forgets details of less important topics at a higher rate compared to details of more important topics.The machine learning model 320 can be configured to generate a set of response tokens based on tokens in the context window 308. The tokens in the context window 308 can include a set of input tokens associated with the input 302, a set of enhanced input tokens generated by the input enhancer 309, a set of summary tokens generated by the summary generator 314, and / or a set of predetermined tokens generated by the data retriever 316. The generated set of response tokens can also be included in the context window 308, such that a subsequent response 304 can be generated based on that set of response tokens. The machine learning model 320 can include an LLM, transformer model, neural network, and / or any other algorithm and / or model configured for natural language processing. Specifically, the machine learning model 320 can take as input the tokens in the context window 308, encode semantic meaning and context of those tokens using an encoder to generate a feature representation of the token, and decode the feature representation using a decoder to generate the set of response tokens. The machine learning model 320 can further normalize the set of response tokens to generate natural language text that can be included in the response 304.
[0062] In some implementations, the machine learning model 320 can be configured to emulate a persona from a plurality of personas in generating the response 304. For example, the machine learning model 320 can include an instruction-following LLM that can receive as input an instruction prompt that includes an indication of the persona (e.g., a therapist, a high school teacher, a college professor, and / or the like). The indication of the persona can be included in a context window of the instruction-following LLM to cause the instruction-following LLM to modify an output based on the persona. In some instances, based on the state (e.g., the predicted level of comprehension for the user), the conversation management application 312 can be configured to cause a transition (e.g., using a state machine) from a first (e.g., initial) persona to a second persona, such that the machine learning model 320 can emulate the second persona in generating subsequent responses. For example, the first persona can be associated with a high school teacher, and as a result of the state indicating a threshold level of knowledge and / or following predefined period of time, the conversation management application 312 can cause a transition to a second persona associated with a college professor persona. As a result, the conversation management application 312 can emulate a graduation experience for the user.11
[0063] In some implementations, the machine learning model can be configured to generate the response 304 based on an indication of a demographic associated with the user. For example, the machine learning model 320 can include an instruction-following LLM that can receive as input an instruction prompt that includes an indication of the demographic. For example, the indication of the demographic can indicate whether the user is elderly, is a child, has a particular education background and / or particular training, and / or the like. Based on the demographic, the machine learning model 320 can cause, for example, terminology appropriate for the user being of the demographic to be included in the response 304. Alternatively and / or in addition, the conversation management application 312 can be configured to generate and / or modify the knowledge graph 318 based on the indication of the demographic. For example, the indication of the demographic can cause the knowledge graph 318 to map to different predetermined data and / or adjust the pacing of delivery of the predetermined data (e.g., by modifying target dates associated with a timeline and specified by the knowledge graph 318), such that the generated response 304 can be tailored to the user being of the indicated demographic.
[0064] FIG. 4 is a schematic diagram illustrating a plurality of interactions 431-436 (e.g., dataflow, transmissions, signals, etc.) between logic components 400 to generate responses, according to an embodiment. The logic components 400 can be associated a compute device that is structurally and / or functionally similar to the respondent compute device 201 of FIG. 2 and / or the respondent compute device 120 of FIG. 1. In some instances, for example, the logic components 400 can be implemented as software stored in memory 210 and configured to be executed via the processor 220 of FIG. 2. For example, at least a portion of the logic components 400 can be included in a conversation management application that is functionally and / or structurally similar to the conversation management application 112 of FIG. 1 and / or the conversation management application 212 of FIG. 2. In some instances, for example, at least a portion of the logic components 400 can be implemented in hardware. The logic components 400 include a tokenizer 406 (which can be functionally and / or structurally similar to, for example, the tokenizer 306 of FIG. 3), a context window 408 (which can be functionally and / or structurally similar to, for example, the context window 308 of FIG. 3), an input enhancer 409 (which can be functionally and / or structurally similar to, for example, the input enhancer 309 of FIG. 3), a state determinator 410 (which can be functionally and / or structurally similar to, for example, the state determinator 310 of FIG. 3), a summary generator 414 (which can be functionally and / or structurally similar to, for example, the summary generator 214 ofFIG. 2 and / or the summary generator 314 of FIG. 3), and a data retriever 416 (which can be functionally and / or structurally similar to, for example, the data retriever 216 of FIG. 2 and / or the data retriever 316 of FIG. 3).
[0065] The tokenizer 406 can receive input text and generate a set of input tokens based on the input text. At 431, the set of input tokens can be sent to the input enhancer 409, which can generate at least one set of enhanced input tokens based on the set of input tokens. The set of input tokens and / or the at least one set of enhanced input tokens can also be sent to the state determinator 410, and the state determinator 410 can determine (e.g., using a machine learning model) a state, importance metric, and / or a topic indication (e.g., classification and / or embedded vector), based on the set of input tokens and / or the at least one set of enhanced input tokens. In some implementations, and although not shown in FIG. 4, the state determinator 410 can determine the state based on the input text rather than and / or in addition to the set of input tokens generated by the tokenizer 406. The set of input tokens and / or the at least one set of enhanced input tokens can also be sent to the context window 408, such that a machine learning model (e.g., the machine learning model 320 of FIG. 3) can generate a response based on the set(s) of input and / or enhanced input tokens.
[0066] At 432, the state, importance metric, and / or topic indication can be sent to the data retriever 416. The data retriever 416 can retrieve predetermined data from a memory based on the state, importance metric, and / or topic indication and as specified by a knowledge graph (e.g., a knowledge graph functionally and / or structurally similar to the knowledge graph 318 of FIG. 3). Subsequently, the data retriever 416 can generate a set of predetermined tokens. Although not shown in FIG. 4, in some instances, the data retriever 416 can retrieve and / or generate the set of predetermined tokens independent of a state, importance metric, and / or topic classification, inferred from the input 302. In such instances, for example, the user may not provide an input and the set of predetermined tokens can be retrieved and / or generated based on context such as, for example, time, location, mood, demographics, scheduled appointments, and / or the like. For example, the data retriever 416 can retrieve and generate a set of timesensitive predetermined tokens based on a predetermined schedule, date, and / or time. At 433, the set of predetermined tokens can be sent to the context window 408, such that a machine learning model can generate a response based on the set of predetermined tokens.
[0067] The summary generator 414 can be configured to condense data from previous conversations within the context window 408. At 435, a set of previous conversation tokenscan be sent from the context window 408 to the summary generator 414. The summary generator 414 can then generate a set of summary tokens, and at 436, the set of summary tokens can be sent to the context window 408, and the previous conversation tokens can be removed from the context window 408. As a result, the machine learning model can generate a response based on the set of summary tokens.
[0068] FIG. 5 is a flowchart showing a method 500 illustrating an example implementation using a system described herein (e.g., the system 100 of FIG. 1). Portions of the method 500 can be implemented using a processor (e.g., the processor 220 of FIG. 2) of any suitable compute device (e.g., the respondent compute device 120 of FIG. 1 and / or the respondent compute device 201 of FIG. 2). The method 500 can include a method of generating a response based on predetermined data.
[0069] The method 500 at 502 includes receiving at a processor, a user input from a user device and generating, via the processor, a set of user input tokens based on the user input. At 504, the method includes providing, via the processor, the set of user input tokens as input to a first machine learning model to generate a set of enhanced input tokens. At 506, a state is determined, via the processor and based on a previous state and at least one of the set of user input tokens or the set of enhanced input tokens. At 508 of the method 500, predetermined data is retrieved, via the processor and from a database based on the state and at least one of the set of user input tokens or the set of enhanced input tokens. The method 500 at 510 includes providing, via the processor, the set of user input tokens and the predetermined data as input to a second machine learning model to generate a set of response tokens. Based on the set of response tokens, at 512, the method 500 includes causing, via the processor, a response to be sent to the user device.
[0070] In an embodiment, a method includes receiving, at a processor, a user input from a user device and generating, via the processor, a set of user input tokens based on the user input. The method also includes generating, via the processor and using a first machine learning model, a set of enhanced input tokens based on the set of user input tokens. A state is determined, via the processor and based on a previous state and at least one of the set of user input tokens or the set of enhanced input tokens. Predetermined data is retrieved, via the processor and from a database based on the state and at least one of the set of user input tokens or the set of enhanced input tokens. The method also includes generating, via the processor and using a second machine learning model, a set of response tokens based on the predetermined data and the userinput. Based on the set of response tokens, the method includes causing, via the processor, a response to be sent to the user device.
[0071] In some implementations, the state can indicate a first predicted level of comprehension for a user associated with the user device, and the previous state can indicate a second predicted level of comprehension for the user and different from the first predicted level of comprehension. Additionally, the predetermined data can include first instructive content, and a knowledge graph can (1) map (a) the state to the first instructive content and (b) the previous state to second instructive content and (2) specify a dependency of the first instructive content on the second instructive content. Additionally, the retrieving can include (1) determining, via the processor and based on the knowledge graph, that the second instructive content was previously conveyed to the user, (2) causing, via the processor, a transition within the knowledge graph based on the state and the previous state, and (3) retrieving, via the processor and in response to the transition, the first instructive content based on the knowledge graph. Additionally, the response can cause the first instructive content to be conveyed to the user.
[0072] In some implementations, the method can further include generating, via the processor, a comprehension progression indication based on the transition. Additionally, the method can further include causing, via the processor and in response to the generating the comprehension progression indication, transmission of a signal indicating the comprehension progression. In some implementations, the determining the state can include predicting, via the processor and using a third machine learning model configured to perform zero-shot classification, a state classification based on at least one of the set of user input tokens or the set of enhanced input tokens. Additionally, the determining the state can further include identifying, via the processor, the state within the knowledge graph based on the state classification and the previous state.
[0073] In some implementations, the predetermined data can include first instructive content, and the determining the state can include predicting, via the processor, using a third machine model, and based on the user input, a mental state of a user associated with the user device, the mental state being associated with at least one of enjoyment of a topic indicated by the user input, interest in the topic, believed importance of the topic, or an emotional state. Additionally, the retrieving can include retrieving, via the processor and based on a knowledge graph, the first instructive content, the knowledge graph specifying a decision (1) determined based on the mental state and (2) between the first instructive content and second instructive content.
[0074] In some implementations, the method can further include receiving, at the processor, an indication of a first date and retrieving, via the processor and based on a knowledge graph specifying an education timeline having at least one goal, time-sensitive predetermined data associated with a goal from the at least one goal, the goal being associated with a second date being within a predefined range that includes the first date. Additionally, the method can further include generating, via the processor and using the second machine learning model, a set of time-sensitive response tokens based on the time-sensitive predetermined data. Additionally, the method can further include causing, via the processor, a time-sensitive response to be sent to the user device based on the set of time-sensitive response tokens.
[0075] In some implementations, the method can further include receiving, at the processor and prior to the receiving the user input, an indication of preliminary content conveyed to a user associated with the user device. The method can also further include determining, via the processor, the previous state based on the indication of the preliminary content. In some implementations, the retrieving can include generating, via the processor and using a third machine learning model, a first embedded vector based on at least one of the set of user input tokens or the set of enhanced input tokens. Additionally, the retrieving can include retrieving, via the processor, the predetermined data from the database based on a distance between the first embedded vector and a second embedded vector, the predetermined data stored at a memory location within the database based on the second embedded vector. In some implementations, the retrieving can include predicting, via the processor and using a third machine learning model configured to perform zero-shot classification, a topic classification based on at least one of the set of user input tokens or the set of enhanced input tokens. The retrieving can also include retrieving, via the processor, the predetermined data from the database based on the topic classification.
[0076] In some implementations, the topic classification can be a high-risk topic classification, and the method can further include causing, via the processor and in response to the predicting the high-risk topic classification, transmission of a signal indicating the high-risk topic classification. In some implementations, the response can include at least one of text data, audio data, image data, or video data. In some implementations, the second machine learning model can be configured to exclude from the response a fact not indicated by the predetermined data. In some implementations, the response can include a citation associated with the predetermined data. In some implementations, the generating the set of response tokens can include concatenating, via the processor, text data associated with at least one document and includedin the predetermined data, to result in concatenated text data. The generating the set of response tokens can also include generating, via the processor, a prompt based on the user input and the concatenated text data and generating, via the processor and using the second machine learning model, the set of response tokens based on the prompt.
[0077] In some implementations, the method can further include including, via the processor, the set of user input tokens and the set of response tokens in a context window and generating, via the processor and using a third machine learning model, a set of summary tokens based on the context window including the set of user input tokens and the set of response tokens. Additionally, the method can further include removing, via the processor, at least one of the set of user input tokens or the set of response tokens, from the context window based on a measure of conversation length. The method can also include including, via the processor, the set of summary tokens in the context window, such that the second machine learning model generates additional responses based on the context window including the set of summary tokens and excluding the at least one of the set of user input tokens or the set of response tokens. In some implementations, the measure of conversation length can include at least one of a period of time, a number of inputs received from the user device, or an amount of data in the context window.
[0078] In some implementations, the first machine learning model can be a first transformer model, and the second machine learning model can be a second transformer model different from the first transformer model. In some implementations, the generating the set of response tokens can be based further on a persona from a plurality of personas. In some implementations, the user input can be a first user input, the state can be a first state, the set of response tokens can be a first set of response tokens, and the persona can be a first persona. Additionally, the method can further include receiving, at the processor, a second user input from the user device and determining, via the processor, a second state based on the second user input. The method can also further include causing, via the processor, a transition from the first persona to a second persona from the plurality of personas and generating, via the processor and using the second machine learning model, a second set of response tokens based on the second persona. In some implementations, the method can further include receiving, at the processor, an indication of a demographic associated with a user of the user device and defining, via the processor, the persona based on the indication of the demographic.
[0079] Examples of computer code include, but are not limited to, micro-code or microinstructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments can be implemented using Python, Java, JavaScript, C++, and / or other programming languages and development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.
[0080] The drawings primarily are for illustrative purposes and are not intended to limit the scope of the subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the subject matter disclosed herein can be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).
[0081] The acts performed as part of a disclosed method(s) can be ordered in any suitable way. Accordingly, embodiments can be constructed in which processes or steps are executed in an order different than illustrated, which can include performing some steps or processes simultaneously, even though shown as sequential acts in illustrative embodiments. Put differently, it is to be understood that such features can not necessarily be limited to a particular order of execution, but rather, any number of threads, processes, services, servers, and / or the like that can execute serially, asynchronously, concurrently, in parallel, simultaneously, synchronously, and / or the like in a manner consistent with the disclosure. As such, some of these features can be mutually contradictory, in that they cannot be simultaneously present in a single embodiment. Similarly, some features are applicable to one aspect of the innovations, and inapplicable to others.
[0082] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the disclosure. That the upper and lower limits of these smaller ranges can independently be included in the smaller ranges is also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0083] The phrase “and / or,” as used herein in the specification and in the embodiments, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements can optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0084] As used herein in the specification and in the embodiments, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the embodiments, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the embodiments, shall have its ordinary meaning as used in the field of patent law.
[0085] As used herein in the specification and in the embodiments, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements can optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with noB present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0086] In the embodiments, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
[0087] Some embodiments described herein relate to a computer storage product with a non- transitory computer-readable medium (also can be referred to as a non-transitory processor- readable medium) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor-readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) can be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein.
[0088] Some embodiments and / or methods described herein can be performed by software (executed on hardware), hardware, or a combination thereof. Hardware modules can include, for example, a processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can includeinstructions stored in a memory that is operably coupled to a processor and can be expressed in a variety of software languages (e.g., computer code), including C, C++, Java™, Ruby, Visual Basic™, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments can be implemented using imperative programming languages (e.g., C, Fortran, etc.), functional programming languages (Haskell, Erlang, etc.), logical programming languages (e.g., Prolog), object-oriented programming languages (e.g., Java, C++, etc.) or other suitable programming languages and / or development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.
Claims
What is claimed is:
1. A method, comprising: receiving, at a processor, a user input from a user device; generating, via the processor, a set of user input tokens based on the user input; providing, via the processor, the set of user input tokens as input to a first machine learning model to generate a set of enhanced input tokens; determining, via the processor, a state based on a previous state and at least one of the set of user input tokens or the set of enhanced input tokens; retrieving, via the processor, predetermined data from a database based on the state and at least one of the set of user input tokens or the set of enhanced input tokens; providing, via the processor, the set of user input tokens and the predetermined data as input to a second machine learning model to generate a set of response tokens; and causing, via the processor, a response to be sent to the user device based on the set of response tokens.
2. The method of claim 1, wherein: the state indicates a first predicted level of comprehension for a user associated with the user device; the previous state indicates a second predicted level of comprehension for the user and different from the first predicted level of comprehension; the predetermined data includes first instructive content; a knowledge graph (1) maps (a) the state to the first instructive content and (b) the previous state to second instructive content and (2) specifies a dependency of the first instructive content on the second instructive content; the retrieving includes: determining, via the processor and based on the knowledge graph, that the second instructive content was previously conveyed to the user, causing, via the processor, a transition within the knowledge graph based on the state and the previous state, and retrieving, via the processor and in response to the transition, the first instructive content based on the knowledge graph; and the response causes the first instructive content to be conveyed to the user.
3. The method of claim 2, further comprising: generating, via the processor, a comprehension progression indication based on the transition; and causing, via the processor and in response to the generating the comprehension progression indication, transmission of a signal that encodes the comprehension progression indication.
4. The method of claim 2, wherein the determining the state includes: providing, via the processor, at least one of the set of user input tokens or the set of enhanced input tokens as input to a third machine learning model configured to perform zeroshot classification, to predict a state classification; and identifying, via the processor, the state within the knowledge graph based on the state classification and the previous state.
5. The method of claim 1, wherein: the predetermined data includes first instructive content; the determining the state includes predicting, via the processor, using a third machine model, and based on the user input, a mental state of a user associated with the user device, the mental state being associated with at least one of enjoyment of a topic indicated by the user input, interest in the topic, believed importance of the topic, or an emotional state; and the retrieving includes retrieving, via the processor and based on a knowledge graph, the first instructive content, the knowledge graph specifying a decision (1) determined based on the mental state and (2) between the first instructive content and second instructive content.
6. The method of claim 1 , further comprising: receiving, at the processor, an indication of a first date; retrieving, via the processor and based on a knowledge graph specifying an education timeline having at least one goal, time-sensitive predetermined data associated with a goal from the at least one goal, the goal being associated with a second date being within a predefined range that includes the first date; providing, via the processor, the time-sensitive predetermined data as input to the second machine learning model to generate a set of time-sensitive response tokens; andcausing, via the processor, a time-sensitive response to be sent to the user device based on the set of time-sensitive response tokens.
7. The method of claim 1, further comprising: receiving, at the processor and prior to the receiving the user input, an indication of preliminary content conveyed to a user associated with the user device; and determining, via the processor, the previous state based on the indication of the preliminary content.
8. The method of claim 1, wherein the retrieving includes: providing, via the processor, at least one of the set of user input tokens or the set of enhanced input tokens as input to a third machine learning model to generate a first embedded vector; and retrieving, via the processor, the predetermined data from the database based on a distance between the first embedded vector and a second embedded vector, the predetermined data stored at a memory location within the database based on the second embedded vector.
9. The method of claim 1, wherein the retrieving includes: providing, via the processor, at least one of the set of user input tokens or the set of enhanced input tokens as input to a third machine learning model configured to perform zeroshot classification, to predict a topic classification; and retrieving, via the processor, the predetermined data from the database based on the topic classification.
10. The method of claim 9, wherein the topic classification is a high-risk topic classification, the method further comprising: causing, via the processor and in response to predicting the high-risk topic classification, transmission of a signal indicating the high-risk topic classification.
11. The method of claim 1 , wherein the response includes at least one of text data, audio data, image data, or video data.
12. The method of claim 1, wherein the second machine learning model is configured to exclude from the response a fact not indicated by the predetermined data.
13. The method of claim 1, wherein the response includes a citation associated with the predetermined data.
14. The method of claim 1, wherein the generating the set of response tokens includes: concatenating, via the processor, text data associated with at least one document and included in the predetermined data, to result in concatenated text data; generating, via the processor, a prompt based on the user input and the concatenated text data; and providing, via the processor, the prompt as input to the second machine learning model to generate the set of response tokens.
15. The method of claim 1, further comprising: adding, via the processor, the set of user input tokens and the set of response tokens to a context window; providing, via the processor, the set of user input tokens and the set of response tokens from the context window as input to a third machine learning model to generate a set of summary tokens; removing, via the processor, at least one of the set of user input tokens or the set of response tokens, from the context window based on a measure of conversation length; and including, via the processor, the set of summary tokens in the context window, such that the second machine learning model generates additional responses based on the context window including the set of summary tokens and excluding the at least one of the set of user input tokens or the set of response tokens.
16. The method of claim 15, wherein the measure of conversation length includes at least one of a period of time, a number of inputs received from the user device, or an amount of data in the context window.
17. The method of claim 1, wherein: the first machine learning model is a first transformer model; and the second machine learning model is a second transformer model different from the first transformer model.
18. The method of claim 1, wherein the second machine learning model generates the set of response tokens based on a persona from a plurality of personas.
19. The method of claim 18, wherein the user input is a first user input, the set of user input tokens is a first set of user input tokens, the state is a first state, the set of response tokens is a first set of response tokens, and the persona is a first persona, the method further comprising: receiving, at the processor, a second set of user input tokens associated with a second user input received from the user device; determining, via the processor, a second state based on the second user input; causing, via the processor, a transition from the first persona to a second persona from the plurality of personas based on the second state; and adding, via the processor, the second set of user input tokens and an indication of the second persona to a context window of the second machine learning model to generate a second set of response tokens.
20. The method of claim 18, further comprising: receiving, at the processor, an indication of a demographic associated with a user of the user device; and defining, via the processor, the persona based on the indication of the demographic.
Citation Information
Patent Citations
Using personalized knowledge patterns to generate personalized learning-based guidance
US20220139245A1
Educational and content recommendation management system
US20230045037A1