Systems and methods for generating state-based responses based on text data and speech data

The system addresses the limitations of existing machine learning models by using a processor to manage input and state tokens, and machine learning models to generate and summarize responses, resulting in improved contextual relevance and reduced hallucination.

WO2025133176A1PCT designated stage expired Publication Date: 2025-06-26COMPASS PATHFINDER LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing machine learning models have limited context, leading to inefficient generation of responses and a propensity for hallucination, where fictional and misleading content is produced due to limited context and failure to update internal states based on conversation progression.

Method used

A system that uses a processor to receive input data, generate input tokens, and determine a state based on these tokens. It then provides the tokens and state to a machine learning model to generate response tokens, which are used to create a response. Additionally, the system uses a second machine learning model to generate summary tokens based on input and response tokens, which are then used to update the context window and improve response generation.

Benefits of technology

The system effectively generates state-based responses with improved context, reducing the likelihood of hallucination and enhancing the relevance and accuracy of responses by incorporating evolving psychological states and summary tokens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087985_26062025_PF_FP_ABST
    Figure EP2024087985_26062025_PF_FP_ABST
Patent Text Reader

Abstract

In an embodiment, a method includes receiving, at a processor, input data from the user device and providing, via the processor, the input data as input to a machine learning model to generate an input token. The method also includes adding, via the processor, a predetermined token different from the input token to the context window based on the input meeting a criterion. The method also includes providing, via the processor, the predetermined token from the context window as input to the machine learning model to generate a response token. A response is sent to the user device based on the response token.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR GENERATING STATE-BASED RESPONSES BASEDON TEXT DATA AND SPEECH DATACROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 613,805, filed December 22, 2023, and titled “SYSTEMS AND METHODS FOR GENERATING STATE-BASED RESPONSES USING MACHINE LEARNING MODELS,” which is incorporated herein by reference in its entirety.FIELD

[0002] One or more embodiments described herein relate to systems and computerized machine learning methods for generating state-based responses.BACKGROUND

[0003] Some known machine learning models can have limited context. As such, it can be desirable to have systems configured to use machine learning models with improved context to generate responses.SUMMARY

[0004] In an embodiment, a method includes receiving, at a processor, input data from a user device, and generating, via the processor, a set of input tokens based on the input data. The set of input tokens is added, via the processor, to a context window, and a state is determined, via the processor, based on the set of input tokens. The method also includes providing, via the processor, (1) the set of input tokens from the context window and (2) the state as input to a first machine learning model to generate a set of response tokens. Additionally, the method includes causing, via the processor, a response to be sent to the user device based on the set of response tokens, and providing, via the processor, the set of input tokens and the set of response tokens from the context window as input to a second machine learning model to generate a set of summary tokens. The method also includes removing, via the processor, at least one of the set of input tokens or the set of response tokens from the context window based on a measure of conversation length. Additionally, the method includes adding, via the processor, the set of summary tokens to the context window, such that the first machine learning model generatesadditional responses based on the context window including the set of summary tokens and excluding the at least one of the set of input tokens or the set of response tokens.

[0005] In an embodiment, a method includes receiving, at a processor, input data from the user device and providing, via the processor, the input data as input to a machine learning model to generate an input token. The method also includes adding, via the processor, a predetermined token different from the input token to the context window based on the input meeting a criterion. The method also includes providing, via the processor, the predetermined token from the context window as input to the machine learning model to generate a response token. A response is sent to the user device based on the response token.

[0006] In an embodiment, a method includes receiving, at a processor and from a user device, first input data having a topic, and providing, via the processor, the first input data as input to the first machine learning model to generate (1) a topic identification associated with the topic and (2) a first psychological metric associated with the topic, the first psychological metric being below a threshold. The method also includes providing, via the processor, the first input data and the first psychological metric as input to a second machine learning model to generate first response data. The method also includes causing, via the processor, the first response data to be conveyed to the user device, and storing, via the processor and at a memory, the first response data based on the topic identification. The method also includes receiving, at the processor and from the user device, second input data, and providing, via the processor, the second input data as input to the first machine learning model to generate the topic identification and a second psychological metric for the topic, the second psychological metric being above the threshold. Additionally, the method includes retrieving, via the processor, the first response data from the memory based on the topic identification and the second psychological metric. The method also includes providing, via the processor, the second input data and the first response data as input to the second machine learning model to generate second response data, and causing, via the processor, the second response data to be conveyed to the user device.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] FIG. 1 is a schematic diagram of a system, according to an embodiment.

[0008] FIG. 2 is a schematic diagram of a compute device included in a system, according to an embodiment.

[0009] FIG. 3 is a schematic diagram of logic components included in a system, according to an embodiment.

[0010] FIG. 4 is a signal diagram showing a plurality of interactions implemented by a system, according to an embodiment.

[0011] FIG. 5 is a flowchart showing a method of using a system to generate a response based on a set of summary tokens, according to an embodiment.

[0012] FIG. 6 is a flowchart showing a method of using a system to generate a response based on predetermined data, according to an embodiment.

[0013] FIG. 7 is a flowchart showing a method of using a system to generate a response based on a previous response, according to an embodiment.DETAILED DESCRIPTION

[0014] Some known machine learning models are configured to generate text. For example, some known large language models (LLMs) can perform as chat agents (e.g., “chat hots”), generating responses based on inputs received from a user. Such known LLMs can, in some instances, use previous conversations (e.g., dialogue) to generate a next response. However, given that context windows (e.g., the tokens and / or text segments used as input to an LLM to generate text) can be limited by known scaling laws, memory constraints, processing resources, and / or the like, the amount of previous conversation data used as context for some known LLMs is limited. Moreover, in some instances, some known LLMs over-emphasize previous conversation data as compared to instruction prompts (e.g., inputs, queries, etc.) received more recently from a user. As a result, previous dialogue can outweigh and / or dominate prompts to these known LLMs, disproportionately influencing responses generated by these known LLMs.

[0015] Additionally, some known LLMs can be prone to hallucination, generating fictional and / or misleading content. Moreover, hallucinations generated by such known LLMs can fail to be mutually consistent, given that context to these known LLMs is limited to the instruction prompt and the conversation history. Some known LLMs also do not determine or update an internal state (e.g., a psychological state, emotional state, mood, etc., to be emulated by the LLM) based on a progression of a conversation. Thus, a need exists for a machine-learning based system that can implement an agent (e.g., a patient agent and / or the like, described herein) having extended conversation context from which to generate a response. A need alsoexists for a machine-learning based system that can generate a response having contextual and / or topical content based on data retrieved from a memory. Additionally, a need exists for a machine-learning based system that can generate a response based on an evolving psychological state that is updated based on, for example, instruction prompts received from a user.

[0016] FIG. 1 is a schematic diagram of a system 100 for generating responses based on received inputs, according to an embodiment. The system 100 includes a user device 110, respondent compute device 120, database 130, and network Nl. The system 100 can include alternative configurations, and various steps and / or functions of the processes described below can be shared among the various devices of the system 100 or can be assigned to specific devices (e.g., the user device 110, the respondent compute device 120, the database 130, and / or the like). For example, in some configurations, a user can provide inputs directly to the respondent compute device 120 rather than via the user device 110, as described herein.

[0017] In some embodiments, the user device 110 and / or the respondent compute device 120 can include any suitable hardware-based computing devices and / or multimedia devices, such as, for example, a server, a desktop compute device, a smartphone, a tablet, a wearable device, a laptop and / or the like. In some implementations, the user device 110 and / or the respondent compute device 120 can be implemented at an edge node or other remote computing facility. In some implementations, each of the user device 110 and / or respondent compute device 120 can be a data center or other control facility configured to run a distributed computing system and can communicate with other compute devices.

[0018] In some implementations, the user device 110 can include a peripheral device(s) (not shown in FIG. 1) configured to receive an input from a user. For example, the user device 110 can include a keyboard that the user can use to generate text data to be used as an instruction prompt by the respondent compute device 120. In another implementation, the user device 110 can include a graphical user interface (GUI) that a user can use to define and / or articulate the instruction response (e.g., by choosing from a list of predefined instruction prompts). In yet another implementation, the user device 110 can include a microphone, camera, and / or video camera, and the user device 110 and / or the respondent compute device 120 can be configured to convert audio data (e.g., speech data), image data, and / or video data to text data and / or tokens (described herein) to be used to generate a response. For example, the user device 110 and / or the respondent compute device 120 can use facial recognition techniques to infer a user’smental state (e.g., apathy, boredom, concern, shock, confusion, anxiety, and / or the like) from a facial expression of the user depicted in video data. This inferred mental state can be communicated to the respondent compute device 120 (e.g., as a set of tokens) to be considered while generating a response.

[0019] The respondent compute device 120 can be configured to generate a response to a user based on input data received from the user (e.g., via the user device 110). For example, the respondent compute device 120 can be configured to execute (e.g., via a processor) a conversation management application 112, which can be functionally and / or structurally equivalent to the conversation management application 212 of FIG. 2, described herein. The conversation management application 112 can be implemented via software and / or hardware and can, for example, act as a participant (e.g., an agent) in a dialogue with a user. Similarly stated, the conversation management application 112 can assume the role of one person in a two-person dialogue. In some instances, this dialogue can occur in the context of training the user. For example, the user can be a trainee undergoing training for a professional role, such as, for example, a crisis hotline counsellor, a doctor practicing bedside manner, an elementary school teacher, a public relations representative, a support guide for psychedelic therapy, and / or the like. To provide a training opportunity to the user, the conversation management application 112 can generate responses that emulate reactions of a patient and / or a person in need of a service to be provided by the user.

[0020] The conversation management application 112 can be further configured to evaluate the performance and / or progression of a user (e.g., a trainee) based on the inputs received from the user. For example, the conversation management application 112 can cause display of a progress bar to the user. The progress bar can indicate, for example, how effective the user’s inputs are at causing the agent to set a simulated intention. If, for example, the user’s input includes a question and / or has a structure, where that question and / or structure is associated with a best practice that a therapist can follow to help a patient set an intention and / or goal, the progress bar can indicate an advancement and / or improvement. Alternatively, if the user’s input is unhelpful for intention setting and / or is not associated with the best practice, the agent can generate a response that is of reduced value to the user. For example, such a response can indicate confusion and / or the like.

[0021] As described herein (e.g., in relation to FIGS. 2-4), the conversation management application 112 can be configured to summarize previous conversations with the user andgenerate subsequent responses based on those summaries. The conversation management application 112 can also determine and / or estimate a state to be emulated and generate responses based on this state. For example, the conversation management application 112 can retrieve predetermined data (e.g., backstory data) from a memory (e.g., the database 130, described herein) and cause that predetermined data to be considered within the context window. Additionally, the conversation management application 112 can retrieve a previously generated response from a memory (e.g., the database 130) based on a current state (e.g., a state indicating increased trust as compared to a previous state associated with the previously generated response). The conversation management application 112 can then generate a response based on this previously generated response. For example, the conversation management application 112 can include in the new response additional information, such that it appears to the user that the simulated patient and / or person in need of service is reflecting on and amending previous responses based on, for example, increased trust in the user.

[0022] In some implementations, the respondent compute device 120 can generate text data that can cause a response to be communicated to the user. For example, in some implementations, the user device 110 and / or the respondent compute device 120 can display the text data to the user (e.g., via a display operably coupled to the user device 110 and / or the respondent compute device 120). Alternatively and / or in addition, the user device 110 and / or the respondent compute device 120 can cause an audio signal to be generated based on the generated text data, such that the response can be sonically conveyed to the user (e.g., via a speaker operably coupled to the user device 110 and / or the respondent compute device 120). In some implementations, the audio signal can indicate a tone of voice (e.g., enthusiastic, pessimistic, etc.) using, for example, a synthesized speech markup language (SSML) and based on a determined state (described herein in relation to, for example, the state determinator 310 of FIG. 3) of the agent. In some implementations, the user device 110 and / or the respondent compute device 120 can be configured to perform speech output throttling to simulate natural (e.g., nominal) speech production rates and / or an anomalous speech production rate (e.g., a slower speech production rate associated with depression, severe depression, and / or the like). As a result, the response (e.g., communicated via the display and / or the audio signal) can have a cadence and / or pacing that depends on the determined state (e.g., a mental state) being emulated by the agent.

[0023] Alternatively and / or in addition, the user device 110 and / or the respondent compute device 120 can cause video data to be generated based on the text data, such that, for example,the user can view (e.g., via the operably coupled display) an animation of a simulated patient and / or person in need of service providing the generated response.

[0024] The database 130 can include at least one memory, repository and / or other form of data storage. The database 130 can be in communication with the user device 110 and / or the respondent compute device 120 (e.g., via the network Nl). In some implementations, the database 130 can be housed and / or included in one or more of the user device 110, the respondent compute device 120, or a separate compute device(s). The database 130 can be configured to store, for example, predetermined data and / or previous response data that can be retrieved or otherwise accessed by one or more compute devices, such as, for example, the respondent compute device 120, to perform at least some of the features (e.g., in relation to the conversation management application 112) described herein.

[0025] The database 130 can include a computer storage, such as, for example, a hard drive, memory card, solid-state memory, ROM, RAM, DVD, CD-ROM, write-capable memory, and / or read-only memory. In addition, the database 130 may include a distributed storage system where data is stored on a plurality of different storage devices, which may be physically located at a same or different geographic location (e.g., in a distributed computing system). In some implementations, the database 130 can be associated with cloud-based / remote storage.

[0026] Database 130 can be networked, via the network Nl, to the user device 110 and / or the respondent compute device 120 directly using wired connections and / or wireless connections. The network Nl can include various configurations and protocols, including, for example, short range communication protocols, Bluetooth®, Bluetooth® LE, the Internet, World Wide Web, intranets, virtual private networks, wide area networks, local networks, private networks using communication protocols proprietary to one or more companies, Ethernet, WiFi and HTTP, cellular data networks, satellite networks, free space optical networks and / or various combinations of the foregoing. Such communication can be facilitated by any device capable of transmitting data to and from other compute devices, such as a modem(s) and / or a wireless interface(s).

[0027] FIG. 2 is a schematic diagram of a respondent compute device 201 of a system, according to an embodiment. The respondent compute device 201 can be structurally and / or functionally similar to, for example, the respondent compute device 120 of the system 100 shown in FIG. 1. The respondent compute device 201 can be a hardware-based computing device, a multimedia device, or a cloud-based device such as, for example, a computer device,a server, a desktop compute device, a laptop, a smartphone, a tablet, awearable device, a remote computing infrastructure, and / or the like. The respondent compute device 201 includes a memory 210, a processor 220, and a network interface 230 operably coupled to a network N2.

[0028] The processor 220 can be, for example, a hardware based integrated circuit (IC), or any other suitable processing device configured to run and / or execute a set of instructions or code (e.g., stored in memory 210). For example, the processor 220 can be a general-purpose processor, a central processing unit (CPU), an accelerated processing unit (APU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic array (PLA), a complex programmable logic device (CPLD), a programmable logic controller (PLC), a remote cluster of one or more processors associated with a cloud-based computing infrastructure and / or the like. The processor 220 is operatively coupled to the memory 210 (described herein). In some embodiments, for example, the processor 220 can be coupled to the memory 210 through a system bus (for example, address bus, data bus and / or control bus).

[0029] The memory 210 can be, for example, a random-access memory (RAM), a memory buffer, a hard drive, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and / or the like. The memory 210 can store, for example, one or more software modules and / or code that can include instructions to cause the processor 220 to perform one or more processes, functions, and / or the like. In some implementations, the memory 210 can be a portable memory (e.g., a flash drive, a portable hard disk, and / or the like) that can be operatively coupled to the processor 220. In some instances, the memory can be remotely operatively coupled with the respondent compute device 201, for example, via the one or more network interface controllers 240. For example, a remote database server can be operatively coupled to the respondent compute device 201.

[0030] The memory 210 can store various instructions associated with processes, algorithms and / or data, including machine learning models (e.g., LLMs, classifiers, tokenizers, embedders, etc., as described herein). Memory 210 can further include any non-transitory computer-readable storage medium for storing data and / or software that is executable by processor 220, and / or any other medium which may be used to store information that may be accessed by processor 220 to control the operation of the respondent compute device 201. For example, the memory 210 can store data associated with the conversation management application 212. The conversation management application 212 can be functionally and / orstructurally similar to the conversation management application 112 of FIG. 1 and / or the conversation management application 312 of FIG. 3, described herein.

[0031] The conversation management application 212 can include a summary generator 214, a data retriever 216, and a response retriever 218. The summary generator 214 can be functionally and / or structurally similar to the summary generator 314 of FIG. 3 and / or the summary generator 414 of FIG. 4, described in further detail herein. The data retriever 216 can be functionally and / or structurally similar to the data retriever 316 of FIG. 3 and / or the data retriever 416 of FIG. 4, described in further detail herein. The response retriever 218 can be functionally and / or structurally similar to the response retriever 318 of FIG. 3 and / or the response retriever 418 of FIG. 4, described in further detail herein. As described herein, the summary generator 214 can be configured generate a summary of a conversation history (e.g., previous dialogue) included in a context window. The conversation management application 212 can then replace, within the context window, the set of tokens representing the conversation history with a set of summary tokens representing the generated summary. As described herein, the data retriever 216 can be configured to retrieve predetermined data (e.g., backstory data) and inject that predetermined data into the context window. As a result, the conversation management application 212 can generate responses based on the predetermined data. The predetermined data can be retrieved based on a state, a topic of conversation, and / or an importance metric associated with the topic of conversation, as described herein. The response retriever 218 can be configured to “reflect” on a previous generated response provided to the user and generate a new response that, for example, modifies this previous response. The modification can be based on, for example, the agent having a different state (e.g., increased trust) as compared to the state when the agent generated the previous response, as described herein.

[0032] The network interface 230 can be configured to connect to the network N2, which can be functionally and / or structurally similar to the network N1 of FIG. 1. For example, network N2 can use any of the wired and wireless short range communication protocols described above with respect to network N1 of FIG. 1.

[0033] In some instances, the respondent compute device 201 can further include a display, an input device, and / or an output interface (not shown in FIG. 2). The display can be any display device by which the respondent compute device 201 can output and / or display data. The input device can include a mouse, keyboard, touch screen, voice interface, and / or any other hand-held controller or device or interface via which a user may interact with the respondent compute device 201. The output module can include a bus, port, and / or other interfaces by which the respondent compute device 201 may connect to and / or output data to other devices and / or peripherals, such as a speaker.

[0034] FIG. 3 is a schematic diagram of logic components 300 for generating a response, according to an embodiment. The logic components 300 can be associated with a compute device (e.g., a compute device that is structurally and / or functionally similar to the respondent compute device 201 of FIG. 2 and / or the respondent compute devices 120 of FIG. 1). In some instances, for example, the logic components 300 can be implemented as software stored in memory 210 and configured to be executed via the processor 220 of FIG. 2. In some instances, for example, at least a portion of the logic components 300 can be implemented in hardware. The logic components 300 include an input 302, a response 304, and a conversation management application 312 (which can be functionally and / or structurally similar to, for example, the conversation management application 112 of FIG. 1 and / or the conversation management application 212 of FIG. 2). The conversation management application 312 includes a tokenizer 306, a context window 308, a state determinator 310, a summary generator 314, a data retriever 316, a response retriever 318, and a machine learning model 320.

[0035] The input 302 can include text data (e.g., string data, phrases, sentences, questions, natural language, etc.) generated (e.g., written, dictated, acted out through gestures and / or sign language, etc.) by a user (e.g., a trainee). The user can provide the text data as an instruction prompt to the conversation management application 312 to solicit a response and / or cause a state transition within the conversation management application 312 (e.g., to an improved state, as described in relation to the state determinator 310 herein). The input 302 can include, for example, a question that is often asked to patients and / or other people in need of service. Such queries can include, for example: “How are you feeling?”, “Has your treatment improved your symptoms?”, “Is anything else bothering you?”, and / or the like. The input 302 can also include, for example, advice, guidance, and / or the like (e.g., treatment suggestions, medical expertise and / or explanations, etc.). In some implementations, the input 302 (e.g., the text data) can be stored in a memory (e.g., the database 130 of FIG. 1), such that a user and / or a supervisor of the user can review the input 302 for further training purposes and / or the like.

[0036] The tokenizer 306 can segment a string of text (e.g., natural language), and the resulting segments (e.g., words, sub-words, characters, punctuation, and / or the like) can be referred toas “tokens.” In some instances, a token can be a numerical representation of a discrete element (e.g., a word, sub-word, etc.) of natural language text. For example, a token can be assigned an integer value that can range from zero to the size of a vocabulary to be considered. This process (referred to herein as tokenization) can permit a machine learning model (e.g., a large language model and / or a transformer) to infer semantic meaning from natural language text based on a decomposed representation of that text. Tokenization can be performed by, for example, a byte pair encoding algorithm and / or any other suitable tokenization method.

[0037] The tokenizer 306 can take as input the text data included in the input 302 and generate a set of input tokens. In some implementations, the tokenizer 306 can be included in the machine learning model 320, as described herein. In some implementations, as described in relation to the data retriever 316 and the response retriever 318, the tokenizer 306 (and / or a tokenizer structurally and / or functionally similar to the tokenizer 306) can generate a set of predetermined tokens and / or a set of previous response tokens. Such sets of tokens can be included in the context window 308, as can a set of summary tokens generated by the summary generator 314, as described herein.

[0038] The context window 308 can be a collection of tokens (e.g., a set(s) of input tokens, a set(s) of response tokens, a set(s) of predetermined tokens, a set(s) of previous response tokens, a set(s) of summary tokens, etc.) stored in a memory (e.g., a memory structurally and / or functionally similar to the memory 210 of FIG. 2) and to be used as context by the machine learning model 320 while interpreting the input 302 and / or generating the response 304. Tokens can be added to and / or removed from such a collection of tokens (e.g., the context window 308) as a conversation between the user and the agent progresses. For example, tokens associated with older dialogue can be removed from the collection of tokens (e.g., the context window 308) in response and / or contemporaneous to tokens associated with more recent dialogue being included in the collection of tokens (e.g., the context window 308). As a result, the conversation management application 312 can infer the meaning of an instruction prompt and / or an appropriate response based on the more recent dialogue (as opposed to the older dialogue).

[0039] Illustrating the context window 308 in use, the machine learning model 320 can interpret (e.g., determine a semantic meaning of) a token (e.g., a target token) included in a sequence of input tokens within the context window 308 based on the remaining tokens within the context window 308. For example, if the input 302 indicates the question “What’s yourfavorite?”, the machine learning model 320 can interpret “favorite” to refer to, for example, “favorite movie” if a token representing “movie” was included in the context window 308 (e.g., as a result of a previous dialogue relating to movies). The machine learning model 320 can then generate a response (e.g., “My favorite movie is Jaws”) based on the context inferred from the tokens in the context window 308. In this example, the machine learning model 320 would not respond with “Dogs are my favorite animal,” for example, even though older dialogue (e.g., dialogue associated with tokens since removed from the context window 308) between the user and the agent may have been related to, for example pets. The context window 308 (e.g., the collection of tokens) can have a fixed size (e.g., a fixed number of tokens) or a variable size (e.g., a number of tokens based on a position of the target token in the sequence).

[0040] The state determinator 310 can determine a state based on the set of input tokens generated by the tokenizer 306 in response to receiving the input 302 (e.g., the user’s questions, prompts, etc.). Alternatively and / or in addition, although not shown in FIG. 3, in some implementations, the state determinator 310 can receive text data (e.g., rather than or in addition to tokens generated by the tokenizer 306) as input to determine a state. In such implementations, the state determinator can receive text data directly from the input 302. The state can include, for example, an emotional state, psychological state, mood, and / or the like, to be emulated by the conversation management application 312 in generating the response 304. Specific examples of a state can include, for example, a degree of interest in and / or attention paid to a topic, a degree of enj oyment of a topic, a degree of comprehension of a topic, an evolving mood state associated with overall contentment, depression, etc., and / or the like. The state can be influenced by, for example, the content, tone, etc., of the user’s prompt included in the input 302. For example, if the user fails to ask follow-up questions and instead includes statements in the prompt, the state determinator can infer that the user lacks empathy, has a condescending tone, etc., and determine a state (e.g., distrust, nervousness, etc.) for the agent in response. As described below, the state can evolve over time, which can permit the user to form a “bond” with the agent (e.g., the virtual patient) simulated by the conversation management application 312, as if the agent were a real human. By determining the state, the state determinator 310 can cause the conversation management application 312 to generate a response 304 that simulates a personality of a human (e.g., a patient, customer, etc.), which can evolve over the course of a conversation and / or relationship with a professional (e.g., a therapist, a customer service professional, etc.).

[0041] In some implementations, the conversation management application 312 can determine a state independent of the input 302 provided by the user. For example, the conversation management application 312 can cause a transition to a state based on a passing of a predefined period of time and / or a number of exchanges in a conversation. As a result, the agent can simulate, for example, a cheerful emotional state during a first day and a depressed emotional state during a second day. Thus, the agent can simulate a human having “good days and bad days” for a reason that might not be attributable to the user. In some implementations, the conversation management application 312 can cause a transition to a state randomly. In some implementations, the conversation management application 312 cause a predetermined transition to a state (e.g., irrespective of the input 302), such that the agent can simulate a response to, for example, drug treatment, therapy, etc., that was effective, ineffective, and / or that had an adverse effect and / or caused reaction. In some implementations, the conversation management application 312 can cause a predetermined transition to a state (e.g., irrespective of the input 302), such that the agent can simulate an emotional reaction to a simulated life event (e.g., a reaction to a loss of a family member).

[0042] In some implementations, the state determinator 310 can classify the instruction prompt included in the input 302 by topic. While the agent is simulating a patient, for example, a topic can include a treatment, symptom, event (e.g. , a death in the family), diagnosis, and / or the like. The state determinator 310 can implement zero-shot topic classification using, for example, a machine learning model, such as BART-MNLI, TARS, GPT, and / or the like. As described herein in relation to the data retriever 316 and the response retriever 318, the topic classification can be used by the conversation management application 312 to locate and retrieve additional data from a memory to add to the context window 308. In some implementations, the state determinator 310 can use a machine learning model to generate an embedded vector based on the input 302. This machine learning model can include, for example, a transformer model and / or a similar model configured to identify relationships in sequential data and / or natural language, such as the machine learning model 320 or a machine learning model separate from the machine learning model 320. This embedded vector can indicate a semantic meaning of an instruction prompt of the input 302 and can be used by the data retriever 316 and the response retriever 318 to locate and retrieve additional data from a memory to add to the context window 308, as described herein.

[0043] In some implementations, the state determinator 310 can generate a psychological metric based on the input 302 and / or the topic. The psychological metric can include, for example, an importance metric (e.g., an importance score) and / or a valence metric (e.g., a valence score). The importance metric can be a subjective (as to the simulated agent) measure of how important and / or significant a topic is (e.g., a believed importance). The valence metric can indicate a degree to which a topic is associated with a positive sentiment, neutral sentiment, negative sentiment, a sentiment therebetween, and / or the like. For example, a topic related to a death of a parent can have a high importance metric and a low valence metric, whereas a topic related to a recently viewed movie can have a low importance metric and a high valence metric.

[0044] In some instances, the state (e.g., the psychological state, the importance metric, the valence metric, etc.) can be represented by a numerical metric (e.g., a continuous variable) and / or a discrete state. The state can have an initial state (e.g., based on configuration data, described herein), and, as a conversation and / or relationship with a user progresses, the state can change and / or evolve. An agent receiving counselling, for example, can deem atopic to be important or unimportant (as indicated by an importance metric), but can change opinion on that topic’s importance over time. In some instances, as part of a training exercise, the user can be tasked with causing the agent to transition to a goal state (e.g., a symptom-free state, a relaxation state, an improved state, etc.).

[0045] In some implementations, a state can be represented as a numerical value, which can be rounded to a nearest discrete value. This discrete value can be used as a key to search for and / or look up (e.g., in a lookup table, described herein in relation to the data retriever 316 and the response retriever 318) and retrieve data (e.g., predetermined natural language text and / or previous response text) indexed according to the key. Alternatively and / or in addition, a state can be represented as a discrete node within a Markov decision process (MDP). The MDP can include, for example, a plurality of nodes (e.g., a state space), where each node can be associated with a different state. A transition from one state (e.g., node) to another state can be caused by, for example, a topic being referenced in the input 302 by the user (e.g., as determined using zero-shot topic classification, described above), an importance and / or valence metric for the topic being above or below threshold value, a specific word and / or phrase being included in the instruction prompt of the input 302, etc.

[0046] The data retriever 316 can be configured to retrieve data from a storage memory (e.g., the database 130 of FIG. 1, the memory 210 of FIG. 2, and / or the like) based on the state determined by the state determinator 310. The retrieved data can include, for example, predetermined data generated by, for example, a human, LLM, etc., before a conversation with the user. The predetermined data can include and / or encode, for example, a text description of a backstory, a childhood recollection, a suppressed recollection, a symptom burden, and / or the like, to be brought up in conversation by the agent. The retrieval of the predetermined data can be caused and / or triggered by a transition to a particular state and / or by a topic being brought up by the user as communicated in the input 302. In this way, the agent can mimic a human who surfaces a recollection in response to a stimulus (e.g., a topic being brought up in conversation, a state of mind, etc.). In some implementations, the conversation management application 312 can be configured to use the tokenizer 306 (or a functionally and / or structurally similar tokenizer) to generate a set of predetermined tokens based on the retrieved predetermined data.

[0047] The storage memory from which the predetermined data is retrieved can include, for example, a database (e.g., a SQL database and / or a searchable vector store) configured for keyvalue search and / or embedded vector (e.g., text embedding) search. To construct the database searchable by key-value pair, the predetermined data can be labelled (e.g., via manual labelling and / or automatic classification using a machine learning model) as being associated with a state, topic, importance metric, etc. The predetermined data can be stored within the database in a location (e.g., at an index) associated with the label. The predetermined data can then be located if, for example, a classification of the input 302 (e.g., a classification of the state, topic, importance metric, etc., associated with the input 302) is associated with, equivalent to, and / or similar to the label. The state determinator 310 can generate the classification using, for example, a machine learning model, as described herein.

[0048] Alternatively and / or in addition, to construct the database searchable by embedded vector, the predetermined data can be associated with an embedded vector that can indicate a semantic meaning (e.g., a triggering state) within an embedding space. To search for the predetermined data, the data retriever 316 can receive an indication of a topic of the input 302 from the state determinator 310. This indication can include, for example, an embedded vector that encodes a semantic meaning of the input 302 and that is associated with the embedding space. Within the embedding space, the position of the embedded vector generated based onthe input 302 and / or received from the state determinator 310 can be compared to the distance of the embedded vector associated with the predetermined data. If this distance is sufficiently small (e.g., below a threshold distance), the associated predetermined data can be located and retrieved.

[0049] Having retrieved the data, the conversation management application 312 can be configured to tokenize the predetermined data (e.g., using the tokenizer 306) and include the resulting set of predetermined tokens in the context window 308. Alternatively, in some instances, the predetermined data can be stored in token form prior to retrieval of the data (e.g., during an initialization and prior to use), such that the set of predetermined tokens can be retrieved and included in the context window 308. As a result of the set of predetermined tokens being included in the context window 308, the machine learning model 320 (described herein) can consider the set of predetermined tokens as context while interpreting the input 302 and / or generating the response 304. For example, the machine learning model 320 can select at least one predetermined token from the set of predetermined tokens for inclusion in a set of response tokens representing the response 304. Such selection of the at least one predetermined token can be based on, for example, a confidence value and / or a probability that the predetermined token(s) could be included in a possible (e.g., relevant, sensible, coherent, etc.) response. If the at least one predetermined token is selected, the agent can reveal at least a portion of the content (e.g., the backstory, childhood recollection, etc.) to the user.

[0050] In some implementations, the conversation management application 312 can retrieve the predetermined data asynchronously to the receiving of the input 302 that triggers the retrieval. For example, using a first machine learning model, the input 302 can be classified and / or embedded into a vector by the state determinator 310 a period of time after a response 304 is generated in response to the input 302. Specifically, the conversation management application 312 can include the set of input tokens associated with the input 302 in the context window 308, generate the response 304 using the machine learning model 320 and based on that context window 308, and then use the input 302 to determine that predetermined data should be retrieved. After retrieving the predetermined data and generating the set of predetermined tokens, the conversation management application 312 can generate a response 304 based on the predetermined data in response to a later input, even if that input (unlike the previous input that caused the retrieval) would not itself be sufficient to cause retrieval of the predetermined data.

[0051] In some implementations, the predetermined data can be retrieved independent of a determined state and / or the input 302. For example, the conversation management application 312 can be configured to retrieve and / or receive (e.g., from a news feed) predetermined data based on a real-world current event (e.g., a championship win by a sports team, a new conflict and / or war, an election, breaking news, etc.), where the predetermined data can include data (e.g., a text description) associated with the current event. In some implementations, the predetermined data can be generated by an LLM agent configured to use a web-scraping tool and / or chain-of-thought reasoning to interpret real world events in the context of a simulated (e.g., predefined) interest(s) of the agent. The category of current event (e.g., sports, politics, entertainment, etc.) can be predetermined based on a simulated interest to be emulated by the agent. As a result of the retrieving and / or receiving, the conversation management application 312 can cause tokens representing the current event data to be injected into the context window 308, such that the conversation management application 312 can generate a response 304 that addresses the current event, models an emotional reaction to the current event and / or simulates interest in the current event.

[0052] The response retriever 318 can be configured to cause the agent to “reflect” on a previous generated response provided to the user and generate a new response that, for example, modifies this previous response. The modification can be based on, for example, the agent having a different state (e.g., increased trust) as compared to the state when the agent generated the previous response. Alternatively and / or in addition, the modification can be based on, for example, a change in psychological metric (e.g., an importance and / or valence metric) associated with atopic of the previous response. For example, based on the progression of a conversation, the importance metric and / or valence metric for the topic can increase above a threshold, causing the agent to revisit the previous response.

[0053] Similar to the storing of the predetermined data described in relation to the data retriever 316, to store a previous response such that it can be later retrieved, the response retriever 318 can be configured to (1) index the text of a previous response within a database based on a label and / or (2) associate the text of the previous response with an embedded vector in an embedding space. The label and / or embedded vector can be associated with the triggering topic and / or state. A machine learning model (e.g., a machine learning model separate from the machine learning model 320) can be used to classify a response 304 generated by the machine learning model 320 (described herein), and this classification can be used as a key to search forthe indexed text and / or to search for a sufficiently near embedded vector within the embedding space.

[0054] Similar to the retrieval of the predetermined data described in relation to the data retriever 316, the conversation management application 312 can retrieve the previous response data asynchronously to the receiving of the input 302 that triggers the retrieval. For example, an input 302 can trigger the agent to (1) generate a first response to the input 302, (2) reflect on a previous response based on the input 302, and (3) generate a new response that modifies that previous response (e.g., generate a reflection response). As described above, to “reflect” on a previous response, the conversation management application 312 can perform a first step of classifying and / or embedding the input 302 and a second step of retrieving the predetermined data. In some instances, to reduce delays perceived by a user that can result from performing these two steps, the conversation management application 312 can generate a first response to the input 302, and, subsequent and / or concurrently to the generation of the first response, perform the classifying and / or embedding the input 302 and the retrieving the predetermined data. Subsequent to the generation of the first response, the conversation management application 312 can generate the reflection response.

[0055] The summary generator 314 can generate a summary of a conversation history (e.g., previous dialogue) included in the context window 308. The conversation management application 312 can then replace, within the context window 308, the set of tokens representing the conversation history with a set of summary tokens representing the generated summary. In some instances, the set of summary tokens can be smaller (e.g., in data size) than the set of tokens previously in the context window 308 and that were replaced by the set of summary tokens.

[0056] The summary generator 314 can use a machine learning model (e.g., the machine learning model 320, a machine learning model separate from the machine learning model 320 and configured for natural language processing, a transformer model, and / or the like) to generate the set of summary tokens based on the tokens in the context window. In some implementations, the summary generator 314 can emphasize and / or include more details about in the generated summary previously discussed topics that have a higher importance as compared to topics that have a lower importance (e.g., as indicated by an importance metric). For example, a conversation history represented in the context window 308 can include a discussion of a significant topic (e.g., a symptom, illness, issue of concern, etc.) having a higherimportance and a discussion of an insignificant topic (e.g., small talk, an exchange of pleasantries, etc.) having a lower importance. The summary generator 314 can generate a summary of this conversation that excludes reference to the insignificant topic or summarizes the insignificant topic more briefly (e.g., with less text and / or data) as compared to the significant topic. In some implementations, a set of summary tokens can be generated for and replace tokens within the context window 308 that are associated with a less important topic (as determined by the importance metric), while tokens within the context window 308 associated with a more important topic can remain in the context window 308 without being summarized and / or replaced (e.g., for a period of time after the set of summary tokens are generated for the less important topic).

[0057] The summary can be generated when the conversation history has exceeded a threshold length and / or time. For example, the threshold can be a number of tokens in the context window 308, a number of inputs 302 and / or responses 304 represented in the context window 308 by (respectively) a set of input tokens and / or a set of response tokens, a length of time elapsed without a summary being generated, etc. The threshold length can be selected based on, for example, memory availability, processing resource availability, and / or the like, such that memory and / or processor usage can be improved. Alternatively and / or in addition, the threshold length can be selected to limit and / or prevent previous dialogue remaining in the context window 308 from outweighing and / or dominating a recently received input 302 (e.g. an instruction prompt). The threshold length can also be selected to limit and / or prevent the machine learning model 320 from hallucinating, such as generating an incorrect, untruthful, fictitious, misleading, and / or irrelevant response. By reducing the context window using the summary, resource usage, prompt domination, and / or hallucinations can be mitigated while context of a previous conversation(s) is still maintained when generating a response.

[0058] Alternatively and / or in addition to the above, the threshold length can also be selected based on a human recollection characteristic. For example, the threshold length can approximate a length of time that information can be retained by a human. In some instances, this length of time can depend on the importance and / or valence of the information and / or the psychological state that the human is in (e.g., a concentrated state, a distracted state, etc.). Thus, the threshold length can be dynamically based on the importance metric associated with the conversation being summarized and / or the state of the agent (as determined by the state determinator 310).

[0059] In some instances, a set of summary tokens can be further summarized after a period of time and / or after a number of exchanges within a conversation. Similarly stated, a set of summary tokens can be generated based on one or more previous summaries of dialogue. For example, an existing set of summary tokens within the context window 308 can represent a summary of multiple topics having a range of importance metrics and / or valence metrics. After a period of time and / or number of exchanges (e.g., based on predetermined thresholds of time and / or exchanges), the summary can be further compacted and / or condensed by generating a new (e.g., smaller) set of summary tokens that excludes, for example, a topic(s) having lower importance relative to a remaining topic(s). As a result of gradually reducing a summary (e.g., by gradually excluding topics based on importance metrics and / or valence metrics), the conversation management application 112 can emulate a human that gradually remembers fewer details over time and / or that forgets details of less important topics at a higher rate compared to details of more important topics.

[0060] The machine learning model 320 can be configured to generate a set of response tokens based on tokens in the context window 308. The tokens in the context window 308 can include a set of input tokens associated with the input 302, a set of summary tokens generated by the summary generator 314, a set of predetermined tokens generated by the data retriever 316, and / or a set of previous response tokens generated by the response retriever 318. The generated set of response tokens can also be included in the context window 308, such that a subsequent response 304 can be generated based on that set of response tokens. The machine learning model 320 can include an LLM, transformer model, neural network, and / or any other algorithm and / or model configured for natural language processing. Specifically, the machine learning model 320 can take as input the tokens in the context window 308, encode semantic meaning and context of those tokens using an encoder to generate a feature representation of the token, and decode the feature representation using a decoder to generate the set of response tokens. The machine learning model 320 can further normalize the set of response tokens to generate natural language text that can be included in the response 304.

[0061] In some implementations, the conversation management application 312 can have an initial state that can be defined based on configuration data. This initial state can represent, for example, an initial psychological state (e.g., a state of depression), level of understanding, level of interest, etc., that the user is tasked with improving through conversation. In some implementations, the initial state can be a classification generated based on a psychometricinstrument and / or a questionnaire (e.g., a questionnaire associated with: the Montgomery- Asberg Depression Rating Scale (MADRS), the Hamilton Depression Rating Scale (HAM-D), the Hamilton Anxiety Rating Scale (HAM-A), etc.). The instrument and / or questionnaire (which can refer to, for example, 50 symptoms, less than 50 symptoms, or more than 50 symptoms) can be completed by a human and / or LLM. In some instances, a symptom can be randomly selected from a plurality of symptoms, and / or a symptom value (e.g., a symptom intensity) can be randomly generated. Based on the instrument, questionnaire, random selection, and / or random generation, symptoms to be emulated by the agent while having the initial state can be identified. For example, a covariance structure of the symptoms to be presented by the agent (as indicated by the instrument and / or questionnaire) can be analyzed to determine relationships between symptoms. Based on these relationships, representations of symptoms (e.g., embedded symptom vectors) can be clustered within an embedding space, such that the symptoms can be collectively classified to determine an initial state (e.g., apparent sadness, pessimistic thoughts, severe depression, etc.). Thus, the covariance structure can be used to create realistic agent states, such as a set of symptoms that typically covary together. In some instances, the psychometric instrument can be used to determine a state subsequent to the initial state and / or in the midst of the conversation. For example, the psychometric instrument can be used to determine a predetermined state and / or a state that is determined independently from the input 302, as described herein.

[0062] FIG. 4 is a schematic diagram illustrating a plurality of interactions 431-436 (e.g., dataflow, transmissions, signals, etc.) between logic components 400 to generate responses, according to an embodiment. The logic components 400 can be associated a compute device that is structurally and / or functionally similar to the respondent compute device 201 of FIG. 2 and / or the respondent compute device 120 of FIG. 1. In some instances, for example, the logic components 400 can be implemented as software stored in memory 210 and configured to be executed via the processor 220 of FIG. 2. For example, at least a portion of the logic components 400 can be included in a conversation management application that is functionally and / or structurally similar to the conversation management application 112 of FIG. 1 and / or the conversation management application 212 of FIG. 2. In some instances, for example, at least a portion of the logic components 400 can be implemented in hardware. The logic components 400 include a tokenizer 406 (which can be functionally and / or structurally similar to, for example, the tokenizer 306 of FIG. 3), a context window 408 (which can be functionally and / or structurally similar to, for example, the context window 308 of FIG. 3), a statedeterminator 410 (which can be functionally and / or structurally similar to, for example, the state determinator 310 of FIG. 3), a summary generator 414 (which can be functionally and / or structurally similar to, for example, the summary generator 214 of FIG. 2 and / or the summary generator 314 of FIG. 3), a data retriever 416 (which can be functionally and / or structurally similar to, for example, the data retriever 216 of FIG. 2 and / or the data retriever 316 of FIG. 3), and a response retriever 418 (which can be functionally and / or structurally similar to, for example, the response retriever 218 of FIG. 2 and / or the response retriever 318 of FIG. 3).

[0063] The tokenizer 406 can receive input text and generate a set of input tokens based on the input text. At 431, the set of input tokens can be sent to the state determinator 410, and the state determinator 410 can determine (e.g., using a machine learning model) a state, importance metric, and / or a topic indication (e.g., classification and / or embedded vector), based on the set of input tokens. In some implementations, and although not shown in FIG. 4, the state determinator 410 can determine the state based on the input text rather than and / or in addition to the set of input tokens generated by the tokenizer 406. The set of input tokens can also be sent to the context window 408, such that a machine learning model (e.g., the machine learning model 320 of FIG. 3) can generate a response based on the set of input tokens. At 432, the state, importance metric, and / or topic indication can be sent to the data retriever 416 and / or the response retriever 418. The data retriever 416 can retrieve predetermined data from a memory based on the state, importance metric, and / or topic indication, and can then generate a set of predetermined tokens. At 433, the set of predetermined tokens can be sent to the context window 408, such that the machine learning model can generate a response based on the set of predetermined tokens. The response retriever 418 can retrieve previous response data from a memory based on the state, importance metric, and / or topic indication, and then generate a set of previous response tokens. At 434, the previous response tokens can be sent to the context window 408, such that the machine learning model can generate a response based on the set of previous response tokens.

[0064] The summary generator 414 can be configured to condense data from previous conversations within the context window 408. At 435, a set of previous conversation tokens can be sent from the context window 408 to the summary generator 414. The summary generator 414 can then generate a set of summary tokens, and at 436, the set of summary tokens can be sent to the context window 408, and the previous conversation tokens can be removed11from the context window 408. As a result, the machine learning model can generate a response based on the set of summary tokens.

[0065] FIG. 5 is a flowchart showing a method 500 illustrating an example implementation using a system described herein (e.g., the system 100 of FIG. 1). Portions of the method 500 can be implemented using a processor (e.g., the processor 220 of FIG. 2) of any suitable compute device (e.g., the respondent compute device 120 of FIG. 1 and / or the respondent compute device 201 of FIG. 2). The method 500 can include a method of generating a response based on a set of summary tokens.

[0066] The method 500 at 502 includes receiving, at a processor, input data from a user device, and generating a set of input tokens, via the processor, based on the input. At 504 of the method 500, the set of input tokens is added, via the processor, to a context window, and a state is determined, via the processor, based on the set of input tokens. The method 500 also includes, at 506, generating a set of response tokens, via the processor, by providing the set of input tokens from the context window and the state as input to a first machine learning model. Additionally, the method 500 at 506 further includes causing, via the processor, a response to be sent to the user device based on the set of response tokens. At 508, the method 500 includes generating a set of summary tokens, via the processor, by providing the set of input tokens and the set of response tokens from the context window as input to a second machine learning model. The method 500 at 510 includes removing, via the processor, at least one of the set of input tokens or the set of response tokens from the context window based on a measure of conversation length. At 512, the method 500 includes adding, via the processor, the set of summary tokens to the context window, such that the first machine learning model generates additional responses based on the context window including the set of summary tokens and excluding the at least one of the set of input tokens or the set of response tokens.

[0067] FIG. 6 is a flowchart showing a method 600 illustrating an example implementation using a system described herein (e.g., the system 100 of FIG. 1). Portions of the method 600 can be implemented using a processor (e.g., the processor 220 of FIG. 2) of any suitable compute device (e.g., the respondent compute device 120 of FIG. 1 and / or the respondent compute device 201 of FIG. 2). The method 600 can include a method of generating a response based on predetermined data.

[0068] The method 600 at 602 includes receiving, at a processor, input data from a user device and generating an input token, via the processor, by providing the input data as input to amachine learning model. At 604, the method 600 includes adding, via the processor, a predetermined token different from the input token to the context window based on the input meeting a criterion. The method 600 at 606 includes generating a response token, via the processor, by providing the predetermined token from the context window as input to the machine learning model. At 608, the method 600 includes causing, via the processor, a response to be sent to the user device based on the response token.

[0069] FIG. 7 is a flowchart showing a method 700 illustrating an example implementation using a system described herein (e.g., the system 100 of FIG. 1). Portions of the method 700 can be implemented using a processor (e.g., the processor 220 of FIG. 2) of any suitable compute device (e.g., the respondent compute device 120 of FIG. 1 and / or the respondent compute device 201 of FIG. 2). The method 700 can include a method of generating a response based on a previous response.

[0070] The method 700 at 702 includes receiving, at a processor and from a user device, first input data, and generating a topic identification and a first psychological metric for the topic, via the processor, by providing the first input data as input to a first machine learning model, the first psychological metric being below a threshold. At 704, the method 700 includes generating first response data, via the processor, by providing the first input data and the first psychological metric as input to a second machine learning model. The method 700 at 706 includes causing, via the processor, the first response data to be conveyed to the user device, and storing, via the processor and at a memory, the first response data based on the topic identification. At 708, the method 700 also includes receiving, at the processor and from the user device, second input data, and generating the topic identification and a second psychological metric for the topic, via the processor, by providing the second input data as input to the first machine learning model, the second psychological metric being above the threshold. The method 700 at 710 includes retrieving, via the processor, the first response data from the memory based on the topic identification and the second psychological metric. At 712, the method 700 includes generating second response data, via the processor, by providing the second input data and the first response data as input to the second machine learning model, and causing, via the processor, the second response data to be conveyed to the user device.

[0071] In an embodiment, a method includes receiving, at a processor, an input from a user device, and generating, via the processor, a set of input tokens based on the input. The set of input tokens is included, via the processor, in a context window, and a state is determined, viathe processor, based on the set of input tokens. The method also includes generating, via the processor and using a first machine learning model, a set of response tokens based on (1) the context window including the set of input tokens and (2) the state. Additionally, the method includes causing, via the processor, a response to be sent to the user device based on the response token, and generating, via the processor and using a second machine learning model, a set of summary tokens based on the context window including the set of input tokens and the set of response tokens. The method also includes removing, via the processor, at least one of the set of input tokens or the set of response tokens from the context window based on a measure of conversation length. Additionally, the method includes including, via the processor, the set of summary tokens in the context window, such that the first machine learning model generates additional responses based on the context window including the set of summary tokens and excluding the at least one of the set of input tokens or the set of response tokens.

[0072] In some implementations, a user of the user device can be a trainee for a role that includes communicating with a person, the state can be associated with a mental state, and the set of response tokens can be associated with an expected reaction, in response to a communication by the trainee, from the person having the mental state. In some implementations, the state can be included in a Markov decision process. In some implementations, the state can be associated with at least one of enjoyment of atopic associated with the input, interest in the topic, comprehension of the topic, or an emotional state. In some implementations, the measure of conversation length can include at least one of a period of time, a number of inputs received from the user device, or an amount of data in the context window. In some implementations, the measure of conversation length can be determined based on a human recollection characteristic. In some implementations, the set of summary tokens can represent a summary of at least one of one or more previous inputs received from the user device, one or more previous responses generated using the first machine learning model, or one or more previous summaries.

[0073] In some implementations, the method can further include receiving, at the processor, configuration data, and generating, via the processor, embedded configuration data based on the configuration data. Additionally, the method can include generating, via the processor, a covariance structure based on the embedded configuration data, and determining, via the processor, an initial state based on the covariance structure. In some implementations, the configuration data can include a representation of at least one symptom associated with a Montgomery-Asberg Depression Rating Scale (MADRS). In some implementations, thegenerating the embedded configuration data can include generating a first embedded symptom vector and a second embedded symptom vector associated with an embedding space, based on the representation of the at least one symptom. Additionally, the generating the covariance structure can include determining, via the processor, a distance, in the embedding space, between the first embedded symptom vector and the second embedded symptom vector.

[0074] In some implementations, at least one of the input or the response can include at least one of text data, audio data, or video data. In some implementations, the input can be a first input, the set of input tokens can be a first set of input tokens, the state can be a first state, the set of response tokens can be a first set of response tokens, and the response can be a first response. Additionally, the method can further include receiving, at the processor, a second input from the user device, and generating, via the processor and using the first machine learning model, a second set of input tokens based on the second input. The method can also include including, via the processor, the second set of input tokens in the context window that includes the first set of input tokens and the first set of response tokens. Additionally, the method can include determining, via the processor, a second state based on the first state and the second set of input tokens, and generating, via the processor and using the machine learning model, a second set of response tokens based on the second state and the context window including the first set of input tokens, the first set of response tokens, and the second set of input tokens. The method can also include causing, via the processor, a second response to be sent to the user device based on the second set of response tokens, and including, via the processor, the second set of response tokens in the context window.

[0075] In some implementations, the method can further include receiving, at the processor, a third input from the user device, and generating, via the processor and using the first machine learning model, a third set of input tokens based on the third input. The method can also include including, via the processor, the third set of input tokens in the context window that includes the set of summary tokens, and determining, via the processor, a third state based on the second state and the third input token. The method can also include generating, via the processor and using the first machine learning model, a third set of response tokens based on the third state and the context window including the set of summary tokens and the third set of input tokens. Additionally, the method can include causing, via the processor, a third response to be sent to the user device based on the third set of response tokens.

[0076] In an embodiment, a method includes receiving, at a processor, an input from the user device and generating, via the processor and using a machine learning model, an input token based on the input. The method also includes including, via the processor, a predetermined token different from the input token to the context window based on the input meeting a criterion. The method also includes generating, via the processor and using the machine learning model, a response token based on the context window including the predetermined token, and causing, via the processor, a response to be sent to the user device based on the response token.

[0077] In some implementations, the predetermined token can be associated with a backstory, and the response can convey at least a portion of the backstory to the user device. In some implementations, the machine learning model can be a first machine learning model, and the predetermined token can be generated by a second machine learning model different from the first machine learning model. In some implementations, the including can include (1) generating, via the processor, an embedded vector based on the input token, (2) retrieving, via the processor, predetermined data from a database based on the embedded vector, and (3) generating, via the processor, the predetermined token based on the predetermined data. In some implementations, the machine learning model can be a first machine learning model, and the including can include predicting, via the processor and using a second machine learning model configured to perform zero shot classification, a classification based on the input token. The including can also include retrieving, via the processor, predetermined data from a database based on the classification, and generating, via the processor and using the first machine learning model, the predetermined token based on the predetermined data.

[0078] In some implementations, the machine learning model can be a first machine learning model, and the including can include generating, via the processor and using a second machine learning model, a first embedded vector based on the input token, and retrieving, via the processor, predetermined data from a database based on a distance between the first embedded vector and a second embedded vector associated with the predetermined data. The including can also include generating, via the processor and using the first machine learning model, the predetermined token based on the predetermined data. In some implementations, the input can meet the criterion if the input causes a transition to a psychological state to be emulated by the machine learning model. Additionally, the including can include determining, via the processor and based on the input token, that the input has caused the transition to the psychological state. The including can also include retrieving, via the processor, predetermined data from adatabase based on the transition to the psychological state, and generating, via the processor and using the machine learning model, the predetermined token based on the predetermined data.

[0079] In an embodiment, a method includes receiving, at a processor and from a user device, first input data, and generating, via the processor and using a first machine learning model, a topic identification and a first psychological metric for the topic based on the first input data, the first psychological metric being below a threshold. The method also includes generating, via the processor and using a second machine learning model, first response data based on the first input data and the first psychological metric. The method also includes causing, via the processor, the first response data to be conveyed to the user device, and storing, via the processor and at a memory, the first response data based on the topic identification. The method also includes receiving, at the processor and from the user device, second input data, and generating, via the processor and using the first machine learning model, the topic identification and a second psychological metric for the topic based on the second input data, the second psychological metric being above the threshold. Additionally, the method includes retrieving, via the processor, the first response data from the memory based on the topic identification and the second psychological metric. The method also includes generating, via the processor and using the second machine learning model, second response data based on the second input data and the first response data, and causing, via the processor, the second response data to be conveyed to the user device.

[0080] In some implementations, the second response data can be responsive to the first input data and different from the first response data. In some implementations, each of the first psychological metric and the second psychological metric can include at least one of an importance metric or a valence metric.

[0081] Examples of computer code include, but are not limited to, micro-code or microinstructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments can be implemented using Python, Java, JavaScript, C++, and / or other programming languages and development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.

[0082] The drawings primarily are for illustrative purposes and are not intended to limit the scope of the subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the subject matter disclosed herein can be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0083] The acts performed as part of a disclosed method(s) can be ordered in any suitable way. Accordingly, embodiments can be constructed in which processes or steps are executed in an order different than illustrated, which can include performing some steps or processes simultaneously, even though shown as sequential acts in illustrative embodiments. Put differently, it is to be understood that such features can not necessarily be limited to a particular order of execution, but rather, any number of threads, processes, services, servers, and / or the like that can execute serially, asynchronously, concurrently, in parallel, simultaneously, synchronously, and / or the like in a manner consistent with the disclosure. As such, some of these features can be mutually contradictory, in that they cannot be simultaneously present in a single embodiment. Similarly, some features are applicable to one aspect of the innovations, and inapplicable to others.

[0084] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the disclosure. That the upper and lower limits of these smaller ranges can independently be included in the smaller ranges is also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0085] The phrase “and / or,” as used herein in the specification and in the embodiments, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements can optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used inconjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0086] As used herein in the specification and in the embodiments, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the embodiments, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the embodiments, shall have its ordinary meaning as used in the field of patent law.

[0087] As used herein in the specification and in the embodiments, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements can optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0088] In the embodiments, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,”“composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

[0089] Some embodiments described herein relate to a computer storage product with a non- transitory computer-readable medium (also can be referred to as a non-transitory processor- readable medium) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor-readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) can be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein.

[0090] Some embodiments and / or methods described herein can be performed by software (executed on hardware), hardware, or a combination thereof. Hardware modules can include, for example, a processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can include instructions stored in a memory that is operably coupled to a processor and can be expressed in a variety of software languages (e.g., computer code), including C, C++, Java™, Ruby, Visual Basic™, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments can be implemented using imperativeprogramming languages (e.g., C, Fortran, etc.), functional programming languages (Haskell, Erlang, etc.), logical programming languages (e.g., Prolog), object-oriented programming languages (e.g., Java, C++, etc.) or other suitable programming languages and / or development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.

Claims

What is claimed is:

1. A method, comprising: receiving, at a processor, input data from a user device; generating, via the processor, a set of input tokens based on the input data; adding, via the processor, the set of input tokens to a context window; determining, via the processor, a state based on the set of input tokens; providing, via the processor, (1) the set of input tokens from the context window and (2) the state as input to a first machine learning model to generate a set of response tokens; causing, via the processor, a response to be sent to the user device based on the set of response tokens; providing, via the processor, the set of input tokens and the set of response tokens from the context window as input to a second machine learning model o generate a set of summary tokens; removing, via the processor, at least one of the set of input tokens or the set of response tokens, from the context window based on a measure of conversation length; and adding, via the processor, the set of summary tokens to the context window, such that the first machine learning model generates additional responses based on the context window including the set of summary tokens and excluding the at least one of the set of input tokens or the set of response tokens.

2. The method of claim 1, wherein: a user of the user device is a trainee for a role that includes communicating with a person; the state is associated with a mental state; and the set of response tokens is associated with an expected reaction, in response to a communication by the trainee, from the person having the mental state.

3. The method of claim 1, wherein the state is included in a Markov decision process.

4. The method of claim 1, wherein: the state is associated with at least one of enjoyment of a topic associated with the input data, interest in the topic, comprehension of the topic, or an emotional state.

5. The method of claim 1, wherein the measure of conversation length includes at least one of a period of time, an amount of input data received from the user device, or an amount of data in the context window.

6. The method of claim 1, wherein the measure of conversation length is determined based on a human recollection characteristic.

7. The method of claim 1, wherein the set of summary tokens represents a summary of at least one of previous input data received from the user device, one or more previous responses generated using the first machine learning model, or one or more previous summaries.

8. The method of claim 1, further comprising: receiving, at the processor, configuration data; generating, via the processor, embedded configuration data based on the configuration data; generating, via the processor, a covariance structure based on the embedded configuration data; and determining, via the processor, an initial state based on the covariance structure.

9. The method of claim 8, wherein the configuration data includes a representation of at least one symptom associated with a Montgomery-Asberg Depression Rating Scale (MADRS).

10. The method of claim 9, wherein: the generating the embedded configuration data includes generating a first embedded symptom vector and a second embedded symptom vector associated with an embedding space, based on the representation of the at least one symptom; and the generating the covariance structure includes determining, via the processor, a distance, in the embedding space, between the first embedded symptom vector and the second embedded symptom vector.

11. The method of claim 1 , wherein at least one of the input data or the response includes at least one of text data, audio data, or video data.

12. The method of claim 1, wherein the input data is first input data, the set of input tokens is a first set of input tokens, the state is a first state, the set of response tokens is a first set of response tokens, and the response is a first response, the method further comprising: receiving, at the processor, second input data from the user device; providing, via the processor and using the first machine learning model, a second set of input tokens based on the second input data; adding, via the processor, the second set of input tokens to the context window that includes the first set of input tokens and the first set of response tokens; determining, via the processor, a second state based on the first state and the second set of input tokens; generating, via the processor and using the first machine learning model, a second set of response tokens based on the second state and the context window including the first set of input tokens, the first set of response tokens, and the second set of input tokens; causing, via the processor, a second response to be sent to the user device based on the second set of response tokens; and adding, via the processor, the second set of response tokens to the context window.

13. The method of claim 12, further comprising: receiving, at the processor, third input data from the user device; generating, via the processor and using the first machine learning model, a third set of input tokens based on the third input data; adding, via the processor, the third set of input tokens to the context window that includes the set of summary tokens; determining, via the processor, a third state based on the second state and the third set of input tokens; providing, via the processor, (1) the set of summary tokens, (2) the third set of input tokens, and (3) the third state as input to the first machine learning model to generate a third set of response tokens; and causing, via the processor, a third response to be sent to the user device based on the third set of response tokens.

14. A method, comprising: receiving, at a processor, input data from a user device; generating, via the processor and using a machine learning model, an input token based on the input data; adding, via the processor, a predetermined token different from the input token to a context window based on the input data meeting a criterion; providing, via the processor, the predetermined token from the context window as input to the machine learning model to generate a response token; and causing, via the processor, a response to be sent to the user device based on the response token.

15. The method of claim 14, wherein: the predetermined token is associated with a backstory; and the response conveys at least a portion of the backstory to the user device.

16. The method of claim 14, wherein the machine learning model is a first machine learning model, and the predetermined token is generated by a second machine learning model different from the first machine learning model.

17. The method of claim 14, wherein the adding includes: generating, via the processor, an embedded vector based on the input token; retrieving, via the processor, predetermined data from a database based on the embedded vector; and generating, via the processor, the predetermined token based on the predetermined data.

18. The method of claim 14, wherein the machine learning model is a first machine learning model, and the adding includes: providing, via the processor, the input token as input to a second machine learning model configured to perform zero shot classification, to predict a classification; retrieving, via the processor, predetermined data from a database based on the classification; and providing, via the processor, the predetermined data as input to the first machine learning model to generate the predetermined token.

19. The method of claim 14, wherein the machine learning model is a first machine learning model, and the adding includes: providing, via the processor, the input token as input to a second machine learning model to generate a first embedded vector; retrieving, via the processor, predetermined data from a database based on a distance between the first embedded vector and a second embedded vector associated with the predetermined data; and providing, via the processor, the predetermined data as input to the first machine learning model to generate the predetermined token.

20. The method of claim 14, wherein: the input data meets the criterion if the input data causes a transition to a psychological state to be emulated by the machine learning model; and the adding includes: determining, via the processor and based on the input token, that the input data has caused the transition to the psychological state, retrieving, via the processor, predetermined data from a database based on the transition to the psychological state, and providing, via the processor, the predetermined data as input to the machine learning model to generate the predetermined token.

21. A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to: receive, from a user device, first input data having a topic; provide the first input data as input to a first machine learning model to generate (1) a topic identification associated with the topic and (2) a first psychological metric associated with the topic, the first psychological metric being below a threshold; provide the first input data and the first psychological metric as input to a second machine learning model to generate first response data; cause the first response data to be conveyed to the user device; store, at a memory, the first response data based on the topic identification; receive second input data from the user device;provide the second input data as input to the first machine learning model to generate the topic identification and a second psychological metric for the topic, the second psychological metric being above the threshold; retrieve the first response data from the memory based on the topic identification and the second psychological metric; provide the second input data and the first response data as input to the second machine learning model to generate second response data; and cause the second response data to be conveyed to the user device.

22. The non-transitory, processor-readable medium of claim 21, wherein the second response data is responsive to the first input data and different from the first response data.

23. The non-transitory, processor-readable medium of claim 21, wherein each of the first psychological metric and the second psychological metric includes at least one of an importance metric or a valence metric.