Persistent expert avatar system with synthesis of multiple domains
The dedicated expert avatar system addresses the limitations of existing expert systems by providing continuous, customized interactions and integrating with domain-specific systems, ensuring personalized and timely responses across various platforms and environments.
Patent Information
- Application Number
- PCT/US2024/033947
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-12-18
AI Technical Summary
Existing expert systems are domain-specific, providing generic responses regardless of user expertise, lack continuity across interactions, and require users to restart conversations due to lack of memory, and do not seamlessly integrate across different platforms.
A dedicated expert avatar (XFF) system that maintains user interactions, preferences, and history, allowing continuous and customized responses across multiple devices and platforms, integrating with domain-specific XAv systems to provide unified guidance.
Enables persistent, continuous, and customized interactions by personalizing responses based on user history and preferences, ensuring seamless integration and timely access to the latest information across different environments and devices.
Smart Images

Figure US2024033947_18122025_PF_FP_ABST
Abstract
Description
PERSISTENT EXPERT AVATAR SYSTEM WITH SYNTHESIS OF MULTIPLE DOMAINSCROSS-REFERENCE TO OTHER APPLICATION
[0001] This application shares some subject matter with PCT Patent Application PCT / US2023 / 013706, filed February 23, 2023, for “Expert-Based Guidance through Virtual Avatars in Augmented Reality and Virtual Reality Environments,” which is hereby incorporated by reference.TECHNICAL FIELD
[0002] The present disclosure is directed, in general, to systems and methods for artificial intelligence (Al) expert systems.BACKGROUND OF THE DISCLOSURE
[0003] Expert systems can use Al techniques and large language models (LLMs) to interact with users in both conversational and technical styles. These systems are typically domain-specific in that they interact as “experts” in a specific field, such as a specific industrial field of use. If a user desires to find information in a different field, she must access a system (whether or not it is an expert system) in that domain to obtain the information. The information obtained from different systems may be incompatible or presented differently, so there is significant difficulty for a user to combine, reconcile, and correlate the outputs of different systems. Also, expert systems interactions with users are currently “generic” in the sense that the same question by different users with different levels of expertise generates the same response, so there is no customization of answers based on user characteristics. Further, current expert systems are “forgetful” in that one interaction with a generative Al system is independent of later interactions after passage of time, so that a user cannot effectively continue a “conversation” later but must reorient the expert system to the previous context before asking for further information. Finally, the interaction across disparate platforms best suited for user environment is not continuousand the user is required to start an interaction from the beginning as they traverse multiple hardware. Improved systems are desirable.SUMMARY OF THE DISCLOSURE
[0004] Various disclosed embodiments include methods for providing persistent, continuous, and / or customized expert interactions with a user, the method performed by a computer system implementing a dedicated expert avatar (DEA), and corresponding systems and computer-readable mediums. A method includes receiving a prompt from a user and receiving user history and user preferences corresponding to the user. The method includes applying the user history and the user preferences to the prompt to produce a plurality of queries. Each of the plurality of queries corresponds to a respective expert avatar (XAv) system. The method includes transmitting the plurality of queries to the corresponding XAv systems and receiving an avatar response from each of the XAv systems. The method also includes validating and customizing the avatar responses by applying the user history and the user preferences to the avatar responses to produce a user- customized response and transmitting the user-customized response to the user as a response to the prompt.
[0005] Note that, as described herein, an “avatar” is not limited to a visual representation in a three-dimensional (3D) environment, but instead refers to the system interacting with the user, whether by text exchanges, audio exchanges, tactile / virtual reality interactions in two-dimensions (2D) or 3D, or otherwise. The disclosed avatar can communicate with its human counterpart in any environment using any device(s) or mode(s) of communication. In various embodiment, the XAvs and the DEA can determine the most appropriate communication method based on their environment, device, and preference. Different modes of interaction are supported in the system; for example, a 3D avatar may be able to interact with the virtual environment as a human might do, while a text- or audio-only avatar may have more limited methods of communication. Note that the discussion herein may refer to such a dedicated XAv as an “XAv Friend Forever” (XFF) or “dedicated expert avatar” (DEA), using those terms interchangeably.
[0006] In various embodiments, the combined response is transmitted for delivery to the user as a textual response (with or without graphical displays), an auditory response (with or without graphical displays), or a virtual or immersive interface such as a 3D avatarresponse in a virtual environment. In various embodiments, the user-customized response combines the substance of the avatar responses and corresponds to previous interactions with the user. In various embodiments, the DEA interacts with the XAv systems via an agent framework. In various embodiments, applying the user history and the user preferences to the prompt to produce a plurality of queries also includes applying company data to the prompt. In various embodiments, applying the user history and the user preferences to the prompt to produce a plurality of queries also includes applying environment data to the prompt.
[0007] In various embodiments, applying the user history and the user preferences to the prompt to produce a plurality of queries includes modifying the prompt using user context and history. In various embodiments, each of the XAv systems is a domain-specific expert avatar system, and each of the queries is directed to an expertise of the corresponding XAv. In various embodiments, the user history and the user preferences are logically separated from a large language model used to produce the user-customized response. In various embodiments, receiving and applying the user history and the user preferences are performed using Retrieval Augmented Generation (RAG). In various embodiments, the user history and the user preferences are stored in separate datastores than company data. In various embodiments, the user history is stored in a vector database that stores a history of the user interactions as pairs of questions and answers.
[0008] Various embodiments also include coordinating communications between different XAv systems. Various embodiments also include transmitting additional queries to the XAv systems to refine the avatar responses. Various embodiments also include analyzing a plurality of interactions with the user, applying a topic analysis to group the plurality of interactions by topic, and vectorizing the plurality of interactions according to the groups.
[0009] Disclosed embodiments include a computer system having a processor and an accessible memory, particularly configured to perform processes as disclosed herein. Disclosed embodiments include a non-transitory computer-readable medium encoded with executable instructions that, when executed, cause one or more computer systems to perform processes as disclosed herein.
[0010] The foregoing has outlined rather broadly the features and technical advantages of the present disclosure so that those skilled in the art may better understand the detailed description that follows. Additional features and advantages of the disclosure will be described hereinafter that form the subject of the claims. Those skilled in the art will appreciate that they may readily use the conception and the specific embodiment disclosed as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Those skilled in the art will also realize that such equivalent constructions do not depart from the spirit and scope of the disclosure in its broadest form.
[0011] Before undertaking the DETAILED DESCRIPTION below, it may be advantageous to set forth definitions of certain words or phrases used throughout this patent document: the terms “include” and “comprise,” as well as derivatives thereof, mean inclusion without limitation; the term “or” is inclusive, meaning and / or; the phrases “associated with” and “associated therewith,” as well as derivatives thereof, may mean to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, or the like; and the term “controller” means any device, system or part thereof that controls at least one operation, whether such a device is implemented in hardware, firmware, software or some combination of at least two of the same. It should be noted that the functionality associated with any particular controller may be centralized or distributed, whether locally or remotely. Definitions for certain words and phrases are provided throughout this patent document, and those of ordinary skill in the art will understand that such definitions apply in many, if not most, instances to prior as well as future uses of such defined words and phrases. While some terms may include a wide variety of embodiments, the appended claims may expressly limit these terms to specific embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] For a more complete understanding of the present disclosure, and the advantages thereof, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, wherein like numbers designate like objects, and in which:
[0013] FIG. 1 illustrates a block diagram of a computer system in which an embodiment can be implemented;
[0014] FIG. 2 illustrates an example of a high-level architecture of a system in accordance with disclosed embodiments;
[0015] FIG. 3 illustrates a high-level overview of an example XFF architecture in accordance with disclosed embodiments;
[0016] FIGS. 4A-4C illustrate examples of possible interactions between a user and an XFF in accordance with disclosed embodiments;
[0017] FIG. 5 illustrates a process in accordance with disclosed embodiments;
[0018] FIGS. 6A-6D illustrate various processes in accordance with disclosed embodiments; and
[0019] FIG. 7 illustrates a flowchart of a process in accordance with disclosed embodiments.DETAILED DESCRIPTION
[0020] FIGS. 1 through 7, discussed below, and the various embodiments used to describe the principles of the present disclosure in this patent document are by way of illustration only and should not be construed in any way to limit the scope of the disclosure. Those skilled in the art will understand that the principles of the present disclosure may be implemented in any suitably arranged device. The numerous innovative teachings of the present application will be described with reference to exemplary non-limiting embodiments.
[0021] A previous patent application for a “System and Methods for Expert Avatars in Industrial Metaverse” described processes in which an “eXpert Avatar” (XAv) captures and transfers the knowledge and actions of an experienced shop-floor worker performing a complex task to other less experienced workers, which might have a different body shape and / or level of experience compared to the expert and might possibly be performing the complex task in different environments. That application also describes MCAD / ECAD use cases by evolving XAvs that are pre-trained in real industrial settings through observations, and then fine-tuned on domain-specific Product Lifecycle Management (PLM) and Electronic Design Automation (EDA) principals in the context of the MCAD / ECAD portfolio to get a deep understanding of the analytics and simulations that are involved in design and manufacturing of complex products.
[0022] However, the techniques described in the previous application are limited in several ways. For example, the guidance provided is generic in that it is not customized for specific users. The guidance may be disjointed in that several agents may provide input that are not connected, particularly if the agents are expert systems for different domains. The guidance could be outdated in that new instructions or information may have been given since pre-trained expert assistance model was created, and the guidance does not have persistence and continuity. That is, as discussed herein, previous digital assistants do not “memorize” previous interactions and cannot continue the “conversation” at later times or across multiple hardware platforms.
[0023] Disclosed embodiments exploit increased capabilities provided by augmented reality (AR), virtual reality (VR), and generative Al technologies to provide real-time domain-specific digital assistance to engineers or other users for comprehending large amounts of data. Disclosed techniques are especially relevant and useful for, but are not limited to, cloud-based assistance to users in performing complex tasks in PLM / EDA contexts including digesting large simulation results and pointing out critical aspects to human engineers.
[0024] Disclosed embodiments provide a dedicated “companion” XAv that can connect and interact with different individual or domain-specific XAvs to provide a unified, persistent, and consistent experience to the user. Preferably, the system implements a personal “companion” XAv exclusively for the use of each user, although of course such an XAv could be dedicated to a group or class of users.
[0025] In various embodiments, an XFF as disclosed herein learns and knows the user, so that it can accommodate the user’s communication preferences, level of expertise in different areas, familiar terminologies, etc.
[0026] In various embodiments, an XFF as disclosed herein maintains a persistent session memory in the sense that it “remembers” its interactions with the user for a long time and can recall and reference that memory when it constructs its questions or provides answers.
[0027] In various embodiments, an XFF as disclosed herein can be instantiated on multiple devices, where each instance communicates with each other, so that the XFF provides an “ever-present” experience for the user. The XFF can be always available in any environment and on any device including desktops, laptops, tablets, phones, VR devices, AR devices, MR devices, etc., and each instance can adjust its behavior and interactions according to the platform on which it is running. For example, in various implementations, the XFF can interact with a user in the form of a virtual 3D avatar in immersive environments because it can then not only provide speech and text interfaces but also support gestures and movements, and show the users what to do or what to look at by actually interacting with the virtual objects within the environment as an expert human would do. The same XFF can then adjust its behavior to use a text-only or audio-onlyinterface depending on changes in the user’s environment, device, or preferences. In text- or audio-only interactions, the XFF may not be able to use gestures, move around, or perform visual demonstrations to the user, but it still interacts with the user and continues its assistance on a more limited basis.
[0028] In various embodiments, an XFF as disclosed herein can use Retrieval Augmented Generation (RAG) to access the latest knowledge, requirements, and regulations or other data from relevant server systems and knowledge bases. This RAG system can augment the LLMs utilized by XAvs and XFF. In various embodiments, an XFF as disclosed herein can communicate with multiple domain-specific XAvs to perform interpretation and synthesis tasks as an interface between the user and the other XAvs. For example, as described herein, the XFF can interpret a user’s prompt or input to determine the appropriate XAvs to query according to the information domains involved; formulate a query for the identified XAvs; receive, combine, and reconcile the output / responses of the identified XAvs; synthesize a response to the user according to the combined responses; formulate the response according to the user’s communication preferences, and present the information in a timely manner.
[0029] A domain-specific XAv can incorporate Generative Al (GenAI) learning techniques from human experts or expert systems to predict guidance and provide assistance to less experienced users as cloud-based services.
[0030] FIG. 1 illustrates a block diagram of a computer system 100 in which an embodiment can be implemented, for example as a computer system particularly configured by software or otherwise to perform the processes as described herein, and in particular as each one of a plurality of interconnected and communicating systems as described herein. The computer system depicted includes a processor 102 connected to a level two cache / bridge 104, which is connected in turn to a local system bus 106. Local system bus 106 may be, for example, a peripheral component interconnect (PCI) architecture bus. Also connected to local system bus in the depicted example are a main memory 108 and a graphics adapter 110. The graphics adapter 110 may be connected to display 111.
[0031] Other peripherals, such as local area network (LAN) / Wide Area Network / Wireless (e.g. WiFi) adapter 112, may also be connected to local system bus 106. Expansion bus interface 114 connects local system bus 106 to input / output (I / O) bus 116. I / O bus 116 is connected to keyboard / mouse adapter 118, disk controller 120, and I / O adapter 122. Disk controller 120 can be connected to a storage 126, which can be any suitable machine usable or machine readable storage medium, including but not limited to nonvolatile, hard-coded type mediums such as read only memories (ROMs) or erasable, electrically programmable read only memories (EEPROMs), magnetic tape storage, and user-recordable type mediums such as floppy disks, hard disk drives and compact disk read only memories (CD-ROMs) or digital versatile disks (DVDs), and other known optical, electrical, or magnetic storage devices. Storage 126 can store any data necessary or useful for performing processes as described herein, including any of the executable instructions or other data and models as described.
[0032] Also connected to I / O bus 116 in the example shown is audio adapter 124, to which speakers (not shown) may be connected for playing sounds. Keyboard / mouse adapter 118 provides a connection for a pointing device (not shown), such as a mouse, trackball, trackpointer, touchscreen, etc.
[0033] Those of ordinary skill in the art will appreciate that the hardware depicted in FIG. 1 may vary for particular implementations. For example, other peripheral devices, such as an optical disk drive and the like, also may be used in addition or in place of the hardware depicted. The depicted example is provided for the purpose of explanation only and is not meant to imply architectural limitations with respect to the present disclosure.
[0034] A computer system in accordance with an embodiment of the present disclosure includes an operating system employing a graphical user interface. The operating system permits multiple display windows to be presented in the graphical user interface simultaneously, with each display window providing an interface to a different application or to a different instance of the same application. A cursor in the graphical user interface may be manipulated by a user through the pointing device. The position of the cursor maybe changed and / or an event, such as clicking a mouse button, generated to actuate a desired response.
[0035] One of various commercial operating systems, such as a version of Microsoft Windows™, a product of Microsoft Corporation located in Redmond, Wash, may be employed if suitably modified. The operating system is modified or created in accordance with the present disclosure as described.
[0036] LAN / WAN / Wireless adapter 112 can be connected to a network 130 (not a part of computer system 100), which can be any public or private computer system network or combination of networks, as known to those of skill in the art, including the Internet. Computer system 100 can communicate over network 130 with server system 140, which is also not part of computer system 100, but can be implemented, for example, as a separate computer system 100.
[0037] Various embodiments include a system that includes an XFF that interacts with a user, providing access to a common set of data sources (either general or domain-specific), continuous multimodal access, applying user- and company-specific data in a persistent manner, and customizing responses based on user characteristics.
[0038] FIG. 2 illustrates an example of a high-level architecture of a system 200 in accordance with disclosed embodiments, that includes XAv services, a common client session library and interface 202, and various XFF clients 204a / 204b.
[0039] XFF clients 204a / 204b can be implemented, for example, on one or more computer systems 100, whether in the form of a desktop or laptop computer, a mobile device such as phone or tablet, an immersive device such as head-mounted display, or other similar device. In particular, XFF clients 204a / 204b can be implemented using one or more server computer systems 100 that interact with user devices such as computers, mobile devices, augmented-reality devices, and others, so that a user’s XFF client can effectively be always connected to the user via any appropriate device, using any appropriate interface including text responses, voice prompts, talking heads, virtual-reality inputs, and others. The user devices provide the interface between the user and the user’s dedicated XFF client. Eachof the devices described herein can be implemented using any suitable computer hardware, and each will typically include at least one or more processors and an accessible memory.
[0040] An XFF client 204a / 204b can include and implement specific interfaces for various domain-specific XAvs. In this example, some domain-specific XAvs can include expert systems such as systems for process engineering simulation (PES), lifecycle services (LCS), simulation tools (STS), integrated electrical services (IES), and other PLM / EDA processes.
[0041] In the example of FIG. 2, a domain-specific PES XAv 206 includes a PES knowledge graph 208, a PES model 210, and a PES XAv interface 212. Domain-specific PES XAv 206 can also interact with knowledge graphs 214 that can include information common to multiple domain-specific XAvs and can be used to train the domain-specific knowledge graphs. In this specific example, knowledge graphs 214 can include product lifecycle management (PLM) and electronic design automation (EDA) knowledge graphs. Each domain-specific XAv can have similar components, as illustrated in FIG. 2.
[0042] Each of the domain-specific XAv interfaces, such as PES XAv interface 212, can interact with the common client session library and interface 202 and with XFF clients 204a / 204b via the common client session library and interface 202. As illustrated in FIG. 2, common client session library and interface 202 acts as an interface between the domainspecific XAvs (and other systems) and the XFF clients 204a / 204b, each of which can implement its own domain-specific XAv interfaces.
[0043] FIG. 2 illustrates a set of domain-specific XAvs, each concentrating on a single aspect of PLM / EDA: engineering, design, simulation, etc. These XAvs can be built using a common architecture, such that while the algorithms and training infrastructure are the same for each XAv, the data provided to each domain-specific XAv is what makes them different.
[0044] One advantage of having separate XAvs for different domains, such as different aspects of PLM / EDA, is that it allows the XAvs to interact with each other as if they were individuals. For example, a PES XAv may interact with the STS XAv to build a simulation,each of which is an expert in their own domain. During interactions, the goals of each XAv, as produced by the XAv, help drive the constraints needed for an appropriate simulation, in a pattern is known as the “Agent Pattern.” This framework allows different Agents to interact with each other - coordinated by another Agent, and in contact with a human - to solve complex tasks. The coordinator allows different Agents to talk directly to each other to define and execute a plan to solve a task. A system with separate XAvs also reduces the overall cost of training and validation.
[0045] Disclosed embodiments enable custom, domain-specific XAvs to coordinate with other more general Agents that may be more suited to other roles.
[0046] A “multimodal” aspect of the disclosed embodiments can provide continuous access to organizational knowledge, guidance, and expertise when and where it is needed, as illustrated in FIG. 1. For instance, a user at home can ask his / her companion XFF to prioritize the task for that day and then start collecting the necessary background information for the first identified task, while she / he gets ready to go to work. That first task may be, for example, a problem in CAE design discovered by the PES XAv and validated by the STS XAv. Once in the car, the user asks XFF to explain the issues and discuss potential solutions through its communications with PES XAv and STS XAv. Once the user gets to the office or the factory floor, the user and its XFF participate in an immersive collaborative meeting with other colleagues so that they can collectively digest the results from a design tool with the help of XAvs, and get guidance on design modifications that corrects the problem.
[0047] The underlying infrastructure, knowledge, and sessions are independent of the way that the user access the XFF; whether the user is immersed, typing, or talking, the same XFF with the same knowledge base and understanding of the user is accessed.
[0048] Information about the user’s environment can be received by the user’s XFF by means of an environment context, such as a key-value mapping that enables the XFF to understand the user’s environment and the mechanisms available to interact with the user. This context tells the XFF if the user is immersed or text-only, the attributes of any objects / users in the environment, and what methods the XFF can use to respond (such asverbally, highlighting / moving objects, moving the display, etc.). In this manner the XFF can respond appropriately in any client environment and generalize to new environments based on capability information provided by the client. For example, a user may be immersed in an XFF-enabled client application that allows users to explode a model and view its constituent parts. When the user communicates with XFF, the client application can attach an environment context that tells XFF what models are loaded, and that it can load or unload models, move them around, and explode them. Then, the XFF can proactively decide to perform these actions, such as exploding a model when the user asks about a component that is in the interior of the model. The client system, in various embodiments, can create the environment context object and provide a mechanism for executing any actions it defines, such as a REST API, so that it can be considered XFF- enabled.
[0049] Disclosed embodiments can be adapted to any available user interface, including virtual 3D avatars or talking heads that can provide relevant assistance or training to any individual performing any task of any type or complexity within the relevant domains. For example, a learning engine can use a semantic database to store captured expert knowledge for performing a task and can generalize it using multiple knowledge graphs, such as knowledge graphs 214.
[0050] For instance, in a product design and engineering context, during conceptual design phase, many design alternatives are considered. By using immersive VR hardwaresoftware combined tools, a system as disclosed herein can help designers, engineers, and manufacturers quickly come to decisions on viable alternatives. Design products such as the NX and Simcenter products offered by Siemens Digital Industries Software, where complex systems must be analyzed, can greatly benefit from the systems and methods disclosed herein. Availability of XAvs in interaction with XFF enables the system to analyze and provide feedback to a user directly within an immersive environment.
[0051] One significant technical advantage of the disclosed systems and methods persistence and integration of previous interactions between the XFF and the user. An XFF as disclosed herein can personalize the guidance and rely on the latest information andprevious experiences / conversations to generate its recommendations. In this way, the XFF system learns the needs and interaction style of the user, and can perform tasks using data and results of previous interactions to provide a more efficient and productive current interaction.
[0052] Further, an XFF as described herein can interact with a user at any available interface level, from a fully-immersive experience to more basic interactions. If users are immersed, then an actionable 3D companion XFF can guide the user through the steps and show him / her what is required. However, if a user is communicating with XFF by classical non-immersive means (e.g., on a desktop, laptop, or smartphone), then the XFF can provide assistance via text, voice, or a talking head (e.g., similar to a chatbot).
[0053] For example, in a fully-immersed environment, the XFF may be able to interact with a user using speech, text, a pointer, the user’s movement / position, gestures, etc., and can visually display or simulate the interactions, results, etc. If the user’s device or interface is more limited, for example to text with a view, text only, or audio only, the XFF interactions may be limited to only those techniques that can be support on the client device. As a user changes devices, however, the XFF can seamlessly continue the session between devices and can add or remove interface capabilities as necessary.
[0054] FIG. 3 illustrates a high-level overview of an example XFF architecture 300 in accordance with disclosed embodiments, for interacting with a user 302 (whether by immersion, text / chat, or audio).
[0055] In XFF architecture 300, an XFF process 304 (which can be implemented on any device as described above) interacts with user 302. XFF 304 stores user-customization information such as user history 306, user preferences 308, and company data 310, which allows the XFF to customize interactions with user 302.
[0056] When interacting with a user, the XFF 304 extracts and applies the user preferences 308 (at 312) and extracts and applies the user history 306 (at 314). As appropriate, these processes can include extracting and applying company data 310. XFF 304 can also identify, extract, and apply environment data (at 316), which can include, for example,determining the means and capabilities by which the user 302 is interacting with XFF 304, determining the location of user 304 (for example if the user 304 is driving, at the workplace or otherwise), determining whether user 304 is in a collaborative session with other users, and determining other physical and virtual environmental factors.
[0057] XFF communicates with a plurality of XAvs 320 via agent framework 318, which can be implemented, for example, using the Autogen software product by Microsoft or by another suitable software system. The XAvs 320 can be domain-specific XAvs, and XAvs 320 interact with data sources 322, which can include both local data sources and remote data sources such as the Internet and other public databases. XAvs 320 can, when appropriate, search for images, data, and other information to respond to queries.
[0058] In this way, user 302 interacts with XFF 304 based on a prompt. “Prompt” is used herein to refer to any statement, command, or request by the user to the XFF for which the XFF should generate a response. By contrast, “query,” as used herein, refers to statements, commands, or requests by the XFF to an XAv. This figure illustrates the prompt 330, the user-customized response 332, the query(ies) 334, and the XAv response(s) 336.
[0059] XFF 304 receives and interprets the prompt, then applies one or more of user history 306, user preferences 308, company data 310, and environmental data to clarify the prompt in context. The XFF 304 generates queries to one or more XAv systems based on the prompt. The XAvs 320 can generate responses to the queries based on the data sources 322 and return these responses to the XFF 304.
[0060] XFF understands how to communicate with XAvs and can modify the prompt from the user by adding context, history, and company data to produce appropriate queries so that the XAvs fully grasp the intent and return responses that are accurate and concise. In other words, in various embodiments, XFF 304 communicates bi-directionally with both software agents (by optimizing and customizing the queries and responses) and humans (by interpreting and validating the prompts and responses) using the most appropriate lingo / vocabulary for each entity.
[0061] When the XFF 304 receives the response(s) to the queries from the XAv(s) 320, the XFF 304 interprets and combines the response(s) according to one or more of the user history 306, user preferences 308, company data 310, and environmental data to produce a response to the prompt, which is delivered to the user.
[0062] In the example of FIG. 3, the XFF architecture 300 saves information relevant to the user or company in three different datastores, which serve different purposes while protecting company IP and privacy of individual workers by separating them from LLMs:
[0063] The User Preferences store 308 can be implemented as a simple key-value store, which stores data specific to the user, such as their experience with different products, the response style they prefer (formal or informal, terse, or step-by-step), their role, etc. Initially this data will be explicitly set by the user, perhaps through an initial interview by the XFF. Over time, the preference values may change; the user might change roles, or gain experience with a particular product. As the XFF has access to user sessions, it can analyze those sessions to determine whether the user preferences should be updated (after verifying with the user that the change should be made).
[0064] The User History store 306 can be implemented as a vector database which stores the history of the user’s interaction with the XFF as pairs of questions and answers. This personalized history can help to augment the limited context of the XFF’s underlying LLMs by finding conversations that may be relevant to the current interaction so they can be included in the context.
[0065] The company data store 310 can be implemented as a second vector database which stores any company relevant information, such as best practices, procedures, regulations, etc. This can help to tailor the responses of the XFF to the company’s specific information.
[0066] Various embodiments can accommodate segregation of data to protect user privacy. For example, in most cases XAvs are companywide or public, so a user might feel that their prompts are being captured and analyzed by the company or third parties. The XFF, however, can be dedicated to a single user, so the XFF only communicates one-on-onewith the user, and the data captured by the XFF is kept private to the user. This can help the user feel that they have control over their data.
[0067] While XAv conversations can be analyzed to understand trends and capture information about what users are asking questions about, the personal nature of the XFF can more easily capture this information from the point of view of the user. This captures the implicit knowledge of the company as understood by the users themselves, which provides a more targeted understanding of the actual tasks that users perform. With anonymization of personal data, this history could even be passed on to someone else to help them understand what the job entails and how best to perform it.
[0068] Conventional AI / LLM systems by nature react only to their given prompt or query, rather than understanding the environment that they are working in. By contrast, the XFF does have that understanding and can proactively perform tasks for the user. For example, if the user has a question about a particular component in an assembly, an XAv can return information about that component, but the XFF can take an action such as opening a link to the component in a textual environment, or automatically loading the assembly model corresponding to the component and highlighting the component in an immersive environment.
[0069] The XFF is shown in Figure 2 as another agent, but with a particular set of skills and interaction methods. The XFF coordinates the conversations between other XAvs to accomplish the goals set out for it by the human user. As the user interacts with the XFF, its “skills” are activated to extract or apply user preferences to the conversation, store or retrieve user’s prior history with the XFF, and provide relevant company data to the conversation. These skills can both understand and modify the conversation as it happens.
[0070] FIGS. 4A-4C illustrate examples of possible interactions between a user and an XFF in accordance with disclosed embodiments. FIG. 4A illustrates an example where XFF automatically provides visual results to an immersed user, and later provides a speech description to an audio user.
[0071] FIG. 4B illustrates an example of a novice user conversation with the XFF to learn to use a function of the NX software product by Siemens Digital Industries Software. In this example, the XFF takes the original prompt at 402, interprets it using user history and user preferences, and produces a query 404 to be passed to one or more XAvs. Upon receiving a response 406 from the XAv(s), the XFF interprets the responses using XFF data to produce a response 408 to the user that is reconciled with the user’s history, preferences, environment, and other factors.
[0072] More specifically, in this example, a novice user is trying to determine how to create a cylinder in Siemens NX. The process of this example is:• The XFF receives a prompt form the user: “How do I create a cylinder in NX?”• The XFF adds the question to the history, and checks whether the question (or a similar one) has been asked before. If it has, then the previous question and answer can be added to the current question’s context. In this case, the XFF notices that part of the answer was covered in a previous conversation “last Thursday” and gets confirmation from the user that they recall that part of the answer before proceeding.• The XFF can now rewrite the question based on its knowledge of the user. In this case, the user is a novice, so step-by-step instructions and a friendly tone are added. Any necessary company data, such as best practices, can also be added to the context.• The question is now passed to the XAvs through the Agent Framework as one or more queries. The XFF coordinates the Agent responses until the question has been appropriately answered.• The response from the XAvs is added to the history.• The answer may be rewritten based on the XFFs knowledge of the user characteristics.• Finally, the response is validated and provided back to the user. The user may continue the conversation, ask another question, etc. at which point the process is repeated.
[0073] FIG. 4C illustrates an example of an advanced user conversation with the XFF showing the use of history. In this example, the XFF takes the original prompt at 410, interprets it using user history and user preferences, and produces a query 412 to be passed to one or more XAvs. Upon receiving a response 414 from the XAv(s), the XFF interprets the responses using XFF data to produce a response 416 to the user that is reconciled with the user’s history, preferences, environment, and other factors.
[0074] As can be seen, the example of FIG. 4C shows an experienced user asking the same question as in the example of FIG. 4B. The same process is used, but now the longer history between the user and XFF comes into play and the response is different. Over time, the company best practices have been updated to recommend using a sketch to create a circle, and then extrude that into a cylinder. However, the XFF associated with this individual recognizes from its history that in the past, the user has successfully used the Cylinder primitive and suggests that in addition to the best practice from the XAvs.
[0075] As noted above, the various data stores, such as user history 306, user preferences 308, company data 310, and others, can be implemented as vector databases, for example using a RAG process. These vector databases can be implemented using open source or commercial offerings, customized for the vectorization of data such as user history and company data. For example, a user session may include an arbitrary number of conversation rounds, where the user provides some input and the XFF provides some response. It is entirely possible that over the course of a longer conversation the topic will change several times. While a standard vectorization algorithm that treats the entire session as a single text input, such treatment may not be able to be effectively matched against a new input from the user.
[0076] FIG. 5 illustrates a process 500 in accordance with disclosed embodiments that can associate the topics with a full conversation. Such a process can be performed, for example,by an XFF system implemented by one or more computer systems 100 or other device as disclosed herein (generically referred to as the “system” below).
[0077] At 502, the system receives a plurality of interactions between the user and the XFF. “Receiving,” as used herein, can include loading from storage, receiving from another device or process, receiving via an interaction with a user, or otherwise. In the case, each interaction may include a user input (whether textual, spoken, gestures, or otherwise) and may include a response from the XFF.
[0078] At 504, for each interaction, the system applies a topic analysis model to identify a topic of that interaction.
[0079] At 506, the system groups similar interactions together by topic to produce topic “documents” for purposes of vectorization.
[0080] At 508, the system vectorizes each topic document using a conventional vectorization algorithm and stores the vectorized documents in a data storage. In this way the system vectorizes interactions according to the topic groups.
[0081] At 510, the system associates the topic documents in the data storage with combined interaction between the user and the XFF. The combined interaction can be the combined interaction of a certain date or session or can be combined over greater periods of time, including the entirety of the user’s interaction with his XFF.
[0082] When the user submits a new request, this process will make it easier for the vector database to return the exact data relevant to the request, while still retaining a relationship to the full historical session to get the full context.
[0083] FIGS. 6A-6D illustrate various processes in accordance with disclosed embodiments that can be performed by the XFF to augment the requests to, and responses from, the underlying LLMs and RAG.
[0084] FIG. 6A illustrates a process for updating a user history store, such as user history 306, in accordance with disclosed embodiments.
[0085] At 602, the system receives one or more interactions in a session.
[0086] At 604, the system vectorizes the user session and its interactions, for example as described above, to produce history-vector pairs 606.
[0087] At 608, the system stores the history-vector pairs in the user history store. Note that in this architecture, user history can be easily deleted from the system if desired.
[0088] FIG. 6B illustrates a process for querying the user history store in accordance with disclosed embodiments.
[0089] At 610, the system receives a user prompt.
[0090] At 612, the system vectorizes the user prompt to produce a vector 614, for example, using a vectorization process as described above.
[0091] At 616, the system queries the user history store using the vector 614 and receives a response 618 to the query. The response 618, in this example, is a set of text-similarity pairs, where the text contains a historical question and answer, and the similarity shows how similar the text is to the input.
[0092] The company data store, or other data storage, can be added to and queried in a similar fashion, potentially with different vectorization approaches appropriate to the exact data being stored.
[0093] FIG. 6C illustrates a process for reformulating a user prompt to include the user preferences, in accordance with disclosed embodiments. In this process, the system can use an LLM to generate a new prompt, providing some instructions to the LLM, the user preferences, and the user’s original prompt. The output is a rewritten prompt that takes the user preferences into account.
[0094] At 620, system receives a user prompt.
[0095] At 622, the system generates a query to an LLM 628. The query can include an LLM prompt 624, user preferences 626, and the original user prompt 620.
[0096] As part of 622, the system uses LLM 628 to generate a rewritten prompt 630, which can then be used as a query to XAvs via an agent framework.
[0097] FIG. 6D illustrates that producing the query that is sent to the agent framework for processing can including the user prompt 620, the user history 306, and other data such as company data 310.
[0098] The final query is used by the Agent Framework to select one or more appropriate Agents to answer the question and generate a response.
[0099] Since all of this data is kept separate, it helps to keep the data up to date. This provides several technical advantages over other approaches. For example, as corporate data (or other data) is updated, the historical sessions still keep track of answers on older data. This can also be helpful to a user trying to track down information from an older session or from a point whether other circumstances applied, such as when laws, guidelines, or practices change.
[0100] If a user leaves a company or opts out of data retention, their XFF history can be deleted independently from the company data. By not mixing user history and company data in the same data store, the system allows greater control over user privacy.
[0101] FIG. 7 depicts a flowchart of a process in accordance with disclosed embodiments that may be performed, for example, by one or more computer systems as disclosed herein implementing a dedicated expert avatar (DEA).
[0102] At 702, the DEA receives a prompt from a user.
[0103] At 704, the DEA receives user history and user preferences corresponding to the user. The DEA can also receive company data.
[0104] At 706, the DEA applies the user history and the user preferences to the prompt to produce a plurality of queries, each of the plurality of queries corresponding to a respective expert avatar (XAv) system. The DEA can also apply the company data to the prompt to produce the plurality of queries.
[0105] At 708, the DEA transmits the plurality of queries to the corresponding XAv systems. Each of the XAvs can use LLMs and RAG to generate avatar responses, which can be domain-specific responses.
[0106] At 710, the DEA receives an avatar response from each of the XAv systems. Steps 706, 708, and 710 can be repeated as the DEA receives avatar responses, refines the queries according to the user history and user preferences, and sends refined queries to the XAv systems to generate further avatar responses. Further, the DEA can coordinate communications between the different XAv systems, refining queries and responses as needed, to ensure that the eventual avatar responses are responsive to the user prompt and accommodate the user history and user preferences.
[0107] At 712, the DEA validates and customizes the avatar responses by applying the user history and the user preferences, and optionally any company data, to the avatar responses to produce a user-customized response that combines the substance of the avatar responses. By applying user history, user preferences, and / or company data, the user-customized response can correspond to previous interactions with the user.
[0108] At 714, the DEA transmits the user-customized response to the user as a response to the prompt. The combined response can be transmitted for delivery to the user as a textual response, an auditory response, a virtual avatar response in a virtual environment, or otherwise.
[0109] The techniques described herein are not limited to the specific XAvs and LLMs described herein, but can be used with any LLM or expert system, including those not yet implemented.
[0110] Another significant advantage is that an XFF as disclosed herein can communicate bi-directionally with both humans and software agents using the most appropriate “language” that each entity understands, resulting in more accurate and concise interactions.
[0111] Any generated response by the XFF is customized for the specific user based on that individual’s knowledge, level of expertise, and preferred style of communication (so- called “user characteristics”).
[0112] Another significant advantage is that responses from various XAvs are combined by the XFF and packaged in a coherent conversation. Another significant advantage is that all information provided by XAVs is evaluated by the XFF to ensure that it is valid, current to all available data (e.g., it includes the latest set of company data / regulations), and can be double-checked for accuracy by the human user. Another significant advantage is that the interaction between XFF and the user is conducted in a continuous manner using the most convenient and / or appropriate hardware device across multiple environments.
[0113] Another significant advantage is that, by preserving the history of previous conversations in a persistent manner, the XFF helps users recall what has been done or learned in the past and thus avoid duplication or redundancy in the conversation.
[0114] An architecture as disclosed herein has the advantage of separating LLMs from company data and user data, thus protecting company intellectual property (IP) and user privacy, and simplifying updates to those data or removing outdated / undesired data from the system. Such an architecture also simplifies anonymizing the data of the experienced users, enabling the use of their expertise for the benefit of the entire enterprise by extracting just the characteristics of their interactions with XFF from the vector databases and including them in the responses.
[0115] An XFF “digital companion” as disclosed herein behaves like an expert colleague and effectively communicates with other agents in order to provide helpful instructions to its human counterpart. Engineers and manufacturers don’t use menus, icons, or cryptic textual guidance to learn and reason; for the most part, engineers explore the industrial environment relying on experts and previous experiences, and in so doing, they gain a rich and robust understanding of the problems and potential solutions. The user-specific XFF helps close the gap between Artificial Intelligence and common sense for industrial and other applications. Humans can exercise common sense, i.e., the ability to use generalknowledge about the world and apply it to any problem or task once guided by one or more experts.
[0116] The architecture as disclosed herein demonstrates how an XFF relies on the information provided by multiple domain-specific expert avatars while at the same time recalling previous conversations and accessing the latest enterprise knowledge and regulations through techniques such as RAG. Furthermore, the fact that the same XFF is always available in any environment and on any device means that a human user is never alone as she / he navigates the complexities of working within a digital industry.
[0117] Of course, those of skill in the art will recognize that, unless specifically indicated or required by the sequence of operations, certain steps in the processes described above may be omitted, performed concurrently or sequentially, or performed in a different order.
[0118] Those skilled in the art will recognize that, for simplicity and clarity, the full structure and operation of all computer systems suitable for use with the present disclosure is not being depicted or described herein. Instead, only so much of a computer system as is unique to the present disclosure or necessary for an understanding of the present disclosure is depicted and described. The remainder of the construction and operation of computer system 100 may conform to any of the various current implementations and practices known in the art.
[0119] It is important to note that while the disclosure includes a description in the context of a fully functional system, those skilled in the art will appreciate that at least portions of the mechanism of the present disclosure are capable of being distributed in the form of instructions contained within a machine-usable, computer-usable, or computer-readable medium in any of a variety of forms, and that the present disclosure applies equally regardless of the particular type of instruction or signal bearing medium or storage medium utilized to actually carry out the distribution. Examples of machine usable / readable or computer usable / readable mediums include: nonvolatile, hard-coded type mediums such as read only memories (ROMs) or erasable, electrically programmable read only memories (EEPROMs), and user-recordable type mediums such as floppy disks, hard disk drives and compact disk read only memories (CD-ROMs) or digital versatile disks (DVDs).
[0120] Although an exemplary embodiment of the present disclosure has been described in detail, those skilled in the art will understand that various changes, substitutions, variations, and improvements disclosed herein may be made without departing from the spirit and scope of the disclosure in its broadest form. The various features, processes, and systems disclosed herein can be combined with one or more of the features, processes, and systems described in the prior application incorporated by reference herein.
[0121] None of the description in the present application should be read as implying that any particular element, step, or function is an essential element which must be included in the claim scope: the scope of patented subject matter is defined only by the allowed claims. Moreover, none of these claims are intended to invoke 35 USC §112(f) unless the exact words "means for" are followed by a participle. The use of terms such as (but not limited to) “mechanism,” “module,” “device,” “unit,” “component,” “element,” “member,” “apparatus,” “machine,” “system,” “processor,” or “controller,” within a claim is understood and intended to refer to structures known to those skilled in the relevant art, as further modified or enhanced by the features of the claims themselves, and is not intended to invoke 35 U.S.C. §112(f).
Claims
WHAT IS CLAIMED IS:
1. A method for providing persistent, customized expert interactions with a user, the method performed by a computer system (100) implementing a dedicated expert avatar (DEA) and comprising: receiving (702) a prompt (330) from a user; receiving (704) user history (306) and user preferences (308) corresponding to the user; applying (706) the user history (306) and the user preferences (308) to the prompt (330) to produce a plurality of queries (334), wherein each of the plurality of queries (334) corresponds to a respective expert avatar (XAv) system (320); transmitting (708) the plurality of queries (334) to the corresponding XAv systems (320); receiving (710) an avatar response (336) from each of the XAv systems (320); validating and customizing (712) the avatar responses (336) by applying the user history (306) and the user preferences (308) to the avatar responses (336) to produce a user-customized response (332); and transmitting (714) the user-customized response (332) to the user as a response to the prompt (330).
2. The method of claim 1, wherein the user-customized response (332) is transmitted for delivery to the user as a textual response, an auditory response, or a virtual avatar response in a virtual environment.
3. The method of claim 1, wherein the user-customized response (332) combines the substance of the avatar responses (336) and corresponds to previous interactions with the user.
4. The method of claim 1, wherein the DEA interacts with the XAv systems (320) via an agent framework.
5. The method of claim 1, wherein applying (706) the user history (306) and the user preferences (308) to the prompt (330) to produce a plurality of queries (334) also includes applying company data (310) to the prompt (330).
6. The method of claim 1, wherein applying (706) the user history (306) and the user preferences (308) to the prompt (330) to produce a plurality of queries (334) also includes applying environment data to the prompt (330).
7. The method of claim 1, wherein applying (706) the user history (306) and the user preferences (308) to the prompt (330) to produce a plurality of queries (334) includes modifying the prompt (330) using user context and history.
8. The method of claim 1, wherein each of the XAv systems (320) is a domain-specific expert avatar system, and each of the queries (334) is directed to an expertise of the corresponding XAv (320).
9. The method of claim 1, wherein the user history (306) and the user preferences (308) are logically separated from a large language model used to produce the user- customized response (332).
10. The method of claim 1, wherein receiving (704) and applying (706) the user history (306) and the user preferences (308) are performed using Retrieval Augmented Generation.
11. The method of claim 1, wherein the user history (306) and the user preferences (308) are stored in separate datastores than company data (310).
12. The method of claim 1, wherein the user history (306) is stored in a vector database that stores a history of the user interactions as pairs of questions and answers.
13. The method of claim 1, further comprising coordinating communications between different XAv systems (320).
14. The method of claim 1, further comprising transmitting (708) additional queries (334) to the XAv systems (320) to refine the avatar responses (336).
15. The method of claim 1, further comprising analyzing a plurality of interactions with the user, applying a topic analysis to group the plurality of interactions by topic, and vectorizing the plurality of interactions according to the groups.
16. A computer system (100) comprising: a processor (102); and an accessible memory (108), the computer system (100) particularly configured to perform processes as in any of claims 1-15.
17. A non-transitory computer-readable medium (126) encoded with executable instructions that, when executed, cause one or more computer systems (100) to perform processes as in any of claims 1-15.
Citation Information
Patent Citations
Artificial intelligence platform with improved conversational ability and personality development
US20190156222A1
Multi-modal dialogue agent
US20200160199A1
Ai guided spectrum operations
US20210117809A1
Cited By
Data analysis agent synthesizing responses from experience data using large language models
US20260236481A1