Data-augmented training of reinforcement learning software agents
By using knowledge base information to train reinforcement learning models and testing them in a sandbox environment, the problem of insufficient performance of software agents in complex environments was solved, and efficient autonomous operation and rapid goal achievement were achieved.
Patent Information
- Application Number
- CN202080082641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-11-27
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-11-27
AI Technical Summary
Existing software agents have limited performance in complex environments, and traditional reinforcement learning paradigms often fail to effectively operate autonomously and learn from past decisions, resulting in poor performance in dynamic environments.
The reinforcement learning model is trained by receiving semi-structured information from the knowledge base and tested in a sandbox test environment. It is trained and deployed using a domain simulator, combined with natural language processing technology and a performance evaluation engine to generate and match policy actions, and dynamically update the model to improve performance.
It improves the performance of reinforcement learning software agents in complex environments, reduces training time, enhances the ability to operate autonomously in dynamic environments, and achieves efficient action execution and goal achievement.
Smart Images

Figure CN114730306B_ABST
Abstract
Description
Background Art Technical Field
[0002] Embodiments of the present invention relate to reinforcement learning software agents, and more particularly to training a reinforcement learning model that supports the reinforcement learning software agent based on information obtained from one or more knowledge stores, and testing the trained reinforcement learning model in an environment with limited connectivity to the external environment. The software agent can be deployed along with the tested and trained reinforcement learning model.
[0003] Related technical discussions
[0004] Software agents or decision agents are deployed in various settings. Examples of software agents include conversational agents or chatbots, online shopping software agents, spam filters, etc. Software agents are typically deployed in dynamic environments. Rather than being programmed to perform a set of tasks, software agents that can receive input from their current environment are configured to act autonomously in order to achieve a desired goal.
[0005] However, software agents, or decision agents, while operating autonomously, may not make optimal decisions. For example, some software agents operate in a seemingly random manner and do not learn from past decisions. Other software agents may use machine learning techniques (e.g., reinforcement learning) as part of their decision-making process, but these software agents still have limited performance in complex situations. Consequently, software agents are often limited in application to simple systems, and traditional reinforcement learning paradigms often fail in complex situations. Summary of the Invention
[0007] According to embodiments of the present invention, methods, systems, and computer-readable media are provided for configuring a reinforcement learning model for a reinforcement learning software agent using data. Access to a knowledge base is received, the knowledge base including data related to topics supported by the software agent. Information from the knowledge base is used to train a reinforcement learning model, which informs the reinforcement learning software agent. The trained reinforcement learning model is tested in a sandbox testing environment having limited connectivity to an external environment. The reinforcement learning software agent is deployed within the environment along with the tested and trained reinforcement learning model to autonomously perform actions to process requests. This method allows for improved performance of reinforcement learning software agents in complex environments, as well as improved training of reinforcement learning software agents for deployment in complex environments.
[0008] In various aspects, information from a knowledge base is in a semi-structured format and includes questions with one or more corresponding answers. Access to a domain simulator is provided to enable a reinforcement learning software agent to perform actions in a test environment. Based on this approach, the present technology is compatible with any type of information that is in a semi-structured format and that provides a domain simulator.
[0009] In other aspects, the knowledge base includes features used by the reinforcement learning model to rank the information provided in the knowledge base. This method allows knowledge from subject matter experts to be automatically processed and accessed and provided as guidance to reinforcement learning software agents in complex environments. In some embodiments, the reinforcement learning model ranks the information in the knowledge base based on user preferences. These techniques can be used to customize specific actions for specific users.
[0010] In various aspects, a query is generated by a reinforcement learning software system. The query is matched to a state in a policy generated by a reinforcement learning model, where each state has a corresponding action. The reinforcement learning software agent executes the action returned from the matched policy. In other aspects, a reward can be received by the reinforcement learning software agent for executing the action when the action achieves a goal or solves the query. This allows the reinforcement learning model to provide ranked information to the reinforcement learning software agent to accelerate training and execute actions based on the ranking to achieve the desired goal. This approach can lead to improved performance of RL (reinforcement learning) software agents in complex environments with high fan-out in available actions.
[0011] In various aspects, the policy generated by the reinforcement learning model is updated at a fixed or dynamic frequency to include questions and answers added to the knowledge base. This allows the reinforcement learning model to be updated and include current or recent information in the knowledge base.
[0012] It should be understood that the present invention summary is not intended to identify the key features or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Generally, the same reference numerals are used to denote the same components throughout the various drawings.
[0014] Figure 1A is a schematic diagram of an example computing environment for a reinforcement learning system according to an embodiment of the present invention.
[0015] Figure 1B According to an embodiment of the present invention, Figure 1A An example computing device of a computing environment.
[0016] Figure 2is a flow chart illustrating example computational commands and operations for a reinforcement learning software agent deployed in a reinforcement learning system according to an embodiment of the present invention.
[0017] Figures 3A to 3B is a diagram illustrating an example environment for a reinforcement learning software agent deployed in a reinforcement learning system according to an embodiment of the present invention. Figure 3A Corresponding to the computing environment. Figure 3B Corresponding to the travel environment.
[0018] Figure 4 is a flow diagram illustrating a sandbox environment for testing reinforcement learning software agents according to an embodiment of the present invention.
[0019] Figure 5 is a schematic diagram of input and output of a test environment according to an embodiment of the present invention.
[0020] Figure 6A is a flowchart illustrating ranking knowledge base data by a reinforcement learning system according to an embodiment of the present invention.
[0021] Figure 6B is a diagram illustrating utilizing a source code developed in a sandbox environment according to an embodiment of the present invention. Figure 6A Flowchart of policy information deployed in a reinforcement learning system.
[0022] Figure 7A and Figure 7B An example of evaluating the performance of a reinforcement learning system according to an embodiment of the present invention is shown.
[0023] Figure 8 A high-level flowchart illustrating the operation of a reinforcement learning system according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0024] A reinforcement learning (RL) model supporting a reinforcement learning (RL) software agent is trained based on information obtained from one or more knowledge bases. The RL model supporting the RL software agent is trained and tested in a setting with limited connectivity to the external environment. The RL software agent can be deployed alongside the tested and trained RL model. The RL software agent can receive positive reinforcement upon completing a sequence of actions and corresponding state changes that lead to reaching a goal. When the same state is reached at a future point in time, the RL software agent retains the memory and can therefore reach the goal by repeating the learned actions.
[0025] An example environment for use with embodiments of the present invention is shown in FIG1 . Specifically, the environment includes one or more server systems 10, one or more client or end-user systems 20, a database 30, a knowledge base 35, and a network 45. The server system 10 and the client system 20 can be remote from each other and can communicate via the network 45. The network can be implemented by any number of any suitable communication media, such as a wide area network (WAN), a local area network (LAN), the Internet, an intranet, etc. Alternatively, the server system 10 and the client system 20 can be local to each other and can communicate via any suitable local communication medium, such as a local area network (LAN), hardwire, wireless link, intranet, etc.
[0026] Client system 20 enables users to interact with various environments, performing activities that can generate queries that are provided to server system 10. As described herein, server system 10 includes RL system 15, which includes RL software agent 105, natural language processing (NLP) engine 110, RL model 115, domain simulator 120, performance evaluation engine 125, and query generation engine 130.
[0027] Database 30 can store various information for analysis, such as information obtained from knowledge base 35. For example, RL system 15 can analyze a knowledge base located in its local environment, or can generate and store a copy of knowledge base 35 as knowledge base 31 in database 30. Knowledge base 35 can include any suitable information in a semi-structured format (such as a question-answer format), including but not limited to online forums, public databases, private databases, or any other suitable repository. In some aspects, multiple answers can be provided for a single question. In other aspects, a question can be generated without any corresponding answer.
[0028] Information from knowledge base 35 may be processed by RL system 15 and stored as processed Q / A data 32. For example, processed Q / A data 32 may include information processed by NLP engine 110 to generate data compatible with training RL model 115 from knowledge base data 31. In various aspects, this information may include mapping a question to each answer, and mapping each answer to one or more features associated with the answer to determine the relevance and quality of the answer.
[0029] Once the RL model 115 has been trained, the output of the trained RL model can be stored as policy data 34. In various aspects, the policy data 34 can be provided in a format compatible with the decision agent, such as a (state, action) format. By matching the state of the RL software agent 105 with the state of the policy data 34, the corresponding action of the policy data can be identified and provided to the RL software agent.
[0030] Database 30 may store various information for analysis, such as knowledge base data 31 obtained from online forums or any other knowledge base, processed Q / A data 32 and policy data 34, or any other data generated through operations involved by the RL system in a sandbox / testing environment or a real environment.
[0031] Database system 30 may be implemented by any conventional or other database or storage unit, may be located local to or remote from server system 10 and client system 20, and may communicate via any suitable communication medium, such as a local area network (LAN), a wide area network (WAN), the Internet, a hardwire, a wireless link, an intranet, etc. The client system may present a graphical user interface (e.g., a GUI, etc.) or other interface (e.g., a command line prompt, a menu screen, etc.) to request information from the user regarding the execution of actions in the environment, and may provide reports or other information, including whether the actions provided by the RL system successfully achieved the goal.
[0032] The server system 10 and the client system 20 may be implemented by any conventional or other computer system preferably equipped with a display or monitor, a base (including at least one hardware processor (e.g., microprocessor, controller, central processing unit (CPU)), etc.), one or more memories and / or internal or external network interface or communication device (e.g., modem, network card, etc.), optional input device (e.g., keyboard, mouse or other input device), and any commercial and custom software (e.g., server / communication software, RL system software, browser / interface software, etc.). By way of example, the server / client includes at least one processor 16, 22, one or more memories 17, 24 and / or internal or external network interface or communication device 18, 26 (such as a modem or network card), and user interface 19, 28, etc. The optional input device may include a keyboard, mouse or other input device.
[0033] Alternatively, one or more client systems 20 may perform RL in a standalone operating mode. For example, the client system stores or accesses data (such as a knowledge base 35), and the standalone unit includes the RL system 15. A graphical user interface or other interface 19, 28 (such as a GUI, a command line prompt, a menu screen, etc.) requests information from the corresponding user, and the state of the system can be used to generate queries.
[0034] RL system 15 may include one or more modules or units to perform the various functions of the embodiments of the present invention described herein. The various modules (e.g., RL software agent 105, natural language processing (NLP) engine 110, RL model 115, domain simulator 120, performance evaluation engine 125, and query generation engine 130, etc.) may be implemented by any combination of any number of software and / or hardware modules or units and may reside in memory 17 of the server for execution by processor 16. These modules are described in more detail below. These components operate to improve the operation of RL software agent 105 in a real-world environment, thereby reducing the time required for the RL software agent to identify and implement one or more actions (actions) to achieve a goal. Multiple answers may exist, resulting in different sequences of actions to achieve the goal.
[0035] RL software agents 105 can provide decision support in a domain as part of RL system 15. Domain simulators, domain models, and knowledge bases can be used to configure such systems. In some aspects, the knowledge base includes data stored in the form of questions and answers corresponding to the domain. RL software agents can be trained using reinforcement learning, where data from the knowledge base provides information related to actions. Before deployment, RL software agents can be tested in a sandbox environment, which is typically a non-production environment using a simulator.
[0036] The NLP engine 110 relies on NLP technology to extract relevant parts of information from the knowledge base 35, which can be in the form of semi-structured text. The NLP engine 110 is able to identify the forum section corresponding to the question, one or more answers to the question, and features associated with each answer, which can be provided to the RL model 115 so that the RL model ranks the answers to the specific question. Features may include author name, author status, upvotes, downvotes, response length, comments on the response by other users indicating that the solution worked, etc., or any combination thereof. In some cases, an "upvote" represents the confidence and approval of the answer by the community and can represent a reward signal to provide motivation for the RL algorithm decision-making process.
[0037] Information retrieval techniques used to crawl the contents of a knowledge base can be used to generate a copy of the knowledge base information and store that information in database 30. Alternatively, RL system 15 can communicate with the knowledge base to obtain relevant information related to the topic (e.g., a set of answers to a particular question). The obtained information can be processed by an NLP engine and stored in database 30 as processed Q / A data 32.
[0038] The RL model 115 includes any suitable algorithm for ranking data associated with a feature set. For example, any suitable algorithm may be used, including but not limited to FastAP, PolyRank, Evolutionary Strategy (ES) ranking, combined regression and ranking, etc. Any suitable method (e.g., pointwise, pairwise, or listwise) may be used to determine the best ranking. Any suitable metric (including average precision, average reciprocal rank, etc.) may be used to evaluate the ranking.
[0039] The domain simulator 120 includes a simulation environment for simulating a real-world environment. The simulation environment can include a sandbox environment, a testbed, or any other environment that operates in a manner similar to a real-world environment, but with limited connectivity to the external environment so that actions performed in the simulation environment are not propagated to the real system. This approach provides the ability to conduct comprehensive training and testing without the potential adverse consequences of performing inappropriate actions. Typically, the domain simulator will include or have the ability to generate a domain model 120 that represents the domain environment.
[0040] The performance evaluation engine 125 may include components for determining the performance of RL software agents in the RL system. For example, in some aspects, the performance evaluation engine 125 may include components that represent lower and upper bounds on performance, such as a random decision agent and / or a planning decision agent. Random agents are the least efficient because the actions chosen are random, while planning agents are the most efficient because the optimal action is known. The performance of the RL system can be benchmarked against these and other decision agents.
[0041] The query generation engine 130 may generate queries based on a state that is used to identify corresponding actions and policies in response to the query.
[0042] The query matching engine 135 can match the generated query with a policy, which can take the form of a (state, action) pair. In other aspects, the query matching engine 135 can be used to match content from the knowledge base with the generated query, which is provided to the NPL engine 110 and the RL model 115 for generating a policy. These and other features are described throughout the specification and drawings.
[0043] The client system 20 and the server system 10 can be connected to each other through any suitable computing device (such as Figure 1BThe computing device 212 shown in FIG. 1 is implemented as a computing device 212 for computing environment 100. This example is not intended to imply any limitation on the functionality or scope of use of the embodiments of the present invention described herein. In any case, the computing device 212 is capable of implementing and / or performing any of the functions set forth herein.
[0044] Among the computing devices, there are computer systems that can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with the computer system include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, etc.
[0045] Computer system 212 may be described in the general context of computer system-executable instructions, such as program modules (e.g., RL system 15 and its corresponding modules), executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types.
[0046] Computer system 212 is shown in the form of a general-purpose computing device. Components of computer system 212 may include, but are not limited to, one or more processors or processing units 155, system memory 136, and a bus 218 that couples various system components including system memory 136 to processor 155.
[0047] Bus 218 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0048] Computer system 212 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system 212, and it includes both volatile and nonvolatile media, removable and non-removable media.
[0049] The system memory 136 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 230 and / or cache memory 232. The computer system 212 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 234 may be provided for reading from and writing to a non-removable non-volatile magnetic medium (not shown and commonly referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), as well as an optical disk drive for reading from and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media). In such cases, each may be connected to the bus 218 via one or more data media interfaces. As will be further depicted and described below, the memory 136 may include at least one program product having a set (e.g., at least one) program modules configured to perform the functions of an embodiment of the present invention.
[0050] By way of example and not limitation, a program / utility 240 having a set (at least one) of program modules 242 (e.g., RL system 15 and corresponding modules, etc.), as well as an operating system, one or more application programs, other program modules, and program data can be stored in memory 136. Each or some combination of the operating system, one or more application programs, other program modules, and program data can include an implementation of a network environment. Program modules 242 generally perform the functions and / or methods of embodiments of the present invention as described herein.
[0051] Computer system 212 can also communicate with one or more external devices 214 (such as a keyboard, pointing device, display 224, etc.); one or more devices that enable a user to interact with computer system 212; and / or any device that enables computer system 212 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication can occur via input / output (I / O) interface 222. Furthermore, computer system 212 can communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via network adapter 225. As depicted, network adapter 225 communicates with other components of computer system 212 via bus 218. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with computer system 212. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archiving storage systems.
[0052] Figure 2 is a flow diagram illustrating example computational commands and operations of an RL software agent for deployment in an RL system according to an embodiment of the present invention.
[0053] The following example illustrates how an RL software agent (RL decision agent) and an RL model can support a task such as opening a text document. At operation 310, a user enters a command to open a document using gedit. At operation 320, the computing environment determines that the command gedit is not found. In response, at operation 325, a query is generated, for example, by the query generation engine 130 of the RL software system. The query is used to identify relevant content from a knowledge base, for example, using the query matching engine 135, where the content may be in the form of policies or in a semi-structured form representing information stored in the knowledge base. In some cases, no exact match may be found for the query. In such cases, results may be returned based on the similarity between the query and the content in the knowledge base.
[0054] In this example, in response to a query match, matching or similar content may be processed by the NLP engine 110 and the RL model 115. However, in other aspects, the knowledge base information may be stored and / or processed at predetermined time intervals to generate ranked strategies stored in the database 30 (e.g., for this embodiment, see Figure 6A-6B ). Either method is applicable to the embodiments provided herein.
[0055] Once the query is matched to content in the knowledge base (e.g., a question-answer pair), the NLP engine 110 and the RL model 115 are configured to generate a policy based on the content and rank the (state, action) pairs of the policies present in the forum (e.g., based on votes, author name, author status, replies, etc.). At operation 335, the RL system identifies the highest-ranked (question, answer) pairs, where the question corresponds to the state and the answer corresponds to the action that can be performed by the RL software agent.
[0056] At operation 340, an action corresponding to the state of the matched policy is provided back to the RL software agent. At operation 345, the computing environment executes the action, in this case, apt-get install gedit. At operation 350, another error is generated, which results in another query being generated. In this case, the system determines that it does not have permission to install gedit. Operations 325-340 are repeated to provide the corresponding action to the RL software agent.
[0057] The computing environment assumes root privileges and gedit is installed at operation 360. The original command is executed, thereby opening the test file at operation 370. At operation 380, support for this particular task by the RL software agent and the RL model ends.
[0058] Figures 3A to 3B is a diagram illustrating an example environment for an RL software agent deployed in a specific domain according to an embodiment of the present invention. A knowledge database (such as an online forum) contains information that can be used to improve the performance of the RL software agent. For example, an online forum may contain a question and answer support forum whose content can be ranked by various parameters within the domain (such as relevance, date, votes / popularity, author status, etc.). This information can be mined by the RL model to identify the best answer to a question (or action for a state).
[0059] Figure 3A Corresponding to the computing environment. Figure 3B Corresponding to the travel environment. In each instance, the user performs an action that generates a specific state. For example, the user may execute a command that requires access to a specific software platform, or may submit a query to an online website that is supported by a decision agent.
[0060] According to the present technique, a query is derived from the user's state and used to identify the policy corresponding to that state. A ranked policy (state, action) pair is selected, derived based on data from a knowledge base that has been processed by the RL model. Thus, the (state, action) pair generated by the RL system 15 determines the appropriate action, leading to the next state. Thus, the question-answer correspondence represents an implicit state transition that implies the correct action / policy for a given state.
[0061] A knowledge base or forum contains information (e.g., structured, unstructured, or semi-structured) from subject matter experts who have solved the same or similar problems. The forum provides pseudo-traces that can be used to bootstrap the learning process of an RL agent. Bootstrapping is performed using the techniques presented in this article, which can be used to train cognitive systems for decision support in various domains, such as software technical support, travel website support, and so on.
[0062] Online forums can be used to assist with programming (e.g., Java, Python, C++, etc.) and operating systems (Windows, Linux, Mac), etc. This method can be used to capture and use grouped data or individual data from the forums.
[0063] These techniques can also be used to adapt actions to specific user preferences.In some cases, a combination of features can be used to adapt actions specific to a given user's preferences.
[0064] Figure 4 A test or sandbox environment for embodiments of the present technology is provided. This architecture represents a general implementation of agent types and environment types. In some aspects, plug-and-play agents can be employed, allowing different types of technologies (e.g., programming, operating systems, travel, scientific literature, etc.) to be evaluated for benchmarking.
[0065] In general, agents 620 may include one or more types of agents. Stochastic agent 605 may randomly perform actions until its goal is achieved. Because actions are performed randomly, this agent can be used to establish a lower bound for evaluating the performance of the learning agent. As the number of next states increases due to the complexity of the system, the performance of the stochastic engine degrades.
[0066] Planning agent 615, a state-based agent, can be used to generate an optimal method or plan for reaching a goal (e.g., through a series of state changes) based on its current state. This domain can be manually coded or learned using computational techniques such as execution tracing. This type of agent executes a plan to achieve a goal and represents optimal efficiency.
[0067] An RL agent 105 learning through reinforcement learning can learn a policy of (state, action) pairs ranked by the RL system, which maps (state, action) pairs to values representing the usefulness of performing the action in the current state. RL agents can include data-driven agents, Q-learning agents, and the like. The RL model can generate a policy that determines the next state of the RL agent 105 based on the current state of the RL agent 105.
[0068] The agent 620 may be provided as part of a package along with the domain simulator 120, which may include a test environment 625 and a simulator 630. The simulator 630 may be deployed in the test environment 625 (e.g., a sandbox) to allow testing without being deployed in a real system. The sandbox environment allows for both learning and benchmarking / performance.
[0069] The simulator can be used to simulate a domain environment using the underlying domain model 120. In various aspects, the system can automatically learn the domain model. For example, the system can learn the domain model of the computing environment based on help pages. Alternatively, the environment can be modeled based on questions from forums, for example, to identify actions that users can perform. The simulator 630 can perform actions 650 and store the results, for example, in the database 30.
[0070] During training, a query can be formulated by the RL software agent 105. The RL software system identifies a policy generated by the RL model, where the state corresponds to or is similar to the state of the RL software agent. Once a match between the current state and the policy state is performed, a corrective action can be selected and executed in the sandbox. If the action solves the query, a reward is received (e.g., successfully solving a user computational problem such as opening a file). On the other hand, if the action fails, the RL software agent will learn that the selected action did not solve the problem and will not select the action for the corresponding state in the future.
[0071] Thus, the simulator 630 receives the policy information 34 generated and ranked by the RL model 115, which has processed the knowledge base data to generate a series of (state, action) pairs. An associated Q-value can be assigned to each policy pair, and the RL software agent, when deployed, will select the highest Q-value to efficiently reach its goal.
[0072] Once trained and tested in the sandbox, the RL software agent can be deployed with the RL model containing ranked knowledge-based data represented as (state, action) pairs.
[0073] Figure 5 The input and output of the sandbox environment are shown. This method can be used in any system where a question answering (Q / A) dataset, a simulator of the environment, and a domain model are available. This embodiment includes any suitable type of automatic model learning for decision support (and execution). This allows the decision support agent to be booted up and trained in the field, as manually coding the decision agent is impractical for large systems with high complexity. The ability to automatically learn the domain model alleviates this complexity.
[0074] A QA dataset typically includes a collection of entries, each of which has a question and zero or more answers. Answers can be structured, semi-structured, or unstructured. Each answer can contain a set of actions and features / parameters related to the question extracted by the RL system 15. Features can include author name, author status, votes in favor, votes against, length of response, comments on the response by other users indicating that the solution worked, etc., or any combination thereof. Typically, features contain information that can be used to evaluate the quality (e.g., efficiency, effectiveness, etc.) of the answer to the question and can serve as a reward signal to the RL system. For example, in some cases, a "yes vote" indicates confidence and recognition of the answer by the community and can represent a reward signal to power the RL algorithm's decision-making process.
[0075] An RL model can be trained based on these features to efficiently determine a sequence of actions that leads to a desired goal or solution. By training the RL model on data that has been at least lightly curated by multiple users, the trained and tested RL model can enable an RL software agent to reach a desired goal more quickly and efficiently than an RL software agent without an RL model. RL systems can be trained based on information in a knowledge base, rather than selecting random or seemingly random actions and determining whether a reward signal is received.
[0076] Figure 6A is a flow chart illustrating ranking knowledge base data by the RL model 115 according to an embodiment of the present invention. Semi-structured data 910 is provided to the NLP 920 to extract forum information, for example, in the form of questions and corresponding answers. The semi-structured data may include questions (Q1, Q2, Q3, ...), where each question has one or more answers (e.g., A1, A2, A3, ...). For each answer, corresponding feature information 930 is extracted. In some aspects, a one-to-one mapping can be performed between the question and each answer, along with the corresponding features for that answer. Any other suitable form of information can be used. This information is provided to the RL model 115. The RL model ranks each answer based on the provided features and generates a policy 950, which may take the form of a ranked list or any other suitable form. When a query is generated by the RL software agent, the state is matched against the policy, and the corresponding action is provided to the RL software agent for execution.
[0077] Among other things, RL techniques can be combined with heuristic techniques to rank (state, action) pairs for a policy.
[0078] Figure 6B is a diagram illustrating utilizing a source code developed in a sandbox environment according to an embodiment of the present invention. Figure 6A The upper portion of the diagram corresponds to training in a sandbox environment. An RL agent 1010 is operating in the sandbox environment, interacting with an environment shell 1020. When a query is generated, a policy is matched based on the state of the RL agent. An action corresponding to the policy is provided to the RL software agent, which executes the action. By executing the action in the environment shell, the state of the RL software agent changes, and feedback regarding the action is provided to the RL software agent. For example, if the action results in reaching the goal, the action can be validated, or a reward can be provided to reinforce the behavior.
[0079] In the real environment, in response to a query generation, the RL system can match the state of the RL software agent with the state in the policy 950 generated by the RL model. The best action (highest ranked) is returned for execution by the RL software agent 1030.
[0080] In this example, the NLP engine and RL model process the knowledge base data before query generation and store the ranked policies in the database.
[0081] Figure 7A-7B Example results for benchmarking and performance evaluation in a sandbox environment are shown. The sandbox provides an environment for RL model training, separate from production-based or real-world environments. Typically, RL model training is performed in the sandbox, while testing is performed in the real world. The performance and learning rate of RL software agents can be benchmarked using a managed domain that includes a random decision agent and an optimal planning agent.
[0082] Figure 7A A graph of the number of epochs versus the length of the epoch is shown. Figure 7A , the RL software agent (without a corresponding RL model) initially has a sharp drop in epoch length, which translates to an increase in learning. The trajectory then levels off and does not appear to converge to an optimal trajectory (e.g., such as from a planning agent). In contrast, the RL software agent (with a corresponding RL model) shows a similar sharp drop in epoch length, and the trajectory continues to decrease until it approaches the optimal trajectory.
[0083] Figure 7B Various agents are shown: a random agent (e.g., a strawman), an optimal agent (e.g., a planning agent), and an RL software agent. As expected with high fan-out, the random agent based on pure self-exploration in the real-world domain does not converge quickly enough (or at all). In contrast, the RL software agent augmented with an RL model trained on a knowledge base has similar performance to the optimal agent. Thus, the RL software agent augmented with external data providing guidance converges or nearly converges to the optimal trajectory.
[0084] Figure 8 is an operational flow diagram illustrating the high-level operation of the techniques presented herein. At operation 810, a computer receives access to a knowledge base related to a topic supported by an RL software agent. At operation 820, the computer uses information from the knowledge base to train an RL model that supports the RL software agent. At operation 830, the computer tests the trained RL model in a test environment that has limited connectivity to an external environment. At operation 840, the RL software agent is deployed within the environment along with the tested and trained RL model to autonomously perform actions to process requests.
[0085] Thus, in various aspects, an RL software agent generates a query based on its current state. The RL system matches this query with a (state, action) pair of policies generated by the RL model based on information obtained from a knowledge base. For example, the RL model, in conjunction with other techniques (e.g., NPL), can process semi-structured information to identify the best action based on multiple answers. The RL software agent executes this action to reach the next state. This process can continue until the goal is reached.
[0086] The present technique improves the operation of RL software agents because the training time of RL software agents can be reduced due to the input from the RL model.Traditionally, training RL software agents involves an iterative, computationally intensive process that may not converge to the goal in complex environments.
[0087] In addition, RL software agents are configured to operate in dynamic environments where executing actions results in state changes. This technology provides the ability to train RL software agents in dynamic, complex test environments, and the ability to deploy RL software agents and trained and tested RL models in real systems.
[0088] Features of embodiments of the present invention include using a knowledge database to enhance and improve RL software agents. RL software agents can be developed in an automated manner based on techniques using domain learning. The performance of decision-support agents using reinforcement learning can be compared to the performance of the best planners used for benchmarking.
[0089] The benefits of these technologies provide instant, scalable and personalized on-site support. Additionally, by using adjudication agents to handle technical support issues, technical support experts are available to resolve more advanced support issues.
[0090] These techniques can be applied to a wide variety of environments. Any environment where simulators and training data are available is suitable for use with embodiments of the present invention. Decision support agents can be used to assist with decisions when implementing and configuring software on an operating system, assist with implementing and configuring networks using network simulators, assist with programming troubleshooting techniques based on information in online programming forums, assist with health-related decisions based on scientific and medical literature, and the like.
[0091] This technology provides a novel, labor-saving approach to building decision support tools that continuously learn from feedback and past data. This approach can be used to scale decision support technology in both enterprise and public environments in an efficient and timely manner.
[0092] In addition, our technique captures the implicit preferences of the user community while enabling continuous lifelong learning for RL software agents and RL models.
[0093] It should be understood that the embodiments described above and shown in the accompanying figures represent only some of the many ways to implement embodiments of the RL software system.
[0094] The environment of embodiments of the present invention may include any number of computers or other processing systems (e.g., client or end-user systems, server systems, etc.) and databases or other repositories arranged in any desired manner, wherein embodiments of the present invention may be applied to any desired type of computing environment (e.g., cloud computing, client-server, network computing, mainframe, stand-alone systems, etc.). The computers or other processing systems employed by the present invention may be implemented by any number of any personal or other type of computers or processing systems (e.g., desktop computers, laptop computers, PDAs, mobile devices, etc.) and may include any commercial operating system and any combination of commercial and custom software (e.g., browser software, communication software, server software, RL system 15, etc.). These systems may include any type of monitor and input device (e.g., keyboard, mouse, voice recognition, etc.) to enter and / or view information.
[0095] It should be understood that the software of embodiments of the present invention (e.g., RL system 15, including RL software agent 105, NLP engine 110, RL model 115, domain simulator 120, performance evaluation engine 125, query generation engine 130, query matching engine 135, etc.) can be implemented in any desired computer language and can be developed by a person of ordinary skill in the computer arts based on the functional descriptions contained in the specification and the flowcharts shown in the accompanying drawings. In addition, any reference herein to software that performs various functions generally refers to a computer system or processor that performs these functions under the control of the software. The computer system of embodiments of the present invention may alternatively be implemented by any type of hardware and / or other processing circuitry.
[0096] The various functions of a computer or other processing system can be distributed in any manner among any number of software and / or hardware modules or units, processing or computer systems and / or circuits, wherein the computers or processing systems can be arranged locally or remotely from each other and can communicate via any suitable communication medium (e.g., LAN, WAN, intranet, Internet, hard wire, modem connection, wireless, etc.). For example, the functions of embodiments of the present invention can be distributed in any manner among various end-user / client and server systems, and / or any other intermediate processing devices. The software and / or algorithms described above and shown in the flowcharts can be modified in any manner to achieve the functions described herein. In addition, the functions in the flowcharts or description can be performed in any order to achieve the desired operations.
[0097] The software of embodiments of the present invention (e.g., RL system 15, including RL software agent 105, NLP engine 110, RL model 115, domain simulator 120, performance evaluation engine 125, query generation engine 130, query matching engine 135, etc.) may be made available on a non-transitory computer-usable medium (e.g., magnetic or optical medium, magneto-optical medium, floppy disk, CD-ROM, DVD, memory device, etc.) on a fixed or portable program product apparatus or device for use with a stand-alone system or systems connected via a network or other communications medium.
[0098] The communication network can be implemented by any number of any type of communication networks (e.g., LAN, WAN, Internet, intranet, VPN, etc.). The computer or other processing system of an embodiment of the present invention can include any conventional or other communication device to communicate on the network via any conventional or other protocol. The computer or other processing system can use any type of connection (e.g., wired, wireless, etc.) for accessing the network. The local communication medium can be implemented by any suitable communication medium (e.g., local area network (LAN), hard wire, wireless link, intranet, etc.).
[0099] The system may employ any number of any conventional or other databases, data repositories, or storage structures (e.g., files, databases, data structures, data or other repositories, etc.) to store information (e.g., RL system 15, including RL software agent 105, NLP engine 110, RL model 115, domain simulator 120, performance evaluation engine 125, query generation engine 130, query matching engine 135, etc.). The database system may be implemented by any number of any conventional or other databases, data repositories, or storage structures (e.g., files, databases, data structures, data or other repositories, etc.) to store information (e.g., knowledge base data 31, processed Q / A data 32, policy data 34, etc.). The database system may be included within or coupled to a server and / or client system. The database system and / or storage structure may be remote from or local to a computer or other processing system and may store any desired data (e.g., knowledge base data 31, processed Q / A data 32, policy data 34, etc.).
[0100] Embodiments of the present invention may employ any number of any type of user interfaces (e.g., graphical user interfaces (GUIs), command lines, prompts, etc.) to obtain or provide information (e.g., knowledge base data 31, processed Q / A data 32, policy data 34, etc.), wherein the interface may include any information arranged in any manner. The interface may include any number of any type of input or actuator mechanisms (e.g., buttons, icons, fields, boxes, links, etc.) arranged in any location to input / display information and initiate desired actions via any suitable input device (e.g., mouse, keyboard, etc.). Interface screens may include any suitable actuators (e.g., links, tabs, etc.) to navigate between screens in any manner.
[0101] The output of the RL 15 may include any information arranged in any manner and may be configurable based on rules or other criteria to provide the user with desired information (e.g., one or more actions, etc.).
[0102] Embodiments of the present invention are not limited to the specific tasks or algorithms described above, but can be used in any application that uses RL software agents to support complex tasks. Furthermore, the approach is generally applicable to provide support in any context and is not limited to any specific application domain, such as manufacturing, health, etc.
[0103] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms as well. It should also be understood that when the terms "comprises", "comprising", "includes", "including", "has", "have", "having", "with" and the like are used in this specification, these terms specify the presence of the features, wholes, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, parts and / or combinations thereof.
[0104] All means or steps plus corresponding structures, materials, actions, and equivalents of functional elements in the following claims are intended to include any structure, material, or action for performing a function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiments are chosen and described in order to best explain the principles of the invention and practical application, and to enable others of ordinary skill in the art to understand the various embodiments of the invention with various modifications suitable for the intended specific use.
[0105] The description of various embodiments of the present invention has been presented for the purpose of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications or technical improvements over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
[0106] The present invention may be a system, method and / or computer program product of any possible degree of technical detail integration. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to execute various aspects of the present invention.
[0107] Computer-readable storage medium can be a tangible device that can retain and store the instruction for use by instruction execution device.Computer-readable storage medium can be, for example but not limited to, electronic storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or any suitable combination of the foregoing.The non-exhaustive list of the more specific example of computer-readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, the mechanical encoding device (such as the convex structure in punch card or groove) of instruction recorded thereon and any suitable combination of the foregoing.Computer-readable storage medium as used herein should not be interpreted as transient signal itself, such as radio wave or other free propagation electromagnetic wave, electromagnetic wave (such as, by optical pulse of fiber optic cable) propagated by waveguide or other transmission medium or the electric signal transmitted by wire.
[0108] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0109] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, configuration data of integrated circuits or source code or object code written in any combination of one or more programming languages, these programming languages include object-oriented programming languages (such as Smalltalk, C++ etc.) and procedural programming languages (such as " C " programming languages or similar programming languages). The computer-readable program instructions can be performed completely on the user's computer, partly on the user's computer, performed as an independent software package, partly on the user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including local area network (LAN) or wide area network (WAN)), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In certain embodiments, the electronic circuit comprising, for example, programmable logic circuit, field programmable gate array (FPGA) or programmable logic array (PLA) can personalize the electronic circuit and perform the computer-readable program instructions by adopting the state information of the computer-readable program instructions, so as to perform various aspects of the present invention.
[0110] Various aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0111] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions direct the computer, programmable data processing device, and / or other equipment to operate in a specific manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0112] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0113] The flow charts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of the possible implementations of the systems, methods and computer program products according to various embodiments of the present invention. To this end, each box in the flow chart or block diagram can represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing a specified logical function. In some alternative implementations, the functions annotated in the box may not occur in the order annotated in the figure. For example, the two boxes shown in succession can actually be completed as a step, simultaneously, substantially simultaneously, in a partially or completely overlapping manner in time, or the boxes can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a system based on dedicated hardware, which performs a specified function or action based on dedicated hardware, or performs a combination of dedicated hardware and computer instructions.
Claims
1. A method for configuring a reinforcement learning software agent using data from a knowledge base, the method comprising: receiving, by a computer, access to the knowledge base related to a subject matter supported by the software agent, wherein the knowledge base includes information about the subject matter from a subject matter expert; training, by the computer, a reinforcement learning model supporting the reinforcement learning software agent using information from the knowledge base; testing, by the computer, the trained reinforcement learning model in a test environment, the test environment including a domain simulator to simulate the domain and not propagating actions performed in the test environment to a production-based environment; validating the actions learned by the reinforcement learning software agent in a test environment using a planning-decision agent and a random software agent, wherein the planning-decision agent produces the best known actions and represents an upper bound on performance, and the random software agent produces random actions and represents a lower bound on performance; and Deploying the reinforcement learning software agent along with the tested and trained reinforcement learning model within the production-based environment to autonomously perform actions to process requests, wherein autonomously performing actions to process requests comprises: generating a content query to obtain content related to the request to execute the command from a knowledge base, wherein the content from the knowledge base includes questions from a forum related to the command and having one or more corresponding answers; generating a policy for processing commands based on content from a knowledge base using a reinforcement learning model, the policy comprising states and corresponding actions, wherein the states correspond to questions about the content and the actions correspond to corresponding answers to the content, the answers indicating actions that can be performed by the reinforcement learning agent to execute the command, wherein the policy ranks the states and actions based on attributes from the forum; generating a policy query based on a state associated with the user and matching the policy query to a state in a policy having a corresponding action; and The reinforcement learning software agent executes the corresponding actions in the policy to carry out the command.
2. The method according to claim 1, wherein The information from the knowledge base is in a semi-structured format.
3. The method according to claim 2, wherein: The knowledge base includes information generated by a plurality of users, and wherein at least a portion of the information is managed by the plurality of users.
4. The method according to claim 1, wherein The knowledge base includes features used by the reinforcement learning model to rank information provided in the knowledge base.
5. The method according to claim 4, wherein The characteristics include one or more of upvotes, downvotes, author name, author title, or author status.
6. The method according to claim 1, wherein The reinforcement learning model ranks the information in the knowledge base based on user preferences.
7. The method according to claim 1, further comprising: A reward is received based on performing the corresponding action by the reinforcement learning software agent when the corresponding action resolves the policy query.
8. The method according to claim 1, further comprising: The policy generated by the reinforcement learning model is updated at a fixed frequency or a dynamic frequency to include questions and answers added to the knowledge base.
9. The method according to claim 1, wherein: A reinforcement learning software agent deployed with a trained reinforcement learning model reaches a desired goal more effectively than a reinforcement learning software agent without a trained reinforcement learning model.
10. A computer system for configuring a reinforcement learning software agent using data from a knowledge base, the computer system comprising: one or more computer processors; one or more computer-readable storage media; Program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising instructions for: receiving access to a knowledge base related to a subject matter supported by the software agent, wherein the knowledge base includes information about the subject matter from a subject matter expert; using information from the knowledge base to train a reinforcement learning model that supports a reinforcement learning software agent; Testing the trained reinforcement learning model in a test environment that includes a domain simulator to simulate the domain and does not propagate actions performed in the test environment to a production-based environment; validating the actions learned by the reinforcement learning software agent in a test environment using a planning-decision agent and a random software agent, wherein the planning-decision agent produces the best known actions and represents an upper bound on performance, and the random software agent produces random actions and represents a lower bound on performance; and Deploying the reinforcement learning software agent along with the tested and trained reinforcement learning model within the production-based environment to autonomously perform actions to process requests, wherein autonomously performing actions to process requests comprises: generating a content query to obtain content related to the request to execute the command from a knowledge base, wherein the content from the knowledge base includes questions from a forum related to the command and having one or more corresponding answers; generating a policy for processing commands based on content from a knowledge base using a reinforcement learning model, the policy comprising states and corresponding actions, wherein the states correspond to questions about the content and the actions correspond to corresponding answers to the content, the answers indicating actions that can be performed by the reinforcement learning agent to execute the command, wherein the policy ranks the states and actions based on attributes from the forum; generating a policy query based on a state associated with the user and matching the policy query to a state in a policy having a corresponding action; and The reinforcement learning software agent executes the corresponding actions in the policy to carry out the command.
11. The computer system according to claim 10, wherein: The information from the knowledge base is in a semi-structured format.
12. The computer system according to claim 11, wherein: The knowledge base includes information generated by a plurality of users, and wherein at least a portion of the information is managed by the plurality of users.
13. The computer system according to claim 10, wherein: The knowledge base includes features used by the reinforcement learning model to rank information provided in the knowledge base, wherein the features include one or more of upvotes, downvotes, author name, author title, or author status.
14. The computer system according to claim 10, wherein: The program instructions also include instructions for the following operations: A reward is received based on performing the corresponding action by the reinforcement learning software agent when the corresponding action resolves the policy query.
15. The computer system according to claim 10, wherein: A reinforcement learning software agent deployed with a trained reinforcement learning model reaches a desired goal more effectively than a reinforcement learning software agent without a trained reinforcement learning model.
16. A computer program product for configuring a reinforcement learning software agent using data from a knowledge base, the computer program product comprising program instructions executable by a computer to cause the computer to: receiving, by a computer, access to a knowledge base related to a subject matter supported by the software agent, wherein the knowledge base includes information about the subject matter from a subject matter expert; using information from the knowledge base to train, by the computer, a reinforcement learning model supporting a reinforcement learning software agent; testing, by the computer, the trained reinforcement learning model in a test environment, the test environment including a domain simulator to simulate the domain and not propagating actions performed in the test environment to a production-based environment; validating the actions learned by the reinforcement learning software agent in a test environment using a planning-decision agent and a random software agent, wherein the planning-decision agent produces the best known actions and represents an upper bound on performance, and the random software agent produces random actions and represents a lower bound on performance; and Deploying the reinforcement learning software agent along with the tested and trained reinforcement learning model within the production-based environment to autonomously perform actions to process requests, wherein autonomously performing actions to process requests comprises: generating a content query to obtain content related to the request to execute the command from a knowledge base, wherein the content from the knowledge base includes questions from a forum related to the command and having one or more corresponding answers; generating a policy for processing commands based on content from a knowledge base using a reinforcement learning model, the policy comprising states and corresponding actions, wherein the states correspond to questions about the content and the actions correspond to corresponding answers to the content, the answers indicating actions that can be performed by the reinforcement learning agent to execute the command, wherein the policy ranks the states and actions based on attributes from the forum; generating a policy query based on a state associated with the user and matching the policy query to a state in a policy having a corresponding action; and The reinforcement learning software agent executes the corresponding actions in the policy to carry out the command.
17. The computer program product of claim 16, wherein: The information from the knowledge base is in a semi-structured format.
18. The computer program product of claim 17, wherein: The knowledge base includes information generated by a plurality of users, and wherein at least a portion of the information is managed by the plurality of users.
19. The computer program product of claim 16, wherein: The program instructions also include instructions for the following operations: A reward is received based on performing the corresponding action by the reinforcement learning software agent when the corresponding action resolves the policy query.
20. The computer program product of claim 16, wherein: A reinforcement learning software agent deployed with a trained reinforcement learning model reaches a desired goal more effectively than a reinforcement learning software agent without a trained reinforcement learning model.