Methods and systems for adapting a learning model
By adapting the behavior model of reinforcement learning-based control systems through calculating an adaptation region and updating actions in response to environmental novelty, the method enhances the system's performance in changing environments.
Patent Information
- Application Number
- PCT/AU2024/051251
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Existing reinforcement learning-based control systems struggle to adapt effectively to environmental novelty, leading to performance drops when faced with unexpected changes or variations in the operating environment.
The method involves adapting the behavior model of the control system by calculating an adaptation region in the embedded state space, determining an adaptation action using an adaptation function, and updating the behavior model to replace compromised actions with the adaptation action.
This approach enables the control system to improve its operational performance in novel environments by effectively updating its behavior model, allowing it to maintain optimal action responses and task performance.
Smart Images

Figure AU2024051251_30052025_PF_FP_ABST
Abstract
Description
"Methods and systems for adapting a learning model"Technical Field
[0001] The present invention relates to adapting a learning model of an agent, such as for example a reinforcement learning neural network, to improve the operation of a control system, such as for example a self-navigating vehicle or a game controller, in response to change or mismatch in the operating environment.Background
[0002] Recent developments in artificial intelligence (Al) have lead to a surge in the use of Al techniques for the purpose of achieving autonomous operation of real-world systems or devices. The nature of the systems being operated may vary depending on the application and the operational environment in which they are deployed. For example, electric, mechanical, or electro-mechanical systems may be operated by Al techniques in self-navigating vehicles, robots, and manufacturing equipment. There is also motivation for achieving Al-based operation of digital control systems that are implemented as hardware or software modules within a computing device (e.g., gaming or simulation engines).
[0003] Autonomous Al-based control is typically achieved via the use of an agent that gathers perception input from the operating environment (e.g., the region around a vehicle) and generates action responses by executing the inputs on a behaviour model. Machine learning techniques provide an Al agent with the ability to represent an optimal or desired behaviour of a control system in an operating environment from unstructured input data, and without manual engineering of a corresponding state space of the control system. For example, deep reinforcement learning (DRL) is able to process complex and noisy inputs as sampled from the environment, and can subsequently produce action responses leading to operation of a control system that matches or exceeds human performance in several domains (see [1, 2, 3]).
[0004] Techniques for autonomous system operation that use networks based on DRL typically excel in operating environments that are stable over time (i.e., where the state space, action space, and transition probability distribution can be accurately assumed to remain unchanged). However, the performance of a DRL based agent has been shown to drop significantly in environments that experience unexpected changes or variations relative to the baseline conditions on which the model of the agent was trained (see [4]). The occurrence of such changes or variations is referred to as “novelty” in the environment.
[0005] Novelty can occur as a result of explicit changes in the operation or state of the control system (e.g., failure of the breaking sub-system of a vehicle controller) or the environment (e.g., a change in road conditions of a vehicle being controlled). Alternatively, novelty can occur implicitly for example in scenarios where the behaviour model of the agent is trained in different conditions to those of the environment in which the control system is deployed. In both cases, the novelty refers to a mismatch or change in the current (i.e., “operating”) environment relative to the baseline environment, where the novelty typically affects the optimal action response desired by the control system to perform a relevant task.
[0006] Accordingly, there is a need to develop approaches that improve the functionality of an intelligent agent in response to environmental novelty, particularly where novelty may occur suddenly and with long-term effects. One solution to improving agent behaviour in response to environmental changes is to collect additional data that is better representative of the changed environment, and to then retrain the behaviour model using the additional data. However, collecting additional data that accurately reflects a novelty experienced at operation time may not always be possible or practically feasible. For example, perception data may not be available to capture the operation of an autonomous vehicle or robot in remote and previously unseen areas (e.g. on Mars).Summary
[0007] There is provided a method comprising: during performance of one or more of episodes of operating a control system in an operating environment, each of the episodes being associated with the control system performing a task in the operating environment based on a behaviour model trained in a baseline environment, and comprising a sequence of states and actions of the control system: in response to a novelty in the operating environment, adapting the behaviour model to the novelty in the operating environment; and using the adapted behaviour model to operate the control system in the operating environment in relation to the task, wherein adapting the behaviour model comprises: (i) calculating an adaptation region in an embedded state space of the operating environment, the adaptation region including one or more compromised model states of the control system, wherein each of the one or more compromised model states is a state where an action of the control system is compromised in the operating environment in relation to the task, and wherein the action is determined by executing the behaviour model; (ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and (iii) updating the behaviour model by replacing the action of each of the one or more compromised model states with the adaptation action, wherein the one or more compromised model states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
[0008] In some embodiments, the behaviour model is a deep reinforcement learning model.
[0009] In some embodiments, updating the behaviour model involves retaining one or more actions of the control system associated with model states that are not within the adaptation region.
[0010] In some embodiments, the one or more compromised model states of the adaptation region include at least a current model state representing, in the embedded state space, a current state of the control system in the operating environment.
[0011] In some embodiments, the embedded state space is a dimensionally-reduced embedded state space.
[0012] In some embodiments, the one or more compromised model states of the adaptation region include at least one other model state additional to the current model state.
[0013] In some embodiments, the adaptation region is based on a cluster formed in the embedded state space at a location of the current model state.
[0014] In some embodiments, the adaptation region is calculated by partitioning the embedded state space based on a spatial relationship between the current model state and other states in the embedded state space.
[0015] In some embodiments, the adaptation region is calculated by constructing a Voronoi cell encompassing the current model state in the embedded state space.
[0016] In some embodiments, the adaptation region includes the current model state and is closed, and wherein the adaptation region is split by: i) selecting the adaptation action associated with the adaptation region; ii) evaluating the adaptation action with respect to the corresponding current state to determine a performance of the adaptation action; and iii) splitting the closed adaptation region into two regions if the performance of the adaptation action is below a threshold level.
[0017] In some embodiments, steps (i) to (iii) are iteratively repeated until the adaptation region is smaller than a predefined threshold, or an overall performance of the control system reaches a predefined threshold.
[0018] In some embodiments, the performance of each adaptation action is determined using an evaluation function that assigns a score to the corresponding action of the state as determined by the behaviour model.
[0019] In some embodiments, determining the adaptation action comprises: selecting the adaptation function from a plurality of candidate functions; and applying the selected adaptation function to the current model state.
[0020] In some embodiments, each adaptation function maps a set of embedded states corresponding to the adaptation region to a set of candidate actions from an action space of the operating environment.
[0021] In some embodiments, the selected adaptation function outputs the adaptation action by evaluating the set of candidate actions in response to an input of the current model state.
[0022] In some embodiments, using the behaviour model comprises: i) determining the current state of the control system in the operating environment; ii) evaluating the behaviour model on the current state to determine an action of the control system; and iii) processing one or more parameters of the action to generate signals to control the control system in the operating environment.
[0023] There is also provided a control system comprising: one or more input components configured to generate input data associated with an operating environment of the control system; one or more operational components configured to, in response to receiving corresponding control signals, cause the control system to perform a task in the operating environment; and an intelligent controller device configured to receive the input data from the one or more input components and transmit the control signals to the one or more operational components, and comprising at least: a data storage medium configured to store a behaviour model; and one or more processors configured to operate the control system based on the behaviour model by executing any of the methods described herein.
[0024] There is also provided a computer-readable storage medium having program code that is executable by a processor device to cause a computing device to perform operations, the operations comprising: during performance of one or more of episodesof operating a control system in an operating environment, each of the episodes being associated with the control system performing a task in the operating environment based on a behaviour model trained in a baseline environment, and comprising a sequence of states and actions of the control system: in response to a novelty in the operating environment, adapting the behaviour model to the novelty in the operating environment; and using the adapted behaviour model to operate the control system in the operating environment in relation to the task, wherein adapting the behaviour model comprises: (i) calculating an adaptation region in an embedded state space of the operating environment, the adaptation region including one or more compromised model states of the control system, wherein each of the one or more compromised model states is a state where an action of the control system is compromised in the operating environment in relation to the task, and wherein the action is determined by executing the behaviour model; (ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and (iii) updating the behaviour model by replacing the action of each of the one or more compromised model states with the adaptation action, wherein the one or more compromised model states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
[0025] There is also provided a method for adapting a behaviour neural network used to operate a control system in an environment, the method comprising: executing the behaviour neural network during one or more episodes, each episode comprising a sequence of states and actions of the control system to perform a task, and in response to a change in the environment during the execution of the behaviour neural network: (i) calculating an adaptation region in an embedded state space of the environment, the adaptation region including one or more compromised network-embedded states of the control system, wherein each of the one or more compromised network-embedded states is a state where an action of the control system is compromised in the environment in relation to the task, and wherein the action is determined by executing the behaviour neural network; (ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and (iii) updating the behaviour neural network by replacing the action of each of the one or more compromisednetwork-embedded states with the adaptation action, wherein the one or more compromised network-embedded states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.Brief Description of Drawings
[0026] Some embodiments are described herein below with reference to the accompanying drawings, wherein:
[0027] Fig. la is a block diagram of a control system operated in an operating environment in accordance with some embodiments;
[0028] Fig. lb is a block diagram of a first example implementation of the control system of Fig. la as a vehicle controller;
[0029] Fig. 1c is a schematic diagram of a vehicle operated by the vehicle controller ofFig. lb;
[0030] Fig. Id is a block diagram of a second example implementation of the control system ofFig. la as a game controller;
[0031] Fig. le is a schematic diagram of a virtual entity operated by the game controller of Fig. Id;
[0032] Fig. 2a is a block diagram of a behaviour model implemented as a neural network in accordance with some embodiments;
[0033] Fig. 2b is a block diagram of an adapted behaviour model that is adapted in response to a change in the environment implemented as a neural network in accordance with some embodiments;
[0034] Fig. 2c is a schematic diagram of an exemplary embedded state space for a behaviour model of an autonomous vehicle in accordance with some embodiments;
[0035] Fig. 2d is a schematic diagram of a control system operating in an environment with a baseline behaviour model that is adapted to produce an adapted behaviour model in accordance with some embodiments;
[0036] Fig. 3 is a flow diagram of a method for operating a control system to perform a task using a behaviour model that is adapted to a novelty in the operating environment in accordance with some embodiments;
[0037] Fig. 4 is a flow diagram of a method for adapting the behaviour model during the method of operating the control system of Fig. 3;
[0038] Fig. 5 is a flow diagram of a method 500 for creating a new adaptation region as a Voronoi cell within the embedded state space during the method for adapting the behaviour model of Fig. 4;
[0039] Fig. 6 is a flow diagram of an exemplary method for determining mappings between behaviour adaptation functions and corresponding adaptation actions in accordance with some embodiments;
[0040] Fig. 7 is a flow diagram of an exemplary method of operating a control system in a post-novelty environment in accordance with some embodiments;
[0041] Fig. 8 is a schematic diagram illustrating the four domains used for an evaluation of the methods of operating a control system while adapting the behaviour model in accordance with some embodiments;
[0042] Fig. 9 is a graph illustrating the overall cumulative rewards on the CartPole domain of the evaluation;
[0043] Fig. 10 is a graph set illustrating the performances of the proposed novelty adaptation approaches using different parameter values on the CartPole domain of the evaluation;
[0044] Fig. 11 is a graph set illustrating the overall cumulative rewards on the Mountain Car domain of the evaluation;
[0045] Fig. 12 is a graph set illustrating the performances of the proposed novelty adaptation approaches using different parameter values on the Mountain Car domain of the evaluation;
[0046] Fig. 13 is a graph set illustrating the performances of the proposed novelty adaptation approaches under different gravity and force values of the evaluation;
[0047] Fig. 14 is a graph set illustrating the overall performance on the CrossRoad domain of the evaluation;
[0048] Fig. 15 is a graph set illustrating the performances of the proposed novelty adaptation approaches under different embedded spaces of the evaluation;
[0049] Fig. 16 is a graph illustrating the overall pass rate of the agents tested in the Angry Birds domain of the evaluation;
[0050] Fig. 17 is a graph set illustrating the performances of the proposed novelty adaptation approaches under different novelties of the evaluation; and
[0051] Fig. 18 is a graph illustrating the average asymptotic pass rate of the proposed novelty adaptation approach compared to other approaches of the evaluation.Description of Embodiments
[0052] Novelty adaptation is the ability of an intelligent agent to adjust its behaviour of operating a control system to perform a task in response to changes in the operatingenvironment. This is an important characteristic of intelligent agents, as it allows them to continue to function effectively in new or unexpected situations that deviate from a baseline environment in which a behaviour model of the agent was developed.
[0053] Performing novelty adaptation remains a critical challenge for many learning paradigms, including for example DRL. Specifically, there are several drawbacks associated with applying existing techniques for the purpose of allowing a DRL based intelligent agent to account for changes in an operating environment of a control system. In one approach, the robustness of the DRL model is improved by incorporating measures of non-stationarity into the underlying decision process (see [5, 6, 7]). While this may be beneficial in environments that change slowly over time, in real-world environments novelty events often occur suddenly and impose long-term effects.
[0054] Another option is to use transfer learning, which involves selecting appropriate source tasks that are related to the target task, learning how they are connected, and transferring knowledge (i.e., determinations or mappings of action responses for a current state of the control system) from the source tasks to the target task. However, this method requires knowledge of the target task in advance, which is referred to as the agent having “known unknowns.” Implementing transfer learning also typically requires the agent to learn multiple action mappings (also referred to as “functions” or “policies”) in the source task, and the action of the agent is chosen from the function that gives the highest evaluated performance value (e.g., Q-value) for a given state for the target task (see [8, 9]).
[0055] Meta reinforcement learning enables intelligent agents to learn new skills quickly by using prior learning experiences. However, typical meta-reinforcement learning techniques require the agent to be trained on multiple distinct tasks to find an appropriate initialization of weights or hyperparameters before the agent can effectively solve new tasks (see
[0010] ).
[0056] Other open-world learning techniques have been developed to tackle the challenge of operating in new and unexpected environments, such as Rapid Motor Adaptation (RMA), HYDRA (see [11, 12]), OpenMIND (see
[0013] ), and RAPid-Learn (see
[0014] ). However, these methods often require an explicit understanding of the environment’s structure, including models of the environment, the a priori identification of possible novelties, and handcrafted predicates and preconditions, which may not always be available in real-world scenarios. Additionally, many of these methods rely on a planning domain definition language (PDDL). More specifically, implementations of these previous approaches involve the use of sets or sequences of the model states, actions, and probabilities to detect novelty and determine an adaptation response. For example, symbolic planning, and / or simulations of the expected sequence of states, may be performed according to a model of the environment that is adjusted to account for the novelty. This process may be computationally intensive when the model state and / or action spaces are large. Accordingly, it is desired to develop an approach to adapting a trained behaviour model, such as, for example, a deep reinforcement learning network, in response to environmental novelty that address one or more of these problems, or other problems, or that at least provide a useful alternative.Overview
[0057] Disclosed herein are embodiments of methods and systems for adapting a behaviour model used to operate a control system in an operating environment. The adaptation is performed by an intelligent agent that operates the control system to perform a task, in response to a change in the environment relative to a baseline environment on which the model is initially trained. The agent is configured to adapt the behaviour model by 1) identifying the configurations (i.e., states) of the control system where the existing model produces an ineffective or undesirable behaviour in the changed environment (referred to as “compromised states”), and 2) adjusting the behaviour of the control system only for the identified states. That is, the actions of the control system as learned from baseline conditions of the environment (i.e., without thenovelty) are retained in all but the configurations of the control system in which there is a need for behavioural modification.
[0058] It is an advantage that the proposed techniques provide for improvements in the operational performance of an autonomous control system in various physical or digital implementations. Examples of the control system include a vehicle configured to perform an autonomous driving task on an urban road, or a game controller configured to control an entity, such as a game character, within a virtual world. The techniques enable an intelligent controller device to rapidly adapt its learning network to environmental changes when the system is carrying out one or more tasks. The intelligent controller is able to use the adapted network to operate the system with improved effectiveness. For example, the operation of a self-navigating vehicle is improved by enabling the vehicle to perform a given navigation task (such as maintaining a safe following distance) more quickly as the described adaptation of its reinforcement learning network takes place more efficiently. Further, the techniques for adaptive learning described herein are iterative enabling the improvements in operational performance of the control system to grow with continued activity of the system in the environment.
[0059] The proposed approach allows the intelligent controller to adapt a network or other model that is trained on the baseline environment (referred to herein as a “baseline model”) to accommodate a novelty in an operational environment by utilizing an embedded representation of the states of the control system. Specifically, the compromised states are localized in a region of the embedded state space (the “adaptation region”) based on a degree of semantic similarity of the control system performing the task over the states.
[0060] Various techniques are proposed to determine the adaptation region in the embedded state space, including partitioning, segmenting or otherwise identifying a subset of the states in the embedded state space based on a spatial relationship between the current model state and one or more other states. The intelligent controller updates the behaviour of the control system by determining one or more adaptation actionsassociated with the adaptation region that, when performed instead of a current action of the behaviour model for the compromised state, improves the performance of the system on the given task.
[0061] Determination of the adaptation region is specific to the type of learning model utilized by the agent. In some examples, the behaviour model is a deep reinforcement learning (DRL) model, where the state space is defined by a neural network (NN) with a plurality of layers representing a forward model of the environment dynamics (e.g., as obtained by supervised training in the pre-novelty baseline environment). It will be appreciated that the techniques described are applicable to any other type of learning model which relies on an embedded representation of model states constructed from training in the baseline environment.
[0062] The intelligent controller learns new functions to determine new (“adaptation”) actions for multiple related states of the system, as captured within a corresponding region of the embedded state space. The use of one or more adaptation actions to adapt the existing behaviour model is referred to as an adaptation principle. Each adaptation principle can be loosely explained in terms of a rule: ”if the agent is in some situation 5 and wanted to do x, it should now do y instead”. That is, each adaptation principle defines a new behaviour of the system (i.e., linking actions performed to a state of the system in the environment) that replaces the underperforming behaviours of the current state and a set of related states that are likely affected by the change in environment. An adaptation principle may be considered open if a number of possible actions may be selected as the adaptation action. The principle is considered closed once an optimal or best action is identified as the adaptation action for the corresponding region, or the intelligent controller determines that adaptation is not possible in this region.
[0063] The proposed techniques provide advantages in enabling an intelligent controller, or agent, to perform a learning process efficiently by altering the behaviour of a control system only when an environmental novelty affects the performance of a current action. This enables the agent to maximize retaining of prior learning andresults in efficient run-time performance while permitting behaviour adjustments to new and unexpected situations. The agent is able to generalize a behavioural adjustment made for one configuration of the system to semantically similar configurations without knowledge of the structure of the changed environment, and without explicitly evaluating the performance of those states (e.g., without simulating the states). Further, the proposed techniques enable the agent to perform data efficient adaptation, particularly when the underlying model or network involves a forward model of the dynamics of the environment, since the amount of data required to learn a useful adjustment is reduced.Control system
[0064] Fig. la illustrates an example of a control system 101 operated in an operating environment 150 according to various embodiments disclosed herein. Operation of the control system 101 is for the purpose of performing a task or activity that is specific to the environment 150. The form of the control system 101 may differ depending on the application to include various physical and digital systems as described herein.
[0065] Control system 101 comprises one or more input components 102a. . . 102n, one or more operational components 104a. . . 104m, and an intelligent controller 112. The intelligent controller 112 is configured to receive the input data from the one or more input components 102a. . . 102n, process the input data to generate control signals for operating the control system 101, and transmit the control signals to the one or more operational components 104a. . . 104m to cause the control system 101 to perform an action related to the task. The form and configuration of the input components 102a. . . 102n and the operational components 104a. . . 104m is dependent on the implementation of the control system 101.
[0066] In the examples of Figs, la-le, the intelligent controller 112 is implemented as a computing device, and comprises a central system bus (not shown), a memory system 108, a processing device 110, and a communications module 106. The processing device 110 may comprise one or more individual processors (e.g., CPUs) ormicroprocessors which perform the execution of sequences of machine instructions, and may have architectures consisting of a single or multiple processing cores such as, for example, a system having a 32- or 64-bit Advanced RISC Machine (ARM) architecture (e.g., ARMvx). The processing device 110 issues control signals to other device components via the system bus, and has direct access to at least some form of the memory system 108.
[0067] The memory system 108 provides internal media for the electrical storage of the machine instructions required to execute the user application. The memory system 108 may include random access memory (RAM), non-volatile memory (such as ROM or EPROM), cache memory and registers for fast access by the processor(s), and high volume storage subsystems such as hard disk drives (HDDs), or solid state drives (SSDs).
[0068] The processes executed by the intelligent controller 112 are implemented as programming instructions of one or more software modules stored on non-volatile storage of the memory system 108. In some other embodiments, the processes may be executed by one or more dedicated hardware components, such as field programmable gate arrays (FPGAs) and / or application-specific integrated circuits (ASICs). Modules and data stored by the memory system 108 include: a behaviour model 120 representing at least one policy (also referred to as a behaviour policy) of operating the control system 101 for performing the task in the environment; control data 124 representing one or more states and / or actions of the control system 101; and observation data 122 representing observations of the environment associated with the states and / or actions. Memory system 108 may also include one or more general application programs providing methods, data structures or other software services that define data or perform functions as required by the intelligent controller 112 (e.g., an operating system). The data and instructions may reside in multiple parts of the memory system 108, including registers, cache, main memory, and high volume storage.
[0069] In some implementations, the communications module 106 comprises one or more I / O device interfaces 105 (not shown). Each VO device interface 105 provides functionality enabling the user to interact with the intelligent controller 112 via one or more VO devices. The VO device interface(s) 105 also provide functionality for the intelligent controller 112 to instruct external components, such as for example displays, audio devices, and / or output peripherals. In some implementations, the communication module 106 is comprised of one or more software components providing an Application Program Interface (API) enabling the components and engines of the intelligent controller 112 to exchange data with other components or engines of a digital control system.
[0070] In some implementations, the communications module 106 includes a modem or transceiver device configured to enable the establishment of a logical connection between the intelligent controller 112 and other devices, such as one or more of the input components 102a. . . 102n and the output components 104. . . 104m, through a wireless or wired transmission media.
[0071] The skilled person in the art will appreciate that many other embodiments may exist including variations in the hardware configuration of intelligent controller 112, and the distribution of program data and instructions to execute the methods described herein.Example: self-navigating vehicle
[0072] Figs, lb and 1c illustrate an example control system 101 in the form of a vehicle controller 101’ for autonomously operating a vehicle 101 A in a navigation environment. The vehicle 101 A may be a ground vehicle and the navigation environment may be a driving environment (e.g., including one or more roads on which the ground vehicle 101 A travels and associated objects and / or conditions). Intelligent controller 112 operates the vehicle 101 A to perform a task in the navigation environment 150, such as monitoring and controlling a distance of the autonomous vehicle 101 A to other vehicles and / or objects.
[0073] In the example of Figs, lb and 1c, input sensors 102a. . . 102n include one or more vision sensors 102a and proximity sensors 102b configured to generate input data associated with the environment 150 and / or the state of the vehicle 101 A within it. In Fig. 1c, sensors 102a, 102b are mounted at fixed positions within an array on opposing lateral sides of the vehicle 101 A. Other configurations may be adopted according to the properties of the vehicle 101 A and / or sensors 102a, 102b.
[0074] Vision sensors 102a may include, for example, one or more monographic cameras, stereographic cameras, and / or laser scanners. A laser scanner includes one or more lasers devices that emit light and one or more sensing components that collect data related to reflections of the emitted light, including 3D scanners having a position sensitive detector (PSD) or other optical position sensor.
[0075] Proximity sensors 102b may include one or more devices configured to detect the presence of objects nearby the vehicle 101 A without any physical contact between the object and the sensors 120b. For example, the sensors 102b may include an emitting component configured to emit an electromagnetic field or electromagnetic radiation, such as for example an infrared beam, and determine a change in the field or return signal as caused by the presence of the object.
[0076] In the example implementations of Figs, lb and 1c, the vision and proximity sensors 102a and 102b are oriented to have a field of view of at least a portion of the driving environment 150 of the vehicle 101 A, such as to enable the generation of input data that is useful for the intelligent controller 112 to automate the operation of the vehicle 101 A for performing the task. In other implementations, additional and / or alternative configurations and / or positionings of the sensors 102a 102b may be utilized from the examples shown.
[0077] In the example of Figs, lb and 1c, the one or more operational components 104a. . . 104m include an acceleration system 104a and a steering system 104b. The acceleration system 104a is configured to receive, from the intelligent controller 112, acceleration control signals indicating a desired degree to which the vehicle 101 A is tobe accelerated or decelerated for the purpose of conducting the task. For example, the acceleration control signals may comprise a real number value quantifying an acceleration, (when positive) or deceleration (when negative), to control a distance of the autonomous vehicle 101 A to other vehicles and / or objects in the driving environment.
[0078] The steering system 104b is configured to receive, from the intelligent controller 112, steering control signals indicating a desired bearing or heading of the vehicle 101 A for the purpose of conducting the task. For example, the steering control signals may comprise an indication of an angle or offset that specifies the heading of the vehicle 101 A. The direction or heading may be specified relative to an axis of the vehicle 101 A, and / or with respect to a relative direction of motion of the vehicle (i.e., forward or reverse).
[0079] Acceleration system 104a and steering system 104b process respective acceleration control signals and steering control signals received from the intelligent controller 112 and generate action response signals. The operational systems 104a and 104b transmit the action response signals to one or more external components 140 to cause the vehicle 101 A to perform an action. For example, acceleration system 104a may process an indication of a desired deceleration provided in the acceleration control signals to generate corresponding action response signals that operate a drive train system 140 (e.g., to control power delivered to one or more wheels of the vehicle 101 A, and / or the engagement of brake calipers for the wheels). In response to receiving the action response signals, the drive train system 140 performs operations that lead to an appropriate action (e.g., a decrease in the speed) of the vehicle 101 A for performing the task. For example, the decrease in the speed of the vehicle 101 A results in adjustment to and control of the distance of the autonomous vehicle 101 A to other vehicles and / or objects in the driving environment.Example: virtual entity in a game world
[0080] Figs. Id and le illustrate an example in which the control system 101 is a digital system 101” for autonomously operating an entity 10 IB in a virtualenvironment. Digital system 101” comprises a game controller 112 that operates a virtual entity 10 IB in a game world 150 (shown as an abstraction in Figs. Id and le) for the purpose of performing an in-game objective or task. For example, the controller 112 may operate a game character 101B in a 2D or 3D maze containing respective hazards and rewards to perform a gameplay related task such as maximizing collection of the rewards in the maze while avoiding the hazards.
[0081] In the example of Figs. Id and le, input components 102a. . . 102n include a position module 102a, an object module 102b, and a world state module 102c configured to generate input data associated with the game environment 150 and / or the state of the character 101B within it. For example, the position module 102a may be configured to provide position data to the controller 112 indicating a relative position of the character 10 IB in the game world 150. Object module 102b may be configured to provide the controller 112 with object data indicating the state or behaviour of objects (rewards and / or hazards) in the game world 150. World state module 102c may be configured to provide the controller 112 with world data indicating conditions or effects of the world 150 not specified by the object or position data.
[0082] In the example of Figs. Id and le, the one or more operational components 104a. . . 104m include a character movement module 104a and a character ability module 104b. For example, the character movement module 104a may be configured to receive control signals from the controller 112 indicating a direction and quantity of movement of the character 10 IB in the game world 150. The character ability module 104b may be configured to receive control signals from the controller 112 indicating commands, abilities or skills to be performed by the character 10 IB in the game world 150.
[0083] The position, object module, and world state modules 102a, 102b, and 102c and character movement and character ability modules 104a and 104b are components within a game engine 144 that maintains the entity 101B within the world 150. In some examples, the game engine 144 is executed by a computing system 142. Computing system 142 may be the same or a component system of the digital system 101”, forexample where the game engine 144 is implemented as a software component maintained in memory system 108. In other examples, the computing system 142 is another computing system in communication with digital system 101” via the APIs of the communications module 106.
[0084] In the described example, the character movement module 104a and character ability module 104b process respective control signals received from the intelligent controller 112 and generate action response signals (e.g., in the case that an action comprises a movement and / or an ability for the character 10 IB). The action response signals are processed by the game engine 144 to cause a corresponding update to the state of the character 10 IB in the environment 150. In some examples, the game engine 144 transmits data to external components 140. External components 140 may include, for example, one or more graphical user interface (GUI) modules 146 and / or other hardware devices 147 configured to process data including signals indicating the state(s) and / or action(s) of the character 10 IB, and generate graphical output representing the same (i.e., to display the action of character 101B to a user 145).Al agent based operation of the control system
[0085] Referring to Fig. la, intelligent controller 112 is configured to act as an agent by conducting a learning process that achieves autonomous control of the system 101 by executing and updating the behaviour model 120. Operation engine 130 processes observation data 122 and control data 124 in one or more control cycles to determine a next action to be performed by the operational components 104a. . . 104m, based on the present state of the control system 101. Adaptation engine 132 is invoked by the operation engine 130 to perform adjustments to the behaviour model 120 in response to a novelty in the environment 150.
[0086] The intelligent controller 112 is configured to perform a plurality of episodes of operating the control system 101 in the environment 150, where each of the episodes are associated with the control system 101 performing the task based on the behaviour model 120. For example, in the application of Fig. lb vehicle 101 A includes a behaviour model 120, for example in the form of a neural network, that represents adeterministic policy or behaviour function for an automated vehicle navigation task (e.g., maintaining a following distance to a nearest object). At the beginning of an episode, a current state of the vehicle 101 A is applied as input to the behaviour model 120 along with a success indicator (e.g., that the perceived following distance is within a range, or at least a minimum value), and an output generated over the model 120 based on the input.
[0087] In a control cycle, the intelligent controller 112 evaluates the model (i.e., executes the network 120 in its adapted form) to determine the behaviour (i.e., next action) of the vehicle 101 A for performing the task, as based on the current state of the vehicle 101 A. The behaviour model 120 output indicating the next action to be performed in a next control cycle of the vehicle 101A is based on the state of the vehicle 101 A in the current cycle. The intelligent controller 112 processes one or more parameters of the determined next action to generate signals to control the vehicle 101 A in the environment 150.
[0088] The vehicle 101 A is then operated by the operational components 104a, 104b (e.g., to steer the vehicle 101 A and / or to control its acceleration and braking functions) as a result of the control signals generated by the intelligent controller 112. The state of the vehicle 101 A after performing the action can then be applied as input to the model 120 along with the success indicator, to generate a further output from the behaviour model 120 in the next control cycle (i.e., based on the new input being the new state of the vehicle 101 A after performing the action). By iteratively conducting control cycles that involve executing the model 120 with the current state of the vehicle 101 A to determine an action, and operating the vehicle 101 A by performing the action, the success indication may be achieved allowing the episode to advance (terminate).
[0089] In some implementations, the adaptation and control cycles are separated where the intelligent controller 112 receives the indication of change(s) in the environment 150, and subsequently invokes adaptation engine 132 to adapt the behaviour model 120, independently from and asynchronously with any control cycles executed by the operation engine 130. In other embodiments, the adaptation and controlcycles are integrated such that the intelligent controller 112 determines novelty of the operating environment 150 as a result of, or synchronously with, the operation of the control system 101. In such embodiments, action(s) of the control system 101 performed in the environment 150 may be evaluated based on one or more changes induced in the environment 150 from the action(s), and thereby creating feedback for adapting the behaviour model 120 (i.e., as a closed-loop).Behaviour model
[0090] In some implementations, the intelligent controller 112 achieves autonomous control of the system 101 via reinforcement learning. In such implementations, the intelligent controller 112 is programmed to maximize a reward function which is a user-provided definition of the task to be performed by the control system 101.
[0091] For the control system 101 having a state stat time t, the intelligent controller 112 chooses and executes action ataccording to a policy 7r(at|st), transitions to a new state st+1according to transition probability p(st+1|st, at), and receives a reward r(st, at). The goal of reinforcement learning is find the optimal policy 7t* which maximizes the expected sum of rewards from an initial state distribution. The reward is determined based on the reward function which is dependent on the task.
[0092] In some implementations, behaviour model 120 is a reinforcement learning model that parameterizes the policy 7r(at|st) for determining an action atof the system 101 based on its current state st. In some implementations, the behaviour model 120 may be determined by model-based reinforcement learning where a forward model of the environment dynamics is explicitly learned during a training process (e.g., via supervised training).
[0093] Alternatively, the behaviour model 120 may be determined by model -free reinforcement learning where forward dynamics of the environment are not explicitly modelled and where the policy is generated by an estimation technique (e.g., dynamic programming or Q-leaming).Deep reinforcement learning network
[0094] Some implementations utilize deep reinforcement learning (DRL) to train the behaviour model 120. Fig. 2a illustrates an example 200 of a behaviour model 120 implemented as a network in DRL (also referred to as a “behaviour network” or “behaviour neural network”). Examples a behaviour neural network include a policy network that, when executed, directly outputs an action associated with a state, and a value network in which an action is inferred from outputs obtained by executing the network (e.g., by selecting the (s, a) pair with the highest Q-value).
[0095] Behaviour network 120 is executed on a current state s of the control system 101 to produce a next action a to be performed by the system 101, as shown in Fig. 2a. A state of the embedded representation of the behaviour network 120, also referred to as a “network-embedded state” for a neural network implementation, is determined by the internal parameters of the behaviour network 120. The behaviour network 120 accepts the current state s as an input 202. State s is represented in the network 120 as a parameterization of weights of input nodes 204. Behaviour policy TT(S) of the network 120 is defined by parameters or weights of the one or more hidden (inner) layers 206 of the neural network. Network 120 produces an output 212 by generating values of weights of output nodes 208 corresponding to the input state 202. The output 212 represents an action a to be performed by the control system 101.
[0096] In some implementations, the behaviour network 120 is a deep neural network with an architecture that encompasses, for example, a multilayer perceptron (MLP), a recurrent neural network (RNN), a transformer, or a convolutional neural network (CNN). In some implementations, the neural network is configured with a single input layer 204 and a single output layer 208, and with multiple hidden layers 206.
[0097] Irrespective of the architecture, the behaviour network 120 implements a function (or “policy”) 210 7r(a|s) which enables the network 120 to determine an action a 212, as represented in an action space, from an input state s 202 defined over acorresponding state space. Policy 210 may therefore be abbreviated as TT(S) = a (i.e., the network 120 defines a function n that maps a state s to an action a).
[0098] The state space of a state s encompasses various information associated with the current state of the control system 101, and / or the current state of one or more components of the control system 101, in the environment 150. For example, the current state of vehicle 101 A may be determined by a position of the vehicle, a speed, a direction of travel, a distance to one or more objects, a current pose(s) of target object(s), and / or time derivatives of any of the aforementioned. Values at input nodes 204 represent a parameterization of the state 202 in response to inputting the state 202 into the behaviour network 120.
[0099] Behaviour network 120 is trained on a set of training data specific to the environment 150 and the task to be performed. Training of the behaviour network 120 in this manner results in a pre-configured or “baseline” network that reflects desired behaviour in an environment with nominal or expected characteristics (i.e., a baseline environment). As a result, the training data, and therefore the baseline network 120, often does not account for changes or variations in the environment 150 in which the control system is deployed. These changes or variations result in a mismatch between the behaviour (i.e., actions) generated by the baseline network 120 and the desired behaviour of the system 101 in the deployed environment 150. This may negatively impact on the ability to effectively operate the control system 101 to carry out the intended task using the baseline network 120 alone.Adapting the network using an embedded space
[0100] Fig. 2b illustrates an example 200’ of a behaviour model 120’ that is adapted in response to a change in the environment 150 relative to the trained baseline environment. When the baseline model is a network such as, for example, a neural network, as depicted in Figs. 2a and 2b, the adapted behaviour model 120’ is referred to as an “adapted behaviour network”. The adapted behaviour network 120’ replaces action a, the output action 212 produced by the baseline network 120, with anadaptation action a1. Each adaptation action a1is determined by a corresponding adaptation function nalp(also referred to as an “adaptation policy”). Intelligent controller 112 is configured to learn one or more adaptation functions nalpdynamically during exploration of the performing the task with control system 101 in environment 150.
[0101] Adapted behaviour network 120’ defines a set 220 of adaptation functionsnap —nap ■ Each of the N adaptation functions in set 220 generates a corresponding action a1... aNto replace output action a generated by the baseline model 120 for a given state s. The adaptation functions naLpare defined over the parameter space defined by the embeddings of the behaviour network 120, which is referred to as the “embedded state space.” That is, the embeddings 214 of network 120 translate a given system state s to a corresponding “model state” ms in the embedded state space of the behaviour model (i.e., the network-embedded state of the behaviour network).
[0102] In general, the embedded state space represents the learning performed by a behaviour model 120 to encapsulate desired behaviour of the system 101 for performing the task. For a behaviour network, the embedded state space refers to the vector space in which data is represented after being processed through one or more layers of the network. One or more dimensionality reduction techniques, such as PC A, or manifold learning techniques, such as t-SNE, or UMAP, may be applied to the processed data to further reduce the dimensionality of the resulting embedded state space (thereby forming a “dimensionally-reduced embedded state space”). As a consequence of the training process, the behaviour network 120 has model states (i.e., network-embedded states) that are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
[0103] For example, with reference to the example of Figs, lb and 1c the state of autonomous vehicle 101 A may be determined by the current speed of the vehicle 101 A, the position of the vehicle 101A in the navigation environment 150, relative positions of another vehicle / object closest to the vehicle 101 A in the navigation environment150, and the speed and / or direction of the other vehicle / object, as determined by sensors 102a... 102c.
[0104] Training a behaviour model 120 as a deep neural network captures relationships between the state parameters as regions in the space of the embedded parameters based on their similarity in effecting how the vehicle 101 A performs the task. For example, a change in the current speed of the vehicle 101A is likely to have a similar semantic effect on the task of controlling the following distance of the autonomous vehicle 101 A to another vehicle / object in the navigation environment, as a decrease in the speed of the other vehicle / object. Accordingly, embedded parameters which are strongly linked to these state parameters will be localized together in the embedded state space (i.e., since changes to either state parameter have a similar effect on the task).
[0105] Each adaptation function nalpis defined by the intelligent controller 112 with respect to a region of the embedded state space referred to as an adaptation region Rl. Embeddings 214 of the behaviour network 120 translate state s into model state ms and a corresponding action a is evaluated to determine whether an adaptation is required (i.e., if action a produced by the baseline model 120, or a corresponding existing adaptation, should be replaced with an adaptation actionFor example, if a new adaptation function nalpis determined for model state ms then the adapted behaviour network 120’ will generate output action aalpinstead of action a. As described herein, adaptation regions Rlare identified, and actions a1of the corresponding adaptation policies naLpare selected, to improve the performance of the system 101 on the task relative to the performance obtained with the baseline network 120.
[0106] Fig. 2c illustrates an exemplary embedded state space for a behaviour model 120 of an autonomous vehicle 101 A. The vehicle 101 A performs a task of controlling its following distance relative to another vehicle / object in the driving environment 150.
[0107] The embedded state space 232 displays the decision boundary of a trained agent (i.e., the intelligent controller 112). The points 231 represent states where the controller 112 chooses accelerate and the points 233 are where the controller 112 applies brakes. In the pre-novelty environment 160, the controller 112 accelerates in response to the vehicle in front of vehicle 101 A travelling slower but being far away, or in response to a vehicle in another lane. The controller 112 will slow down the vehicle 110A when the other vehicle is close to it.
[0108] After the ’’Reduced Braking Power” novelty occurs, in environment 150 the controller 112 applies brakes earlier; therefore, the controller 112 learns to adapt the behaviour model given the presence of a slow car ahead of the vehicle 101 A, instead of continuing to accelerate (i.e., the controller 112 applies the brakes instead). The localization of the states in the embedded state space 232 enables the controller 112 to adjust its behaviour in response to changes in environment 150 by altering the behaviour only in the region 240. In this example, the controller 112 compensates for the reduced braking power novelty allowing the controller 112 to continue to navigate the vehicle 101 A safely in its environment, even when faced with the unexpected change.
[0109] Fig. 2d is a schematic illustration 250 of control system 101 operating in an environment 150 with a baseline behaviour model 120 that is adapted to novelty to produce adapted behaviour network 120’. Control system 101 receives observation data including an indication of environment conditions and determines its current state stat time t. The current state stis applied to the baseline model 120 to generate the baseline action atand model state mst. Intelligent controller 112 generates adapted behaviour model 120 in response to receiving a change or variation signal 254 specific to the operating environment 150.
[0110] Intelligent controller 112 makes an adaptation decision 256 by processing the baseline action atand the model state ms and the change or variation signal 254. In response to determining that an adaptation is required, the intelligent controller 112 may utilize one or more adaptation functions 258 to determine the new actions(s)(adaptation actions). The adaptation functions 258 can be pre-existing functionsnap —nap that apply to the model state mstbased on its locality within a corresponding adaptation region R1... RN. The adaptation functions 258 may comprise new functions generated by the intelligent controller 112 on determination that the model state mstrepresents behaviour of the control system that is compromised in the environment 150 in relation to the task. A model state mstwith a previously adapted behaviour may be further adapted by the determination of a new adaptation function generating a new adaptation action. This is advantageous in enabling the performance of the model to be improved dynamically as the behaviour of the control system 101 is explored over time.
[0111] Alternatively, the intelligent controller 112 may decide 256 not to perform adaptation of the baseline action at. Intelligent controller 112 generates updated model parameters 262 replacing the baseline action atwith the new adaptation action a1, or retaining the baseline action at. Action atoutput from the adapted model 120’ is used to provide control signals to external components 140 of the entity controlled by the control system 101 (e.g., a drive train of autonomous vehicle 101A). Operation of the control system 101, and therefore the external components 140, may influence the next state st+1of the system 101.Intelligent control with novelty adaptation
[0112] Fig. 3 illustrates a method 300 for operating control system 101 to perform a task using a behaviour model 120 that is adapted to a novelty in the operating environment 150.Baseline model
[0113] At step 302, behaviour model 120 is configured as a baseline behaviour network using DRL. For example, configuration of the baseline behaviour network 120 may be performed using supervised learning. In such implementations, a set of training data is collected with a set of states and corresponding actions {s, a} of the controlsystem 101. The training data is specific to performing the task in a baseline environment.
[0114] The training data is processed to generate an embedded state space representation for the network 120. With reference to Fig. 2b, embeddings 214 are generated by a training process conducted on the training data, such as for example backpropagation using a squared error based loss function. In some implementations, such as those where the network 120 is a deep neural network with a large embedded parameter space, an unsupervised pre-training process is applied to build a deep feature hierarchy for one or more of the hidden layers 206.
[0115] In some implementations, the training process is performed by the intelligent controller 112 in an off-line mode and / or prior to use of the intelligent controller 112 to operate the control system 101. In other implementations, the baseline network 120 is trained with another computing device, and the trained network is stored in the memory system 108 of the intelligent controller 112 as a baseline model / network 120.Adaptation of the baseline model
[0116] Adaptation of the (baseline) behaviour network 120 occurs in response to a novelty in environment 150 relative to the baseline environment used to train the baseline network 120. With reference to Fig. 3, steps 304 and 306 are performed during an adaptation cycle 303 executed by the intelligent controller 112.
[0117] In some examples, the intelligent controller 112 determines a change in the environment 150 in response to generating, or otherwise obtaining, observation data comprising observations about the control system 101 and / or environment 150. Observation data may comprise input data from one or more input components 102a. . . 102n. For example, observation data may include imaging data generated by vision sensors 102a and / or proximity data generated by proximity sensors 102b. Such embodiments perform active novelty detection by receiving and processing data to detect changes in conditions in the operating environment 150 relative to the baseline environment.
[0118] In other implementations, the intelligent controller 112 does not actively or explicitly receive information about the operating environment 150, but instead implicitly perceives a novelty occurrence based on the performance of the control system 101. For example, the intelligent controller 112 may be configured to evaluate an action a determined by the baseline network 120 (or adapted behaviour network 120’) for a state s, to determine that (further) adaptation of the network 120 (or adapted behaviour network 120’) is required without detecting or measuring the environment 150.
[0119] Fig. 4 illustrates an exemplary method 400 performed by the intelligent controller 112 to adapt the behaviour network 120 including: (i) calculating an adaptation region Rlin an embedded state space of the operating environment (i.e., at step 402); (ii) determining, using an adaptation function nalp, an adaptation action aapassociated with the adaptation region Rl(i.e., at step 404); and (iii) updating the behaviour model 120 by replacing the action (a) of each compromised state with the adaptation action aap(i.e., at step 406).
[0120] Each adaptation region Rlincludes one or more compromised model states over the embedded state space defined as embedded state set MS^p. Formally, the zth adaptation function associated with region Rland embedded state set MS^pis given as naLp: MSaP-> A where A is the set of the available actions of the intelligent controller 112. Therefore, assuming there are N adaptation regions, the adapted behaviour policy is
[0121] That is, the behaviour network 120 is updated with a function that retains one or more actions of the control system associated with model states that are not within any adaptation region (i.e., the actions that are not compromised).
[0122] In response to determining that the current model state mstis compromised (e.g., by evaluation of the current action at, which may be the baseline action a or an existing adaptation aap) the intelligent controller 112 creates a new adaptation function, defined over a corresponding adaptation region, to improve on the compromised behaviour by replacing the action of compromised model state mst.
[0123] Each new adaptation region Rlis created to include the model state mstcorresponding to compromised action at. The region Rlmay be defined to include at least one other model state additional to the current model state mst, such as for example model states adjacent to current state mstin the embedded space. This enables the intelligent controller 112 to efficiently adapt behaviour model 120 for the current state of the control system 101 by generalizing the determined adapted behaviours to other states semantically similar to the current model state.Calculating the adaptation region
[0124] The adaptation region R includes the set of one or more compromised model states MS of the control system 101. Compromised model states are those states in the embedded space where an action atdetermined by the behaviour model 120, for the control system 101 in the operating environment 150, is compromised in relation to the task (i.e., the action leads to an undesired, or sub-par, performance outcome).
[0125] At step 402 in response to determining that an adaptation is required for action atof model state mstcorresponding to the current state st, the intelligent controller 112 calculates an adaptation region Rlfor the model state mst.
[0126] For implementations in which the behaviour model 120 is a deep reinforcement learning network, the calculation of the adaptation region is enabled by the existence of an embedded space that can be partitioned, segmented or otherwise divided to allow the last output layer 208 to specify the desired action a. That is, semantically similar states are close to each other in the embedded space of behaviour network 120. In such implementations the adaptation region is calculated by capturinga spatial relationship between the current model state mstand other states in the embedded state space.
[0127] Intelligent controller 112 may utilize one or more different techniques to calculate a new adaptation region R for model state mstaccording to various exemplary implementations. In some implementations, the new adaptation region is initially formed at a location of the current model state mstin the embedded space and expanded or moved therefrom in subsequent processing.
[0128] In one example, the embedded space is uniformly partitioned into a plurality of predetermined sections along each dimension. The predetermined section including model state mstis determined as the new adaptation region R. This approach is advantageous in that all possible adaptation regions may be calculated prior to realization of the corresponding model state, thereby reducing the number of computations that need to be performed by the intelligent controller 112 during runtime.
[0129] Uniform division of the embedded state space is most effective when the model state density within the each resulting section is sufficiently high to produce reliable predictions. However, the number of sections required to achieve reasonable precision increases exponentially with the dimensions of the embedding space.
[0130] In other implementations, a new adaptation region is formed by dynamically partitioning, segmenting or otherwise dividing the embedded state space sequentially over time with successive determinations of desired behaviour adaptations.
[0131] In one example, each new adaptation region is based on a cluster formed in the embedded state space. The intelligent controller 112 is configured to use one or more density-based, distribution-based, centroid-based, or hierarchical-based clustering algorithms to form the adaptation region at a location of the current model state mst. For example, unsupervised k-means clustering may be applied by the intelligent controller 112 to assign observations in the embedded state space into k clusters, one ofwhich has a centroid corresponding to the location of the current model state mst. Assignment of one or more candidate model states mscto the cluster of the adaptation region R is based on whether the respective candidate model state is located closer to the current model state mstor to the centroid of another cluster.
[0132] Alternatively, the intelligent controller 112 is configured to determine the adaptation region R by applying a classifier or pattern recognition model to data of the embedded state space. The classifier or pattern recognition model may be a machine learning model, such as a GMM or SVM trained using pre-labelled data or on unsupervised data.
[0133] In another example, the new adaptation region is calculated by constructing a Voronoi cell encompassing the current model state in the embedded state space. A Voronoi cell is a partitioned area in a plane of like cells, where each cell includes a given set of points. Inversely, for each point there is a corresponding region, the Voronoi cell, which consists of all points of the plane closer to that point than to any other. Therefore, the set of compromised states MS with similar semantic characteristics are represented as a Voronoi cell created by the model state mst. The intelligent controller 112 creates the Voronoi cells dynamically by only including, in any given cell, the model states that share the compromised behaviour and that can therefore be modified to produce the same desired adapted behaviour.
[0134] The intelligent controller 112 is therefore configured to learn new functions to adapt behaviour network 120 on a per model state and a per action basis. This improves run-time performance by only processing the states that are semantically similar to model state mstfor adaptation, thereby limiting the number of parameters of the network 120 that are required to be adjusted in a given episode.
[0135] Fig. 5 illustrates a method 500 performed by the intelligent controller 112 to create a new adaptation region as a Voronoi cell within the embedded state space.
[0136] At step 502, the new adaptation region R is constructed by forming a region initially including the current model state mst, and then the region is updated until closed with an update to the behaviour model (i.e., using step 406 of method 400). The formed region R is divided into sub-regions by executing steps 504 to 508. At step 504, the intelligent controller 112 selects the adaptation action (i.e., as determined at step 404 of method 400).
[0137] At step 506, the intelligent controller 112 evaluates the adaptation action with respect to its corresponding current state to determine a performance of the adaptation action. Then, at step 508, the intelligent controller 112 splits the region R into two regions if the performance of the adaptation action, as determined from its evaluation, is below a pre-determined level (as determined at step 507).
[0138] The performance of each adaptation action is evaluated using an evaluation function Eval(. ) that assigns a score, for example as a real -valued number, to the corresponding action of the candidate state as determined by the behaviour network 120 (i.e., ac= TT(SC)) . For example, for a digital control system 101” operating the Angry Birds game, the function Eval(-) returns the reward of an action, which is the number of pigs that were destroyed. The threshold may be set to ‘ 1’ such that mstis included among the existing model states to create a new Voronoi cell if the action at has not destroyed any pig. In other examples, a trained Q-value function may be defined as an alternative, or in addition to, the manually defined threshold function.
[0139] The intelligent controller 112 is configured to iteratively repeat steps 504 to 508 until the adaptation region is smaller than a predefined threshold, or until an overall performance of the control system reaches a predefined threshold.Determining adaptation actions
[0140] Following the determination of the adaptation region R for a model state mst(at step 402), at step 404 the intelligent controller 112 determines the adaptation action aap(omitting the ‘t’ for simplicity of notation) for the region Rl. In someimplementations, determining the adaptation action comprises: (i) selecting an adaptation function nalpcorresponding to the region Rlfrom a plurality of candidate functions; and (ii) applying the selected adaptation function to the current model state mstto generate the action aap.
[0141] Each adaptation function nalpmaps a set of embedded statesMS1corresponding to the adaptation region Rlto a set of candidate actions from an action space A of the operating environment 150. In some implementations, a selected behaviour adaptation function nalpoutputs the adaptation action aalpE A by evaluating the set of candidate actions in response to an input of the current model state mst(e.g., by determining a performance of one or more of the candidate actions in the set).
[0142] If the adaptation region R is newly created for the model state mst(i.e., in step 402), then intelligent controller 112 generates a new behaviour adaptation function naLpto output the adaptation action aap.
[0143] For compromised model states MS1, the intelligent controller 112 initializes a set of available actions from the action space as the “candidate action” set Aalpfor Ttalp, and the adaptation principle becomes an open adaptation principle. For model state mstin the region Rl, the intelligent controller 112 chooses an action aapfrom Aalpand evaluates the action using the Eval( ) function. Intelligent controller 112 initializes a variable BestScore to store the best action values of each mstused to create the corresponding regions { / ?} in the embedded space. To decide if iap(mst)= aap replaces 7r(st) = a, the intelligent controller 112 compares the action value of aapagainst both the T and BestScore values. If the value of aapis greater than or equal to both the value of a and T, then aapis considered a good action and the evaluation value of aapupdates the BestScore for mst.
[0144] Otherwise, the aapis removed from the available action set, and the intelligent controller 112 proceeds with one or more other candidate actions. The adaptationprinciple becomes a closed adaptation principle when only one action remains in the available action set or the intelligent controller 112 fail to find an adaptation action, in which case the baseline action will be used. In some cases, the Eval st, aap, st+1) may reach the highest value an action can get. For those cases, the adaptation principle is closed directly with just aapas no other action will evaluate to a higher value.
[0145] Fig. 6 illustrates an example method 600 performed by the intelligent controller 112 to determine mappings between behaviour adaptation functions and corresponding adaptation actions.
[0146] Assuming a trained agent (intelligent controller 112) with 3 available actions a1, a2, and a3, represented by blue, yellow, and grey regions 602, 604, 606, respectively, in the trained embedding space 601. Once a novelty is introduced, the agent 112 process the post-novelty environment state stto assess the model state mst. As this is the first model state, there is no associated adaptation function. Hence, as the model state falls in the blue region 602, and the behaviour of the baseline network 120 is retained to give 7r(st) = a1.
[0147] The intelligent controller 112 then performs a control cycle with action at= a1and obtains the next environment state st+1The intelligent controller 112 evaluates, at step 609, if atis a good action by comparing the Eval st, at, st+1) with the predefined threshold and the best scores value. If atis a good action, the intelligent controller 112 retains the action for mst. Otherwise, the action at(and therefore the current model state mst) is considered compromised and an adaptation is required (at step 610) to generate an adaptation action. In this case, as no adaptation function is currently associated with the model state mst, the intelligent controller 112 initializes the region R1with state set MS^pto cover the whole space of a1(i.e., the intelligent controller 112 has learned that the action a1leads to compromised behaviour).
[0148] In response to the initialization of the adaptation function, the function is defined by an open adaptation principle. The corresponding action space is initializedto contain all possible actions as candidate actions for the function, except the action that the baseline agent has already selected, e.g., the candidate action set is {a2, a3}. The intelligent controller 112 will sample one action from the candidate action space each time the adaptation region is activated and will remove the action from the candidate action space if the sampled action is inadequate.
[0149] At step 612, the intelligent controller 112 sets the next action at+1to a3. In response to control system 101 assuming the next state st+2, the intelligent controller 112 evaluates if at+1is a good action. In response to determining at+1to be a good action and Eval(st+1, a3,st+2) reaches the maximum possible value. In that case, the state of the adaptation function becomes closed as the intelligent controller 112 stops searching for other actions because at+1is already the best possible action.
[0150] Otherwise, the intelligent controller 112 tries another available action when a model state falls in the adaptation region 602. Consider a next state st+2and a model state mst+2that still falls in the region 602 but is considered an undesired action. In such a situation, the intelligent controller 112 partitions the region 620 using Voronoi Tessellation with mstand mst+2. A new adaptation region and corresponding function (open) is learned by the intelligent controller 112, and the prior (closed) adaptation function is stored.
[0151] The partitioning of the embedded space by the intelligent controller 112 is enabled by the localization of conceptually related states in the embedded space of the network 120. In response to an insufficiency of model states to decide the exact boundary of regions that require adapting the network 120, the intelligent controller 112 assumes that states close to the existing model state should have the same behaviour (i.e., their actions should be determined by the same adaptation function). The intelligent controller 112 applies the same adaptation function until the function fails to generate adequate actions. In response, the intelligent controller 112 creates a new adaptation region around the model state that the prior function could not improve upon. The intelligent controller 112 continues this process until a termination criteria is reached, or until no further adaptation can be performed.Updating the model
[0152] With reference to Fig. 4, at step 406 the intelligent controller 112 updates the behaviour model 120 by replacing the action (at) of each compromised state with the determined adaptation action (uap). Updating the behaviour model 120 may involve preserving some or all parameters of the behaviour model 120 (i.e., where the model 120 is retained as a baseline model 120) and generating new parameters (e.g., as update data) representing behaviour adaptation functions TT and actions aapto form the new adapted behaviour model (network) 120’. In the described implementations, generating the adapted behaviour network 120’ involves retaining one or more actions of the control system 101 (i.e., of the baseline network 120) that are associated with model states that are not within any adaptation region ( / ?l).
[0153] For example, the intelligent controller 112 may be configured to generate update data in memory system 108 to represent N adaptation functions nalpi = 1 ... N. The update data is generated by the adaptation engine 132 during the one or more adaptation cycles performed by the controller 112 and accessed by the operational engine 130 to replace action outputs that are generated by the baseline model 120.
[0154] Referring to Fig. 3, at step 308 the operation engine 130 of the intelligent controller 112 uses the adapted behaviour model 120’, defined by the baseline model 120 in combination with the derived adaptation functions nalpi = 1 ... N, to operate the control system 101 in the operating environment 150. An exemplary process performed by the intelligent controller 112 for using the behaviour model 120’ comprises: i) determining the current state of the control system 101 in the operating environment 150; ii) evaluating the behaviour model 120’ on the current state to determine an action of the control system 101; and iii) processing one or more parameters of the action to generate signals to control the control system 101 in the operating environment 150.Adaptation during episode-based operation
[0155] Method 400 is performed by the intelligent controller 112 during performance of one or more episodes, each of the episodes being associated with the control systemperforming the task in the environment 150 based on the network 120. Each episode has one or more steps consisting a sequence of states ({st}) and actions ({at}) of the control system 101. The environment E is represented by the intelligent controller 112 as a Markov Decision Process (MDP) {S, A, Pr, V }, where S is the set of environment states, A is the set of available actions of the agent, Pr(st+1\at, st) is the transition probability distribution that returns the probability of st+1given atand st, and V is the value or reward function that returns the immediate reward after the transition from stto st+1with action at.
[0156] In each episode, the intelligent controller 112 determines a state stG S, selects an action atG A, and then receives the reward vt+ 1 for the action atand the resulting state st+1until the episode ends. Novelties are considered as transformations of the underlying environment E via a transformation function (|>, such that a postnovelty environment is represented as (p E) = {S<p ,A<p, P^, V^}. In other words, a novelty may or may not change the state space, action space, transition probability, and the reward function of the environment.
[0157] Fig. 7 illustrates an example method 700 of operating a control system 101 in a post-novelty environment 150 while adapting the underlying behaviour network. The control system 101 may be a physical system including one or more components that control a robot, vehicle, or other mechanical, electrical or electro-mechanical system, such as the autonomous vehicle 101A of Figs, lb and 1c.
[0158] Alternatively, the control system 101 may be a digital system including one or more components of a module, engine, or simulation that control a digital entity, such as the game character 10 IB of Figs. Id and le. Moreover, while operations of method 700 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0159] At step 702, intelligent controller 112 starts a new episode to perform a task.
[0160] At step 704, the intelligent controller 112 initializes a random or pseudorandom process for task exploration. For example, the intelligent controller 112 may cause game character 10 IB to move to a random starting position in game world 150. As another example, the intelligent controller 112 may cause the vehicle 101 A to assume a particular state (e.g., particular position, velocity, and / or acceleration) in the navigation environment 150.
[0161] At step 706, the intelligent controller 112 initializes parameters of adapted behaviour network 120’ used by the system 101 to generate actions. Initialization is based on the parameters of the baseline network 120, and any updated parameters from prior episodes if such are available. For example, the intelligent controller 112 may generate and / or retrieve update data specifying the replacement of one or more actions of the baseline network 120 with one or more recently determined adaptation actions generated based on the method 400 of Fig. 4.
[0162] In some implementations, replacing one or more actions of the baseline network 120 comprises replacing values or parameters of nodes the baseline network model 120 with other values to form adapted behaviour network 120’. In other implementations, the baseline network 120 is fully retained and update data provides additional information to supplement values or parameters of the baseline network 120, with the update data and baseline model data together forming the adapted behaviour network 120’. In some embodiments, first execution of step 706 involves initializing the adapted behaviour network 120’ with the baseline network 120.
[0163] At step 708, the intelligent controller 112 identifies a current state stof control system 101. The current state stcan include the current system state and / or the current state of one or more objects in environment 150. The current state of the objects can be determined based data received from input components 102. For example, the current state of an object can be based on sensor data from one or more sensors, such as vision sensors 102a of a vehicle 101 A when the task is to maintain a distance between the vehicle 101 A and other objects. The current state of an object in a navigation environment 150 can be based on vision sensor data captured by the vision sensor 102a(e.g., a current position of an object relative to the vehicle 101 A). At the first iteration of step 708 the current state stwill be an initial state s0. For example, an initial state of a vehicle 101 A may be determined by the current state of one or more components of the vehicle 101 A such as its position, velocity, and / or acceleration.
[0164] At step 710, the intelligent controller 112 selects an action atto perform based on the current state stand based on the adapted behaviour network 120’. For example, the system may apply the current state stas input to a reinforcement learning policy TT defined by adapted behaviour network 120’ and generate, over the model based on the input, output that indicates an action at. In some implementations, the output includes values to generate control signals to set a state of one or more operational components 104 of the control system 101, and selecting the action can comprise selecting those values as the action.
[0165] At step 712, the intelligent controller 112 evaluates the selected action atto determine if an adaptation is required. Evaluation of the current action atis performed using a call to a predetermined function Eval(st, at, st+1) and a corresponding performance threshold T. A separate performance threshold T may be defined for each task and / or environment 150. The model state mstand action atare considered as compromised in response to the evaluation returning a value less than the threshold T.
[0166] In response to determining that the current action atis compromised, at step 714 the intelligent controller 112 selects an adaptation function to adapt the adapted behaviour network 120’. The selected adaptation function may be an existing adaptation function nalpthat corresponds to an open region for example in response to the model state mstfalling within a corresponding existing adaptation region Rl. Alternatively, the intelligent controller 112 defines a new adaptation region R and corresponding adaptation function n using the method 500 of Fig. 5. According to this approach, a new adaptation region, and a corresponding new adaptation function, is defined in response to a current action being below the performance threshold, unless the current action results from an adaptation function for an open region (i.e., when theintelligent controller 112 is still in the progress of resolving the best adaptation action to use for the region).
[0167] At step 716, the intelligent controller 112 determines an adaptation action aapto replace current action based on the selected or new adaptation function. For example, the intelligent controller 112 may assign adaptation action aapfrom a set of candidate actions of the selected policy for an open principle, using the method 600 of Fig. 6.
[0168] At step 718, the intelligent controller 112 adapts the behaviour of the control system 101 by setting the current action atto the adaptation action aap. This may involve replacing the control signal values of the current action with corresponding values of the adaptation action, for example to cause the control system 101 to operate one or more operational components differently (i.e., setting at= aapover the output parameter space).
[0169] In response to determining that the current action atis not compromised, then at step 713 the intelligent controller 112 proceeds with the current action atwithout parameter or value replacement.
[0170] At step 720, the intelligent controller 112 executes the action and observes the reward vtand the subsequent state st+1that results from the action. For example, the intelligent controller 112 may generate one or more control signals to cause the acceleration system of the vehicle 101 A to change the speed of the vehicle 101 A according to the action. The intelligent controller 112 may observe the reward based on a reward function V and optionally based on a success signal provided to the intelligent controller 112. For example, the reward function for a task of maintaining a safe distance to a nearest object in the navigation environment may be based on a difference between a position of the vehicle 101 A and that of the nearest object. The subsequent state st+1can include a subsequent state of the vehicle 101 A and / or one or more environmental objects. For example, the subsequent state of vehicle 101 A may include the state of the vehicle 101A as a result of the action, such as positions, velocity, and / or acceleration of the vehicle 101 A. In some implementations, at step 720 the intelligentcontroller 112 additionally or alternatively identifies other observations as a result of the action, such as for example vision sensor data captured by a vision sensor 102a of the vehicle 101 A after implementation of the action.
[0171] At step 722, the intelligent controller 112 determines if the current episode commenced at step 702 has ended based on whether success or other criteria has been met. For example, the intelligent controller 112 may determine success if the reward vtobserved at step 720 satisfies a threshold. In some implementations, the intelligent controller 112 determines success based on criteria such as a threshold amount of time and / or a threshold number of iterations of steps 708 to 720 having been performed during the episode.
[0172] If the intelligent controller 112 determines that the episode has ended, the intelligent controller 112 proceeds to step 702 and starts a new episode. According to method 700, in response to commencing a new episode, at step 706 of the new episode, the intelligent controller 112 initializes the adapted behaviour network 120’ with parameters that are updated relative to the adaptations performed in the immediately preceding episode (as a result of generating update data to represent adaptation actions determined in method 600 of Fig. 6 and / or other methods). For example, the initialization of the adapted behaviour network 120’ may involve generating or retrieving update data representing parameters of newly defined (i.e., in the previous episode) adaptation functions, in addition to data of the baseline network 120. In these and other manners, each episode of method 700 can utilize a behaviour model that is updated based on the most recently performed actions. Multiple episodes may be performed by the intelligent controller 112 utilizing method 700 until the task is completed and / or until some other signal is received (e.g., an error occurs).
[0173] In response to the intelligent controller 112 determining that the episode has not ended (at step 722), the intelligent controller 112 proceeds to step 708 and performs an additional iteration of steps 708 to 720.Evaluation
[0174] An evaluation of the proposed techniques was performed in the context of a game controller operating as an agent in four action domains, each with different reward structures and task types. The results demonstrate that the proposed methods enable a reinforcement learning-based agent to rapidly and effectively adjust to a wide range of novel situations across all tested domains. Further, the proposed techniques outperform standard machine learning approaches (i.e., Online learning and finetuning) and the current state-of-the art open-world learning agents by a large margin.Domains and protocol
[0175] Evaluation domains included: CartPole (see
[0015] ) and MountainCar (see [16, 17]) (classic control); CrossRoad (see
[0018] ) (path-finding), and Angry -Birds using Science Birds Novelty (see
[0019] ) (physical reasoning). Fig. 8 is a schematic diagram 800 that illustrates the four domains used for the evaluation.
[0176] Evaluation follows the open-world learning trial setting discussed in (see
[0020] ) and used in (see [12, 11, 14, 19]) to evaluate the proposed methods, which is summarized as follows. A trial consists of a sequence of episodes, which starts from episodes drawn from the pre-novelty environment E. After several pre-novelty episodes, a novelty is introduced, and the environment E is transferred to the post novelty environment <|)(E). After that, all the remaining episodes are drawn from <|)(E). In each trial, there is one and only one transformation applied to the pre-novelty environment. The agent’s objective is to maximize the cumulative reward in the trial. In this work, a trial consists of 80 episodes. The first 40 episodes are sampled from the pre-novelty environment, and the other 40 are from the post-novelty environment. The novelty is always introduced in episode 0. Therefore all the pre-novelty episodes have a negative index from -40 to -1. The same settings were applied for all domains, except for Angry Birds where the number of pre-novelty episodes in a trial is a random number.Novelty detection and generation
[0177] Novelty detection is performed using a simple heuristic in each domain to determine if the novelty has been introduced. That is, the agent reports novelty when the 5-episode rolling average reward is less than a domain dependent threshold. The novelty detection heuristic is activated from the episode index 5 (the 45th episode) in every trial in CartPole, MountainCar and CrossRoad. The novelty status is provided to the agent in AngryBirds.
[0178] For CartPole, Mountain Car and CrossRoad, post-novelty environments are created by varying the environment’s physical or spatial parameter values using the novelties presented in the 2022 AIBirds Competition (see
[0021] ) Novelty Track for the Angry Birds domain. Further details are provided as follows.CartPole
[0179] The OpenAI Cartpole environment is a classical RL benchmark based on the cart-pole problem described by
[0015] , In this environment, a pole is connected to a cart. The task is to maintain the pole upright by moving the cart either left or right. The reward of +1 is given for each step that the pole remains upright. The episode ends if the pole is greater than 12 degrees from the vertical position or if the cart is greater than 2.4 units from the center (i.e. the cart reaches the edge of the display). In this evaluation, the maximum possible reward is truncated at 200, so the total reward for an episode is between 1 and 200.Baseline Agents
[0180] The proposed methods were evaluated using a pre-trained DQN and a pretrained PPO from
[0022] , The two agents achieve perfect performance in the pre-novel environment. For each agent, performance was evaluated on the trials with four settings: no learning (-Baseline), online learning (-Online), fine-tuning (-FineTune), and Principle Adaptation (i.e., the proposed methods).Novelty
[0181] The parameters in CartPole were varied to create novel situations for the agents. The parameters that were varied in CartPole are: 1) length of the pole, 2) gravity, 3) mass of the cart, 4) mass of the pole, and 5) the magnitude of the pushing force, each with the default value of 0.5, 9.8, 1, 0.1 and 10 respectively.
[0182] To explore the limits of the proposed method, the parameter values were sampled from an extensive range; the lower range being set to be the default parameter value divided by ten and the upper to the default value multiplied by 10 of the default value for all parameters. Therefore, the ranges of the parameters are:• Length: the length of the pole from 0.05 to 5 unit;• Gravity: gravity range from 0.98 to 98 unit;• MassCart: the mass of the cart from 0.1 to 10 unit;• Masspole: the mass of the pole from 0.01 to 1 unit; and• Force mag: the magnitude of the pushing force from 1 to 100 unit.
[0183] For each trial, the novel environment’s parameter values were sampled uniformly from the parameter range given above. The number of trials run per agent was 5000. All agents used the same detection heuristic: a novelty is detected if the five- rolling average episodic reward is less than 150.CartPole Results and Discussion
[0184] Fig. 9 is a graph 900 showing the overall cumulative rewards on CartPole for all the models evaluated. The proposed methods are referred to as Novelty Adaptation Principles Learning (“NAPPING”) methods for the purposes of the discussion below.
[0185] The median is used instead of the mean to alleviate the impact of the results from unsolvable / challenging novel environments. From the figure, we notice that noneof the baseline, online learning, and fine-tuning agents showed any adaptation. In contrast, the PPO-NAPPING agent and the DQN-NAPPING reacted quickly and effectively to the novelties. Significantly, the median cumulative rewards of PPO- NAPPING reach 200 in less than ten episodes. DQN-NAPPING, although showing substantial adaptation progress, adapts less efficiently when compared to PPO- NAPPING. It suggests the PPO learned a more suitable embedding space that allows NAPPING to identify practical adaptation principles quickly.
[0186] Fig. 10 is a graph set 1000 demonstrating the DQN-NAPPING and PPO- NAPPING performances by different parameter values for only the post-novelty episodes. As we start to detect novelties from episode 5, the first five episodes illustrate the performance of the DQN-Baseline and PPO-Baseline. The results confirm that the performance of both baseline agents dropped significantly even when the task became more manageable, e.g., when the gravity became lower, and the pushing force became larger. As shown in the figures, NAPPING enables standard DQN and PPO to adapt to almost all the novelties, except those adamant ones, e.g., when the length of the pole or the pushing force becomes too small, and when the gravity becomes too large.MountainCar
[0187] OpenAI MountainCar is a popular control domain that is based on the deterministic MDP problem described in
[0017] , In the domain, an underpowered car moves along a curve and attempts to reach a goal state at the top of the ’’mountain” by selecting between three actions on every time step, Forward, Stay, and Backward. The Forward action accelerates the car in the positive x direction. Backward accelerates the car in the negative x direction. The agent’s state is described by two state variables: the horizontal position, x, and velocity, x . The agent receives a reward of -1 per time step. Thus the goal of the agent is to reach the goal state as soon as possible. An episode ends when the agent reaches the goal state or 500 time steps has passed, whichever comes first. Therefore, the reward for an episode is between 0 to -500.Baseline Agents
[0188] In order to examine NAPPING’s performance when the baseline agents have different performance levels, two versions of PPO were included, namely the PPO- Weak and PPO-Strong. Also included is the DQN agent from
[0022] for evaluation. PPO- Strong and DQN solve the pre-novelty environment with an average episodic reward of around 120 and -100, respectively. PPO-Weak fails in the pre-novelty environment with -500 average rewards but still manages to work in some less demanding postnovelty environments. PPO-Strong has a superior generalization ability, with solving even the most novel situations. This agent was included to serve as an empirical maximum performance for this domain so that we can evaluate how much NAPPING can help DQN and PPO-Weak to adapt more intuitively.Novelty
[0189] To create novelty situations in MountainCar, the pushing force (default:©.001) and the gravity (default:©.0025) of the environment were varied. As extreme values in MountainCar can quickly render the game unsolvable, e.g., small pushing force and massive gravity, values of pushing force from [0.0001, 0.02] and gravity from [0.0001, 0.005] were selected to create novelties.
[0190] Like CartPole, each trial in MountainCar consists of 80 episodes. The first 40 episodes came from the pre-novelty environment, and the rest 40 from the novel environment. For each trial, the novel environment’s parameter values were sampled uniformly from the parameter listed range. The number of trails ran per agent was 2000. Like the agents in CartPole, all agents detect novelty from episode 5 (the 45th episode) using the same detection heuristic: a novelty was detected if the five-rolling average episodic reward is less than -120 or greater than -80.MountainCar Results and Discussion
[0191] Fig. 11 is a graph set 1100 displaying the overall cumulative rewards on the Mountain Car domain. Unlike CartPole, the cumulative rewards in MountainCar reflect how fast an agent reaches the goal state; therefore, it is not always possible for theagent to return to the pre-novelty performance level when the environment becomes harder to solve. It is clear from the graphs of Fig. 11 that all -NAPPING agents adapt to novel situations rapidly, with a significant upward jump in the performance in episode 5 when the agents learn adaptation principles. It suggests that in the MountainCar domain, NAPPING achieves significant adaptation performance within the very first episode in the post-novelty environment. Of further note is that although DQN performs better in the pre-novelty environment, it has lower average rewards when a novelty is introduced compared to the PPO-Strong agent.
[0192] While PPO-Strong-NAPPING seems more robust than DQN-NAPPING, they converge quickly to the same performance level. The performance of DQN-NAPPING, although it starts much lower than PPO-Strong, quickly picks up after episode 5 and converges to the same level as PPO-Strong’ s performance. This indicates that with NAPPING, DQN is able to reach the empirical maximum performance level. Although the PPO-Weak agent fails in the pre-novelty environment, the baseline still solves some of the novelties in MountainCar. The PPO-Weak-NAPPING result shows that even with a weak baseline agent, our method still manages to help the agent react to unknown changes with a significant jump in the performance in the first novel episode. However, we also notice that the asymptotic performance of PPO-Weak is lower than that of PPO-Strong-NAPPING and DQN-NAPPING, indicating that the baseline performance affects the performance of the proposed methods.
[0193] Fig. 12 is a graph set 1200 displaying the average performance of -NAPPING agents on the Mountain Car domain by different novel parameter values. There are several notable features of the results. First, unlike PPO-Strong, where the performance seems more consistent with the task’s difficulty, DQN-NAPPING’ s performance dropped across all the gravity values. On the other hand, PPO-Strong achieves better performance when the environment becomes more straightforward (e.g., larger pushing force or lower gravity values). Second, NAPPING assists DRL agents to adapt to not only the novelties that make the task harder (post-novelty baseline performance lower than per-novelty), but also those that make the environment easier (post-novelty baseline performance higher than per-novelty).
[0194] Fig. 13 is a graph set 1300 displaying the performance of the proposed - NAPPING agents under different gravity and force values. Each dot represents an evaluated trial with novel environment parameter values, with the x axis showing the value of force and y showing gravity. The color of the dots represents the average accumulative reward achieved in the environment. The larger the reward, the deeper the color and the larger the size of the dot. The first row / column shows the performance of the baseline model, while the second shows the performance of the same DRL agent but with NAPPING. The red dots represent the environment where the agent fails, i.e., having -500 rewards.
[0195] In general, the dots are darker in the bottom left direction and are lighter and eventually turn to red in the top left direction. The novel task becomes easier to solve with a larger pushing force value and a smaller gravity level. While the difficulty of the environment increases quickly in the top left direction, resulting in possibly unsolvable environments around the top left comer. The orange cross is the default parameter value of the domain. With reference to Fig. 13, the proportion of failed tasks reduced significantly for both the NAPPING based PPO-Weak and DQN agents, witnessing 59.4% and 71.45% reduction in the number of failed tasks, respectively. Further, the performances of PPO-Strong-NAPPING and DQN-NAPPING are very similar. It indicates that with NAPPING, the DQN agent is able to reach the empirical maximum performance level.CrossRoad
[0196] CrossRoad is a toy domain inspired by OpenAI Freeway. Compared to OpenAI Freeway, CrossRoad allows the user to fully control the speeds and locations of moving objects and it allows the agent to move not just in upward and downward directions, but also in left and right directions. It has 8 ’’cars” (red boxes) moving horizontally at different speeds and directions. The goal is to cross the road while navigating the ’’player” (green box) and avoiding colliding with the cars. The agent can either move left, right, up, down, or stay in a place. The episode starts with the player being placed on the bottom line and ends if the player collides with a car, crosses theroad, or reaches the episode length of 100 steps. A +1 reward is given if the agent crosses the road, -1, if the agent collides with a car or 100 steps are passed. This evaluation assumed that the agent only has partial observable state space; the agent sees only its location and the adjacent eight cells, each with the x and y coordinates. Therefore, the state is an 18 dimensional vector. The reward agent can get in each episode is therefore either -1 or 1.Baseline Agents
[0197] A DQN and a PPO agent were trained using the default setting described in
[0022] , The trained agents perform perfectly in the pre-novelty environment. Like previous domains, evaluations involved the offline (-Baseline) version, online learning (-Online) version, and fine tuning (-FineTune) version and a comparison of these with the performance of NAPPING (-NAPPING) for both the DQN and PPO agent.Novelty
[0198] Eight novelty bases were created in CrossRoad, representing eight different novel layouts of the game environment. These bases are:• Super Slow Speeds: the cars are still moving in the same direction, but the speed is much slower;• Super Fast Speeds: the cars are still moving in the same direction, but the speed is much faster;• New Speeds: the cars are moving in the same direction but at different speeds;• Opposite Direction: the cars are moving at the same speeds but in opposite directions;All to the Left: All cars are moving to the left;• All to the Right: All cars are moving to the right;• Shift Speeds: speeds of the cars are shifted;• Reverse Cars: the order of the cars is reversed, i.e. the first car from the top now become the last car at the bottom.
[0199] The initial position and the speed of cars for each novelty base was further varied by adding noise on both the initial position and the speed of cars. For each car’s initial position, a uniformly sampled noise was added from [-1, 1], and for speed, the noise is sampled from [-10, 10], This noise was also added to the pre-novelty environment. In total, nine different novelties were produced. A novelty was detected if the five-rolling average episodic reward is less than 1. The trials per novelty per DRL agent was set to 100, totalling 1800 trials.CrossRoad Results and Discussion
[0200] Fig. 14 is a graph set 1400 showing the overall performance of the evaluated agents. Similar to the results in previous domains, only agents with NAPPING showed adaptation behavior, with the DQN-NAPPING and PPO-NAPPING reaching around 0 cumulative rewards, indicating a 50% of solving rate of novel tasks. In comparison, all other agents fail to learn any helpful policy to react to the novelties. When comparing PPO-NAPPING and DQN-NAPPING, Fig. 14 illustrates that although PPO-NAPPING has a slightly more robust baseline model and adapts faster than DQN-NAPPING, the asymptotic performance of PPO-NAPPING is slightly lower than that of DQN- NAPPING.
[0201] Fig. 15 is a graph set 1500 displaying the performance of the DRL agents with NAPPING. It is observed that different DRL agents with different embedded spaces show different adaptation performances. For example, DQN-NAPPING responds to the New Speed novelty much faster than PPO-NAPPING and has a much higher asymptotic performance. While PPO-NAPPING performs slightly better in All to theRight and Super Fast Speeds novelties and much better in Shifted Speeds and Super Slow Speeds novelties.Angry Birds
[0202] Angry Birds is a simple and intuitive game with realistic physical simulation providing a testing domain for physical reasoning (see [23, 24]). The goal of Angry Birds is to destroy particular in game entities (i.e., pigs) by shooting birds from a slingshot. Pigs are normally protected by physical structures made of blocks with varied sizes, shapes, and materials. Some birds have special powers that can be activated after being released from the slingshot. The only actions available to the players are to select a bird trajectory by pulling the bird back in the slingshot to the release coordinates (x, y) and then tapping the screen at time t after release to activate the special power.
[0203] In this evaluation, the proposed NAPPING methods are tested using the novelties used in the most recent AIBIRDS competition (see
[0021] ), which focused on continuous action spaces and environments in that the agent does not have complete knowledge about the physical parameters of objects.AIBIRDS Novelty Track competition
[0204] In the AIBIRDS Novelty Track competition (see
[0019] ), an agent can request two types of state representations: the screenshot and the ground truth. A screenshot state representation is a 480 x 640 colored image, and the ground truth representation is in JSON format containing all foreground objects in a screenshot. Each object in the ground truth representation is represented as a polygon of its vertices (provided in order) and its respective color map containing a list of 8-bit quantized colors that appear in the game object with their respective percentages. An agent can request screenshots or ground truth representations of the game level at any time while playing.
[0205] In the evaluation, the action space of Angry Birds is set to three discrete actions: objects to shoot, high trajectory or low trajectory, and target offsets. Morespecifically, using the trajectory planner provided in the AIBIRDS Competition Novelty Track, the agent first decides which object to shoot at. Next, depending on whether the bird should reach the object from the top left side or just the left side, the agent needs to decide whether the high or low trajectory should be used. To increase the freedom of the agent’s strategy and compensate for the noise in the trajectory planner, the agent was also allowed to choose an offset for the target object from 7 different locations starting from no changes to 15 / 25 / 35 pixels vertical variations. The action space was further simplified by assuming that the agent uses only the full strength of the slingshot, i.e. the bird is pulled to the furthermost point from the slingshot before being released. When NAPPING samples the adaptation action in Angry Birds, it always starts with the (object, trajectory, offset) triplet purposed by the baseline agent and searches for a triplet that with a maximum Eval value by firstly varying objects, then trajectory and finally offset.Baseline Agents
[0206] For the Angry Birds domain, a DQN (see [25, 26]) agent was trained and the DQN-Rel agent that contains a relational module was used to evaluate the performance of NAPPING. Both agents were trained on a set of pre-novelty tasks provided by the competition organizer following the same setup used in
[0024] , The pre-novelty task set contains 2450 game levels that are generated from seven task templates. Both agents were trained on in total of 20,000 episodes with over 95% of training pass rates.
[0207] The evaluation compares the proposed approach with a heuristic adaptation agent Naive Adaptation, and two current state-of-the-art open world learning agents HYDRA (see [11, 12]) and OpenMIND (see
[0013] ). OpenMIND was the official champion of the 2022 AIBIRDS Competition Novelty Track. The Naive Adaptation is built on top of the Naive Agent, which shoots only at the pigs. Naive Adaptation uses the strategy of the Naive Agent in the pre-novelty game levels.
[0208] After receiving the novelty signal, it searches for a combination of (objects, trajectories, and delays) that solve a game level and keeps a record of each triplet tried.Once a solution triplet (e.g., a solution triplet can be (pig 1, high trajectory, delay 5 seconds)) is found for a trial, the Naive Adaptation will keep using the triplet until it doesn’t work anymore, where the agent starts to search for another triplet. Also included were another Anonymous open-world learning agent and a non-adaptation agent DataLab (the 2014 and 2015 AIBIRDS competition winner; see
[0027] ) to examine the effect of the novelties on standard agents for comparison.Novelty
[0209] There are several novelties available in the 2022 AIBIRDS competition Novelty Track:• Green Egg: A green rectangular object that has the same color as pigs but a different shape from the pigs;• Bone: An object that needs to be destroyed to pass the game level;• Taller Slingshot: Slingshot is now twice high as normal - agent needs to adjust the release point accordingly;• Pig Color: All the objects’ colors changed to pig’s color;• Cicumcircle: circumcircle of objects instead of shapes in the ground truth;• Floating Wood: Wood objects can float on the ice objects; and• Bird Likes Pig: Red Bird cannot cause damage to pigs.
[0210] All the agents were evaluated following the trial setting, but unlike in previous domains, the number of pre-novelty episodes is a random number between each trial as that was the setting used in the competition. There are 40 novel episodes, and 40 trials were run for each novelty.Angry Birds Results and Discussion
[0211] Fig. 16 is a graph 1600 showing the overall pass rate of the agents tested in the Angry Birds domain. Unlike in the previous domains, in Angry Birds the number of pre-novelty episodes is random; therefore, we set the episode where novelty is introduced to be 0, and pre-novelty episodes are, therefore, the negative ones. DQN- NAPPING and DQNRel-NAPPING have the strongest novelty adaptation performance among all the agents evaluated.
[0212] Averaging just above 0.9 in the pre-novelty instances, the pass rates of the - NAPPING agents drop to below 0.4 right after a novelty is presented and then quickly recover to around 0.8 before the end of the trial. The Naive Adaptation agent, although it responds quickly to some of the novel tasks in just 1 episode, has a lower asymptotic performance compared to that of DQN-NAPPING and DQNRel-NAPPING.OpenMIND, as one of the state of-the-art open-world learning agents, achieved above 0.8 pass rate in the pre-novelty and above 0.4 in the post-novelty environment. Both HYDRA and the Anonymous agents have similar pre-novelty performance, Anonymous is more novelty resilient than HYDRA with a less significant reduction in the performance after a novelty is introduced. However, Anonymous seemed to struggle to find the strategy to adapt to novelties, as the average pass rate gradually drops after novelty is introduced. Despite being the previous champion in the standard AIBIRDS Competition track, the non-adapting DataLab has the lowest overall performance among all the agents, with a pre-novelty pass rate of around 0.3 to below 0.2 in the post-novelty episodes.
[0213] Fig. 17 is a graph set 1700 that shows the pass rate of the agents by different novelties. It is observed that the -NAPPING agents respond to all novelties, returning to the pre-novelty performance except for the Bone novelty, where the agents need to understand the new object Bone is another type of pig and needs to be destroyed to pass the game level. The -NAPPING agents are also robust to the Pig Color novelty, as the agents (also the Naive Adaptation agent) use the object’s label directly instead of thecolor maps. Nevertheless, the Naive Adaptation agent fails to adapt to the Cicumcircle and the Floating Wood novelty.
[0214] The results illustrate the advantages of the proposed NAPPING approach over Naive Adaptation. Naive Adaptation relies on an input state assumption that objects are represented by vertices which is in appropriate for Cicumcircle. Further, Naive Adaptation is not well suited to providing varying trajectories over the different game levels in response to the Floating Wood novelty. For example, to deal with this novelty effectively a high trajectory must be used when objects are closer to the slingshot, and a lower trajectory is required if the high trajectory path is blocked and when the game objects are far away from the slingshot. On the other hand, the -NAPPING agents do not rely on assumptions about the input states since the agent operates directly over the embedding space of the underlying DRL models. Also, NAPPING agents can identify the different actions for tasks with different spatial arrangements and timing requirements, as these differences are assumed to have different representations in the embedded space of the DRL agent.
[0215] Fig. 18 is a graph 1800 illustrating the average asymptotic (last ten episodes) pass rate of the proposed approach compared to other approaches on the AIBIRDS Competition Novelty Track. The NAPPING agents have similar asymptotic performance levels, with DQNRel being slightly better in the Bird Likes Pig and Cicumcicle novelty and slightly worse in other novelties. NAPPING outperforms the Anonymous agent which produces higher pass rates higher than the non-adapting agent DataLab. The Naive Adaptation has a relatively high pass rate if the heuristic utilized is effective for the specific novelties introduced, and Open-MIND is effective in adapting the Green Egg and Taller Slingshot novelties. However, neither of these approaches produce consistently high pass rates for all the novelties evaluated.
[0216] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The presentembodiments are, therefore, to be considered in all respects as illustrative and not restrictive.References[1] O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al., Grand master level in starcraft ii using multi-agent reinforcement learning, Nature 575 (7782) (2019) 350-354.[2] D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., Mastering the game of go with deep neural networks and tree search, nature 529 (7587) (2016) 484-489.[3] A. Lazaridis, A. Fachantidis, I. Vlahavas, Deep reinforcement learning: A state- of-the-art walkthrough, Journal of Artificial Intelligence Research 69 (2020) 1421-1471.[4] S. Witty, J. K. Lee, E. Tosch, A. Atrey, K. Clary, M. L. Littman, D. Jensen, Measuring and characterizing generalization in deep reinforcement learning, Applied Al Letters 2 (4) (2021) e45.[5] K. Khetarpal, M. Riemer, I. Rish, D. Precup, Towards continual reinforcement learning: A review and perspectives, arXiv preprint4 arXiv:2012.13490 (2020).[6] S. Padakandla, P. KJ, S. Bhatnagar, Reinforcement learning algorithm for non- stationary environments, Applied Intelligence 50 (11) (2020) 807 3590-3606.[7] W. C. Cheung, D. Simchi-Levi, R. Zhu, Reinforcement learning for nonstationary markov decision processes: The blessing of (more) optimism, in: International Conference on Machine Learning, PMLR, 2020, pp. 811 1843— 1854.[8] F. Fernandez, M. Veloso, Probabilistic policy reuse in a reinforcement learning agent, in: Proceedings of the fifth international joint conference on Autonomous agents and multiagent systems, 2006, pp. 720-727.[9] A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. P. van Hasselt, D. Silver, Successor features for transfer in reinforcement learning, Advances in neural information processing systems 30 (2017).
[0010] T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, S. Levine, Metaworld: A benchmark and evaluation for multi-task and meta reinforcement learning, in: Conference on robot learning, PMLR, 2020, 815 pp. 1094-1100.
[0011] R. Stem, W. Piotrowski, M. Klenk, J. de Kleer, A. Perez, J. Le, S. Mohan, Model-based adaptation to novelty in open-world ai, Proceedings of the 32nd International Conference on Automated Planning and Scheduling, Bridging the Gap Between AI Planning and Reinforcement Learning.
[0012] M. Klenk, W. Piotrowski, R. Stern, S. Mohan, J. de Kleer, Model-based novelty adaptation for open-world ai, in: International Workshop on Principles of Diagnosis (DX), 2020.
[0013] D. J. Musliner, M. J. Pelican, M. McLure, S. Johnston, R. G. Freedman, C. Knutson, Openmind: Planning and adapting in domains with novelty, Proceedings of the Ninth Annual Conference on Advances in Cognitive Systems (2021).
[0014] S. Goel, Y. Shukla, V. Sarathy, M. Scheutz, J. Sinapov, Rapid-leam: A framework for learning to recover for handling novelties in open-world environments, arXiv preprint arXiv: 2206.12493 (2022).
[0015] A. G. Barto, R. S. Sutton, C. W. Anderson, Neuronlike adaptive elements that can solve difficult learning control problems, IEEE Transactions on Systems, Man, and Cybernetics SMC-13 (1983) 834-846.
[0016] G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, W. Zaremba, Openai gym, ArXiv abs / 1606.01540 (2016).
[0017] A. W. Moore, Efficient memory-based learning for robot control, Tech, rep., University of Cambridge (1990).
[0018] nikonkate, Novelty domains, https: / / github.com / nikonkate / novelty-domains (2022).
[0019] C. Xue, V. Pinto, P. Zhang, C. Gamage, E. Nikonova, J. Renz, Science birds novelty: An open-world learning test-bed for physics domains, Proceedings of the AAAI Conference on Artificial Intelligence, Designing Artificial Intelligence for Open Worlds (2022).
[0020] V. Pinto, J. Renz, C. Xue, P. Zhang, K. Doctor, D. W. Aha, Measuring the performance of open-world Al systems, Proceedings of the AAAI Conference on Artificial Intelligence, Designing Artificial Intelligence for Open Worlds (2022).
[0021] AIBIRDS, Angry birds Al competition [cited 22.11.2022], URL http: / / aibirds.org /
[0022] A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Emestus, N. Dormann, Stable-baselines3 : Reliable reinforcement learning implementations, Journal of Machine Learning Research 22 (268) (2021) 1-8. URL http : / / ; ml r.org / papers / v22 / 20- 1364 ,htm 1
[0023] J. Renz, X. Ge, M. Stephenson, P. Zhang, Ai meets angry birds, Nature Machine Intelligence 1 (7) (2019) 328-328.
[0024] C. Xue*, V. Pinto*, C. Gamage*, E. Nikonova, P. Zhang, J. Renz, Phy-q as a measure for physical reasoning intelligence., Nature Machine Intelligence 5 (2023) 83-93, *equal contribution.
[0025] Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, N. Freitas, Dueling network architectures for deep reinforcement learning, in: International conference on machine learning, PMLR, 2016, pp. 1995-2003.
[0026] H. Van Hasselt, A. Guez, D. Silver, Deep reinforcement learning with double q-leaming, in: Proceedings of the AAAI conference on artificial intelligence, Vol. 30, 2016.
[0027] J. Renz, X. Ge, R. Verma, P. Zhang, Angry birds as a challenge for artificial intelligence, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 30, 2016.
Claims
CLAIMS:
1. A method comprising: during performance of one or more of episodes of operating a control system in an operating environment, each of the episodes being associated with the control system performing a task in the operating environment based on a behaviour model trained in a baseline environment, and comprising a sequence of states and actions of the control system: in response to a novelty in the operating environment, adapting the behaviour model to the novelty in the operating environment; and using the adapted behaviour model to operate the control system in the operating environment in relation to the task, wherein adapting the behaviour model comprises:(i) calculating an adaptation region in an embedded state space of the operating environment, the adaptation region including one or more compromised model states of the control system, wherein each of the one or more compromised model states is a state where an action of the control system is compromised in the operating environment in relation to the task, and wherein the action is determined by executing the behaviour model;(ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and(iii) updating the behaviour model by replacing the action of each of the one or more compromised model states with the adaptation action,wherein the one or more compromised model states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
2. The method of claim 1, wherein the behaviour model is a deep reinforcement learning model.
3. The method of claim 2, wherein updating the behaviour model involves retaining one or more actions of the control system associated with model states that are not within the adaptation region.
4. The method of any of claims 2 to 3, wherein the one or more compromised model states of the adaptation region include at least a current model state representing, in the embedded state space, a current state of the control system in the operating environment.
5. The method of any of claims 2 to 4, wherein the embedded state space is a dimensionally-reduced embedded state space.
6. The method of any of claims 4 to 5, wherein the one or more compromised model states of the adaptation region include at least one other model state additional to the current model state.
7. The method of claim 6, wherein the adaptation region is based on a cluster formed in the embedded state space at a location of the current model state.
8. The method of claim 6, wherein the adaptation region is calculated by partitioning the embedded state space based on a spatial relationship between the current model state and other states in the embedded state space.
9. The method of claim 8, wherein the adaptation region is calculated by constructing a Voronoi cell encompassing the current model state in the embedded state space.
10. The method of any of claims 8 to 9, wherein the adaptation region includes the current model state and is closed, and wherein the adaptation region is split by: i) selecting the adaptation action associated with the adaptation region; ii) evaluating the adaptation action with respect to the corresponding current state to determine a performance of the adaptation action; and iii) splitting the closed adaptation region into two regions if the performance of the adaptation action is below a threshold level.
11. The method of claim 10, wherein steps (i) to (iii) are iteratively repeated until the adaptation region is smaller than a predefined threshold, or an overall performance of the control system reaches a predefined threshold.
12. The method of any of claims 10 to 11, wherein the performance of each adaptation action is determined using an evaluation function that assigns a score to the corresponding action of the state as determined by the behaviour model.
13. The method of any of claims 1 to 12, wherein determining the adaptation action comprises: selecting the adaptation function from a plurality of candidate functions; and applying the selected adaptation function to the current model state.
14. The method of claim 13, wherein each adaptation function maps a set of embedded states corresponding to the adaptation region to a set of candidate actions from an action space of the operating environment.
15. The method of claim 14, wherein the selected adaptation function outputs the adaptation action by evaluating the set of candidate actions in response to an input of the current model state.
16. The method of any of claims 1 to 15, wherein using the behaviour model comprises: i) determining the current state of the control system in the operating environment; ii) evaluating the behaviour model on the current state to determine anaction of the control system; and iii) processing one or more parameters of the action to generate signals to control the control system in the operating environment.
17. A control system comprising: one or more input components configured to generate input data associated with an operating environment of the control system; one or more operational components configured to, in response to receiving corresponding control signals, cause the control system to perform a task in the operating environment; and an intelligent controller device configured to receive the input data from the one or more input components and transmit the control signals to the one or more operational components, and comprising at least: a data storage medium configured to store a behaviour model; and one or more processors configured to operate the control system based on the behaviour model by executing the method of any of claims 1 to 16.
18. A computer-readable storage medium having program code that is executable by a processor device to cause a computing device to perform operations, the operations comprising: during performance of one or more of episodes of operating a control system in an operating environment, each of the episodes being associated with the control system performing a task in the operating environment based on a behaviour modeltrained in a baseline environment, and comprising a sequence of states and actions of the control system: in response to a novelty in the operating environment, adapting the behaviour model to the novelty in the operating environment; and using the adapted behaviour model to operate the control system in the operating environment in relation to the task, wherein adapting the behaviour model comprises:(i) calculating an adaptation region in an embedded state space of the operating environment, the adaptation region including one or more compromised model states of the control system, wherein each of the one or more compromised model states is a state where an action of the control system is compromised in the operating environment in relation to the task, and wherein the action is determined by executing the behaviour model;(ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and(iii) updating the behaviour model by replacing the action of each of the one or more compromised model states with the adaptation action, wherein the one or more compromised model states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
19. A method for adapting a behaviour neural network used to operate a control system in an environment, the method comprising:executing the behaviour neural network during one or more episodes, each episode comprising a sequence of states and actions of the control system to perform a task, and in response to a change in the environment during the execution of the behaviour neural network:(i) calculating an adaptation region in an embedded state space of the environment, the adaptation region including one or more compromised network- embedded states of the control system, wherein each of the one or more compromised network-embedded states is a state where an action of the control system is compromised in the environment in relation to the task, and wherein the action is determined by executing the behaviour neural network;(ii) determining, using an adaptation function, an adaptation action associated with the adaptation region; and(iii) updating the behaviour neural network by replacing the action of each of the one or more compromised network-embedded states with the adaptation action, wherein the one or more compromised network-embedded states are localized in a region of the embedded state space based on a degree of semantic similarity of the control system performing the task.
Citation Information
Patent Citations
Optimizing control actions of a control system via automatic dimensionality reduction of a mathematical representation of the control system
US20210350049A1
Agent behavior model for simulation control
US20210370972A1
Task performing agent systems and methods
US20220075383A1
Automated cruise control system
US20220219691A1
Process controller with META-reinforcement learning
US20220291642A1
Cited By
Path planning method based on deep reinforcement learning-fast exploration random tree
CN121028798A