Action Pruning by Logical Neural Networks
The integration of logical neural networks in reinforcement learning enables action pruning by evaluating logical inferences and calculating action probabilities, enhancing efficiency and reducing search space complexity.
Patent Information
- Application Number
- JP2023531653
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-24
- Filing Date
- 2021-10-22
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2041-10-22
AI Technical Summary
Existing reinforcement learning methods lack effective integration of logical rules and do not utilize logical neural networks (LNNs) for action pruning, which is crucial for reducing the branching factor of the search space in Markov decision processes.
A method and system utilizing a logical neural network (LNN) to evaluate logical inferences, calculate upper and lower bounds for actions, and determine their probabilities, thereby pruning actions that violate a policy in reinforcement learning.
Enhances reinforcement learning by dynamically pruning irrelevant actions, improving efficiency and reducing the search space complexity through logical inference and probability calculation.
Smart Images

Figure 0007698385000009 
Figure 0007698385000010 
Figure 0007698385000011
Abstract
Description
Technical Field
[0001] The present invention generally relates to artificial learning, and more particularly to action pruning by a logical neural network.
Background Art
[0002] A logical neural network (LNN) can be trained using logical functions such as NOT, AND, OR. An LNN has weights, activation functions, backward functions, and gradients. Such a structure is a state-of-the-art neural network for this type of training. Logical information seems to be useful for reinforcement learning in general, but there is no method of integration.
[0003] Action pruning (AP) is a technique that dynamically removes actions from a Markov decision process (MDP) and reduces the branching factor of the search space. However, there is no AP method that includes the following items: defining logical rules as human input for AP; calculating the probability of action candidates by an additional network generated from human input; using these probabilities as Q-values or weights of random actions.
Summary of the Invention
[0004] According to an aspect of the present invention, a computer-implemented method for action pruning in reinforcement learning is provided. The method includes receiving a current state of the environment. The method further includes evaluating a logical inference based on the current state of the environment using a logical neural network (LNN) structure. The method also includes outputting an upper bound and a lower bound for each action from a set of possible actions of an agent in the environment in response to the evaluation of the logical inference. The method further includes calculating a probability for each pair of a possible action of the agent in the environment and the current state of the environment by using the upper bound and the lower bound. Each of the calculated probabilities indicates a respective priority of each action. The method further includes obtaining a policy in reinforcement learning for the current state of the environment by using the calculated probabilities. The method also includes pruning one or more actions from the set of actions as ones that violate the policy such that the one or more actions are ignored.
[0005] According to another aspect of the present invention, there is provided a computer program product for action pruning in reinforcement learning. The computer program product includes a non-transitory computer-readable storage medium having program instructions embodied therein. The program instructions are executable by a computer and cause the computer to perform a method. The method includes receiving a current state of the environment. The method further includes evaluating a logical inference based on the current state of the environment using a logical neural network (LNN) structure. The method also includes outputting an upper bound and a lower bound for each action from a set of possible actions of an agent in the environment in response to the evaluation of the logical inference. The method further includes calculating a probability for each pair of a possible action of the agent in the environment and the current state of the environment by using the upper bound and the lower bound. Each of the calculated probabilities indicates a respective priority of each action. The method further includes obtaining a policy in reinforcement learning for the current state of the environment by using the calculated probabilities. The method also includes pruning one or more actions from the set of actions as violating the policy such that the one or more actions are ignored.
[0006] According to yet another aspect of the present invention, a computer processing system for safe reinforcement learning is provided. The computer processing system includes a storage device for storing program code. The computer processing system further includes one or more hardware processing units for executing program code that receives the current state of the environment. The one or more hardware processing units further execute program code that evaluates logical inferences based on the current state of the environment using a logical neural network (LNN) structure. The one or more hardware processing units also execute program code that outputs upper and lower limits for each action from a set of possible actions of an agent in the environment in response to the evaluation of the logical inference. The one or more hardware processing units further execute program code that calculates probabilities for each pair of possible actions of an agent in the environment and the current state of the environment by using the upper and lower limits. Each of the calculated probabilities indicates the respective priority of each action. The one or more hardware processing units further execute program code that obtains a policy in reinforcement learning for the current state of the environment by using the calculated probabilities. The one or more hardware processing units also execute program code that prunes one or more actions from the set of actions as violating the policy such that the one or more actions are ignored.
[0007] These and other features and advantages will become apparent from the following detailed description of its exemplary embodiments, read in conjunction with the accompanying drawings.
[0008] In the following description, details of the preferred embodiments are provided with reference to the following figures.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0010] Embodiments of the present invention are directed to action pruning by a logical neural network.
[0011] One or more embodiments of the present invention define a more complex structure for actions that are worse than the structures of the prior art.
[0012] Also, in one or more embodiments of the present invention, a logical neural network (LNN) that can be trained from a predetermined trajectory is used (other logical frameworks cannot be trained in that way).
[0013] A logical neural network (LNN) is a new representation method that simultaneously provides the main characteristics of both neural networks (learning) and symbolic logic (reasoning). LNN can incorporate domain knowledge and support composite first-order logical formulas. One or more embodiments of the present invention employ LNN to standardize knowledge induction, knowledge representation, and reasoning.
[0014] In one embodiment, the LNN is implemented as a form of a recurrent neural network that corresponds one-to-one to a set of logical formulas in any of various systems of weighted real-valued logic, and the evaluation performs logical inference. The features that distinguish the LNN from other neural networks are: (1) the neural activation functions are constrained to implement the truth functions of the logical operations they represent (i.e., ∧, ∨, ¬, →, ∀ and ∃ in FOL), (2) the results are represented at the boundaries of truth values to distinguish known, almost known, unknown, and contradictory states, and (3) bidirectional inference is allowed (e.g., in addition to being able to prove y when x is given, or similarly ¬x when ¬y is given, x→y is evaluated as usual.). The nature of the modeled logical system depends on the family of activation functions selected for the neurons of the network that implement the various atoms and operations of the logic.
[0015] FIG. 1 is a block diagram showing an exemplary computing device 100 according to one embodiment of the present invention. The computing device 100 is configured to perform action pruning (AP) by a logical neural network (LNN).
[0016] Computing device 100 may be implemented as any type of computing device or computer device capable of performing the functions described herein, including, but not limited to, a computer, server, rack-based server, blade server, workstation, desktop computer, laptop computer, notebook computer, tablet computer, mobile computing device, wearable computing device, network device, web device, distributed computing system, processor-based system, or consumer electronic device, or combinations thereof. Additionally, or alternatively, computing device 100 may be implemented as one or more compute threads, memory threads, or other components of a rack, thread, computing chassis, or physically disassembled computing device. As shown in FIG. 1, computing device 100 illustratively includes a processor 110, an input / output subsystem 120, a memory 130, a data storage device 140, and a communication subsystem 150, or other components and devices commonly found in a server or similar computing device, or combinations thereof. Of course, computing device 100 may include other or additional components in other embodiments, such as those commonly found in a server computer (e.g., various input / output devices). Further, in some embodiments, one or more of the exemplary components may be incorporated into or form part of another component. For example, memory 130, or a portion thereof, may be incorporated into processor 110 in some embodiments.
[0017] Processor 110 may be implemented as any type of processor capable of performing the functions described herein. Processor 110 may be implemented as a single processor, multiple processors, one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more single or multi-core processors, one or more digital signal processors, one or more microcontrollers, or other one or more processors or one or more processing / control circuits.
[0018] Memory 130 may be implemented as any type of volatile or non-volatile memory or data storage capable of performing the functions described herein. In operation, memory 130 may store various data and software used during the operation of computing device 100, such as an operating system, applications, programs, libraries, and drivers. Memory 130 may be communicatively coupled to processor 110 via I / O subsystem 120 and may be implemented as a circuit or component or both to facilitate input / output operations between processor 110, memory 130, and other components of computing device 100. For example, I / O subsystem 120 may be implemented as a memory controller hub, an input / output control hub, a platform controller hub, an integrated control circuit, a firmware device, a communication link (e.g., a point-to-point link, a bus link, a wire, a cable, a light guide, a printed circuit board trace, etc.), or other components and subsystems for facilitating input / output operations, or a combination thereof, or may otherwise include them. In some embodiments, I / O subsystem 120 may form part of a system-on-chip (SOC) and may be incorporated onto a single integrated circuit chip together with processor 110, memory 130, and other components of computing device 100.
[0019] The data storage device 140 may be implemented as any one or more types of devices configured for short-term or long-term storage of data, such as, for example, a memory device and circuitry, a memory card, a hard disk drive, a solid state drive, or other data storage devices. The data storage device 140 can store program code for action pruning (AP) by a logical neural network (LNN). The communication subsystem 150 of the computing device 100 may be implemented as any network interface controller or other communication circuitry, device, or collection thereof that can enable communication between the computing device 100 and other remote devices via a network. The communication subsystem 150 may be configured to effectuate such communication using any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, InfiniBand®, Bluetooth®, Wi-Fi®, WiMAX, etc.).
[0020] As illustrated, the computing device 100 may also include one or more peripheral devices 160. The peripheral devices 160 may include any number of additional input / output devices, interface devices, or other peripheral devices or combinations thereof. For example, in some embodiments, the peripheral devices 160 may include a display, a touch screen, a graphics circuit, a keyboard, a mouse, a speaker system, a microphone, a network interface, or other input / output devices, interface devices or peripheral devices or combinations thereof, or combinations thereof.
[0021] Of course, computing device 100 may also include other elements (not shown) and may omit certain elements, as would be readily contemplated by those of ordinary skill in the art. For example, various other input devices or output devices or both may be included in computing device 100 depending on the particular implementation, as would be readily understood by those of ordinary skill in the art. For example, various types of wireless or wired or both input devices, or output devices, or both may be used. Additionally, additional processors, controllers, memories, etc. of various configurations may also be utilized. Further, in another embodiment, a cloud configuration may be used (see, for example, FIGS. 9-10). These and other variations of processing system 100 would be readily contemplated by those of ordinary skill in the art in view of the teachings of the invention provided herein.
[0022] As used herein, the terms “hardware processor subsystem” or “hardware processor” can refer to a processor, memory (including RAM, (one or more) caches, etc.), software (including memory management software), or combinations thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements can be included in a central processing unit, a graphics processing unit, or a separate processor or arithmetic element-based controller (e.g., logic gates, etc.), or combinations thereof. The hardware processor subsystem can include one or more on-board memories (e.g., caches, dedicated memory arrays, read-only memories, etc.). In some embodiments, the hardware processor subsystem can include one or more memories (e.g., ROM, RAM, basic input / output system (BIOS), etc.) that can be on-board or off-board or dedicated for use by the hardware processor subsystem.
[0023] In some embodiments, the hardware processor subsystem can include and execute one or more software elements. The one or more software elements can include an operating system, or one or more applications or specific code or both, for achieving a specified result.
[0024] In other embodiments, the hardware processor subsystem can include a dedicated special circuit that executes one or more electronic processing functions to achieve a specified result. Such a circuit can include one or more application-specific integrated circuits (ASICs), FPGAs, or PLAs, or a combination thereof.
[0025] These and other variations of the hardware processor subsystem are also contemplated in accordance with embodiments of the present invention.
[0026] FIG. 2 is a block diagram showing an exemplary LNN graph structure 200 to which the present invention can be applied, according to an embodiment of the present invention.
[0027] The LNN graph structure reflects the mathematical formula it represents.
[0028] TIFF0007698385000001.tif13164
[0029] For example, when the beard is TRUE, the upper and lower limit values have a high value such as ~1.0. When the beard is FALSE, the upper and lower limit values have a low value such as ~0.0.
[0030] FIG. 3 is a block diagram showing exemplary pseudo-code 300 of a first algorithm for an upward pass, according to an embodiment of the present invention.
[0031] The pseudo-code 300 corresponds to an upward pass for inferring formula truth value boundaries for sub-logical formula boundaries. The pseudo-code 300 includes upward propagation of boundaries from leaves, negation, multi-input discrete, and tightening of existing boundaries.
[0032] FIG. 4 is a block diagram showing exemplary pseudo-code 400 of a second algorithm for a downward pass according to an embodiment of the present invention.
[0033] The pseudo-code 400 corresponds to a downward pass for inferring a formula truth value boundary for a partial formula boundary. The pseudo-code 400 includes negation, multi-input discrete, and downward boundary propagation to leaves.
[0034] FIG. 5 is a block diagram showing exemplary pseudo-code 500 of a third algorithm for a recursive inference procedure according to an embodiment of the present invention.
[0035] The pseudo-code 500 corresponds to a recursive inference procedure involving a traversal of a recursive directed graph. The pseudo-code 500 includes a loop until convergence, traversing from leaves to roots visiting the roots of all formulas in order, and traversing from roots to leaves.
[0036] Here, the logical neural network (LNN) will be further described. All neurons return a pair of values in the range 0 to 1 representing the lower and upper bounds of the truth values of the corresponding partial formulas and propositions. To facilitate the interpretation of the boundaries, a true threshold 1 / 2 < α < 1 is defined such that a continuous truth value is considered true if it is greater than α and false if it is less than 1 - α. The boundary values identify one of the four main states that a neuron can take, while the secondary state gives an interpretation more true or more false.
[0037] FIG. 6 is a block diagram showing an exemplary architecture 600 and corresponding signals according to an embodiment of the present invention.
[0038] The architecture 600 includes a semantic parser 610, a reinforcement learning element 620, a logical neural network (LNN) 630, an LNN action pruning element 640, and an environment 650.
[0039] The semantic parser 610 semantically analyzes the input agent state.
[0040] The reinforcement learning element 620 may be a Long Short Term Memory-Deep Q Network (LTSM-DQN). The reinforcement learning element 620 is a base reinforcement learning method that predicts candidate actions for safe reinforcement learning.
[0041] The LNN 630 is for understanding the logical function of safety restrictions.
[0042] The LNN action pruning element 640 is for avoiding unnecessary actions.
[0043] The environment 650 is the place where actions by the agent are performed.
[0044] The following signal definitions apply.
[0045] s t represents the state of the agent.
[0046] s t ' represents the semantically modified state of the agent.
[0047] act t represents the action at time t.
[0048] reward t represents the reward at time t.
[0049] TIFF0007698385000002.tif13167
[0050] TIFF0007698385000003.tif8145
[0051] Figures 7-8 show an exemplary method 700 according to an embodiment of the present invention.
[0052] In block 710, configure one or more hardware processing devices as a logical neural network (LNN) structure having a plurality of neurons and connection edges. The plurality of neurons and connection edges of the LNN structure correspond one-to-one with a system of logical formulas and execute a method for performing logical inference.
[0053] In block 720, for each corresponding logical connection in each formula of the system of logical formulas, configure at least one of the plurality of neurons. One neuron further has one or more link connection edges that provide information including input information containing the operands of the logical connection and parameters configured to implement the truth function of the logical connection. Each of the at least one neuron of the corresponding logical connection has a corresponding activation function for providing a calculation, and the calculation of the activation function returns a pair of values indicating the upper and lower limits with respect to the formula of the system formula, or returns the truth value of a proposition. The system formula is different from the logical formula in that the system formula has logical neurons and activation functions that are not in the logical formula.
[0054] In block 730, for each corresponding proposition of the formula of the system formula, configure at least one other neuron of the plurality of neurons. The at least one other neuron has one or more link connection edges corresponding to the formula that provide information for proving the boundary regarding the truth value of the corresponding proposition, and the information further includes parameters configured to aggregate the tightest limits. The term "aggregating the tightest limits" means collecting the tightest limits for a given action.
[0055] In block 740, receive the current state of the environment.
[0056] In block 750, use the logical neural network (LNN) structure to evaluate logical inferences based on the current state of the environment.
[0057] In block 760, upper and lower limits of each action are output from a set of possible actions of an agent in the environment according to the evaluation of logical inferences.
[0058] In block 770, for each pair of a possible action of an agent in the environment and the current state of the environment, probabilities are calculated by using the upper and lower limits, and each of the calculated probabilities indicates the respective priority of each action. In this specification, the term "priority" means a value for prioritizing taking the target action.
[0059] In block 780, a policy in reinforcement learning for the current state of the environment is obtained by using the calculated probabilities.
[0060] In block 790, one or more actions are pruned from the set of actions as ones that violate the policy so that the one or more actions are ignored (not executed by the agent in the environment).
[0061] Next, the definition of probability from the LNN according to an embodiment of the present invention will be described.
[0062] For an action a t defined or trained by a human or from a pair of an input state and an action, t and a predetermined state s, a probability is calculated from a logical neural network (LNN).
[0063] TIFF0007698385000004.tif74167
[0064] TIFF0007698385000005.tif20168
[0065] TIFF0007698385000006.tif31153
[0066] Next, the possible requirements of the LNN used according to the present invention will be described.
[0067] Each neuron (represented by a proposition) needs to have an upper limit value and a lower limit value, and these are logically connected to logical conjunction operators (AND, OR, IMPLY gates) by activation functions and weight values.
[0068] The input layer has several propositions for state inputs (in the present invention, the logical states for each environmental state input), the hidden layer has logical operators (which have several weight values), and the output layer has several propositions for actions.
[0069] The output layer does not need to set action values from the output of reinforcement learning. The output of the output layer (which is the output of the LNN) is used for action selection (including not only rejection but also recommendation).
[0070] In order to calculate the action value, propositions for all actions need to be set.
[0071] The parameters (that is, weights and bias values) are trainable during execution.
[0072] Next, action pruning according to an embodiment of the present invention will be further described.
[0073] TIFF0007698385000007.tif42167 Here, a is the action to be targeted for calculating the probability, and A is all actions. The value v(a;s t ) represents the level of the truth value of the proposition after discounting the conflicting value.
[0074] TIFF0007698385000008.tif57167
[0075] This disclosure includes a detailed description regarding cloud computing, but it is understood that the implementations of the teachings described herein are not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented with any other type of computer environment now known or developed in the future.
[0076] Cloud computing is a service delivery model for enabling convenient and on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), where the resources can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0077] The characteristics are as follows.
[0078] On-demand self-service: A cloud consumer can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider.
[0079] Broad network access: Computing capabilities are available over a network and can be accessed via standard mechanisms, thereby facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).
[0080] Resource pooling: The computing resources of a provider are pooled and provided to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated according to demand. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the provided resources. However, consumers may be able to specify a location at a higher level of abstraction (e.g., country, state, data center).
[0081] Rapid elasticity: Computing capabilities can be prepared quickly and flexibly, so in some cases, it can automatically scale out immediately and be released promptly to scale in immediately. To consumers, the computing capabilities available for preparation often seem unlimited, and they can purchase any quantity at any time.
[0082] Measured service: Cloud systems utilize a measurement function at a certain level of abstraction suitable for the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. It is possible to monitor, control, and report resource usage to provide transparency to both the provider and the consumer of the service being utilized.
[0083] The service model is as follows.
[0084] Software as a Service (SaaS): The function provided to consumers is that they can use the provider's applications running on cloud infrastructure. The applications can be accessed from various client devices via a client interface such as a web browser (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, and even individual application functions. However, this does not apply to limited settings of user-specific application configurations.
[0085] Platform as a Service (PaaS): The function provided to consumers is to deploy the applications created or obtained by consumers to the cloud infrastructure using the programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, and storage, but can control the deployed applications and, in some cases, also control the configuration of the hosting environment.
[0086] Infrastructure as a Service (IaaS): The function provided to consumers is to prepare processors, storage, networks, and other basic computing resources that allow consumers to deploy and run any software that may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but can control the operating systems, storage, and deployed applications, and in some cases, can also partially control some network components (e.g., host firewalls).
[0087] The deployment models are as follows.
[0088] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0089] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0090] Public Cloud: This cloud infrastructure is provided to an unspecified number of people or large industry groups and is owned by an organization that sells cloud services.
[0091] Hybrid Cloud: This cloud infrastructure is a combination of two or more cloud models (private, community, or public). It retains the entities specific to each model but is bound by standard or individual technologies to achieve data and application portability (e.g., cloud bursting for load distribution between clouds).
[0092] The cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0093] Referring to FIG. 9, an exemplary cloud computing environment 950 is depicted. The cloud computing environment 950 includes one or more cloud computing nodes 910. In contrast, local computer devices used by cloud consumers (e.g., PDA or mobile phone 954A, desktop computer 954B, laptop computer 954C, or automotive computer system 954N or combinations thereof, etc.) can communicate. The nodes 910 can communicate with each other. The nodes 910 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, the private, community, public, or hybrid clouds described above or combinations thereof. Thereby, the cloud computing environment 950 can provide infrastructure, platform, software, or combinations thereof as a service, and cloud consumers do not need to maintain resources on local computer devices. Note that the types of computer devices 954A - N shown in FIG. 9 are merely exemplary, and it should be understood that the computing nodes 910 and the cloud computing environment 950 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.
[0094] Referring to FIG. 10, a set of functional abstraction model layers provided by the cloud computing environment 950 is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 10 are merely exemplary, and the embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided.
[0095] The hardware and software layer 1060 includes hardware components and software components. Examples of hardware components include mainframe 1061, a server 1062 based on a reduced instruction set computer (RISC) architecture, server 1063, blade server 1064, storage device 1065, and network and network components 1066. In some embodiments, the software components include network application server software 1067 and database software 1068.
[0096] The virtualization layer 1070 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual server 1071, virtual storage 1072, virtual network 1073 including a virtual private network, virtual applications and operating systems 1074, and virtual client 1075.
[0097] As an example, the management layer 1080 can provide the following functions. Resource preparation 1081 enables the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 1082 enables cost tracking when resources are utilized within a cloud computing environment, and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only the protection of data and other resources, but also the identification and authentication of cloud consumers and tasks. The user portal 1083 provides access to the cloud computing environment for consumers and system administrators. Service level management 1084 enables the allocation and management of cloud computing resources so that the required service level is met. Planning and fulfillment of service quality assurance (SLA) 1085 enables the advance arrangement and procurement of cloud computing resources that are expected to be needed in the future according to the SLA.
[0098] The workload layer 1090 provides examples of functions available in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 1091, software development and lifecycle management 1092, delivery of virtual classroom education 1093, data analysis processing 1094, transaction processing 1095, and secure reinforcement learning by LNN 1096.
[0099] The present invention can be a system, method, computer program product, or a combination thereof integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0100] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. As a more specific example of a computer-readable storage medium, there can be a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punch card, a mechanically encoded device that records instructions in a raised structure in a groove, and suitable combinations thereof. A computer-readable storage device as used herein should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.
[0101] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network or a combination thereof). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers or a combination thereof. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0102] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or an object-oriented programming language such as SMALLTALK® or C++, and a procedural programming language such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executable entirely on the user's computer as a stand-alone software package, or partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize for carrying out aspects of the present invention.
[0103] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0104] These computer-readable program instructions can be provided to a general-purpose computer, a processor of a special-purpose computer, or other programmable data processing apparatus to generate a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus produce means for implementing the functions / operations specified in one or more blocks of a flowchart, a block diagram, or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing apparatus, or other devices that function in a particular manner or a combination thereof, such that the computer-readable program instructions stored therein constitute one of the manufactured articles that include instructions for implementing the aspects of the functions / acts specified in one or more blocks of a flowchart, a block diagram, or both.
[0105] Like instructions that execute functions / acts specified in one or more blocks of a flowchart, a block diagram, or both on a computer, other programmable apparatus, or other device, the computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device and generate a computer-implemented process.
[0106] The flowcharts and block diagrams in the figures illustrate the structure, functionality, and operation of the implementation that can be executed by systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which constitutes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions shown in the blocks may occur in a different order than shown in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may be executed in the reverse order depending on the related functions. It should also be noted that each block of the block diagram or flowchart diagram, or both, and combinations of blocks of the block diagram or flowchart diagram, or both, can be implemented by a special purpose hardware-based system that executes the specified function or operation, or a combination of special purpose hardware and computer instructions.
[0107] As used herein, references to "one embodiment", "an embodiment", and other variations of the present invention mean that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" and any other variations throughout this specification do not necessarily all refer to the same embodiment.
[0108] It should be understood that any use of the following, namely, “ / ”, “and / or”, “at least one of”, e.g., in “A / B”, “A and / or B”, “at least one of A and B”, is intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of “A, B, and / or C” and “at least one of A, B, and C”, such expressions are intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This can be extended for any number of listed items, as will be readily understood by those having ordinary skill in the art of this and related arts.
[0109] Preferred embodiments of the systems and methods (which are intended to be illustrative and not limiting) have been described, but it is noted that modifications and variations can be made by those skilled in the art in light of the above teachings. Accordingly, it will be understood that changes can be made within the scope of the invention as outlined by the appended claims in the disclosed specific embodiments. Thus, while the aspects of the invention have been described with the particularity and detail required by patent law, what is claimed and desired to be protected by patent is set forth in the appended claims.
Claims
1. A computer-implemented method for action pruning in reinforcement learning, comprising: receiving a current state of an environment; evaluating logical inferences based on the current state of the environment using a logical neural network (LNN) structure; outputting upper and lower bounds for each action from a set of possible actions of an agent in the environment in response to the evaluation of the logical inferences; calculating probabilities for each pair of a possible action of the agent in the environment and the current state of the environment by using the upper and lower bounds, wherein each calculated probability indicates a respective priority of each action; obtaining a policy in reinforcement learning for the current state of the environment by using the calculated probabilities; pruning one or more actions from the set of actions as violating the policy such that the one or more actions are ignored. A computer-implemented method as described above.
2. The computer-implemented method according to claim 1, wherein each pair of the possible action of the agent in the environment and the current state of the environment is defined by a human.
3. The computer-implemented method according to claim 1, wherein each pair of the possible action of the agent in the environment and the current state of the environment is trained from input state-action pairs.
4. The computer-implemented method according to claim 1, wherein the probabilities are calculated by further using logical rule conflict values in addition to using the upper and lower bounds, and each logical rule conflict value represents a level of conflict for each of a plurality of logical rules associated with the LNN.
5. The computer-implemented method according to claim 4, wherein a conflict includes having a lower bound value higher than an upper bound value.
6. The computer-implemented method according to claim 1, further comprising performing exploration in the environment in response to the policy.
7. The computer-implemented method according to claim 1, further comprising assisting boundary interpretability by using a true threshold 1 / 2 < α < 1 such that continuous truth values are considered true when the continuous truth values are greater than α and considered false when the continuous truth values are less than 1 - α.
8. Configuring one or more hardware processing devices as the LNN structure having a plurality of neurons and connection edges, wherein the plurality of neurons and connection edges of the LNN structure correspond one-to-one with a system of logical formulas and execute a method for performing the logical inference, further comprising, at least one neuron of the plurality of neurons is related to a corresponding logical conjunction in each formula of the system of logical formulas, and the at least one neuron further comprises input information including operands of the corresponding logical conjunction and information including parameters configured to implement a truth function of the corresponding logical conjunction, and has one or more link connection edges providing the information, and each of the at least one neuron has a corresponding activation function for providing a calculation, and the calculation of the activation function returns a pair of values indicating an upper limit and a lower limit regarding a formula of a system formula, or returns a truth value of a proposition of the formula of the system formula, The computer-implemented method according to claim 1.
9. The computer-implemented method according to claim 8, wherein at least one other neuron of the plurality of neurons is related to the proposition, and the at least one other neuron has one or more link connection edges corresponding to a mathematical formula providing information for proving upper and lower limits regarding a truth value of the corresponding proposition and information including parameters configured to aggregate the tightest limits.
10. A computer program for action pruning in reinforcement learning, the computer program including program instructions executable by a computer to cause the computer to execute a method that receives a current state of an environment, evaluates a logical inference based on the current state of the environment using a logical neural network (LNN) structure, outputs an upper limit and a lower limit of each action from a set of possible actions of an agent in the environment according to the evaluation of the logical inference Calculating probabilities for each pair of possible actions of the agent in the environment and the current state of the environment by using the upper limit and the lower limit, where each of the calculated probabilities indicates the respective priority of each action; Obtaining a policy in reinforcement learning for the current state of the environment by using the calculated probabilities; Pruning one or more actions from the set of actions as those that violate the policy so that the one or more actions are ignored; A computer program including the above.
11. The computer program according to claim 10, wherein each pair of the possible actions of the agent in the environment and the current state of the environment is defined by a human.
12. The computer program according to claim 10, wherein each pair of the possible actions of the agent in the environment and the current state of the environment is trained from input state-action pairs.
13. The probability is calculated by further using logical rule conflict values by using the upper limit and the lower limit, and each of the logical rule conflict values represents the level of conflict for each of a plurality of logical rules related to the LNN. The computer program according to claim 10.
14. The computer program according to claim 13, wherein the conflict includes having a lower limit value higher than the upper limit value.
15. The computer program according to claim 10, performing exploration in the environment in response to the policy.
16. The computer program according to claim 10, further including assisting boundary interpretability by using a true threshold 1 / 2 < α < 1 such that continuous truth values are regarded as true when the continuous truth values are greater than α and are regarded as false when the continuous truth values are less than 1 - α.
17. Configuring one or more hardware processing devices as the LNN structure having a plurality of neurons and connection edges, wherein the plurality of neurons and connection edges of the LNN structure correspond one-to-one with a system of logical formulas and execute a method for performing the logical inference; further including At least one of the plurality of neurons is related to a corresponding logical conjunction in each formula of the system of logical formulas, and the at least one neuron further includes input information including operands of the corresponding logical conjunction and information including parameters configured to implement a truth function of the corresponding logical conjunction, and has one or more link edges that provide the information, and each of the at least one neuron has a corresponding activation function for providing a calculation, and the calculation of the activation function returns a pair of values indicating an upper limit and a lower limit regarding the formula of the system formula, or returns a truth value of a proposition of the formula of the system formula. The computer program according to claim 10.
18. A computer processing system for safe reinforcement learning, a storage device for storing program code, one or more hardware processing units for executing the program code, and the program code includes receiving a current state of the environment, evaluating a logical inference based on the current state of the environment using a logical neural network (LNN) structure, outputting an upper limit and a lower limit of each action from a set of possible actions of an agent in the environment according to the evaluation of the logical inference, for each pair of a possible action of the agent in the environment and the current state of the environment, calculating a probability by using the upper limit and the lower limit, wherein each of the calculated probabilities indicates a respective priority of each action, acquiring a policy in reinforcement learning for the current state of the environment by using the calculated probabilities, pruning one or more actions from the set of actions as those violating the policy so that the one or more actions are ignored, A computer processing system that executes.
19. The computer processing system according to claim 18, wherein each pair of the possible action of the agent in the environment and the current state of the environment is defined by a human.
20. The computer processing system according to claim 18, wherein each pair of the possible actions of the agent and the current state of the environment in the environment is trained from the input pair of state and action.
Citation Information
Patent Citations
Method of improving motion of robot
JP2016196079A
Reinforcement learning device
JP2020034994A