Multiple prediction network

Through multiple prediction networks and shared weight technology, the control and modular problems of artificial intelligence agents in skills and knowledge training are solved, learning speed and computing efficiency are improved, and efficient prediction of different states and skills are achieved.

CN113228063BActive Publication Date: 2025-07-08SONY CORP OF AMERICA +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202080007396.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-04
Filing Date
2020-01-02
Publication Date
2025-07-08
Estimated Expiration
2040-01-02

AI Technical Summary

Technical Problem

In the prior art, artificial intelligence agents lack control ability in skills and knowledge training, and are difficult to modularly and layered to learn higher-level skills and knowledge, and fail to effectively predict empirical characteristics.

Method used

Multi-head prediction network, multi-input prediction network, multi-skill prediction network and parameterized skill prediction network are adopted to improve training efficiency and generalization ability by receiving environmental status information and additional inputs.

Benefits of technology

It realizes efficient prediction of different states and skills, improves the learning speed and computing efficiency of artificial intelligence agents, and enhances the generalization ability of unseen inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113228063B_ABST
    Figure CN113228063B_ABST
Patent Text Reader

Abstract

Methods and systems for training and / or operating an artificial intelligence agent can use multi-input and / or multi-prediction networks. A multi-prediction is a computational construct where the network weights that are shared can be used to compute multiple related predictions, typically but not necessarily neural networks. This allows for more efficient training based on the amount of data and / or experience desired, and in some cases, allows for more efficient computation of those predictions. There are several related and sometimes combinable schemes for multi-prediction networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present invention generally relate to intelligent artificial agents. More particularly, the present invention relates to training intelligent artificial agents through multi-forecasts and / or methods for making forecast calculations more efficient. Background Art

[0002] The following background information may present examples of specific aspects of the prior art (e.g., but not limited to, solutions, facts, or common knowledge), and although these examples are expected to help further educate the reader about other aspects of the prior art, they should not be construed as limiting the present invention or any of its embodiments to anything stated or implied or inferred therefrom.

[0003] Forecasts are predictions that are useful in many kinds of artificial intelligence (AI) systems. A forecast is a prediction of an outcome based on the state of the world and conditioned on the skills or behaviors performed by an agent. Forecasts can be used to predict the outcome of a current behavior in the current state or to make hypothetical predictions conditioned on hypothetical behaviors for planning purposes. Examples of forecasts include the distance to the termination of a skill, the time to the termination of a skill, the value of a state feature at the termination of a skill, etc.

[0004] Currently known systems for training artificial agents exhibit various problems. In many cases, users lack the ability to control the skills and knowledge learned by the agent, or such learned skills and knowledge may be items that the user does not consider as important as other desired skills and knowledge. Moreover, conventional systems may lack the ability to hierarchically organize skills and knowledge in a modular way for learning higher-level skills and knowledge. Also, in conventional systems, artificial agents may not learn a specific form of knowledge, namely, the prediction of empirical features during skill execution.

[0005] In view of the foregoing, there is a need to improve the training of skills and knowledge in artificial intelligence agents. Summary of the Invention

[0006] Embodiments of the present invention provide a multi-head prediction method for creating artificial intelligence in machines and computer-based software applications, the method comprising: receiving an input from an environment as state information; and outputting a plurality of forecasts, each of the plurality of forecasts corresponding to a different state information feature.

[0007] Embodiments of the present invention also provide a multi-input prediction method for creating artificial intelligence in machines and computer-based software applications, the method comprising: receiving an input from the environment as state information; receiving additional input from at least one of a prediction ID, a skill ID, and a parameter value; and outputting a prediction for each additional input.

[0008] Embodiments of the present invention also provide a prediction network method for creating artificial intelligence in machines and computer-based software applications, the method comprising: receiving an input from the environment as state information; receiving additional input from at least one of a prediction ID, a skill ID, and a parameter value; embedding the additional input into a learned reduced vector representation before the additional input is input into the prediction network; and outputting a prediction for each learned reduced vector representation.

[0009] These and other features, aspects, and advantages of the present invention will be better understood with reference to the following drawings, description, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Some embodiments of the present invention are illustrated by way of example and are not limited by the figures of the drawings, in which like reference numerals may indicate similar elements.

[0011] Figure 1A Illustrates a multi-head prediction network according to an exemplary embodiment of the present invention;

[0012] Figure 1B Illustrates an example of weighting of input nodes of a neural network;

[0013] Figure 2 Illustrates a multi-input prediction network according to an exemplary embodiment of the present invention;

[0014] Figure 3 Illustrates a multi-skill prediction network according to an exemplary embodiment of the present invention;

[0015] Figure 4 Illustrates a parameterized skill prediction network according to an exemplary embodiment of the present invention;

[0016] Figure 5 Illustrates a hybrid skill ID and multi-prediction network according to an exemplary embodiment of the present invention; and

[0017] Figure 6 Illustrates embedding using a prediction ID in a multi-prediction network according to an exemplary embodiment of the present invention.

[0018] Unless otherwise indicated, the illustrations in the various figures are not necessarily drawn to scale.

[0019] Now, the present invention and its various embodiments can be better understood by turning to the following detailed description of the illustrated embodiments. It should be clearly understood that the illustrated embodiments are merely set forth by way of example and are not a limitation of the invention as finally defined in the claims. Detailed Description

[0020] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the invention. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well as the singular forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components and / or groups thereof.

[0021] All terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs, unless otherwise defined. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0022] In describing the present invention, it will be understood that many techniques and steps are disclosed. Each of these techniques and steps has its own advantages, and each can also be used in combination with one or more, or in some cases all, of the other disclosed techniques. Therefore, for clarity, this description will avoid repeating each possible combination of the individual steps in an unnecessary manner. However, the specification and claims should be read with the understanding that such combinations are fully within the scope of the present invention and the claims.

[0023] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be clear to one of ordinary skill in the art that the present invention may be practiced without these specific details.

[0024] Unless otherwise expressly stated, devices or system modules that at least generally communicate with each other need not communicate with each other continuously. Additionally, devices or system modules that at least generally communicate with each other can communicate directly or indirectly through one or more intermediaries.

[0025] The description of embodiments having several components communicating with each other does not imply that all such components are required. Instead, various alternative components are described to illustrate the many possible embodiments of the present invention.

[0026] As is well known to those skilled in the art, when designing the best configuration for the commercial implementation of any system, especially embodiments of the present invention, many careful considerations and compromises must typically be made. A commercial implementation in accordance with the spirit and teachings of the present invention can be configured according to the needs of a particular application, whereby any (one or more) aspects, (one or more) features, (one or more) functions, (one or more) results, (one or more) components, (one or more) solutions, or (one or more) steps of the teachings associated with any described embodiment of the present invention can be appropriately omitted, included, adapted, mixed and matched, or improved and / or optimized by those skilled in the art using their average skills and known techniques to achieve a desired implementation that meets the needs of the particular application.

[0027] "Computer" can refer to one or more devices and / or one or more systems that are capable of accepting structured input, processing the structured input according to specified rules, and producing a processing result as output. Examples of computers can include: computers; fixed and / or portable computers; computers having a single processor, multiple processors, or multi-core processors that can operate in parallel and / or not in parallel; general-purpose computers; supercomputers; mainframes; superminicomputers; minicomputers; workstations; microcomputers; servers; clients; interactive televisions; web appliances; telecommunications devices having Internet access; hybrid combinations of computers and interactive televisions; portable computers; tablet personal computers (PCs); personal digital assistants (PDAs); portable telephones; dedicated hardware for emulating a computer and / or software, such as, for example, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), one chip, multiple chips, systems-on-a-chip, or chip sets; graphics processing units (GPUs); data acquisition devices; optical computers; quantum computers; biometric computers; and generally devices that can accept data, process the data according to one or more stored software programs, generate results, and typically include input, output, storage, arithmetic, logic, and control units.

[0028] Those skilled in the art will recognize that, in appropriate circumstances, some embodiments of the present disclosure may be practiced in a network computing environment having many types of computer system configurations, including personal computers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. In appropriate circumstances, embodiments may also be practiced in a distributed computing environment where tasks are performed by local and remote processing devices linked through a communication network (either hardwired, wirelessly, or by a combination thereof). In a distributed computing environment, program modules may be located in local and remote memory devices.

[0029] "Software" can refer to the prescribed rules for operating a computer. Examples of software can include segments of code in one or more computer-readable languages; graphical and / or textual descriptions; applets; precompiled code; interpreted code; compiled code; and computer programs.

[0030] The example embodiments described herein can be implemented in an operating environment including computer-executable instructions (e.g., software) installed on a computer, in hardware, or in a combination of software and hardware. Computer-executable instructions can be written in a computer programming language or can be implemented in firmware logic. If written in a programming language conforming to recognized standards, such instructions can be executed on various hardware platforms and can be used to interface with various operating systems. Although not limited thereto, computer software program code for performing aspects of the present invention can be written in any combination of one or more suitable programming languages, including object-oriented programming languages and / or conventional procedural programming languages, and / or programming languages such as, for example, HyperText Markup Language (HTML), Dynamic HTML, Extensible Markup Language (XML), Extensible Stylesheet Language (XSL), Document Style Semantics and Specification Language (DSSSL), Cascading Style Sheets (CSS), Synchronized Multimedia Integration Language (SMIL), Wireless Markup Language (WML), Java.TM., Jini.TM., C, C++, Smalltalk, Python, Perl, UNIX Shell, Visual Basic or Visual Basic Script, Virtual Reality Markup Language (VRML), ColdFusion.TM or other compilers, assemblers, interpreters, or other computer languages or platforms.

[0031] Computer program code for performing the operations of aspects of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc. and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). The program code may also be distributed among a plurality of computing units, where each unit processes a part of the overall computation.

[0032] Aspects of the present invention are described below with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or one or more block diagrams.

[0033] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified (one or more) logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may in fact be executed substantially simultaneously, or may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware for performing the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0034] These computer program instructions may also be stored in a computer-readable medium, which can direct a computer, other programmable data processing apparatus, or other devices to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture, the article of manufacture including instructions for implementing the functions / acts specified in the flowchart and / or one or more block diagrams.

[0035] In addition, while process steps, method steps, algorithms, etc. may be described in a sequential order, such processes, methods, and algorithms may be configured to work in an alternative order. In other words, any sequence or order of steps that may be described does not necessarily indicate that the steps are required to be performed in that order. The steps of the processes described herein may be performed in any practical order. Additionally, some steps may be performed simultaneously.

[0036] It is clear that the various methods and algorithms described herein may be implemented by, for example, appropriately programmed general-purpose computers and computing devices. Generally, a processor (e.g., a microprocessor) will receive instructions from a memory or similar device and execute those instructions, thereby performing the processes defined by those instructions. Additionally, various known media may be used to store and transmit programs that implement these methods and algorithms.

[0037] As used herein, the term "computer-readable medium" refers to any medium that participates in providing data (e.g., instructions) that can be read by a computer, processor, or similar device. Such media may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks and other permanent storage. Volatile media includes dynamic random access memory (DRAM), which typically constitutes main memory. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a system bus coupled to a processor. Transmission media may include or convey acoustic waves, light waves, and electromagnetic radiation, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, DVD, any other optical medium, punched cards, paper tape, any other physical medium with hole patterns, RAM, PROM, EPROM, FLASH EPROM, EEPROM, or any other memory chip or cartridge, a carrier wave as described below, or any other medium from which a computer can read.

[0038] Carrying an instruction sequence to a processor may involve various forms of computer-readable media. For example, the instruction sequence (i) may be transferred from RAM to the processor, (ii) may be carried on a wireless transmission medium, and / or (iii) may be formatted according to various formats, standards, or protocols, such as Bluetooth, TDMA, CDMA, 3G.

[0039] Embodiments of the present invention may include apparatus for performing the operations disclosed herein. The apparatus may be specially constructed for the desired purpose or may include a general-purpose device selectively activated or reconfigured by a program stored in the device.

[0040] Embodiments of the present invention may also be implemented in one or a combination of hardware, firmware, and software. They may be implemented as instructions stored on a machine-readable medium that can be read and executed by a computing platform to perform the operations described herein.

[0041] More specifically, as those skilled in the art will recognize, aspects of the present invention may be implemented as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may generally be referred to herein as a "circuit", "module", or "system". Additionally, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0042] In the following description and claims, the terms "computer program medium" and "computer-readable medium" may be used to generally refer to media such as, but not limited to, removable storage drives, hard disks installed in hard disk drives, etc. These computer program products may provide software to a computer system. Embodiments of the present invention may be directed to such computer program products.

[0043] Embodiments within the scope of the present disclosure may also include tangible and / or non-transitory computer-readable storage media for carrying or having computer-executable instructions or data structures stored thereon. Such non-transitory computer-readable storage media may be any available media accessible by a general-purpose or special-purpose computer, including the functional design of any of the special-purpose processors discussed above. By way of example and not limitation, such non-transitory computer-readable media may include RAM, ROM, EEPROM, CDROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code components in the form of computer-executable instructions, data structures, or processor chip designs. When information is transmitted or provided to a computer through a network or other communication connection (wired, wireless, or a combination thereof), the computer appropriately views the connection as a computer-readable medium. Thus, any such connection is appropriately referred to as a computer-readable medium. The above combinations should also be included within the scope of computer-readable media.

[0044] Although non-transitory computer-readable media include, but are not limited to, hard disk drives, optical discs, flash memory, volatile memory, random access memory, magnetic memory, optical memory, semiconductor-based memory, phase change memory, optical memory, memory that is refreshed periodically, etc.; however, non-transitory computer-readable media do not include pure transient signals; that is, the medium itself is transient.

[0045] Here, an algorithm is generally considered to be a self-consistent sequence of actions or operations that lead to a desired result. These include physical manipulations of physical quantities. Usually, although not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. For mainly general reasons, it has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. However, it should be understood that all of these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities.

[0046] Unless otherwise specifically stated, and as will be apparent from the following description and claims, it should be recognized that throughout the specification, descriptions using terms such as "processing", "computing", "calculating", "determining", etc. refer to actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or transform data represented as physical (such as electronic) quantities within the registers and / or memory of the computing system into other data similarly represented as physical quantities within the memory, registers, or other such information storage devices, transmission, or display devices of the computing system.

[0047] In a similar manner, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory or can be transmitted to an external device to cause a physical change or actuation of the external device. A "computing platform" can include one or more processors.

[0048] The term "robot" or "agent" or "intelligent agent" or "artificial agent" or "artificial intelligence agent" can refer to any system directly or indirectly controlled by a computer or computing system that emits actions or commands in response to sensing or observation. The term can refer to, but is not limited to, traditional physical robots with physical sensors (such as cameras, touch sensors, distance sensors, etc.), or simulated robots existing in a virtual simulation, or "bots" such as email bots or search bots that exist as software on a network. It can, but is not limited to, refer to any limb robot, walking robot, industrial robot (including, but not limited to, robots for automated assembly, painting, repair, maintenance, etc.), wheeled robot, vacuuming robot or lawn mowing robot, personal assistant robot, service robot, medical or surgical robot, flying robot, driving robot, aircraft or spacecraft robot, or any other robot, vehicle, or otherwise real or simulated robot that can operate under substantially autonomous control, including stationary robots such as smart home or workplace appliances.

[0049] Many practical embodiments of the present invention provide components and methods for efficiently performing activities by an artificial intelligence agent.

[0050] In some embodiments, a "sensor" may include, but is not limited to, any source of information about the agent's environment and, more particularly, how control may be directed to an end point. In a non-limiting example, sensory information may come from any source, including but not limited to, sensory devices such as cameras, touch sensors, distance sensors, temperature sensors, wavelength sensors, sound or voice sensors, proprioceptive sensors, position sensors, pressure or force sensors, speed or acceleration or other motion sensors, etc., or from compiled, abstract, or situational information (e.g., the known position of an object in space) that may be compiled from a collection of sensory devices in combination with previously held information (e.g., the most recent position of an object), location information, location sensors, etc.

[0051] The term "(one or more) observations" refers to any information received by the agent about its environment or itself in any way. In some embodiments, the information may be a signal or sensory information received through sensory devices such as, but not limited to, cameras, touch sensors, distance sensors, temperature sensors, wavelength sensors, sound or voice sensors, position sensors, pressure or force sensors, speed or acceleration or other motion sensors, location sensors (e.g., GPS), etc. In other embodiments, the information may also include, but is not limited to, compiled, abstract, or situational information compiled from a collection of sensory devices in combination with stored information. In a non-limiting example, the agent may receive abstract information about its own or other objects' locations or characteristics as an observation. In some embodiments, the information may refer to a person or customer or to their characteristics such as purchasing habits, personal contact information, personal preferences, etc. In some embodiments, an observation may be information about the internal parts of the agent, such as, but not limited to, proprioceptive information or other information related to the agent's current or past actions, information related to the agent's internal state, or information that has been calculated or processed by the agent.

[0052] The term "action" refers to any component of an agent that is used to control, affect, or influence the agent's environment, the agent's physical or simulated self, or the internal functions of the agent, which can ultimately control or influence the agent's future actions, action selection, or action preferences. In many embodiments, an action can directly control a physical or simulated servo or actuator. In some embodiments, an action can be an expression of a preference or set of preferences that ultimately intends to influence the agent's selection. In some embodiments, information about the agent's (one or more) actions can include, but is not limited to, the probability distribution of the agent's (one or more) actions, and / or output information intended to influence the agent's ultimate action selection.

[0053] The term "state" or "state information" refers to any collection of information related to the state of the environment or the agent, which can include, but is not limited to, information about the agent's current and / or past observations.

[0054] The term "policy" refers to any function or mapping from any complete or partial state information to any action information. A policy can be hard-coded or can be modified, adjusted, or trained using any suitable learning or teaching method (including, but not limited to, any reinforcement learning method or control optimization method). A policy can be an explicit mapping or can be an implicit mapping, such as, but not limited to, a mapping that may result from optimizing a particular measurement, value, or function. A policy can include associated additional information, features, or characteristics, such as, but not limited to, start conditions (or probabilities) reflecting under what conditions the policy can start or continue, and termination conditions (or probabilities) reflecting under what conditions the policy can terminate.

[0055] The term "distance" refers to any monotonic function. In some embodiments, distance can refer to the space between two points on a surface as determined by a convenient metric (such as, but not limited to, Euclidean distance or Hamming distance). When the distance between two points or coordinates is small, they are "close" or "nearby".

[0056] Broadly, embodiments of the present invention provide methods and systems for training and / or operating an artificial intelligence agent. A multi-prediction is a computational construct where the shared network weights can be used to calculate multiple related predictions, typically but not necessarily a neural network. This allows for more efficient training based on the required amount of data and / or experience, and in some cases, more efficient computation of those predictions. For multi-prediction networks, there are several related and sometimes combinable schemes. The following discussion describes these schemes with reference to the associated figures.

[0057] In FIGS. 1 to Figure 6In each of them, f(x) refers to a prediction, where x can be a state, a prediction id, a skill id, a parameter value, or a combination thereof; s refers to a state; g refers to a prediction id; k refers to a skill id; and p refers to a parameter value.

[0058] Reference Figure 1A , which shows a multi-head prediction network. Here, a single network has multiple outputs, and each output is a prediction of a different feature. As shown in Figure 1, the input to the network is the current state, represented by multiple state inputs S. The weights / parameters of the network in all layers except the last layer of the network are shared among different predictions. Figure 1B Figure shows a simple example of the weights w1 to w4 of a set of inputs 1, x1, x2, and x3 of a single activation node in a single hidden layer of a neural network. As can be appreciated, without sharing the weights, calculations may be involved among different predictions, especially when the number of hidden layers and activation nodes in the neural network grows. Therefore, such sharing has three benefits. First, such sharing can enable faster learning of predictions. Second, since the calculations are shared in the lower layers of the network, such sharing may result in a lower computational cost for calculating multiple predictions. Third, such sharing can lead to generalization of state features.

[0059] For example, a single multi-head prediction network can predict the distance, color, shape, and weight of the nearest object based on a given state. The agent can receive inputs from sensors, etc., as state input data, and can generate predictions that determine that there is a blue circular 3-ounce ball located four feet away at 40 degrees forward. These predictions are indicated as f1(s), f2(s), f3(s), and f4(s) in Figure 1.

[0060] Now refer to Figure 2 , which shows a multi-input prediction network. Here, a single network is capable of calculating the values of several different predictions. In addition to the current state S, it also takes prediction IDs (g1 to g4) as inputs. For example, a single network may be able to predict the distance to any one of a red, green, blue, or yellow block. It can be indicated to the network which one of the four you want to predict by supplying a vector of g values, where only one of the g values is "turned on". As shown in the accompanying figure, in the case of g2 = 1, you can ask the network to calculate the distance to the green ball based on the remaining state information.

[0061] The output of the multi-input prediction network is the corresponding predicted value f(s, g) of the prediction ID supplied as an input. The network is shared, which means that the weights / parameters are common among multiple predictions. Compared with multi-head prediction, parametric prediction has a significant advantage, that is, parametric prediction can be generalized to new or untrained predictions, because the neural network has the ability to generalize to unseen inputs through sufficient training.

[0062] For example, such a multi-input prediction network may be able to predict the distance, color, shape, or weight of an object from an image. The user supplies a flag as input, which tells the network which value should be calculated.

[0063] Reference Figure 3 , shows a multi-skill prediction network. This network is able to calculate the same type of prediction for different skills. In addition to the state S, the prediction network also takes the skill IDs (k1 to k4) as input and outputs a prediction value f(s,k). The multi-skill prediction network is able to generalize predictions based on skills that share some common state dependencies.

[0064] For example, a multi-skill prediction network can be used to calculate the duration of one of the following skills: run-to-door, walk-to-door, skip-to-door, or crawl-to-door, all depending on how far the agent is from the door. Here, as Figure 2 shown in [], the [0,1] layer is designed to represent the "one-hot" nature of the supplied input. In the figure, by setting the second skill (walk-to-door) flag equal to 1 and the rest of the flags to zero, you are asking the network to calculate the prediction if you performed the walk-to-door skill.

[0065] Reference Figure 4 , shows a parameterized skill prediction network. This network is able to predict state features or other predictions based on variable input parameters that affect behavior. For example, the prediction f(s,p) can predict how far the ball will roll when kicking the ball, where the input parameter p is the difficulty of the kick or all the joint angles for the kicking motion plan.

[0066] Reference Figure 5 , shows a hybrid network. In the example shown, this network combines Figure 1A the multi-head predictions of [] with one or more of the skill-conditioned networks (such as the skill-conditioned networks shown in Figure 3 or 4). For example, for a set of similar skills, such as run-to-door, walk-to-door, skip-to-door, or crawl-to-door, a single network may be able to calculate three output predictions, such as the distance traveled, the duration, and the knee pain. The input will include normal state information as well as the encoding of the skill ID.

[0067] Reference Figure 6 , embedding is a technique for forcing even more generalization across inputs. Embedding can be used with any conditional input. In Figure 6 , the conditional input is first embedded into a learned reduced vector representation to form the input for a parameterized prediction.

[0068] For example, a network that needs to predict the duration of running to a door, walking to a door, jumping to a door, or climbing to a door can learn to cluster running and jumping into one category, and cluster climbing and walking into a second category, and then make predictions conditional on these two categories.

[0069] It should be noted that many combinations of these networks are possible. For example, a combination Figure 2 and Figure 3 of networks can have a prediction network that is conditional on both a skill ID and a prediction ID. Or a combination Figure 1A 、 Figure 3 and Figure 4 of networks can be combined to obtain a network that can predict several state variable predictions for multiple skills using a common actual value input parameter (such as the amount of force).

[0070] For example, a network can be established that predicts the distance, duration, and knee pain experienced for four different skills (running, walking, jumping, and climbing) and an "effort" input parameter.

[0071] In view of and in accordance with the teachings of the present invention, those skilled in the art will readily recognize that, depending on the needs of a particular application, any of the foregoing steps may be appropriately replaced, reordered, removed, and additional steps may be inserted. In addition, in accordance with the foregoing teachings, any entity and / or hardware system that those skilled in the art will readily know to be suitable may be used to implement the specified method steps of the foregoing embodiments. For any method step described in this application that can be executed on a computer, a typical computer system can be used as a computer system in which these aspects of the present invention can be implemented when appropriately configured or designed. Therefore, the present invention is not limited to any specific tangible implementation.

[0072] Unless otherwise expressly stated, all features disclosed in this specification, including any accompanying abstract and drawings, may be replaced by alternative features serving the same, equivalent, or similar purpose. Therefore, unless otherwise expressly stated, each feature disclosed is only an example of a series of equivalent or similar features.

[0073] Specific embodiments of the intelligent artificial agent may vary depending on a particular context or application. By way of example and not limitation, the foregoing intelligent artificial agent is mainly directed to two-dimensional embodiments; however, similar techniques may alternatively be applied to higher-dimensional embodiments, and embodiments of the present invention are considered to be within the scope of the present invention. Therefore, the present invention will cover all modifications, equivalents, and alternatives falling within the spirit and scope of the appended claims. It should be further understood that not all disclosed embodiments in the foregoing specification must meet or achieve each objective, advantage, or improvement described in the foregoing specification.

[0074] The claim elements and steps of this document may have been numbered and / or lettered solely for ease of reading and understanding. Neither any such numbering nor letters themselves are intended to and should not be used to indicate the order of elements and / or steps in the claims.

[0075] Without departing from the spirit and scope of the present invention, many changes and modifications can be made by those of ordinary skill in the art. Therefore, it must be understood that the illustrated embodiments are set forth only for purposes of example and should not be regarded as limiting the invention as defined by the following claims. For example, although the fact that the elements of the claims are set forth in a certain combination is presented below, it must be clearly understood that the present invention includes fewer, more, or different other combinations of the disclosed elements.

[0076] The words used in this specification to describe the present invention and its various embodiments should be understood not only in the sense of their common definition, but also by special definition in this specification to include the general structure, material, or acts of a single species.

[0077] Therefore, the definitions of the words or elements in the following claims are defined in this specification to include not only combinations of elements literally set forth. Thus, in this sense, any one element in the following claims can be equivalently replaced by two or more elements, or two or more elements in the claims can be replaced by a single element. Although the elements may be described above as acting in certain combinations and even initially claimed as such, it should be clearly understood that in some cases, one or more elements from the claimed combination can be excised from the combination, and the claimed combination can be directed to a sub - combination or a variant of the sub - combination.

[0078] Therefore, the claims should be understood to include what is specifically shown and described above, what is conceptually equivalent, what can be obviously substituted, and also what incorporates the basic idea of the present invention.

Claims

1. A predictive network method for creating artificial intelligence in machines and computer-based software applications, the method comprising: Receiving an input from the environment as state information to input into a multiple predictive network, wherein the input from the environment includes information from sensors, including information from at least one of the following: a camera, a touch sensor, a distance sensor, a temperature sensor, a wavelength sensor, a sound or voice sensor, a proprioceptive sensor, a position sensor, a pressure or force sensor, a speed or acceleration motion sensor, and / or a location sensor; Receiving an additional input from at least one of a predictive ID, a skill ID, and a parameter value; Before the additional input is input into the multiple predictive network, embedding the additional input into a learned reduced vector representation, wherein embedding the additional input into a learned reduced vector representation includes clustering the actual values of the additional input into one or more categories and using the one or more categories as inputs for prediction; And Outputting a prediction for each learned reduced vector representation based on the received input from the environment as state information, wherein outputting a prediction includes outputting a plurality of predictions, each prediction among the plurality of predictions corresponding to a different state information feature, and wherein the weights or parameters of the multiple predictive network in all layers except the last layer of the multiple predictive network are shared among each prediction among the plurality of predictions.

2. The predictive network method according to claim 1, further comprising inputting at least one of a plurality of skill IDs and a plurality of predictive IDs to provide a hybrid network, wherein the plurality of predictions are output respectively based on the plurality of skill IDs and the plurality of predictive IDs for a set of similar skills or similar predictions.

3. The predictive network method according to claim 1, wherein a plurality of predictive IDs are included in the additional input, and wherein the output prediction is a predicted value of the predictive ID supplied as an input.

4. The predictive network method according to claim 3, wherein the additional input includes a plurality of skill IDs.

5. The predictive network method according to claim 4, further comprising generalizing the prediction based on skills sharing a common state dependence.

Citation Information

Patent Citations

  • Target tracking method of multi-scale expression based on convolutional neural network

    CN106651915A

  • Apparatus and system for vehicle classification and verification

    CN107430693A

  • Method and system for adaptive forecast of energy resources

    EP2688015A1

  • Computer-Implemented Systems And Methods For Large Scale Automatic Forecast Combinations

    US20130024167A1

  • Method for creating predictive knowledge structures from experience in an artificial agent

    US20160012338A1