Memory within embedded agents
Patent Information
- Application Number
- KR1020227003741
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-08
- Filing Date
- 2020-07-08
- Publication Date
- 2026-09-21
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure 112022012470642-PCT00064_ABST
Abstract
Description
Technology Field
[65535] The embodiments described herein relate to the field of Artificial Intelligence (AI) and to systems and methods for implementing and using memory in embedded agents. More specifically, but not exclusively, the embodiments relate to unsupervised learning. Background Technology The goal of Artificial Intelligence (AI) is to build computer systems that possess human-like capabilities, including human-like learning and memory. Most modern machine learning techniques rely on "offline" learning, where AI systems are provided with data ready for learning and complete data limited to specific domains. Creating AI systems that experience the world's objects and events in a human-like manner and learn from implemented interactions remains a significant challenge in conventional technology. Thanks to their implementations and sensorimotor feedback loops with their environment, such AI agents can influence and guide their own learning. Such agents will understand streams of multimodal data from the world and retain information in a meaningful and useful manner. An additional important challenge is creating flexible AI-embodied agents that not only learn from their own experiences but can also possess memories created or modified by external sources (e.g., human users). Hierarchical Temporal Memory (HTM) is an approach to replicating human memory based on a computational structure with multiple registers as analogues to cortical layers. HTM is configured to replicate patches of the cerebral cortex. Nevertheless, HTM fails to provide memory in embedded agents that enables them to learn and unfold from sensorimotor experiences in real time. Brief explanation of the drawing Fig. 1: Schematic diagram of the Convergence Divergence Zone (CDZ) architecture. Fig. 2: Associative Self-Organizing Map (ASOM). Fig. 3: Qualification signals for different modalities. Fig. 4: Method by which a qualification trace generates a qualification window. Fig. 5: Steps of a learning event. Fig. 6: User interface for setting qualifications for different modalities. Fig. 7: Display of ASOM training. Fig. 8: Query viewer input field. Fig. 9: User interface for identifying query patterns. Fig. 10: Display of LTM and STM. Fig. 11: Working Memory System (WM System). Specific details for implementing the invention Computational structures provide embedded agents with memory that can be populated and / or written from experience in real time. Embedded agents (which may be virtual objects, digital entities, or robots) are provided with one or more experience memory stores that influence or direct the behavior of the embedded agents. The experience memory store may include a Convergence-Divergence Zone (CDZ), which simulates the human memory's ability to represent external reality in the form of mental imagery or simulations that can be re-experienced during recall. The memory database is created in a simple, writable manner, allowing experiences to be learned or written during the live operation of the embedded agents. Eligibility-Based Learning determines which elements from streams of multimodal information are stored in the experience memory store. Experience memory storageIn one embodiment, experiences experienced by an agent are stored in one or more experience memory stores. "Experience" should be broadly interpreted as something that an embedded agent can detect or perceive, such as objects, events, emotions, observations, actions, or any combination thereof. Experience memory stores / stores may store dimensionally reduced representations of experiences as neural network weights. Convergence-Divergence Zone CDZIn one embodiment, the experience memory store is implemented as a Convergence-Divergence Zone (CDZ). A CDZ is a network that receives convergent projections from sites where activities are recorded and returns divergent projections to the same sites. Patterns within the CDZs maintain 'dispositions' to act in response to or to complete partially presented perceptual patterns. Hierarchically upstream associative memory associates combinations of activities from lower sensory and / or motor maps to form implicit memory (e.g., cohesive characteristics of objects), which enables downstream reconstruction of component characteristics. For example, an experience memory store that stores experiences of objects that can be used for object classification can be implemented using CDZs as follows: each unimodal object classification path is a hierarchy of CDZs where explicit maps of objects are constructed during perception and reconstructed during recall. Activating a pattern in any single lower-level modality, once learned, can trigger a pattern in a higher-level multimodal CDZ. Subsequently, this activity triggers a 'top-down' flow of activity to other CDZs, allowing the experience memory store to activate patterns that have been learned and are associated with the initial pattern. Figure 1 illustrates an example of CDZ 1. Multimodal CDZs are situated above each higher-level single-mode CDZ. Associations between two modalities X and Y are maintained in a distinct region ("convergence region") Z that is linked independently to both X and Y, rather than through direct links from X to Y. Representations converge from multiple regions into region Z. Declarative representations in the convergence region store associations between stimuli in sub-regions.Patterns are explicitly activated to reveal associations. When a convergence zone representation is activated, it reveals the activity of associated patterns within a set of sub-regions and thus functions as a 'divergence zone' that spreads activity from a single region to a certain range of regions. Convergence-divergence zones can be implemented using maps capable of receiving inputs from multiple modalities, which can then be activated by any of those modalities. These maps can be Associative Self-Organizing Maps (ASOMs), which associate different inputs by taking activation maps from low-level maps and associating simultaneous activations. An ASOM receives an input vector having the size and number of input fields corresponding to neuron weight vectors, where each input field represents a different modality or input type. Once trained on many inputs, the map learns topological groupings of similar inputs. ASOMs can operate using both "bottom-up" construction methods and "top-down" reconstruction methods. ASOMs can generate predictions that can be compared against incoming information. Lowest-level (unrelated) sensory, motor, or other activity sites may be implemented as maps such as Self-Organizing Maps (SOMs) or in any other suitable manner. Figure 3 illustrates the mapping between low-level SOMs and higher-level ASOMs, where ASOMs are the structural input fields of Convergence Zones (CDZs). Low-level SOMs include sensorimotor inputs corresponding to visual, audio, touch, neurochemical (NC), and positional modalities. In a hierarchically structured set of CDZs, lower-order CDZ SOMs provide inputs to higher-level CDZ SOMs.Higher-level ASOMs that serve as convergence-divergence regions include VAT (Visual-Audio-Touch), VM (Visual-Motor), and VNC (Visual-Neurochemical). The VAT ASOM convergence-divergence region is associated with the location aspect in the higher-level VAT-Activity-Location ASOM. CDZs enable real-time learning systems to store multimodal memory and emotion memory. This is how, for example, when an embedded agent "imags" a dog or hears the word "dog," one or more neurons in the high-level ASOM representing a dog are activated. The higher-level ASOM holds pointers to lower-level sensory maps of vision (showing an image of a dog), audio (hearing a dog barking), and even emotion state maps that reproduce the emotions the embedded agent experienced when it first encountered the dog. Aspects"Aspect" should be interpreted broadly as one aspect of something existing, including its representation, manifestation, or experience. Objects and / or events may be experienced in different aspects, including but not limited to visual, audio, touch, motor, and neurochemical. In one embodiment, each aspect input is represented and / or learned by individual SOMs. An architecture including maps associated with each aspect may be used so that when two or more aspects are experienced simultaneously, the combination is stored in higher-level (associative) maps as pointers to each of the two senses in their original lower-order maps. Associative maps may be activated by an input corresponding to any of the aspects to which they are associated. If an input is received from only one of the aspects, corresponding representations from the other aspects may be predicted. Visual input may be streamed to the embedded agent in any suitable manner. In one embodiment, visuals are provided to the embedded agent via a camera capturing the real-world environment. Images may be transmitted from a screencast of a user interface or otherwise from a computer system. The image of the embedded agent may therefore be of the real world, which may include viewing a human user through a camera, or of a "virtual world" or computer representation (e.g., a screen representation or a VR / AR system representation), or any combination of both. Since both "real world" and "interface" views may be presented to the embedded agent, the embedded agent has two distinct views. Each view may have an associated saliency map that controls attention. In one embodiment, only a single salient region across these two maps is always selected for attention: thus, when it becomes an attention routine, the two views can be treated as a single view having two parts.A subregion of the camera input can be automatically mapped to a virtual "fovea," that is, a smaller area of the video input corresponding to where the embedded agent's eyes are directed. The fovea image subregion can be further processed in modules, for example, an effect classifier and / or an object classifier. This allows the embedded agent to focus on only a small part of the camera input, thereby reducing the dimensionality. In one embodiment, a 28 x 28 RGB fovea image is provided. An ambient image may also be processed, but at a much lower resolution. The audio input is transmitted through a microphone to capture the waveform processed by the auditory system. In one embodiment, acoustic features are analyzed using FFT and other techniques to generate a spectrogram, which is used as input to the auditory SOM (e.g., a 20 x 14 (fxt) spectrogram). The auditory SOM learns the tonotopic map of the audio input. Alternatively and / or additionally, digital audio input, such as from audio files, or streaming from a computer system may be delivered to the embedded agent. Acoustic signals can be analyzed through a deep neural network, which provides vectors of values corresponding to the incoming words. These are fed to a second, independent auditory SOM that learns word mappings. The sound phase map and word map may be further integrated by a higher-level auditory ASOM, which is the final representation of the audio modality. Touch sensations may be provided to the embedded agent based on the interaction between the embedded agent and the virtual environment. For example, whenever a part of the embedded agent's body "crosses" with another object within the embedded agent's environment, the object crossing may trigger a touch sensation in the embedded agent.Such touch sensations may be associated with the proprioceptive map of the embedded agent's body, the map of the embedded agent's environment, and / or any other aspect. When the embedded agent touches specific "touchable" objects in the virtual world, a collision is detected, and activity is triggered at mechanoreceptors on the embedded agent's effectors (e.g., fingers). Touch sensations may be provided to the embedded agent via computer input devices such as a mouse, keyboard, or touchscreen. For example, "touching" the screen is projected onto the part of the embedded agent's body that the fingers (on the touchscreen) or mouse cursor "contacts" on the mechanoreceptor map. Symbolic inputs (e.g., keyboard inputs) may be mapped to any touch sensations, such as object textures. For example, a haptic object type SOM may map different object textures. The shapes of objects can also be registered through a haptic system involving both touch and motor movements. The “location” modality can represent the foveal location, including the x and y coordinates of the embedded agent’s fovea. The coordinates can be directly converted into a 10 x 10 activation map via a location-to-activity SOM. The interoceptive sense is the embedded agent’s perceptual sense of the internal state of the embedded agent’s body. The interoceptive state space map is formed by taking inputs from signals representing the body’s instantaneous state, for example, in relation to hunger, thirst, fatigue, heart rate, pain, and aversion. Neurochemical parameters represent physiological internal state variables that are part of the emotional system. The interoceptive map represents the state space of the embedded agent.Examples of neuromodulators that can be modeled include acetylcholine for motor function, cortisol as a stress marker, and oxytocin for social bonding. Basic expressions of primary emotions can be mapped into a high-dimensional neurochemical space, which regulates behavioral responses and provides a mapping from a continuous series of instinctively felt states to distinct psychological categories. Intrinsic sensations can contribute to embedded agent decisions as events are associated with the body's emotional neurochemical states, causing the recalled emotion of the imagined event to become a factor in making decisions. The proprioceptive system provides the embedded agent with perceptual awareness via proprioceptors regarding the composition of the embedded agent's body, including the positions of the agent's effectors (e.g., limbs, head, and the composition of the agent's torso). The proprioceptive map may include information regarding the angles of each joint transmitted from the skeletal model of the embedded agent's body. In more detailed biomechanical models of the embedded agent's muscular system, proprioceptive maps may also contain information regarding muscle elongation and tension. Motor patterns can be used to map types of movements. Individual words can be associated with representations of objects, movements, events, or concepts through recorded words, auditory phonemic representations, and / or other symbols. One or more symbols associated with a representation of a concept can be stored as modalities to be associated with sensory modalities representing the concept. Any other suitable modal (or virtual representations of similar ones), such as taste or smell, can be implemented. Specific aspects of modalities can be modeled as modalities in themselves. For example, an image modal can be subdivided into several modalities, including light modal, color modal, and form modal. Internal sensations such as temperature, pain, hunger, or calmness can be modeled. Direct creation of the experience memory storeIt is possible to store a trained neural network (e.g., SOM) along with its post-training weights in an embedded agent that has not directly experienced the weights. In this way, a "blank" embedded agent can be provided with knowledge (e.g., of objects) embedded in the neural network weights of its experience memory stores / stores. Memory database (memory files)In one embodiment, representations of experiences may be stored in a memory database in addition to the experience memory store. The memory database may be automatically populated and / or created through the experiences of the embedded agent. A user or an automated system may retrieve memories stored in the memory database and / or create new memories in the memory database and / or delete memories. Raw data corresponding to representations in each experienced aspect is stored in the memory database and may be associated with the corresponding experience. For example, components of memory associated with visual aspects may be linked to image files (e.g., JPEG, PNG, etc.), and components associated with auditory aspects may be linked to audio files (e.g., MP3). The memory database may be implemented in any suitable manner, for example, as a folder and / or database storing a collection of files. In one embodiment, the memory database is a CSV file storing experiences. CSV entries may contain or point to representations of raw data associated with the experience corresponding to the entry. Storing memories as other raw data corresponding to associated images or raw inputs enables experiences to be reproduced / processed by the agent. The embedded agent can learn them as if the embedded agent were experiencing those inputs. In one embodiment, during the live operation of the embedded agent, the embedded agent simultaneously stores memories of the experience in both the experience memory store and the memory database.For example, the experience of a dog barking sound may be stored as a multimodal memory stored in an experience memory store, and may also be stored in a memory database as attributes of an entry corresponding to the experience, including other related multimodal data such as images, sounds, emotional valence, and text / speech utterances. Storing the experience in files may also involve storing additional data about the experience, such as metadata, or the time the event occurred (timestamp), the GPS location of the event, or any other contextual information related to the experience. Filling memories through experience In one embodiment, memories stored in a memory database are populated from the real-time experiences of the embedded agent during the course of the embedded agent's live operation. The agent interacts with sensory streams from the real world and / or virtual world as described in New Zealand provisional patent application NZ744410, which is titled "Machine Interaction," assigned to the assignee of the invention, and incorporated herein by reference. As described herein, the embedded agent may selectively learn new, emotional, or user-signed experiences through experience. In an embedded agent where the experience memory store is implemented as a CDZ, memories are stored in the CDZ. Whenever a new memory of an experience is stored in the CDZ, representations from lower-level SOMs are stored as attributes and / or files in a new entry to the memory database. Training of experience memory storage through memory databasesA memory database can be used to train an experience memory store. Entries within the memory database are provided to the experience memory store as training inputs during consolidation. Memories encoded in the experience memory store enable the agent to recognize objects, concepts, and events and make predictions. For example, a user can create a set of input files for specific learning domains. For instance, an agent can become a "dog expert" without experiencing the dogs during live operation by being provided with a memory database containing images of different dog breeds according to associated aspects, such symbols include the names of the dogs, spectrograms of their barking sounds, and the emotional responses evoked by the dogs. In one embodiment of the CDZ, entries within the memory database are used to retrain the CDZ by changing the weights of the underlying convergence / divergence regions (e.g., SOMs / ASOMs). During training, the raw files / data corresponding to the entries are reread by the experience memory store one experience at a time. Taking an example of an object learning event, raw data corresponding to visual, auditory, and touch modalities is loaded and triggers learning events. Long-term memory (LTM) learning events occurring during memory enrichment can take place on a much faster time scale than real-time learning, as discussed under the section titled "Memory Enrichment." In one embodiment, the raw files used to "train" the agent may be displayed to simulate the agent's "dreaming" as the agent "re-experiences" or "reimags" past experiences. Memory ReconstructionEntries within a memory database can be reread to reconstruct memories: for example, they can train a short-term memory (STM) experience memory store to generate "virtual events" or train a long-term memory experience memory store during memory reinforcement. As raw sensory inputs are stored as weights of neurons in low-level maps, it may be possible to reconstruct the raw sensor input (e.g., an image) that triggered the learning event from the experience memory store. However, because potentially multiple different input vectors can modify the weights of a single neuron, the resulting weights in a neural network may be a mixture of multiple input instances. By explicitly storing individual input vectors and their constituent input fields as distinct entries with associated attributes, the memory database provides a method for accurately reconstructing individual experiences. Modification or deletion of memoriesMemories may be selectively modified by the user, for example, by modifying entries within the memory database (explicit modification, e.g., changing the significance of an object) or by deleting all entries. By deleting all entries, the entire memory of the embedded agent may be deleted, leaving a blank slate. In one embodiment, at each reinforcement, the experience memory store is wiped and completely refilled by training using an updated memory database (which may include the edited or deleted entries). In experience memory stores that are SOMs, the wipe of the experience memory store may be achieved by randomizing all neuron weights. In other embodiments, rather than wiping the entire experience memory store, updated or modified experiences may be placed in the experience memory store, and specific data points may be selectively deleted from the experience memory store by "unlearning" them. In the "forgetting" model, experiences are time-stamped or otherwise marked to indicate the recency of memory, and older events can be "forgotten" by deleting them from the experience memory store and / or memory database. Writing of memories Instead of agents requiring new experiences to generate new memories, memory entries corresponding to those experiences can be directly "injected" into the agent's memory. This creates an instructible, artificially manipulable embedded agent. For example, the agent can be programmed to have directed autonomous responses to experiences (e.g., verbal responses to specific stimuli). Thus, entries within memory databases can not only be "written" by external tools but can also be learned directly through the embedded agent's real-time sensorimotor experience. Writing using a Text CorpusMemory creation can be performed in relation to a text corpus. An example of a markup text corpus for creating event memory is as follows: [Timestamp] The red car (image, sound) drove (action) to the left (place). I<didn't> like it (emotion) Real-time sensorimotor context can reflect word choices (like, dislike) and the directive and emotional states of the embedded agent. This can be achieved by providing lookup tables for raw inputs (e.g., images, sounds, feelings, etc.) associated with symbols such as words. This allows for the rapid generation of inputs to learn events through sentences. Data matching corresponding words within the lookup tables is extracted to train the experience memory store and / or generate detailed entries associated with raw data within the memory database. In embedded agents possessing prior knowledge regarding objects, actions, and emotions, events can be constructed by associating event components using syntactic structures. Memories can be classified, labeled, or tagged in a manner that allows individual memories to be easily located, modified, and / or deleted. A user interface may be provided to facilitate users from viewing and editing the memories of the embedded agents. Implementation example using Self-Organizing Maps (SOMs) Self-organizing maps Both aspects and convergence-divergence regions can be represented using unsupervised learning-based memory structures, also known as Self-Organizing Maps (SOMs) or Kohonen Maps. To provide a discretized / quantized representation of this data, an SOM (which can be 1, 2, or 3... or n-dimensional) is trained on the dataset. Subsequently, this discretization / quantization can be used to classify new data within the context of the original dataset. Weighted-distance functionIn conventional SOMs, the buoyancy between an input vector and a neuron's weight vector is calculated using a simple distance function (e.g., Euclidean distance or cosine similarity) over the entire input vector. However, in some applications, it may be desirable to weight certain parts of the input vector (corresponding to different input fields) more highly than others. In one embodiment, an associative self-organizing map (ASOM) is provided for a multimode memory, where each input field corresponding to a subset of the input vector contributes to a distance function by a term called the ASOM Alpha Weight. The ASOM calculates the difference between a set of input fields and a neuron's weight vector by first dividing the input vector into input fields (which may correspond to different attributes recorded in the input vector). The differences between vector components within different input fields contribute to the total distance by different ASOM Alpha Weights. A single generated activity of ASOM is calculated based on a weighted distance function, where different parts of the input vector may have different semantics and their own ASOM alpha weight values. Thus, the total input to ASOM includes whatever the inputs are associated with, such as different modalities, activities of other SOMs, or other things. Figure 2 illustrates the architecture of ASOM integrating inputs from multiple modalities. Input to ASOM Is It consists of input fields (32). Each input field is for vectors of neurons The input field (32) may be: direct one-hot coding of sensory input; a 1D probability distribution, a 2D matrix of activities of a low-level self-organizing map, or any other suitable representation. The ASOM (3) of FIG. 2 is It consists of neurons, and each neuron is a weight vector corresponding to the entire input having, Partial weight vectors for of It is divided into input fields. Input If this is provided, each ASOM neuron first calculates the input field-expressed distance between the input and the neuron's weight vector: Here Is It is the bottom-up mixing coefficient / gain (ASOM alpha weight) of the i-th input field. is an input field-specific distance function. Any suitable distance function or function may be used, including but not limited to the following: Euclidean distance, KL divergence, cosine-based distance. In one embodiment, the weighted distance function is based on Euclidean distance as follows: Here K is the number of input fields, and α i is the corresponding ASOM alpha weight for each input field, and D i Is i It is the dimension of the i-th input field, and x j (i) or w j (i) are each i of the nth input field j It is the nth component or the corresponding neuron weight. In some embodiments, ASOM alpha weights may be normalized. For example, when a Euclidean distance function is used, the active ASOM alpha weights are usually made to sum to 1. However, in other embodiments, ASOM alpha weights are not normalized. Not normalizing can lead to more stable distance functions (e.g., Euclidean distances) in certain applications, for example, where multiple input fields or high-dimensional ASOM alpha weight vectors in ASOMs dynamically change from sparse to dense. Methods for sampling memory: Dreaming and inhibition-of-return (IOR)It may be desirable to randomly reconstruct items stored in the SOM. This occurs, for example, in the construction of pseudo-training items during the reinforcement of long-term memories or in the random generation of motor movements during motor babbling. In these cases, the training records of the SOM induce a probabilistic selection of SOM neurons to be reconstructed. Sampling can be combined with an Inductive Repression (IOR) process to sample from the entire set of trained values. Training records When reconstructing from the entire activity of the SOM, all neurons contribute in proportion to the similarity of their weight vectors to the input vector, regardless of whether these neurons represent meaningful hypotheses or were trained to contain initial random noise. To provide a more complete reconstructed output, more weights may be given to neurons that have received more training (ignoring untrained neurons). The adaptation amount of each neuron can be recorded as a value from 0 to 1, accessible in the SOM parameter "training history." The training history is an additional scalar weight of each neuron initialized to 0 and connected to a fixed input of 1. Thus, whenever this particular neuron is trained, the neuron's training history increases in proportion to the current learning rate because it is a winner or a neighbor of a winner (potentially adapted due to the goodness of the match). This means that during the training process, the training record increases toward 1. The average of the training record values of all neurons in the map ("map occupancy") represents the map's free capacity to learn new inputs without overwriting old inputs. A "maximum occupancy" of 1 indicates a "full / clustered" map (no free capacity), and a "maximum occupancy" of 0 indicates an untrained map. The training record is the value of the activation mask in the activity calculation of the SOM (term m i It can serve as ). In Bayesian terms, uniform m i Instead of using equal prior probabilities (where all hypotheses are equally probable), this is equated with adopting the a priori based on observational frequencies: that is, the resulting probability distribution is based on the assumption that the input is one of the inputs seen prior to training the SOM. conditioning Training records can decay over time, which means that while the search avoids regions of recent training, if a region has not been reactivated for a long time, it can be reused for new memories. The way training records store the history of training can be tuned through the Training Record Decay parameter. A Training Record Decay value of 1 means no decay. Training Record Decay values less than 1 mean that the training records will reflect only the most recent training (which has recency determined by a value between 0 and 1). Content inspection by top-down reconstruction of weights In a hierarchy of connected SOMs where the activity of a lower-level SOM provides input to a higher-level SOM, the flow of activation can be reversed during top-down reconstruction: the reconstructed input from the higher-level SOM provides a top-down signal to the lower SOM, i.e., the expected activation pattern. This signal can be supplied from the top-down bias field of the lower SOM. It can be combined with the activation pattern derived from the lower SOM's own inputs. The raw contents of "memories" stored in the neurons can be retrieved, where the memory is identical to individual events or a mixture of multiple events, depending on the training environments and also the SOM parameters (e.g., small sigma, high learning rate = "sharp" individual memories, while larger sigma and lower learning rate result in generalized and mixed memories). ASOM Configuration for Fast LearningWhile backpropagation-based learning methods require slow learning, the local characteristics of SOM neuron representations through small weight updates allow them to learn input patterns very quickly, even in a single exposure. The problem faced in learning inputs "quickly" (after only a few presentations) in conventional SOMs is the overwriting of previously encoded inputs. Distinct training items (or at least training items considered distinct for the purpose, e.g., members of different classes) must be kept distinct in the SOM by encoding them in distinct neurons or regions. At the same time, items that are sufficiently similar to one another must be encoded in the same neuron or region. Unlike previous attempts to associate using slow-learning SOMs, the ASOMs described herein can learn "quickly." The ASOM can be configured to learn quickly by selecting high learning constant / learning frequency values, allowing a given input to be encoded by a single SOM neuron (or region) in a single exposure. However, changing the learning constant is not sufficient to allow for truly fast learning of large sets of items. It can be determined whether an input is "new" or not, and if the match is not close enough, the "winning neuron" is not overwritten; instead, a different neuron is selected. A "best matching threshold" parameter can be defined to control whether an item presented to the ASOM is considered "new" or "old." The "best matching threshold" is a threshold for the (raw denormalized) activity value of the SOM neuron that responds most strongly to the input item. If this value falls below the "best matching threshold," the item is considered "new"; otherwise, the item is considered "old."New entries are stored in the SOM as separate patterns; old entries update existing patterns. When encountering a new entry, the "search method" parameter determines which neuron to assign to encode the new input. Any suitable search method may be used. Examples include the following: Noise for input search: Add random noise to the current input and find a new winner based on the Gaussian activation function applied to the distance to this modified input. Noise for activation search: A new winner is selected from a composite activation map, which is a mixture of the original activation map and a secondary map filled with random noise. The mixing coefficient for the secondary map is called compare_noise, and it determines how much the original map will be distorted. Small values of compare_noise will cause local search near the original winner. Instead of mixing activity with noise, the secondary map can be set to something that encodes values biased toward or away from specific regions of the SOM—for example, values that inversely reflect how often and how recently each neuron has been active—to ensure that the previous winning neuron is avoided and to promote more even filling of the SOM. A particularly useful method is to track the amount of training each neuron has received (either collectively or recently)—the so-called training history—and to suppress the selection of winners from trained regions (engaging previously untrained neurons / dead neurons). The history of each neuron / region of the map, its training level, and competition for the winning neuron still depends on similarity to the input, but is biased away from regions that have undergone extensive training. The network becomes evenly populated, and "dead neurons" (which are never trained due to poor initial weights) are reduced. Using the inverse of the training history as activation noise ensures that if unused neurons exist, they will be assigned first. To maintain the topographical organization of the SOM and place new winners near original winners, "comparison-noise" can be set to small values. If comparison noise is small, original activations will still have a strong influence, and thus there is a possibility that new winners will emerge near old winners.Next, it will be trained with the current input, and the original winner will encode what it previously had, and the new input will not overwrite it but rather be represented by nearby neurons. Establishing a pattern as a map of values that inversely reflects how often and how recently each neuron has been active (which promotes more even filling of the SOM by engaging previously unused neurons and ensuring that the first winner is not selected) can be calculated for SOM isomorphic vectors using the following pseudocode: If an item is considered "old," it is stored in the SOM in the region where some learning has already occurred. In a standard SOM, if the same item is presented repeatedly, the region representing these neurons will expand and potentially grow in size, eventually dominating the entire SOM. This is an inefficient use of the SOM. To control this effect, the "Best Matching Learning Multiplier" parameter adjusts the learning frequency of the SOM based on the activity of winning neurons. If the "Best Matching Learning Multiplier" is set to 0, exactly repeated items will not induce any new learning in the SOM. If it is set to 1, there is no adjustment to the learning frequency of the original SOM. The multiplier M for the learning frequency can be calculated using the following equation: M = 1 - Raw Wiener Activity * (1 - Best Matching Learning Multiplier)Even in the case of a perfect match, a non-zero value lower than 0 for the best match learning multiplier may be desirable because some training is appropriate. This is because more neurons within the neighborhood of the perfect match can adapt toward that value, and the reconfigured "soft" output will also reflect the frequency with which different values were encountered. As previously discussed, a problem faced in fast-learning SOMs is that large learning frequencies increase the risk of overwriting neurons. At small learning frequencies, weights are averaged rather than completely overwritten. Depending on the values of its parameters, the SOM can be configured for slow or fast learning. Slow learning (as described by standard Kohonen SOMs, and similar to cortical learning in the brain where individual memories can be generalized / mixed) is characterized by: smaller learning frequencies, higher values of the neighborhood size sigma, and disabled novelty detection by setting best_match_threshold=0. Fast learning is characterized by a maximum learning frequency, a very small sigma, and a best_match_threshold set to high; it is similar to hippocampal learning in the brain and can behave like a probabilistic lookup table (having a high degree of representing individual experiences individually / orthogonally and accurately). Because the ranges of the above parameters are continuous, a mix of fast and slow learning is achievable in the SOM. It is possible to adaptively lower the learning frequency when the SOM is crowded (in terms of its map occupancy, as described under the section "Training History")—then the SOM automatically switches to a standard slow learning SOM (since continuing to learn fast across the entire map would mean overwriting / forgetting old knowledge). When the learning frequency is reduced, new memories will be mixed with the most similar old ones.In one embodiment, the "speed" of learning depends on the SOM capacity. In a SOM with sufficient capacity, the SOM can be configured to learn individual memories quickly (even in a single shot) and very accurately. A transition to more gradual learning may occur as the SOM approaches its full capacity (to mix them rather than completely replace old memories with new ones). To monitor how much capacity remains (i.e., unused neurons that can be trained without overwriting older memories), map occupancy can be defined as the average value of training records per neuron, i.e., sum_i(Training Record[i]) / map_size. A value of 0 signifies an empty / untrained map, and a value of 1 signifies a full map. To transition from a fast map type to a slow map type, parameters, namely training frequency, sigma, and best matching threshold, can be gradually adapted while increasing map occupancy. Alternatively, a separate switch may occur when the map occupancy exceeds a predetermined threshold, e.g., 90% (0.9). Targeted oblivion "Forgetting" everything learned by SOMs can be achieved by replacing all neuron weight vectors with random noise (by the same method in which SOMs are initialized). However, there are situations where "targeted forgetting" is useful, for example: "Undo" the most recently learned experience that was learned by mistake Forgetting all memories of a specific kind, namely all images associated with the sound of a gunshot Forgetting infrequent memories (assuming they are experiences that occurred by chance and were repeatedly faced, and are not of very good quality). Forgetting very old memories (assuming that training was unstable at the start and that representations from that point onward are of low quality). Targeted forgetting is controlled by a mask called the "Reset Mask" (similar to an activation mask). The Reset Mask is isomorphic to the SOM (i.e., has one mask value for each SOM neuron). When replacing the weight vectors of neurons with noise, only those neurons with Reset Mask=1 will be reset, while the others (with Reset Mask=0) will be retained. Alternatively, the Reset Mask values can be between 0 and 1; in this case, the original weight vector will be mixed with random noise having a mixing coefficient determined by the Reset Mask value as follows: During a reset, the training records are updated, allowing the training records of the reset neurons to be erased (i.e., for distinct reset masks, Training Record:=0 for those neurons where Reset Mask=1). For successive mixing (memory blurring): Appropriate reset masks can be set according to the following requirements: The reset mask is set to the most recent activation map of the SOM (the activity of the entire SOM immediately after training on the experience to be returned to). This causes partial forgetting-blurring proportional to the magnitude of the activity. Alternatively, distinct reset masks can be generated; for example, for probabilistic SOMs, the reset mask is set to 1 for all neurons whose activation exceeds the reset threshold and to 0 for the rest. Or, for non-probabilistic SOMs, the mask values are set to 1 for the winning neuron and to 0 for all other neurons. Stimuli associated with memories to be forgotten are input. In the above example, a gunshot is provided on the audio input field, and the ASOM alpha weight for video is set to 0 (to extract all videos associated with the gunshot). The generated activation map can be used directly as a reset mask. Alternatively, a separate reset mask may be generated, for example, by setting the reset mask to 1 for all neurons whose activation exceeds a set threshold and to 0 for the rest. Or, it may be set to 1 for the winning neuron and to 0 for all other neurons. The reset mask is set to the 1-training record, or its discretized version (ResetMask[i]=1 if trainingRecord[i] < threshold, and 0 otherwise). Training record decay is set to a value < 1 during training. This will cause the training record to decay to 0 over time for those neurons whose training record is not "refreshed" by new training. Subsequently, the reset mask is set to the 1-training record, or its discretized version (ResetMask[i]=1 if trainingRecord[i] < reset threshold, and 0 otherwise). ASOM VisualizationSOMs can be used as tools for visualizing multidimensional data. Figure 7 illustrates a display of an ASOM associating five input fields (digit bitmap, even, less5, mult3, color) during ASOM training. The visualization shows how ASOMs can be queried to display the configuration of ASOM weights during training and where data satisfying the query is represented in the query. Training data is specified, where each data includes (binary) flags following the digit to indicate whether the digit is (in order) even, less than 5, a multiple of 3, and (arbitrary) color. The ASOM is trained on the data in an arbitrary fitting manner. The neighborhood size and learning rate can be incrementally increased. Figure 7 illustrates static views of the input pattern (where the digit itself is represented as a 20x20 bitmap), the reconstructed output pattern, flags indicating when the network is plastic / trained, and the weights. Since ASOM associates five input fields (digit bitmap, even, less5, mult3, color), the weight matrix is decomposed into input field weight matrices. When binary information is represented, the color white represents 0 / false and black represents 1 / true. When bitmaps and color maps are represented, the colors represent their natural meanings. Once ASOM is trained, it is possible to formulate queries and dynamically view where the regions that best satisfy those queries are located on the map. Query views are displayed in the columns of the dynamic query viewer input field, as illustrated in Fig. 8. Each query is independent of the others and can be manipulated using sliders on their respective tabs, as illustrated in the screenshot in Fig. 9.Queries are displayed side by side, allowing the user to visually compare regions corresponding to different queries. To generate a query, the user or the automated system may specify one or more query patterns (e.g., as illustrated in FIG. 9). The strengths of the influences for each defined pattern may also be specified (as alpha / input field weights for each input field). The strengths may be binary (0 or 1), or sequential / fuzzy mixed queries may be supported. The map of each view may show the regions of the ASOM that best correspond to the query, and the output shows the reconstructed data that best approximates the query. By combining patterns, it is possible to request questions such as 'what are even multiples of three less than five?' or 'which digits are shades of blue?'. It is possible to increase or decrease the matching strictness, that is, the activation sensitivity of the ASOM (as illustrated by the matching strictness variable in FIG. 9). In this example, if the map is completely white or the output bitmap is completely black, it may be desirable to decrease the strictness of the matching, and if the map is too dark or the bitmap is too blurry, it may be desirable to increase the strictness. Each view has two copies of the master ASOM, one for visualizing its activities and the other for reconstructing the output. While ASOM requires activities to be normalized to sum to 1 to calculate the output as a weighted combination of activities, the activity map showing the ASOM must display raw activities without normalization to determine the actual range in which the weights of each neuron satisfy the query.Examples of ASOMs include: VAT (visual / audio / touch), VM (visual / motor), VNC (visual / NC), VATactivityL (VAT / location), HC, or action-result (V1 / V2 / M / L1 / L2 / NC). Cross-mode object representation SOMCross-mode object representations can be learned in a SOM that associates different sensory modalities of an object. In one embodiment, it serves as a cross-mode object representing a SOM that associates visual, audio, and touch inputs and learns modal-integrated representations of object types. It takes inputs from three SOMs that learn monomodal representations of object types: a visual object type SOM, an auditory object type SOM, and a tactile object type SOM. A signal detection process can be implemented by providing each input field of the CDZ SOM to the associated signal detection process, which looks for the start of a signal in such fields and triggers an eligibility trace for such fields when the start occurs. In the case of the cross-mode object representation SOM, these signals can come from attention systems in three distinct modalities. Learning in the cross-mode object representation SOM, and in its input SOMs, is driven by low-level events detected by attention systems in three distinct modalities. For example, an event for an auditory stimulus may be a sound louder than a predetermined threshold. Learning also requires some degree of match between these different signals, as implemented using qualifying traces. Qualifying traces (encoded by leaky integrate-and-fire neurons) are initiated in each modality when an event is detected. When traces for two modalities are active simultaneously, learning occurs in the type SOMs for these modalities and also in the cross-mode object representation SOM. When an event is detected in only one modality, a different connection mode is triggered. The pattern activated in the modality registering the event is passed as input to the cross-mode object representation SOM, which activates the modality-independent object representation.Subsequently, these representations are used to reconstruct representations in other single-mode type SOMs, allowing patterns in missing modalities to be inferred from the top down. In this model, the cross-mode object type SOM is activated by stimuli in a single modality and by stimuli in multiple modalities (provided when they are congruent). Emotional Object Association SOMEmotional states can be associated with input stimuli. For example, pairing loud, sudden, and / or fearful noise with the presentation of an object can induce emotional conditioning in an embedded agent. When the agent subsequently encounters that object, it will induce a fear reaction. Such associations are learned in a Convergence Zone (CDZ) ASOM called the Emotional Object Associations ASOM. In one embodiment, this is achieved by an ASOM that associates neurochemical aspects with visual aspects (VNC SOM). This ASOM takes inputs from a Visual Object Type SOM and a Neurochemical SOM that holds the agent's emotional state. Each input field of the Emotional Object Associations ASOM has an associated signal detection process, which looks for the onset of a signal in that field and triggers an eligibility trace for that field when the onset occurs. The signal that triggers eligibility for the Object Type SOM is the selection of a new important region in the importance map, which can be triggered, for example, by significant movement within the field of view. Events triggering eligibility for the emotion state medium are calculated from 'phasic' signals associated with the emotion state vector, which signal abrupt changes in these vectors. By the principle of start-dependent learning, the SOM is allowed to learn associations between representations in its input fields only when eligibility traces for these fields are simultaneously active exceeding their respective thresholds. In such cases, this principle ensures that emotional associations are learned when a newly significant object is recognized that occurs temporally simultaneously with the sudden onset of a given emotion. After learning emotional associations with a given object of type O, the presentation of O as a newly significant object automatically activates the associated emotion. This occurs through the principle of reconstructing missing inputs. Voluntary learningOperant conditioning is a process in which motor actions generated by an agent in a given context become associated with a reward stimulus that arrives a short time after the action is executed. Circuits that learn these associations run continuously in the embedded agent to suggest actions that lead to a reward in a given context. The Action Result ASOM is a convergence region that is hierarchically constructed on previous convergence regions. The Action Result ASOM learns the association between a perceptual context stimulus occurring at a given time (visual object type T1 appearing at location L1), a motor action performed after a short time, and a reward stimulus occurring at a slightly longer time thereafter. The reward stimulus is associated with an object of a different type T2 appearing at a different location L2. The Action Result SOM needs to store representations of the perceptual context stimulus because they will disappear until the reward stimulus appears. The T1 and L1 inputs to the Action Result SOM hold copies of the previous location selected in the importance map and the previous object type to be invoked in the object type SOM. The T2 and L2 inputs are the currently selected importance position and the currently active object type. Therefore, the operation result SOM learns the association between the remembered object and the currently perceived object. Eligibility windows can be adjusted to accommodate associated events occurring at different times. Behaviors based on reconstruction memoryAll embedded agent behaviors may be influenced by reconfiguring memory. The use of a neural behavioral modeling framework for creating and animating an implemented agent or avatar is also disclosed in US10181213B2 assigned to the assignee of the present invention and is incorporated herein by reference. Within a neural behavioral model such as that described in US10181213B2, reconfiguring inputs from different aspects may alter the internal state of the embedded agent and, accordingly, modify the agent's behavior. Creating emotional memories enables the autonomous triggering of emotional expressions in the embedded agent. In customer service avatars, experience memory stores can be used to efficiently program responses in the avatars. For example, a brand loyalty customer service avatar can be programmed to have a positive emotional response to all trademarks associated with the brand, including visual trademarks and word trademarks. For example, hearing the brand name "Soul Machines" may be associated with a "happy" feeling that alters the agent's neurochemical state accordingly. When that word is heard, a neurochemical state of happiness is predicted, causing the avatar to smile. In one embodiment, direct responses to experiences or events can be "injected" into the agent via writable memory. For example, there is memory that associates the brand "Soul Machines" with states of caution / alertness and wide-eyedness—that is, a state where the visual aspect is consistently qualified, and other aspects indirectly associated with the visual aspect via ASOMs receive top-down inputs to generate predictions. In toy applications, embedded agents, such as avatars or virtual characters, can be programmed by users to exhibit behaviors.For example, a child playing with a virtual friend can "inject" a memory into the virtual friend that a specific object is unpleasant, either through experience (presenting an object in front of the virtual friend and generating negative facial expressions and / or words) or through an interface. The personality or character of an embedded agent can be created by creating memories, including the agent's likes and dislikes. By creating emotional states associated with objects and / or events, it is easy to develop an agent with a specific personality: multiple objects can be presented and associated with emotional states. For example, an avatar can be programmed to have the personality of an animal lover by creating file-based memories of various different animals—each associated with the emotion of "happy." Similarly, objects can be associated with emotions of anger, sadness, or neutrality. Workout plansA motor memory system may be provided to separately store motor movements and / or motor plans that the agent can activate to execute corresponding motor movements. Examples of motor movements include, but are not limited to, pressing, dragging, and pulling. Each movement is characterized by a specific spatiotemporal pattern, such as a sequence of proprioceptive joint positions in the case of movements, or a sequence of visually perceived joint positions in the case of observed movements. The temporal dimension may be implicitly represented by recurring projections, where the ASOM associates the current input with its own activity in a previous computational time step. Embedded agents may have motor control systems that enable the embedded agents to intentionally move body parts, such as limbs or other effectors. Information regarding the angles of each joint may be conveyed from the skeletal model of the agent's body. The agent may have hand-eye coordination capabilities to reach specific points in visual space. A Self-Organizing Map Model (ASOM) enables an agent to learn hand-eye collaboration, allowing the agent to interact with a surrounding 3D virtual (or real in the case of VR / AR) space that changes with realistic eye movements and hand-reaching motions. Once trained, the ASOM is used for inverse kinematics and can return joint angles when a target location is presented. Kinematic actions can be individual movements (e.g., reaching a certain point in space to touch an object) or sequential movements. For example, object-directed kinematic actions such as grasping, slapping, and punching are sequential movements in which the agent's hand (or other effector) moves toward a target object along different trajectories and / or velocities.For example, in the case of striking and punching motions, the trajectory is faster than the approach; the trajectory may also involve pulling the hand back. The trajectory of the fingers of the hand may also be described. Grasping involves extending the fingers of the hand and then folding them. Punching and striking involve configuring the hand into a specific shape before contacting an object. Multiple motions are each associated with targets and may be instructed to generate motion plans. A system for generating plans is described in provisional patent application NZ752901, titled "SYSTEM FOR SEQUENCING AND PLANNING," which is also owned by the applicant and incorporated herein by reference. Motions and / or motion plans may be associated with episodes (as WM motions), objects (as affordances of objects), or any other experience or aspect. Motor actions and / or motor plans may be associated with labels or other symbols identifying the motor actions and / or motor plans in the experience memory store and / or memory database, and / or working memory. Examples of motor plans include: playing a song on a keyboard, drawing an image on a touchscreen, and opening a door. As described in New Zealand provisional patent application NZ744410, which is titled "Machine Interaction" and incorporated herein by reference, motor plans may be associated with user interface events and may trigger events on an application or computer system with which the agent is interacting. For example, double-tapping a target may be modified to a "double-click" on the user interface (in this case, the motor action of double-tapping a "button" triggers a double-click of the button on the user interface). Perception and maintenance of memoriesEvent-driven cognition concerns which events are perceived (and thus communicated to other subsystems of the embedded agent), and biologically realistic reinforcement schemes manage which events are retained. Memory is constructed from events; however, humans remember things that differ from expectations more strongly. Applying this principle to embedded agents allows them to be driven by events rather than being continuously triggered by sensory input. This is a form of time-compression that reduces the computation required for agents to respond to "events." Snapshots of time associated with events can be retained based on importance, volume exceeding a threshold, motion, novelty, contextual information, or any other appropriate metric. Events contribute to memory storage by triggering qualification traces in modalities that generate qualification windows suitable for learning. Qualification-based learning can be used to determine which "events" the agent retains. The occurrence of an event triggers the qualification trace (19) of the modality. Each input channel has its own qualification trace. For example, if a bottom-up event (e.g., a sufficiently loud sound) is present, the input channel for "sound" is open during the qualification window and closes after a certain period of time. Input types (modality) can be associated with unique qualification neurons. A neuron receives the input if the event occurs in its corresponding input channel. The input channel is qualified for a duration when the neuron's voltage exceeds a threshold. Leaky Integrator (LI) neurons can implement qualification traces to facilitate qualification-based learning. The activity of LI neurons is initiated at a certain level and decays over time.A “qualification window” is defined while the activity of a given LI neuron exceeds a predetermined threshold: during this period, some associated circuits are qualified for learning. FIG. 4 illustrates how a qualification trace (19) in a certain aspect creates a qualification window (18) between the event triggering and the voltage threshold of the leakage integrator neuron. In a neural behavioral model as described in US10181213B2, which is also assigned to the assignee of the present invention, qualification traces may operate over entire networks rather than individual synapses. The steps involved in creating qualification traces and qualification windows for each low-level input may be performed in the “connectors” of the programming environment described in US10181213B2. Qualification windows may be used within the CDZ to control how and when learning occurs, and to control how activity spreads through the system of convergence zones. For example, in a simple perceptual convergence zone implemented by a SOM with two input fields, each input field has an associated signal detection process that detects the onset of a signal in that field and triggers an eligibility trace for that field when the onset occurs. Learning and activity in the SOM are now controlled by several general principles that operate across all convergence zone SOMs. The eligibility window for each input field of the convergence zone SOM can be adjusted based on context parameters. For example, certain parameters of the agent's emotional state can cause certain windows to become longer or shorter: thus, frustration can make certain windows shorter, and relaxation can make them longer. In the case of a CDZ SOM that takes input directly from perceptual or motor signals, the signal detection processes associated with its input fields capture the onset of sensory or motor stimuli.When a CDZ SOM takes its input from another CDZ SOM, the signal detection process can identify the onset of a clear signal in the lower CDZ SOM. This is read from a measure of change in the activity pattern of the lower SOM, signaling that it indicates something new. This measure of change can be combined with a measure of the entropy of the lower SOM. If the SOM pattern can be interpreted as a probability distribution, its entropy can be measured. The change must result in a low-entropy state that conveys that the SOM is confidently representing its input pattern. If the lower SOM is configured to learn slowly, the upper SOM will not learn until the encodings of its own inputs in the lower SOM become sufficiently clear. Activity can flow downward from the upper CDZ SOM to the lower CDZ SOM through a "top-down activation" field. If the qualification of a low-level SOM is high, its activity activates the upper associated SOMs, which subsequently provides top-down signals to other connected low-level SOMs in real time. These can serve as real-time top-down guides for bottom-up processes that compute these inputs. In timing-parameterized qualification-based learning, two inputs to an association network must occur within a specified time of each other to allow associations to be learned within these networks. Some ASOMs may be allowed to learn associations between representations in their input fields only when qualification traces for these fields are simultaneously active (exceeding their respective thresholds). This ensures that learning occurs only when new signals arrive simultaneously or with a certain temporal coincidence, preventing the learning of associations between random or noise signals. Therefore, learning. SimultaneousQualification windows are required. For example, to learn the association between visual and tactile stimuli, simultaneous qualification windows must exist for visual and tactile representations. These simultaneous qualification windows model the simultaneous occurrence of different low-level multimode 'events'. Figure 3 illustrates qualification signals for different modalities. Figure 5 illustrates the flags and phases of a learning event. A learning event is triggered by two simultaneous events (0 and 1). While the event is occurring, the agent can acquire more information by, for example, saccading or waiting for an audio sequence to complete. This delay can be implemented using leaky integrator neurons. The length of the delay can be changed by altering the input frequency constant and / or membrane frequency constant of the leaky neurons. All illustrated periods are controlled by separate leaky neurons. At the end of the event, plasticity exists at 2 and 3 for "1st-order" SOMs, and at 4 for 2nd-order SOMs. This is a general sequence of learning events where, if multiple layers of CDZs exist in the hierarchy, the learning phases can be expanded accordingly to accommodate the hierarchy. If the eligibility trace is active for only one input field, the SOM is placed in a mode where its activity is driven solely by that field. (i.e., the ASOM alpha weights for other input fields are set to 0.) Subsequently, activity in other input fields is reconstructed from the active SOM pattern, which is discussed under "Examining content by top-down reconstruction of weights." The reconstructed value provides a useful top-down bias when the missing input field is passed by the lower-level classification process.Reconstruction also provides a simple model of perceptual 'filling-in,' and associated information missing by it is imagined. It is possible for a reconstructed (or predicted) input to arrive from the bottom up while the qualification trace for the first field is still active. In this case, the SOM will perform some additional learning to strengthen the associations used for prediction. On the other hand, if an unpredicted signal arrives in the second field while the qualification trace for the first field is still active, new associations will be learned by the SOM. In one embodiment, a time-dependent dopamine plasticity model is implemented to determine which (and to what extent) events observed by the embedded agent are learned / retained. The amount of learning that occurs, i.e., the "plasticity" of the relevant system, is influenced by several factors. The level of the relevant qualification signal(s) is one factor. Another important factor is the strength of the matching 'reward' signal. Reward signals can be implemented as neurotransmitter levels, particularly dopamine levels. Maps such as ASOMs, probabilistic SOMs, and mixtures thereof can be associated with a plasticity variable that determines when the weights of the ASOM are updated. To prevent overtraining, plasticity can be dynamically turned on at distinct moments or time intervals (e.g., when a new input arrives). Before updating the weights, if a good match exists between the input and the winning neuron, the learning frequency constant is reduced to prevent the neurons from overlearning. Memory enhancementExperience memory stores may contain two independent and competing SOMs, namely Short-term Memory (STM) and Long-term Memory (LTM). The STM can be configured for rapid online / fast learning, which can be trained in a 1-shot learning style with a high LFC and poor topographical arrangement of the data. The STM can act as a buffering system for learning that has not yet been reinforced. The STM memory can be erased after each reinforcement. The LTM can be configured for slow offline learning with a learning frequency constant of low and time decay, which can be configured to result in good topographical groupings. The LTM is trained during memory reinforcement (which can be represented as "sleeping" in the avatar). Training data from the STM can be simulated or replayed into the LTM SOM. During reinforcement, the STM can activate trained units (e.g., self-organizing map neurons) randomly or systematically to regenerate object types and images, and provide this as training data for the LTM. LTMs are trained on regenerated data pairs or tuples, and training on new data can be interleaved with their own training data. Alternatively, instead of regenerating objects using STM SOMs, raw data files from entries in an in-memory database may be provided to train LTMs. For example, LTM and STM object classifiers can compete continuously for visual recognition. Both LTM and STM object classifiers can map representations in visual space (e.g., pixels) to a common one-hot encoding of object types. If there is no sufficiently good match in the LTM, the system assumes a certain match in the STM. An LTM match is satisfied when entropy is below a threshold and winner activity is above a threshold.Therefore, STM and STM classifiers represent object types collectively and separately (where STM represents object types after the last reinforcement, and LTM represents object types learned before the last reinforcement). When faced with a new object, LTM can first check whether the object exists. If the object does not exist, the object is learned in STM. Figure 10 illustrates the displays of LTM and STM. The leftmost window shows the current fovea input for the visual system. The second window indicates which system (STM / LTM) has a more sufficient or better match for the fovea (bottom-up) input through the purple rectangle. Green indicates which system is being influenced by the top-down influence. The upper and lower halves of these windows correspond to STM and LTM, respectively. For the remaining displays, the upper half belongs to STM, and the lower half belongs to LTM. The four squared windows (highlighted by solid lines) show the SOM training history with the input image, output image, SOM weights, and SOM winner overlay (clockwise). The two windows on the right are for duplicates of the visual SOM, displaying the next prediction in the sequence image. Training from a memory database Experiences from a memory database can be used to train LTMs during sleep. Many iterations from the memory database can be provided (randomly or systematically). languageConnecting the memory systems described herein to a linguistic system can be based on meaning for the embedded agents by modeling the embedded agents in environments in which they can interact. Instead of providing a symbolic knowledge base, the embedded agents create their own meaning from the sensory inputs they receive and the actions they generate. By connecting the memory systems described herein to a linguistic system, relevant syntactic structures extracted from specific languages can capture cross-linguistic generalizations. Memories of EpisodesAn agent may experience episodes representing events in the world that can be reported as simple sentences. Episodes are events expressed as sentence-sized semantic units centered on actions and their participants. Different objects play different "semantic roles" or "thematic roles" in episodes. For example, a WM agent is the cause or initiator of an action, and a WM patient is the target or experiencer of an action. Episodes may involve agents performing actions, perceiving actions performed by other agents, planning or imagining events, or remembering past events. Like other experiences, episodes can be stored in an experience memory store and a memory database. Representations of episodes can be stored and processed in a working memory system, which processes episodes as preparatory sequences / regularities encoded as distinct actions. The WM system (40) links low-level object / episode perception with memory, (high-level) behavior control, and language. FIG. 11 illustrates a working memory system (WM system) (40) configured to process and store episodes. The WM system (40) includes WM episodes (42) and WM entities (41). WM entities (41) define entities characterized by episodes. WM episodes (42) include all elements containing episodes, including WM entities and actions. In a simple example of a WM episode (42) including individual WM agents and WM patients: the WM agents, WM patients, and WM actions are processed sequentially to fill the WM episode. Individual storage / medium (46) can be used to store WM entities and to determine whether an entity is new or a re-focused entity.Individual storage / mediums may be implemented as SOM or ASOM, where new objects are stored with the weights of newly recruited neurons, and re-focused objects update the neurons representing the re-focused objects. In one embodiment, the individual storage / medium is a convergence zone of CDZ that stores unique combinations of attributes of objects, such as location, number, and features, as distinct objects. The ASOM preferably has a high learning rate and a nearly zero neighborhood size, can learn objects immediately (in one shot), and does not have a topography configuration (thus, so that representations of different objects do not affect each other). The features of different objects are stored with the weights of different neurons; for this purpose, if the activity of a winning neuron is below a novelty threshold, a new unused neuron is recruited, otherwise the winner's weights are updated. Locations, numbers, and features are received sequentially from individual buffers (48) one at a time. The individual storage / medium (46) system is queried each time with non-zero alphas for the already filled components, and thus it can predict numbers and characteristics based on such location, for example, if an object was recently seen at a certain location. However, adaptability in the system is turned on only when the location-number-characteristic sequence is successfully completed. Subsequently, the individual storage / medium (46) system goes through its learning cycle, and when it is completed, it is kept there for a while so that the object can be copied (along with its old / new state) to its respective agent / patient buffer in the episode buffer. The episode storage / medium (47) stores WM episodes. The episode storage / medium can be implemented as a SOM or ASOM that is trained on combinations of objects and actions. In one embodiment, the episode storage / medium is a convergence zone of the CDZ that stores unique combinations of episode elements.The episode storage / medium (47) can be implemented as an ASOM having three input fields—an agent, a patient, and a behavior that take inputs from their respective WM episode slots. The mixing coefficient (alpha) for the input fields is non-zero only when the input of the input fields has been successfully processed. This means that as the input fields are progressively filled, the ASOM conveys predictions regarding the remaining input fields, such as which episodes these agents typically accompany. Individual buffers (48) sequentially acquire the attributes of the objects. When the sequence is completed (the holding gates of all buffers are closed), the adaptability in the episode storage / medium (47) is turned on, and the episode storage / medium (47) can store this particular combination of location, number, and attributes as a new object (or update the re-attention). When the episode storage / medium (47) completes its processing, the entire cycle starts again. The episode buffers sequentially acquire the elements of the episodes. Adaptability in the episode repository / media system is turned on only when the episode sequence is successfully completed. This ensures that inaccurate representations are not learned when attention mechanisms guess incorrect episode participants. Reinforcement learning components If a perceived episode brings a specific reward to the agent, it can be associated with the episode in the episode store / medium (47) as an additional input field. During episode perception, the episode store / medium (47) having a zero ASOM alpha weight as a reward will bring a prediction of the expected reward associated with the currently perceived episode. During action execution, the reward input can be used to prepare the medium to preferentially activate episodes associated with a specific value of the reward. EmotionsIf a perceived episode is associated with a specific felt emotion or emotional value, it may be associated with an episode in the episode repository / medium (47) as an additional input field. During episode perception, the episode repository / medium (47), having a zero ASOM alpha weight for the emotion, will bring the emotion associated with the predicted episode. During action execution, the emotional value may be used to prepare the medium to preferentially activate episodes associated with similar emotions. Details and Variations of Embedded Agents Top-down and bottom-up neurobehavioral modelingEmbodiments of the present invention improve artificial intelligence by combining low-level modeling capable of emergent behavior with high-level abstract models that are less biologically based but faster and more effective for given tasks. An example of an embedded agent having an architecture capable of emergent behavior is disclosed in US10181213B2, which is also assigned to the assignee of the present invention and is incorporated herein by reference. A highly modular programming environment allows for top-down cognitive architectures having interconnected high-level "black boxes" (modules (10)). Each "black box" ideally contains a collection of interconnected biologically valid low-level models, but may just as easily include abstract rules or logical descriptions, access to knowledge bases / databases, conversational engines, processing inputs using traditional machine learning neural networks, or any other computational techniques. The inputs and outputs of each module (12) are exposed as "variables" of the module that can be used to drive behavior (and consequently animation parameters). Connectors communicate the variables between the modules (12). In the simplest way, a connector copies the value of one variable to another module (12) at each time step. These high-level symbolic processes are integrated with behaviors arising from models of low-level neural circuits. Emergent behaviors interact with the high-level processes in a natural manner. The circuits performing the computation execute sequentially in parallel without any central control point. The programming environment can hardcode this principle by preventing any single control script from executing a sequence of commands for the modules (12). The programming environment supports the control of cognition and behavior through a set of neurologically valid and distributed mechanisms. Leakage integrators have three main parameters controlling timing: IFC, MFC, and voltage thresholds.When modifying parameters to control timing, the easiest way is to adjust the MFC—increase it for faster decay, and vice versa. A typical voltage threshold in CDZ can be 0.1. Emotions Emotions are modeled as regulated brain-body states in which experienced or felt components and behavioral responses exist. An emotional neuroscience-based approach is modeled in which behavioral circuits are modulated by physiological parameters. Physiological modulation alters the interaction of sensory, cognitive, and motor states. Virtual neurotransmitters are generated in response to stimuli, which can map to emotions and guide behavioral responses. For example, a "threatening stimulus" triggers the release of virtual norepinephrine and cortisol, which release energy for a fight-or-flight response, and induces fear. A smiling face or a soft voice (as evaluated by certain functions) can trigger virtual oxytocin and dopamine, which map to positive valence states and distinct emotions such as happiness, generate smiling facial expressions, and reduce agitational behaviors. AdvantagesTherefore, a real-time learning network architecture is provided—most prominent machine learning algorithms learn offline and require large amounts of data. Agents can learn autonomously (by interacting with their environment), be taught (by the user presenting specific stimuli), or be fully user-controlled (by injecting memories). The architecture is a general architecture for learning capable of learning different types of objects. Furthermore, the real-time learning network architecture is not a black box, as the causes of emergent behavior can be understood. It is possible to trace back through the paths that cause the behavior. SOMs can accept any form of input, such as one-hot vectors, RGB images, feature vectors from deep neural networks, or others. Additionally, the architecture is hierarchically stackable—low-level inputs are integrated, and further integrated with other associated regions. This allows discontinuous aspects to be indirectly related, which can lead to complex behaviors. Maps as disclosed herein enable agents to flexibly encode events and retrieve stored information during live operations. In the process of experiencing the world, a new event to be encoded is presented to the map representing the remembered events. However, while the embedded agent is experiencing these events, this same map is used in 'query mode,' where it is presented with parts of the events experienced so far and is asked to predict the remaining parts, so that these predictions can serve as a top-down guide for sensorimotor processes. SOMs provide an alternative way to construct HTM-type systems, but they have the advantage of being self-constructed as topography and consequently clustering information better.Unlike conventional deep networks that must be trained slowly and offline, SOMs support fast, one-shot learning. SOMs facilitate the learning of generalizations for the input patterns they receive. SOMs can store their memories in the weight vectors of each neuron within the map. This allows for dual representation: SOM activity represents a probability distribution across multiple options, but the content of each option is stored in the weights of each neuron and can be reconstructed top-down. ASOMs can flexibly associate inputs from different sources / modules and provide them with dynamically changeable attention / importance. The flow of activation can be reversed—ASOMs support both bottom-up (from input to activation) and top-down (from activation to reconstructed input) processing, as well as combinations thereof. SOMs can reject noisy inputs, reconstruct missing parts, return prototypes, or highlight parts where the input and prototype differ. All of these work across multiple levels of the hierarchy of SOMs. In a summary embodiment: a method for animating an embedded agent comprises: receiving a sensory input corresponding to a first representation of an experience in a first aspect; querying an experience memory store to retrieve a second representation of an experience in a second aspect; and animating the embedded agent using the second representation in the second aspect.In another embodiment: a system for storing memory for an embedded agent comprises an experience memory store filled with experiences experienced in the course of operation of the embedded agent, wherein each experience is associated with a plurality of representations of the experience in different aspects, and the experience memory store stores the representations of the experiences as neural network weights; and a memory database storing copies of the experiences stored in the experience memory store, wherein the memory database stores raw data corresponding to the representations of the experience in different aspects. In another embodiment: a method for selectively storing experiences experienced by an agent in the course of live operation of the embedded agent comprises the steps of receiving representations of inputs from a plurality of input streams for receiving inputs in a plurality of aspects, wherein each input stream is associated with at least one condition for generating an eligible trace in the input stream; and detecting simultaneous eligible traces of two or more input streams ("eligible" input streams). and includes the step of storing and associating representations of inputs from qualified input streams. In another embodiment: a method for training a SOM comprising a plurality of neurons—each neuron associated with a weight vector and a training record—comprising the steps of: receiving an input vector; determining whether the input vector is "new"; if the input vector is not new: selecting a first winning neuron that favors a higher similarity between the input vector and the winning neuron, and modifying the weight vector of the first winning neuron toward the input vector; if the input vector is new: selecting a second winning neuron that favors neurons with lower training records, and modifying the weight vector of the second winning neuron toward the input vector.In another embodiment: a method-implemented system for selectively storing experiences experienced by an agent in the course of operation of an embedded agent, comprising the steps of: receiving representations of input from a plurality of input streams for receiving input in a plurality of aspects—each input stream being associated with at least one condition for generating an eligible trace in the input stream—; detecting simultaneous eligible traces of two or more input streams ("eligible" input streams); and storing and associating representations of input from the eligible input streams.
Claims
Claim 1 A method for animating an embedded agent, comprising: receiving a sensory input corresponding to a first representation of an experience in a first modality; querying an experience memory store to retrieve a second representation of said experience in a second modality — said experience memory store being implemented as a Convergence Divergence Zone (CDZ) to store multimodal experiences previously experienced by said embedded agent in the course of real-time operation of said embedded agent as neural network weights, while simultaneously copying and storing the raw data of said experiences in a memory database —; and animating said embedded agent using said second representation in the second modality. Claim 2 delete Claim 3 A method according to claim 1, wherein the embedded agent is animated using a neurobehavioral model. Claim 4 delete Claim 5 A method according to claim 1, wherein the plurality of aspects comprises one or more of the group including visual, audio, touch, motor, neurochemical, or positional aspects. Claim 6 A method according to claim 1, wherein the experiences in the experience memory storage include created experiences. Claim 7 A method according to claim 1, wherein the second expression in the second aspect corresponds to the internal state of the embedded agent, and the animation of the embedded agent is directed by modifying its current internal state according to the extracted internal state. Claim 8 In claim 7, the second expression in the second aspect corresponds to the emotional state of the embedded agent. Claim 9 A method according to claim 1, wherein the second expression in the second aspect is re-experienced by the embedded agent during the recall. Claim 10 A method according to claim 1, wherein the experience memory storage is a collection of Self Organizing Maps (SOMs). Claim 11 In paragraph 10, the above SOMs are arranged as Convergence Divergence Zones (CDZs), a method. Claim 12 A system for storing memory for an embedded agent, comprising: an experience memory store implemented as a Convergence Divergence Zone (CDZ) filled with multimodal experiences previously experienced by the embedded agent during the course of the real-time operation of the embedded agent, wherein each experience is associated with multiple representations of the experience in different aspects, and the experience memory store stores the representations of the experiences as neural network weights; and a memory database storing copies of the experiences stored in the experience memory store, wherein the memory database stores raw data corresponding to the representations of the experiences in different aspects. Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete Claim 28 delete Claim 29 delete Claim 30 delete Claim 31 delete Claim 32 delete Claim 33 delete Claim 34 delete Claim 35 delete Claim 36 A system in which a method for directing the behavior of an embedded agent is implemented, the method comprising: receiving a sensory input corresponding to a first representation of an experience in a first aspect; querying an experience memory store to retrieve a second representation of the experience in a second aspect — wherein the experience memory store is implemented as a Convergence Divergence Zone (CDZ) to store multimodal experiences previously experienced by the embedded agent in the course of the real-time operation of the embedded agent as neural network weights, while simultaneously copying and storing the raw data of the experiences in a memory database —; and directing the behavior of the embedded agent using the second representation in the second aspect. Claim 37 delete
Citation Information
Patent Citations
Unsupervised behavior learning system and method for predicting performance anomalies in distributed computing infrastructures
US20150074023A1
System for neurobehaviorual animation
US20190172242A1