Neuromorphic systems for learning spatial and temporal patterns and associated methods
Patent Information
- Application Number
- PCT/US2025/033278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-12
- Filing Date
- 2025-06-12
- Publication Date
- 2026-01-15
AI Technical Summary
Existing neuromorphic spatial and temporal learning systems face challenges in handling high-dimensional input spaces, require extensive parameter tuning, struggle with prediction ambiguities in similar temporal sequences, operate primarily in unsupervised learning, and lack flexibility in integrating spatial and temporal components, leading to inefficiencies in processing complex real-world datasets.
A spatial pooler with supervised and unsupervised learning capabilities, using distance-based overlap measurement and distributed threshold-based winner selection, along with batched learning and directed initialization, enhances computational efficiency and adaptability. A temporal memory system integrates spatial and temporal learning flexibly, and a language model improves predictive text generation.
The system achieves more accurate classification, reduced computational complexity, and improved adaptability to diverse datasets, enabling efficient processing of large and complex datasets in resource-constrained environments.
Smart Images

Figure US2025033278_15012026_PF_FP_ABST
Abstract
Description
NEUROMORPHIC SYSTEMS FOR LEARNING SPATIAL ANDTEMPORAL PATTERNS AND ASSOCIATED METHODSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to US Provisional Application No. 63 / 659,200, titled “Neuromorphic Models of Learning for Classification Tasks” and filed on June 12, 2024, which is incorporated by reference herein in its entirety.
[0002] This application is related to U.S. Application No. 17 / 531 ,576, titled “Neural Processing Units (NPUs) and Computational Systems Employing the Same” and filed on November 19, 2021 , and US Application No. 17 / 877,724, titled “Artificial Intelligence (Al) System for Learning Spatial Patterns in Sparse Distributed Representations (SDRs) and Associated Methods” and filed on July 29, 2022, each of which is also incorporated by reference herein in its entirety.TECHNICAL FIELD
[0003] Various embodiments concern supervised and unsupervised learning algorithms for spatial and temporal pattern recognition, as well as neuromorphic computing systems capable of employing the same.BACKGROUND
[0004] Neuromorphic computing systems aim to emulate the structure and function of biological neural networks to process information. These systems typically use artificial neural networks that include interconnected nodes (referred to as “neural computational units,” “neural processing units,” “neurons,” or “cells”) that can learn and adapt based on input data. Neuromorphic computing uses Sparse Distributed Representations (SDRs) in which information is encoded using a smaller subset of active neural computational units within a larger population. In this context, SDRs help improve computational efficiency and robustness by reducing redundancy and focusing processing power on the most relevant features of the input data.
[0005] Spatial and temporal pattern recognition is often used in neuromorphic systems. Spatial patterns refer to the arrangement or distribution of features in data, while temporal patterns involve sequences or time-dependent relationships. Existingneuromorphic spatial learning methods can be used to extract relevant features and identify recurring patterns in input data, but only in an unsupervised manner. Existing neuromorphic temporal learning systems typically work by forming connections between neurons that represent sequential patterns, enabling the systems to predict future states based on learned temporal sequences and context.
[0006] As a result, existing neuromorphic spatial and temporal learning systems face several limitations. Traditional spatial poolers can struggle with handling highdimensional input spaces and can require extensive parameter tuning to achieve optimal performance across different datasets. Existing temporal learning implementations frequently encounter difficulties in distinguishing between similar temporal sequences that diverge only in later steps, leading to prediction ambiguities. Additionally, both systems typically operate in purely unsupervised learning paradigms, limiting their applicability in scenarios where labeled data is available and could enhance learning efficiency. The integration between spatial and temporal components in current implementations is often rigid, making it challenging to adapt these systems to domains with varying spatial and temporal dependencies. Furthermore, computational efficiency remains a concern, particularly when scaling to complex real-world applications with high-dimensional data and long temporal sequences. As a result, there is a growing need for more flexible and scalable neuromorphic architectures that can seamlessly integrate spatial and temporal learning, leverage supervised learning when appropriate, and efficiently process large- scale, real-world datasets.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Detailed descriptions of implementations of the present technology will be described and explained through the use of the accompanying drawings.
[0008] Figure 1 includes a diagrammatic illustration of a spatial pooler having 5 cortical columns.
[0009] Figure 2 includes a diagrammatic illustration of operations performed by an unsupervised spatial pooler.
[0010] Figure 3 includes a diagrammatic illustration of an example temporal memory having 5 cortical-columns of 3 neural computational units each.
[0011] Figure 4 illustrates the distal synapses of a neural computational unit in the temporal memory shown by Figure 3.
[0012] Figure 5 includes a diagrammatic illustration of operations performed by a temporal memory over two successive time-steps.
[0013] Figure 6A includes a diagrammatic illustration of “burst” columns of a temporal memory.
[0014] Figure 6B includes an illustration of a distal connectivity map of a temporal memory.
[0015] Figure 7 includes a diagrammatic illustration of active neural computational units at a first time step in the temporal memory of Figure 6A.
[0016] Figure 8 includes a diagrammatic illustration of the neural computational units in the temporal memory of Figure 6A, in which certain neural computational units are predicted to lie in winner columns.
[0017] Figure 9 includes a diagrammatic illustration of the neural computational units in the temporal memory of Figure 6A, in which certain neural computational units are determined to be active at a second time step.
[0018] Figure 10 includes a diagrammatic illustration of operations performed by a spatial pooler operating in an unsupervised mode.
[0019] Figure 11 includes a diagrammatic illustration of operations performed by an unsupervised spatial pooler.
[0020] Figure 12 includes a diagrammatic illustration of pipelined operations performed by the spatial pooler of Figure 10.
[0021] Figure 13 includes a diagrammatic illustration of a compact representation of operations performed by the spatial pooler of Figure 10 in the unsupervised mode.
[0022] Figure 14 includes a diagrammatic illustration of operations performed by the spatial pooler of Figure 10 in a supervised mode.
[0023] Figure 15 includes a diagrammatic illustration of two spatial poolers used to emulate the functions of a temporal memory.
[0024] Figure 16 includes a diagrammatic illustration of autoregressive encoder operations of a language model.
[0025] Figure 17 includes a diagrammatic illustration of autoregressive encoder operations for input sequence encoding during an inference phase.
[0026] Figure 18 includes a diagrammatic illustration of autoregressive decoder operations for output sequence generation during the inference phase.
[0027] Figure 19 includes a diagrammatic illustration of input processing stages of a pipeline architecture.
[0028] Figure 20 includes a flow diagram illustrating a process for selecting implementations of the pipeline architecture.
[0029] Figure 21 includes a diagrammatic illustration of an output token generator being trained in accordance with a first implementation.
[0030] Figure 22 includes a diagrammatic illustration of output text sequence generation in accordance with a second implementation.
[0031] Figure 23 includes a diagrammatic illustration of output text sequence generation in accordance with a third implementation.
[0032] The technologies described herein will become more apparent to those skilled in the art from studying the Detailed Description in conjunction with the drawings. Embodiments or implementations describing aspects of the invention are illustrated by way of example, and the same references can indicate similar elements. While the drawings depict various implementations for the purpose of illustration, those skilled in the art will recognize that alternative implementations can be employed without departing from the principles of the present technologies. Accordingly, while specific implementations are shown in the drawings, the technology is amenable to various modifications.DETAILED DESCRIPTION
[0033] Neuromorphic computing systems aim to emulate the structure and function of biological neural networks for information processing tasks. These systems typically use artificial neural networks composed of interconnected nodes (referred to as “neural computational units,” “neural processing units,” “neurons,” or “cells”) thatcan learn and adapt based on input data. SDRs are often employed in such systems, where information is encoded using a small subset of active neural computational units within a large population, thereby improving computational efficiency and robustness by reducing redundancy and focusing on salient features.
[0034] Spatial and temporal pattern recognition can be implemented using neuromorphic systems. Spatial patterns refer to the arrangement or distribution of features in data, while temporal patterns involve sequences or time-dependent relationships. Learning algorithms enable neuromorphic systems to extract relevant features and identify recurring patterns in input data. Until now, only unsupervised learning - which finds structure without explicit labels - could be applied to develop pattern recognition capabilities in neuromorphic architectures, especially those employing Hebbian learning. Many existing learning systems face challenges in efficiently handling both spatial and temporal patterns, especially when dealing with noisy or incomplete data. Additionally, some systems struggle to generalize well from limited training examples or adapt to new patterns over time, which can significantly limit their effectiveness in dynamic environments or real-world environments.
[0035] Many current approaches also face difficulties in scaling to handle large and complex datasets while maintaining computational efficiency. The ability to process information quickly with low power consumption is crucial for practical applications of neuromorphic systems, such as in mobile phones, tablet computers, edge computing devices, and Internet of Things (loT) devices where energy and computational resources are limited. Furthermore, integrating spatial and temporal learning capabilities in a unified architecture that can flexibly handle different types of input patterns remains an ongoing challenge. Addressing these limitations could enable more robust and versatile neuromorphic systems for a wide range of pattern recognition and prediction tasks, including speech recognition, anomaly detection, and autonomous control.
[0036] The present disclosure describes a spatial pooler that can implement supervised and unsupervised learning capabilities for efficient spatial pattern recognition. As further discussed below, the spatial pooler can use overlap measurement based on distance determination and a distributed threshold-based winner selection mechanism. The disclosed pooler architecture enables moreaccurate classification while reducing computational complexity compared to existing architectures. The disclosed spatial pooler can also incorporate batched learning and directed initialization techniques to improve convergence and prediction accuracy. These enhancements allow the spatial pooler to handle high-dimensional input spaces more effectively and adapt to different datasets with reduced parameter tuning, addressing key limitations of many current spatial pooler designs and supporting broader applicability across diverse domains.
[0037] Building on the spatial pooler, architectures for a temporal memory system are disclosed for learning both spatial and temporal patterns. The architectures disclosed use two synchronized supervised spatial pooler components to emulate the functionality of a temporal memory system while enabling more efficient hardware implementation. The disclosed architectures enable the temporal memory system to better distinguish between similar temporal sequences that diverge in later steps, reducing prediction ambiguities encountered by many existing temporal learning systems. The temporal memory system’s ability to integrate spatial and temporal learning in a flexible manner addresses the challenge of adapting to domains having varying spatial and temporal dependencies, such as video analysis and sequential decision-making tasks.
[0038] A disclosed language model extends the capabilities of the temporal memory system to enable predictive text generation tasks. The language model incorporates an encoder-decoder architecture and introduces mechanisms for handling context-dependent token representations and end-of-sequence prediction. These enhancements enable the language model to generate more coherent and contextually appropriate text sequences compared to many current language modeling approaches. In addition, the language model’s neuromorphic design offers potential advantages in computational efficiency and adaptability for natural language processing tasks, such as machine translation, summarization, and conversational artificial intelligence (Al).
[0039] Integrating the components described above, a pipeline architecture is disclosed that provides an end-to-end solution for sequential pattern recognition applications such as predictive text generation, time series analysis, video processing, network security, or industrial monitoring. The pipeline architecture leverages thedisclosed language model as its learning algorithm and incorporates flexible embedding and tokenization stages. The pipeline architecture enables more efficient processing of both spatial and temporal patterns in input data while maintaining the ability to scale to large and complex datasets. The neuromorphic design of the pipeline architecture offers potential improvements in processing speed and energy efficiency compared to many existing processing systems, addressing the need for computational efficiency in practical applications of neuromorphic computing and supporting deployment in resource-constrained environments.
[0040] The methods disclosed herein can be executed entirely in software, for example, on conventional processing units, such as central processing units (CPUs) and graphics processing units (GPUs), or using specialized processing units, such as neural processing units (NPUs). Accordingly, the approaches introduced here could be implemented through the execution - by a conventional processing unit and / or a specialized processing unit - of instructions in a non-transitory medium. This flexibility allows for deployment across a wide range of hardware platforms, from general- purpose computers to dedicated neuromorphic integrated circuits (also called “chips”). Note that while embodiments may be described in the context of software, features of those embodiments may be similarly applicable to firmware and hardware, enabling further optimization depending on whether there is a preference for speed, power consumption, or computational resource consumption.
[0041] The description and associated drawings are illustrative examples and are not to be construed as limiting. This disclosure provides certain details for a thorough understanding and enabling description of these examples. One skilled in the relevant technology will understand, however, that the invention can be practiced without many of these details. Likewise, one skilled in the relevant technology will understand that the invention can include well-known structures or features that are not shown or described in detail, to avoid unnecessarily obscuring the descriptions of examples.Terminology
[0042] References in the present disclosure to "an embodiment" or "some embodiments" mean that the feature, function, structure, or characteristic being described is included in at least one embodiment. Occurrences of such phrases donot necessarily refer to the same embodiment, nor are they necessarily referring to alternative embodiments that are mutually exclusive of one another.
[0043] Unless the context clearly requires otherwise, the terms "comprise," "comprising," and "comprised of" are to be construed in an inclusive sense rather than an exclusive or exhaustive sense. That is, in the sense of "including but not limited to." The term "based on" is also to be construed in an inclusive sense. Thus, the term "based on" is intended to mean "based at least in part on."
[0044] The terms "connected," "coupled," and variants thereof are intended to include any connection or coupling between two or more elements, either direct or indirect. The connection or coupling can be physical, logical, or a combination thereof. For example, elements may be electrically or communicatively coupled to one another despite not sharing a physical connection.
[0045] The term "module" may refer broadly to software, firmware, hardware, or combinations thereof. Modules are typically functional components that generate one or more outputs based on one or more inputs. A computer program may include or utilize one or more modules. For example, a computer program may utilize multiple modules that are responsible for completing different tasks, or a computer program may utilize a single module that is responsible for completing all tasks.
[0046] When used in reference to a list of multiple items, the word "or" is intended to cover all of the following interpretations: any of the items in the list, all of the items in the list, and any combination of items in the list.
[0047] As used herein, the term “tokenizer” may refer to a component configured to convert any sequence of items into discrete units called tokens. For example, the tokenizer may break down input text into words, subwords, characters, or other linguistic units based on predetermined rules or learned patterns. The tokenizer may generate tokens corresponding to entries in a predefined dictionary, where each token may be assigned a unique identifier.
[0048] As used herein, the term "embedder" may refer to a component configured to convert tokens into numerical feature vectors of predetermined dimensions. The embedder may generate context-dependent embeddings where the same token may have different numerical representations based on its usage indifferent linguistic contexts. The embedder may utilize pretrained models or may be trained to capture semantic relationships between tokens.
[0049] As used herein, the term "SDR generator" may refer to a component configured to convert input data into sparse binary vectors where only a small subset of bits are active within a larger population. The SDR generator may process embeddings or other numerical representations to create sparse bit patterns that encode information in a distributed manner, enabling efficient pattern recognition and learning in neuromorphic computing systems.
[0050] As used herein, the term "language model" may refer to a computational system configured to learn patterns in text sequences and generate predictions for subsequent tokens or text. A language model may process input sequences to establish temporal relationships between tokens and may utilize neuromorphic architectures with spatial poolers to perform predictive text generation tasks.
[0051] As used herein, the term "label generator" may refer to a component configured to create labels for input tokens based on predetermined criteria or learned patterns. A label generator may generate single labels, multi-labels, or pseudo-labels depending on the operational mode, and may process embeddings or token identifiers to produce labels for supervised or semi-supervised learning.
[0052] As used herein, the term "classifier" may refer to a component configured to categorize input data into predefined classes or categories. A classifier may process SDRs, embeddings, or other feature vectors to generate predictions with associated confidence scores, and may translate internal representations into output tokens or class labels for decision-making purposes.
[0053] Each of a tokenizer, an embedder, an SDR generator, a language model, a label generator, and a classifier may be implemented using hardware, software, firmware, or a combination thereof.Introduction to Neuromorphic Learning
[0054] Neuromorphic computing systems aim to emulate the structure and function of biological neural networks to process information. These systems can use spatial poolers to process SDRs for learning and recognizing spatial patterns in input data.
[0055] Figure 1 includes a diagrammatic illustration of an example spatial pooler 100 having five cortical columns, co / o-4. A spatial pooler is a component that processes input data to create sparse distributed representations, learning spatial patterns by adjusting synaptic connections between the input data space and columns of neural computational units through competitive learning mechanisms. The spatial pooler 100 can be used to implement a neuromorphic model for spatial pattern learning and recognition and can operate in an unsupervised mode (sometimes referred to as an “unsupervised spatial pooler”). The spatial pooler 100 can process an SDR 104 using its cortical columns. Each cortical column includes a neural computational unit and feed-forward (proximal) synapses. The proximal synapses connect a column to specific offsets in the iSDR 104. A synapse is a connection between two neural computational units or between a neural computational unit and an input bit, allowing signals to be transmitted. The synapses represent binary-weighted connections between the neural computational units of the spatial pooler 100.
[0056] Figure 2 includes a diagrammatic illustration of operations performed by an unsupervised spatial pooler. The solid lines shown by Figure 2 represent input and output dataflow to and from the unsupervised spatial pooler. The dotted lines represent internal dataflow of the unsupervised spatial pooler. In Operation 1.1 .1 , an overlap score for each column of the unsupervised spatial pooler is determined by measuring the overlap of its proximal synapses with the 1 -bits in an iSDR. A boosted overlap score can be determined by adding a boost-score to the overlap score. The boostscore is added to the overlap score to mimic the biologically-observed behavior of homeostasis. The boost-score is inversely proportional to the duty-cycle (winning frequency) of each column. The boost-score incentivizes neurons having a lower-duty cycle to win more and participate in the learning of the unsupervised spatial pooler.
[0057] In Operation 1.1.2, one or more columns having the highest boosted overlap score are identified as “winner columns” for the iSDR. As used herein, a winner column is a cortical column selected based on certain predetermined criteria (e.g., overlap score) for a given input pattern. The output produced by Operation 1 .1 .2 is a bit-vector called an output SDR (oSDR). Each bit in the oSDR corresponds to a column in the unsupervised spatial pooler. If and only if the column is a winner column, then the corresponding bit is set to 1 . The number of winner columns per iSDR is typically low and therefore the oSDR is a sparse bit-vector.
[0058] During training of the unsupervised spatial pooler, the strength or permanence of the synapses of the winner columns can be adjusted, e., using Hebbian learning. Permanence refers to the strength value of a synaptic connection that determines whether the synapse is connected or disconnected based on threshold criteria. Each column is assigned a fixed set of offset locations in the iSDR to process. This set is called its potential field. A synaptic connection may be created between a cortical column and each location in its potential field. However, a synapse is connected only when its associated synaptic strength is greater than a connection threshold. After the unsupervised spatial pooler is initialized, the unsupervised spatial pooler learns to grow or decay the synaptic strengths (and synaptic connections) based on processing training iSDRs using Hebbian learning in an unsupervised manner. As used herein, a "training iSDR" refers to an iSDR used during the learning phase to train synaptic connections. For example, in Operation 1.2.1 , the synaptic strengths of the synapses of the winner columns are updated based on the associated bits in the iSDR. If the synaptic strength of a synapse exceeds the connection threshold, the connection is accordingly updated in Operation 1.2.2.
[0059] After training is complete, the unsupervised spatial pooler learns to pool an iSDR into an oSDR by retaining semantically important information and discarding noise. As a result, the generated oSDRs are easier to cluster or classify.
[0060] A spatial pooler can be used to implement or emulate a temporal memory as shown by Figure 3. A temporal memory is a structure that can learn and recognize sequences of patterns over time, using multiple interconnected neural computational units within the temporal memory’s columns to predict future states based on current and past inputs, enabling context-dependent processing. The diagrammatic illustration of the example temporal memory 300 shown by Figure 3 has 5 columns of 3 neural computational units each. The temporal memory shown by Figure 3 is also sometimes referred to as a “temporal pooler.” Figure 3 also shows the shared proximal synapses 312 of column co / o (labeled 308). Each neural computational unit is assigned two identifiers. The identifier (J,k) to the right of each neural computational unit indicates that the neural computational unit is theneural computational unit of the 7thcolumn. The identifier to the left of each neural computational unit is determined by the expression c*j + k, where c denotes the number of neural computational units per column. In the example temporal memory of Figure 3, c equals three. Each columnin the temporal memory can process input data (e.g., iSDR 304) using its proximal synapses. In this aspect, the operation of the temporal memory is similar to that of a spatial pooler, i.e., the proximal synapses operate at the column level. For the ISDR 304, a set of winner columns are determined from the columns of the temporal memory based on their proximal overlap scores. The strengths of the proximal synapses of the winner columns are then updated based on Hebbian learning.
[0061] Figure 4 illustrates the distal synapses of a neural computational unit having the identifier 6 in the temporal memory 300 shown by Figure 3. Distal synapses are directional connections from one neural computational unit in a neuromorphic model (such as a temporal memory) to another neural computational unit that convey activation signals between successive iSDRs. A neural computational unit may be connected to another neural computational unit, including to neural computational units in the same column, or even to itself. The incoming distal synapses of the neural computational unit having the identifier 6 are shown by Figure 4 using solid lines and the outgoing distal synapses are shown using dashed lines.
[0062] A temporal memory typically computes one iSDR per time-step. A neural computational unit in the temporal memory becomes “active” when it receives sufficient input to exceed its activation threshold, typically through its synaptic connections. Active neural computational units represent recognized patterns and can influence other neural computational units through their connections. If a neural computational unit is active in a time step, the active neural computational unit sends out a predictive signal to each neural computational unit connected to its outgoing distal synapses. At the beginning of the next time step, each neural computational unit accumulates the incoming predictive signals from the previous time step as its predictive score. If the predictive score of a neural computational unit is greater than a predictive threshold, the neural computational unit enters a predicted state in which the neural computational unit and its associated column are predicted to become winners in the current time step.
[0063] If a column of the temporal memory that was predicted to become a winner column truly becomes a winner column in the current time step, then the prediction was correct, and the strengths of the associated distal synapses of the column are increased. Otherwise, the prediction was incorrect, and the strengths ofthe associated distal synapses are decreased. The temporal memory thus learns to predict which columns will become winner columns in the next time step in the temporal context of the current winner columns. The multiple neural computational units in a column enable the column to learn higher-order sequences, i.e. , sequences having a length greater than one.
[0064] Figure 5 includes a diagrammatic illustration of operations performed by a temporal memory over two successive time steps, / and / +1 . The input and output dataflows to and from the temporal memory are shown in solid lines and the internal dataflows are shown in dotted lines. The dotted lines from Operations 1 .2 to 1 .1 , and Operations 2.2.1 and 2.2.2 to 2.1 depict dataflows between the two successive time steps.
[0065] Operations 1.1 and 1 .2. are related to spatial-pooling, in which winner columns are selected from among the columns of the temporal memory based on an iSDR and the proximal synapses of the columns. Operation 1 .1 of Figure 5 includes the combined function of Operations 1 .1.1 and 1.1.2 shown by Figure 2. Similarly, Operation 1 .2 of Figure 5 includes the combined function of Operations 1 .2.1 and 1 .2.2 shown by Figure 2.
[0066] Temporal pooling is a process that can learn and recognize sequences of patterns over time by adjusting connections between neural computational units across multiple time steps, enabling prediction of future states based on current and past inputs. The operations shown by Figure 5 that are related to temporal pooling are numbered beginning with 2 and are shown in Figure 5. The predicted neural computational units for time step / +1 shown by Figure 5 are based on the active neural computational units (determined in Operation 2.2.1 ) and the updated distal synapses (determined in Operation 2.2.2) of time step / . A predicted neural computational unit refers to a neural computational unit that is predicted to become active in the next time step based on the current temporal context and / or the neural computational units that are active in the current time step.
[0067] At Operation 2.1 of time-step / +1 , a predictive overlap score of each neural computational unit in the temporal memory is determined. The neural computational units having a predictive score greater than a predictive threshold are determined to be predicted neural computational units. Based on the predicted neural computationalunit, the predicted columns may be optionally determined in Operation 2.1.1 to generate a predicted oSDR. In the predicted oSDR, the bits located at offsets corresponding to predicted winner column identifiers are set to 1. The predicted column identifiers are determined by selecting the values of j, where (J,k) is the identifier of a predicted neural computational unit in time-step / +1.
[0068] In Operation 2.2.1 , the predicted neural computational units determined in Operation 2.1 are synchronized with the winner columns determined in Operation 1.1. The resulting active neural computational units for time step / +1 are derived from either (1 ) a predicted neural computational unit belonging to a winner column or (2) a representative neural computational unit from a winner column that has no predicted neural computational unit. In the latter scenario, the winner column is referred to as having “burst.” The representative neural computational unit can be selected by identifying the neural computational unit in the burst column that has the highest predictive score. If there is a tie, the neural computational unit having the lowest number of input distal synaptic connections is selected from among the neural computational units that have the highest predicted score. If a tie persists, it can be broken by a random selection between the tied neural computational units.
[0069] The strengths of input distal synapses of the predicted neural computational units are adjusted in Operation 2.2.2 as shown by Figure 5. For example, the strengths of the input distal synapses of active neural computational units at time step / +1 that are connected to active neural computational units of time-step / are increased. The strengths of other input distal synapses to such neural computational units are decreased. Conversely, the predicted neural computational units that do not transform into active neural computational units at time step / +1 are labeled as incorrectly predicted neural computational units. The strengths of the input distal synapses to the incorrectly predicted neural computational units from the active neural computational units of time step i are decreased.Illustrative Example
[0070] Figure 6A includes a diagrammatic illustration of burst columns of a temporal memory 600. As described with reference to Figure 5, burst columns occur when a winner column in a temporal memory system has no predicted neural computational units. In this case, all the neural computational units in the columnbecome active, representing a novel or unexpected input pattern. The operations illustrated with reference to Figure 5 are described in more detail with respect to two example time steps as follows. At a first time step, assume that there are no active neural computational units from a previous time step. Consequently, there are no predicted neural computational units at the end of Operation 2.1 . Assume further that coh and co were selected as winner columns after processing the first iSDR (iSDRi) in Operation 1.1 as shown by Figure 5. In Operation 2.2.1 of Figure 5, both these columns are determined to have burst as shown by Figure 6A. The predictive scores of neural computational units having the identifiers 3, 4, 5, 6, 7, and 8 are determined and are found to be tied at 0. The number of incoming distal connections of these neural computational units are then determined as shown by Figure 6B.
[0071] Figure 6B includes an illustration of a distal connectivity map of the temporal memory 600 shown by Figure 6A. The distal connectivity map shows which neural computational units in the temporal memory have distal synaptic connections to other neural computational units, enabling the learning and prediction of temporal sequences. The distal connectivity map shown by Figure 6B is a binary matrix having n rows and n columns. Here, n corresponds to the total number of neural computational units in the temporal memory. In the example of Figure 6A, n equals 15. A “1 ” value in a position (c,d) of the distal connectivity map shown by Figure 6B means that there is a connected distal synapse from neural computational unit c to neural computational unit d.
[0072] Figure 7 includes a diagrammatic illustration of active neural computational units in the temporal memory 600 at the first time step after the processing described with reference to Figure 6A is performed. The number of incoming distal connections into the neural computational unit 3, neural computational unit 4, and neural computational unit 5 are 5, 9 and 7, respectively, with neural computational unit 3 having the least incoming distal connections. Therefore, neural computational unit 3 is selected as the active neural computational unit to represent the burst column, coh . Similarly, the number of incoming distal connections into neural computational unit 6, neural computational unit 7, and neural computational unit 8 are 5, 9, and 2 respectively, with neural computational unit 8 having the least incoming distal connections. Therefore, neural computational unit 8 is selected as the active neural computational unit to represent the burst column, co / 2. Both neuralcomputational unit 3 and neural computational unit 8 are shown shaded in dark grey by Figure 7.
[0073] Table I shows the predictive scores for the neural computational units in the temporal memory illustrated by Figure 7 during the second time step. In the second time step, Operations 1.1 and 2.1 shown by Figure 5 are performed in parallel. In Operation 2.1 , the predictive score of all the neural computational units in the temporal memory are determined. Because neural computational unit 3 and neural computational unit 8 were the predicted neural computational units at the end of the first time step, the predictive scores of the neural computational units at the second time step can be determined as shown by Table I.Table I. Predictive scores for the neural computational units (cells) in the temporal memory.
[0074] The predicted neural computational units are identified based on a predictive score threshold PSmin. In the example of Figure 7 and Table I, PSmin is 2. Therefore, neural computational unit 9, neural computational unit 10, and neural computational unit 12 are determined to be predicted neural computational units. These neural computational units are shown as hatched in Figure 8. Figure 8 includes a diagrammatic illustration of the neural computational units in the temporal memory 600, in which neural computational unit 9, neural computational unit 10, and neural computational unit 12 are predicted to lie in winner columns in the second time step.
[0075] Figure 9 includes a diagrammatic illustration of the neural computational units in the temporal memory 600 of Figure 6A, in which certain neural computational units are determined to be active at a second time step. Without loss of generality, assume that the columns coin and coh are winner columns after the second iSDR (iSDR +i) has been processed in Operation 1 .1 shown by Figure 5. At Operation 2.2.1 (shown by Figure 5), neural computational unit 9 and neural computational unit 10 belong to the winner column cola and therefore transform into active neural computational units. Because there were no predicted neural computational units incolumn colo, column colds determined to have burst. As a result, neural computational unit 2 is selected as a representative active neural computational unit for column colo because neural computational unit 2 has the highest predictive score among the neural computational units in column colo as shown by Figure 9.
[0076] The permanence values of the distal synapses connecting neural computational unit 3 and neural computational unit 8 - which are active neural computational units at the end of the first time step - to neural computational unit 2, neural computational unit 9, and neural computational unit 10 - which are active neural computational units at the second time step - are increased. The permanence values of the other incoming synapses of neural computational unit 2, neural computational unit 9, and neural computational unit 10 are decreased. Note that the predicted neural computational unit 12 belongs to column co which was not determined to be a winner column at the end of the second time step. Hence, neural computational unit 12 was incorrectly predicted to become active. As a result, the synaptic strengths of the distal synapses connecting neural computational unit 3 and neural computational unit 8 - which are active at the end of the first time step - to neural computational unit 12 are decreased.
[0077] Further, if the permanence of a potential distal synapse increases beyond the connection threshold, then the potential distal synapse is transformed into an active distal synapse. A potential distal synapse is an inactive connection between neural computational units whose permanence value is below the connection threshold but could become active. On the other hand, if the permanence of an active distal synapse decreases below the connection threshold, then the active distal synapse is disconnected and transformed into a potential distal synapse. An active distal synapse is a connected temporal link between neural computational units whose permanence value exceeds the connection threshold, enabling predictive signal transmission.Selected Embodiments of Spatial Poolers
[0078] Introduced here are spatial poolers that can implement supervised, unsupervised, and semi-supervised clustering and classification systems that can operate on completely unlabeled, partially labeled, or fully labeled datasets. In some embodiments, a spatial pooler uses distance-based overlap measurement todetermine overlap between synaptic connections and iSDRs, where distances are established for multiple columns containing neural computational units. The spatial pooler can employ a distributed threshold-based winner selection mechanism that compares established distances to column-specific thresholds, enabling decentralized column selection without requiring global synchronization. The disclosed spatial pooler may incorporate directed initialization, batched learning, and / or ensemble learning to enhance prediction accuracy and reduce computational complexity compared to existing spatial learning implementations.A. Overlap Measurement
[0079] As shown by Figure 2, to determine overlap between the synaptic connections of a spatial pooler and an iSDR, the overlap of the proximal synapses of each column with the iSDR are determined. The synaptic connections may be represented as a binary connection vector having the same length as the iSDR, where bit positions corresponding to connected synapses are set to 1 . The overlap score may be determined by counting bit positions where both the iSDR and connection vector are set to 1 , establishing a distance between the iSDR and synaptic connections. Alternatively, bitwise operations or different distance metrics may be employed to determine the relationship between vectors.
[0080] With conventional overlap measurement methods, more densely connected columns may have higher probability of selection as winner columns because greater active connections do not negatively impact overlap scores. Consequently, learning capacity may concentrate in densely connected neural computational units while sparsely connected units receive reduced participation. The disclosed distance-based overlap measurement addresses this limitation by establishing distances that account for connection density differences between columns. In some embodiments, Hamming distance, Minkowski distance, Euclidean distance, Manhattan distance or Jaccard distance can be used to determine the overlap.
[0081] Various distance metrics may be evaluated to determine optimal performance while maintaining computational efficiency. Distance-based measurements may be calculated by determining differences between the iSDR and connection vector of each column. The distance may be reduced when vectors havesimilar density characteristics. If a column becomes more densely connected, automatic penalties through increased distance values can reduce selection probability. Greater density differences may result in higher penalties. In the Hebbian learning embodiments used herein, columns modify connection strengths through winner column selection, providing automatic regulation of winning frequency and connection density.B. Threshold-Based Winner selection
[0082] In conventional spatial poolers, columns with the highest overlap scores are typically selected as winner columns. This selection process may require overlap scores from all columns to be collected, sorted, and ranked. Such a selection process may require synchronization and can incur computational and communication overhead. Additionally, there is no guarantee of minimum overlap between a winner column's connection vector and the iSDR. Consequently, the synaptic strengths of a winner column may be modified for an iSDR with which it has minimal overlap, potentially reducing learning efficiency and pattern recognition accuracy. In contrast, the spatial pooler disclosed herein uses distributed threshold-based winner column selection where each column independently compares its distance to a columnspecific threshold, avoiding synchronization overhead and guaranteeing minimum overlap.
[0083] Figure 10 includes a diagrammatic illustration of operations performed by a spatial pooler operating in an unsupervised mode. As used herein, the term "unsupervised mode" may refer to a learning mode where the spatial pooler processes input data without ground truth labels, using pseudo-labels or unknown labels for pattern recognition and clustering operations. As used herein, a ground truth label refers to the correct, known classification or category associated with a training sample for supervised learning. A pseudo-label refers to a placeholder label assigned to unlabeled data, e.g., unknown, enabling semi-supervised or unsupervised learning operations.
[0084] As shown in Figure 10, the distributed overlap calculation for a column with respect to an iSDR is performed in Operation 1.1.1. The calculated overlap scores may include distance measurements between the iSDR and synaptic connections for each column. Operation 1.1.2 may employ threshold-based winner selectionmechanisms. The distance established for a column may be compared with columnspecific thresholds, and one or more columns meeting the threshold criteria may be identified as winner columns for that iSDR. In some embodiments, when the distance between a column's synaptic connections and the iSDR is determined to be less than the column's associated threshold, that column is identified as a winner column. Alternatively, ranking algorithms, probabilistic selection methods, or competitive learning mechanisms may be employed to identify the one or more winner columns. The threshold-based approach may enable decentralized processing where each column can independently evaluate its selection status, or parallel processing architectures may be utilized to reduce computational dependencies. Additionally, the threshold may provide minimum overlap guarantees between the connected proximal synapses of winner columns and the iSDR.
[0085] An oSDR including a sparse bit-vector is produced where each bit position corresponds to a specific column in the spatial pooler. When a column is identified as a winner column through threshold-based selection, the corresponding bit at that column's offset position is set to one, otherwise it remains zero. The resulting oSDR maintains sparsity with only a small number of active bits representing the winner columns, creating a distributed encoding that retains semantically important spatial pattern information while discarding noise for subsequent temporal processing or classification operations.C. Batched Learning
[0086] In conventional spatial poolers, synaptic strength updates and connection updates may be performed after processing each iSDR. In some cases, the resulting learning rate may be too aggressive. In the disclosed spatial pooler, batched learning approaches can lead to improved and more generalizable results. Additionally, batched operations enable parallel processing of operations across multiple time steps and enable improved computational scheduling that can improve overall system efficiency and throughput.
[0087] Figure 11 includes a diagrammatic illustration of operations performed by an unsupervised spatial pooler for two successive time steps. The requirement to update the proximal synaptic connections at the end of one time step (in Operation 1 .2.2) before calculating overlap scores in a subsequent time step (in Operation 1.1.1 )may create computational dependencies that can limit parallel processing capabilities or reduce opportunities for concurrent execution of operations across time steps. This dependency is addressed in the disclosed spatial pooler by implementing batched learning approaches where synaptic connection updates are deferred until predetermined intervals.
[0088] Figure 12 includes a diagrammatic illustration of pipelined operations performed by a spatial pooler for multiple time steps corresponding to a batch size. As used herein, a batch refers to a predetermined group of input samples processed together before updating synaptic connections, enabling pipelined operations and improved efficiency. As shown in Figure 12, Operation 1 .2.2 may be configured to accumulate synaptic strength updates throughout a batch for each column. At the batch boundary, Operation 1 .2.3 may calculate updated synaptic connections based on the accumulated synaptic strengths. As used herein, a batch boundary (or a boundary of a batch) refers to a predetermined point where accumulated synaptic strength updates are applied to connections and accumulators are reset before processing continues. The operations shown by Figure 12 may be implemented using parallel processing architectures or may utilize decentralized computation where processing occurs locally at individual columns without requiring global synchronization.
[0089] Batched operations can reduce the computational dependencies between consecutive time steps within a batch, with potential synchronization requirements limited to accumulator management. Therefore, operations across successive time steps of a batch may enable parallel processing or pipelined execution. The computational dependency between the final operation of one batch and the initial operation of a subsequent batch may be managed through appropriate scheduling mechanisms. The synaptic strength accumulators may also be reset at batch boundaries. As used herein, the term "batch boundary" refers to a predetermined point in processing where accumulated operations are executed, parameters are updated, and accumulators are reset before beginning the next processing batch.
[0090] Figure 13 includes a diagrammatic illustration of a compact representation of operations performed by the spatial pooler of Figure 10 in the unsupervised mode. As shown by Figure 13, Operation 1.1 depicts the combined functions of Operations1.1.1 and 1 .1.2 (shown by Figure 10) performed sequentially. Winner column identifiers are determined for an iSDR in Operation 1 .1 based on the proximal synaptic connections of the columns. Operation 1.2 represents the combined functions of Operations 1 .2.1 , 1 .2.2, and 1 .2.3 (shown by Figure 10). In Operation 1 .2, iSDRs and oSDRs are received for multiple time steps and synaptic strength updates are accumulated over a batch. At the batch boundary, the synaptic connections are updated in Operation 1 .2.D. Semi-Supervised Hebbian Learning
[0091] Figure 14 includes a diagrammatic illustration of operations performed by the spatial pooler of Figure 10 in a supervised mode. As used herein, the term "supervised mode" may refer to a learning mode where the spatial pooler processes input data with accompanying ground truth labels, enabling classification by comparing column labels with training labels. Conventional spatial poolers typically update proximal synaptic strengths and connections in an unsupervised manner. Consequently, conventional systems may be unable to learn class boundaries as defined by ground truth labels of training samples for supervised classification purposes. Addressing these limitations, the disclosed spatial pooler incorporates supervised learning capabilities whose operations are depicted by Figure 14. Each column in the disclosed spatial pooler may be assigned a column label. Each iSDR used for training the spatial pooler may be accompanied by a set of ground truth labels Li. Although a single label can be used, certain tasks may require the handling of multiple labels.
[0092] When a column is determined to be a winner column, its label may be compared to the ground truth label(s) of the training iSDR. If the column label matches one of the ground truth label(s) accompanying the training iSDR, then it is determined that the column is a “correct winner column.” The strengths of the winner column’s synapses are modified to better align (align more closely) to the training iSDR. Synaptic strengths are adjusted to increase the overlap between the column's connection vector and the input Sparse Distributed Representation (iSDR), such that the strengths of the winner column’s synapses better align (align more closely) to the training iSDR. This process involves incrementing the permanence values of synapses corresponding to active bits in the iSDR and decrementing those corresponding toinactive bits. Synapses with permanence values exceeding a connection threshold become connected, while those falling below can become disconnected.
[0093] The rules of synaptic strength updates in this case are similar or identical to the Hebbian-based rules employed by unsupervised spatial poolers. Alternatively, if the column label does not match any label(s) accompanying the training iSDR, then it is determined that the column is an "incorrect winner column." The strengths of the column's synapses may be modified to diverge from the iSDR in Operation 1 .2 shown by Figure 14. The synaptic strength updates may be reversed from those in the unsupervised approach, as illustrated by the supervised learning pathway in Figure 14. This may allow columns near decision boundaries to be assigned a label and modify the learning such that the column aligns more strongly with iSDRs having the same label and diverges from iSDRs having dissimilar labels. The minimum-overlap guarantee provided by the threshold-based winner selection in Operation 1 .1 ensures that these columns are positioned appropriately relative to the decision boundary.
[0094] For classification, the predicted label, Pi for an ISDR, ISDR, can be computed through various methods. The most frequently occurring winner column label may be determined as the predicted classification. As used herein, "predicted classification" refers to the output classification determined by the spatial pooler based on winner column labels and confidence scores. Alternatively, a weighted score using the overlap score of the winner columns may be determined. In some embodiments, a confidence score is determined based on the distribution of the winner column labels and / or their overlap scores. The ability to predict output class label(s) with or without confidence scores can prevent the need for a subsequent classification engine that may be required for unsupervised spatial poolers. As used herein, "confidence score" refers to a numerical measure indicating the certainty or reliability of the predicted classification.
[0095] Further, the disclosed spatial pooler retains the ability to perform unsupervised clustering as a special case. In this case, all the pooler columns are assigned an unknown label. The training iSDRs may also be associated with an unknown label. The resulting winner columns are correct winner columns because their labels match the unknown label accompanying the training iSDRs.
[0096] The disclosed spatial pooler may also be configured in a semi-supervised learning mode where training labels are available only for a subset of the training set. As used herein, the term "semi-supervised mode" may refer to a learning mode where the spatial pooler processes input data with labels available for only a subset of training samples, using unknown labels for unlabeled data. The unknown label is associated with the training samples without a ground truth label. Initialization of the columns and the learning performed by the spatial pooler in the semi-supervised mode is similar to or the same as in the supervised mode. A distinction between operation in the semisupervised mode and in the supervised mode lies in the handling of incorrect winner columns when the iSDR is associated with the unknown label. Since the iSDR's label is currently not determined, appropriate updates are made to the synaptic strengths of the winner columns, given the ambiguity of the label in the semi-supervised mode.E. Directed Initialization
[0097] In some embodiments, one or more candidate iSDRs may be identified for each column and the columns synaptic strengths and connections may be initialized using this set of candidate iSDRs. A label may also be assigned to each column based on the labels of the set of candidate iSDRs. This directed initialization approach may improve training efficiency, reduce sensitivity to input ordering, or enhance overall system performance for various classification tasks.F. Ensemble Learning
[0098] In some embodiments, the multiple columns are not initialized at the beginning of training. The spatial pooler may be built incrementally as a collection of ensembles of columns, where an ensemble includes a subset of neural computational units and may perform as a predictor. For example, the multiple columns in the spatial pooler are arranged among multiple ensembles, wherein the spatial pooler is incrementally constructed across multiple iterations by training at least some of the ensembles in each iteration. At the beginning of training, a spatial pooler with a single ensemble is initialized using a selection of training iSDRs. After training this spatial pooler, its performance may be evaluated for samples in the training dataset. For each training sample / , an individual training loss, may be calculated based on the distances from all its winning columns. For example, a measurement can be taken that indicatesthe level of representation for an ISDR within the pooler. This measurement can be incorporated into the logic for the next ensemble construction.
[0099] If no columns are selected as winner columns for an observation, then the individual training loss for that observation may be assigned a high value. The observations may be ranked based on descending order of their individual training loss (loss!), and the observations with the highest training loss may become candidates for subsequent ensemble creation.
[0100] At the beginning of each iteration, a new ensemble may be added to the spatial pooler. Multiple options may exist for which ensembles are trained during the iteration, wherein the training may be applied to the latest-added ensemble, a window of the latest-added ensembles, or all the ensembles. After the training is complete for the iteration, the entire combined spatial pooler may be used to determine the individual loss of the observations. The addition of the ensembles may be terminated based on criteria, wherein the criteria may include reaching a desired spatial pooler size or determining that the addition of more ensembles does not result in a reduction in the overall training loss. The overall training loss may include the sum of all the individual losses over all the training observations. This method of spatial pooler construction may lead to focused learning that may prioritize input regions where the spatial pooler has no representation, followed by observations with low or incorrect representation, or a combination thereof in the current spatial pooler.Selected Embodiments of Temporal Memories
[0101] A temporal memory is disclosed herein that has similar features and operations to the disclosed spatial pooler. For example, the input and output in both cases may follow the SDR format, the calculation of the proximal overlap score of the column may be the same as the distal predictive score of a neural computational unit, the selection of the winner columns and the predicted neural computational units may be similar, the Hebbian growth and decay of the proximal and distal synaptic strengths may be similar, and the connection criteria for proximal and distal synapses may be similar. A computational system that is representative of a temporal memory may be implemented using a first spatial pooler and a second spatial pooler that is synchronized with the first spatial pooler, along with auxiliary operations.
[0102] Figure 15 includes a diagrammatic illustration of two spatial poolers used to emulate the functions of a temporal memory. The first spatial pooler has cortical columns containing neural computational units that process input SDRs and generate output SDRs. The second spatial pooler has columns containing one or more neural computational units each that handle cellular activities and distal synaptic connections, wherein the spatial poolers are synchronized through auxiliary operations for temporal pattern learning. The operations of the first spatial pooler are shown in white boxes in Figure 15, while the operations of the second spatial pooler are shown using grey boxes. The auxiliary operations are shown in boxes with a grey hatched background. Operations 1.1 and 1.2 of the second spatial pooler are renamed as Operations 2.1 and 2.2 for ease of reference and clarity of description.
[0103] The first spatial pooler performs operations related to the cortical columns and proximal synapses. The second spatial pooler performs operations related to cellular activities and distal synapses. The first spatial pooler and the second spatial pooler are synchronized through the output of Operation 1 .1 of the first spatial pooler providing one of the inputs to Operation 2.2 of the second spatial pooler. The first auxiliary operation, Operation 2.3.1 facilitates this synchronization. The other two auxiliary operations, Operation 2.3.2 and Operation 2.3.3 are configured to generate the predicted spatial patterns for subsequent time steps.
[0104] The setup and operations of the first spatial pooler are similar to or the same as the operations of the spatial pooler disclosed above. In some embodiments, the first spatial pooler is configured to generate an oSDR by performing spatial pattern recognition on a received iSDR. For example, the first spatial pooler receives an iSDR with an accompanying label and processes it using Operation 1.1 for winner column selection and Operation 1.2 for proximal synaptic updates, thereby generating an oSDR that encodes spatial patterns for subsequent temporal processing by the synchronized second spatial pooler.
[0105] The second spatial pooler is configured to establish a temporal pattern by performing temporal sequence learning on the oSDR. For example, the second spatial pooler determines active neural computational units at a first time step based on predicted neural computational units and the winner columns. Using distal synaptic connections, the second spatial pooler calculates predictive overlap scores for eachneural computational units based on incoming signals from active neural computational units of the previous time step, identifying predicted neural computational units with scores greater than the predictive threshold. The temporal pattern is established by updating distal synaptic strengths between active neural computational units of the first time step and predicted neural computational units of the second time step, strengthening connections for correct predictions and weakening connections for incorrect predictions, thereby learning sequential relationships that enable future state prediction based on the current temporal context.
[0106] For a computational system representative of a temporal memory with n columns, the first spatial pooler has n columns. The spatial patterns for each time step / are presented as iSDRi to the first spatial pooler with the accompanying label During training, the output of the first spatial pooler at each time step is represented by oSDRi. If this is the last time step of the batch, then proximal synaptic connection updates may be passed to Operation 1 .1 of the first spatial pooler for time step / +1 .
[0107] If the emulated temporal memory has n columns and c neural computational units per column, then the second spatial pooler emulates the nxc neural computational units in the temporal memory using nxc columns. The iSDR at time-step / includes the active neural computational units and corresponds to the oSDR of the previous time-step / -1. The incoming distal synapses of the neural computational units in the temporal memory are emulated by the proximal synapses of the corresponding columns of the second spatial pooler. The proximal synaptic strengths and connections are initiated through assignments. Given a neural computational unit identifier (j,k), the column label of the corresponding column in the second spatial pooler is set to j, thereby establishing correspondence between neural computational units and their respective columns.
[0108] Given the active neural computational units in time-step / -1 , the predicted neural computational units may be calculated by Operation 2.1 of the second spatial pooler. These predicted neural computational unit identifiers are converted to predicted spatial class labels using Auxiliary Operations 2.3.2 and 2.3.3. In Operation 2.3.2, the predicted winner columns for a predicted neural computational unit with identifier (j,k) is j. In Operation 2.3.3, the column labels of the corresponding columns of the first spatial pooler are determined to predict spatial class labels. The outputpredicted spatial class labels are labeled PPi in Figure 15. After processing a subsequence that is common to multiple learned sequences, the predicted oSDR may include the composite of all the oSDRs that appear in the next time step in those sequences.
[0109] Auxiliary operation 2.3.1 provides synchronization between the first spatial pooler and the second spatial pooler. Auxiliary operation 2.3.1 receives the identifiers of predicted neural computational units for time-step / from operation 2.1. Auxiliary operation 2.3.1 receives the column identifiers of the winner columns as a multi-label from Operation 1 .1 . As used herein, a "multi-label" refers to a set of multiple labels assigned to a single input sample for classification tasks. From the above inputs, the list of active neural computational units and consolidated neural computational units for time-step / are determined. The active neural computational units may be determined from one of the following: a predicted neural computational unit belonging to a winner column, or a representative neural computational unit from a bursting column, wherein the bursting column includes a winner column determined in Operation 1.1 that had no predicted neural computational units from Operation 2.1 . The active neural computational unit identifiers are represented as an oSDR for the second spatial pooler operations for time step / +1 , wherein this serves as the iSDR for Operation 2.1 in the next time step / +1 .
[0110] The consolidated neural computational units from Operation 2.3.1 and the winner column identifiers from Operation 1.1 serve as the second part of the synchronization of the first spatial pooler and the second spatial pooler for time step / . The consolidated neural computational units that were also active neural computational units will have their column labels in the oSDR generated by Operation 1.1. Therefore, when these labels are encountered as accompanying ground truth labels in Operation 2.2, the active neural computational units are determined as correctly predicted and their synaptic strengths are updated accordingly. The consolidated neural computational units that were not active, wherein they were incorrectly predicted to be active, will not have their column labels passed to Operation 2.2 of the second spatial pooler. Therefore, their synaptic strengths will be adjusted accordingly. If this is the last time step of the batch for the second spatial pooler, then the updated distal synaptic connections are passed to Operation 2.1 of the secondspatial pooler for time step / +1 . The batch size of the first spatial pooler and the second spatial pooler may be different.A. Ensemble Learning
[0111] The second spatial pooler can be constructed incrementally using ensembles of neural computational units. For example, a computational system can evaluate the performance of establishing temporal patterns by the second spatial pooler by calculating an individual training loss for each observation based on distances from winning columns, ranking observations by descending training loss, and identifying candidates with the highest loss for ensemble creation. Ensembles of neural computational units can be added to the set of columns of the second spatial pooler shown by Figure 15 when the desired pooler size is reached or when ensemble addition fails to consistently reduce overall training loss, enabling incremental construction and focused learning on input regions with inadequate representation.
[0112] At first, the number, c, of neural computational units per column may be kept relatively small. The performance of the second spatial pooler in predicting the next token may be assessed using (i) column labels that are often mis-predicted or not predicted or (ii) neural computational units that are most overloaded with the combination of incoming and outgoing synaptic connections. In either of these cases, additional neural computational units may be added to the columns of the second spatial pooler or existing neural computational units may be split into two or more neural computational units. Addition of neural computational units in the second spatial pooler alters the iSDR of the second spatial pooler. Therefore, the bit-offsets of the existing neural computational units across the iSDRs before and after the addition of the ensemble are maintained. New bit-offsets corresponding to the new neural computational units in the ensemble are added to the iSDR.Selected Embodiments of Language Models
[0113] Figure 16 includes a diagrammatic illustration of autoregressive encoder operations of a language model implemented by a computational system. The encoder-decoder architecture disclosed herein can implement a learning and prediction algorithm used for predictive text generation. The disclosed language model is derived from the temporal memory described above with three main modifications.First, the language model uses an autoregressive encoder-decoder architecture that encodes a sequence of tokens as an internal state of the encoder-decoder architecture and then uses this state as the basis for generating at least one token in an autoregressive manner. Second, the language model may generate an End of Sequence (“<EOS>”) token to end the output text produced based on a context of the input text and the output text. The language model is trained using training text that includes <EOS> tokens to predict occurrence of the EOS tokens. Third, the language model can generate multiple representations of a sequence of tokens based on multiple contexts of the input text, wherein the output text is predicted based on the multiple representations of the sequence of tokens.
[0114] During training, the encoder-decoder architecture operates in an autoregressive encoder mode to learn from the training sequences. The training operations are similar to or the same as those of the temporal memory disclosed above and are shown by Figure 16. In some embodiments, the <EOS> tokens are inserted at the beginning and end of the training sequences. During training, the language model learns to begin processing when it encounters an <EOS> token at the beginning of a sequence and learns to predict an <EOS> token as the last token in the sequence. During inference, the processing of the input text begins when no neural computational units are active and the <EOS> token is the first token in the input sequence. After acquiring the tokens in the input sequence, the encoder-decoder architecture switches to a decoder mode using the encoded internal state as the initial state. The generation of the output in the decoder continues until the <EOS> token is predicted.
[0115] As shown by Figure 16, at time-step / , an iSDR corresponding to the / thtoken in the sequence of tokens is passed to Operation 1 .1 , and the corresponding label Li is passed to Operation 1.2. The iSDR and its processing in Operation 1.1 is the same as that of the temporal memory. However, the labels passed to Operation 1 .2 and its processing can be different from that of the temporal memory. The nature of the labels passed can be different for different implementations. In one implementation, the first spatial pooler is configured to work in the unsupervised mode and Li is always equal to unknown for the tokens in the input sequence. In another implementation, the first spatial pooler is configured to work in the supervised mode with Li as a single label for all the tokens in the sequence. In yet another implementation, the first spatial pooler is configured to work in the supervised modeand Li is a multi-label for the tokens in the sequence. The multi-label may include a set of e labels, one for each dimension in the embedding of token / . To accommodate these different implementations (illustrated and described in more detail with reference to Figures 20-23), Li is always treated as a set of labels. As used herein, "e labels" refers to a set of labels corresponding to each dimension of an embedding vector, and "e-dimension" refers to the predetermined number of dimensions in the embedding vector representation. As used herein, an "embedding" refers to a numerical feature vector representation of predetermined dimensions that encodes semantic information about tokens.
[0116] Figure 17 includes a diagrammatic illustration of autoregressive encoder operations for input sequence encoding during an inference phase. During inference, the encoder-decoder architecture encodes the sequence of tokens using the operations shown in Figure 17. These operations are the same as the encoding functions during training except the synaptic update operations 1 .2 and 2.2 are foregone. After processing the tokens in the sequence of tokens, the active neural computational units of the second spatial pooler denote the internal state of the encoder-decoder architecture. This internal state serves as the initial state for the decoder mode used in generating at least one output token.
[0117] Figure 18 includes a diagrammatic illustration of autoregressive decoder operations for output sequence generation during the inference phase. When the decoder operations begin, the sequence of tokens has been passed through the first spatial pooler. For each time step / , the predicted neural computational unit identifiers are generated in Operation 2.1 based on the context of the active neural computational units from the previous time step. The corresponding predicted labels correspond to the identifiers of the predicted winner columns. However, the generation of the active neural computational units for time step / at Operation 2.3.1 waits until at least one token for this time step is predicted. The predicted winner columns from Operation 2.1 are used to generate the output oSDR or output class prediction for time step i in Operations 2.3.2 and 2.3.3. The nature of the output PPi for the encoder-decoder architecture is based on the implementation selected. The / thoutput token is produced based on the confidence scores generated in Operation 2.3.3. Once this output token is produced, the columns of the first spatial pooler corresponding to this output token are passed to Operation 2.3.1 as the winner columns. The generation of the activeneural computational units based on these winner columns and the predicted neural computational units from Operation 2.1 is identical to those of the temporal memory. The generation of output tokens may be stopped once the <EOS> token is predicted for a time step.Selected Embodiments of Pipeline Architectures
[0118] Figure 19 includes a diagrammatic illustration of input processing stages of a computational system that can implement a pipeline architecture 1900. The pipeline architecture 1900 represents an end-to-end pipeline for sequential pattern recognition applications such as predictive text generation, time series analysis, video processing, network security, or industrial monitoring. For example, given the input sequence Mary had a, the output sequence little lamb is generated. The pipeline architecture 1900 disclosed herein uses the language model described above as a learning algorithm and a processing stage. The pipeline architecture 1900 includes multiple stages which may be configured based on the labels used to train the language model 1904. In some embodiments, the pipeline architecture 1900 is referred to as a language pipeline.
[0119] The pipeline architecture includes a tokenizer 1908 and an embedder 1912 as shown by Figure 19. The tokenizer 1908 can convert input text sequences into discrete tokens arranged in temporal order, where tokens may be word-aligned or subword units depending on the tokenizer implementation. The tokens can correspond to entries in a precomputed dictionary of the language. The input text sequence includes a series of tokens from the precomputed dictionary. For example, the input sequence Mary had a little lamb, may be broken down into the tokens <EOS>, Mary, had, a, little, lamb, ., <EOS>. Although the tokens in the example above are word- aligned, this might not always be the case and is dependent on how the tokenizer is implemented. The position of a token in the dictionary serves as its token identifier. The dictionary also contains the special token <EOS> that denotes the beginning and end of a sequence.
[0120] The embedder 1912 is configured to compute numerical feature vectors of predetermined dimensions for the sequence of input tokens, generating context- dependent embeddings where the same token has different numerical representations based on linguistic context. The embeddings take into consideration the context ofother tokens found in the sequence of input tokens. For example, the token novel found in the context of innovation would have a different embedding from the one used in the context of a book. The embedder 1912 may be pretrained to generate the embeddings for the sequence of input tokens.
[0121] Figure 20 includes a flow diagram illustrating a process for selecting implementations of the pipeline architecture. A token may have different embeddings based on the context in which it appears. In the pipeline architecture shown by Figure 19, the different embeddings can be processed using a label generator 1916, the language model 1904, and an output handling stage. Figure 20 shows three potential implementations for processing the different embeddings beginning at the label generation and language model stages. The pipeline implementations referenced by Figure 20 are based on the types of labels associated with the tokens for training the language model 1904. These implementations are based on the semi-supervised learning capability of the first spatial pooler used in the language model.
[0122] The process shown by Figure 20 can be performed by the pipeline architecture illustrated and described in more detail with reference to Figure 19. At 2004, the pipeline architecture evaluates whether labels are used with tokens in the sequence of input tokens. If labels are not used with tokens, Implementation I is selected at 2012, where the language model operates in the unsupervised mode with unknown labels for all tokens. If labels are used with tokens, the pipeline architecture determines whether the number of columns in the language model is greater than or equal to the dictionary size at 2008. At 2008, if the number of columns is greater than or equal to the dictionary size, Implementation II is selected at 2016, where token identifiers are used as labels in the supervised mode. If the number of columns is not greater than or equal to the dictionary size, Implementation III is selected at 2020, where embedding bucket identifiers are used as multi-labels in the supervised mode.A. Implementation I
[0123] The first implementation of the pipeline architecture uses the first spatial pooler in the unsupervised mode. Hence, the input labels to the language model are set to unknown for all input tokens in the sequence of input tokens.
[0124] Figure 21 includes a diagrammatic illustration of an output token generator 2104 being trained in accordance with Implementation I. The language model outputs the oSDR for each time step. The tokens in the language dictionary have no corresponding output label. Hence, a classifier is used to implement the output token generator 2104 as a separate stage to translate the predicted oSDRs to token identifiers during output sequence generation. Once the first spatial pooler of the language model is sufficiently trained, and the oSDRs corresponding to a token have stabilized, the output token generator 2104 is trained. The token identifiers and their corresponding oSDRs serve as the labels and features for training, respectively. If multiple sequence predictions are desired, then a classifier that provides multiple likely predictions with confidence scores is used. For example, a single-hidden layer based perceptron model or a supervised spatial pooler with lower runtime complexity can be used as a classifier.
[0125] When the language model and the output token generator 2104 are trained, the pipeline architecture is ready to generate output tokens. Given an oSDR from the language model for a time step, the output token generator predicts a token identifier, which is used to look up the token in the dictionary. Using the generated tokens, the classifier produces a natural language response by concatenating the predicted tokens in temporal order to form coherent output text sequences that follow the input text sequence.B. Implementation II
[0126] Figure 22 includes a diagrammatic illustration of output text sequence generation in accordance with Implementation II. In Implementation II, the input token identifiers are used as the labels. The language model is configured in the supervised mode with more columns than the number of tokens in the language dictionary. The language model is initialized such that at least one column is dedicated to each token identifier. Effectively, all embeddings of the same token become members of a class, with the token identifier as its class label. The language model addresses the different embeddings of the same token. In each time step, a respective token identifier associated with the likely tokens in the output sequence is predicted with corresponding confidence scores. The token identifiers can be used to look up eachtoken from the dictionary. The output sequence produced results from a concatenation of the tokens in order.C. Implementation III
[0127] Figure 23 includes a diagrammatic illustration of output text sequence generation in accordance with Implementation III. In Implementation III, the pipeline architecture learns to predict the embeddings of the output tokens. The embeddings of the tokens may be pretrained and static or may be derived during the training process. Pre-trained static embeddings can be obtained, e.g., from contemporary open-source language models. The embedding vector for a token can have values in the continuous range between -1 to 1 for each of its e-dimensions. In the pipeline architecture, this range is discretized into b intervals. Interval j from dimension / is assigned the unique label / xb + j. Therefore, there are a total of bxe possible labels. Each input token in the sequence of input tokens is assigned a multi-label which is a set of e labels, one each for each dimension. For example, if b equals 4 and the intervals for all dimensions are uniform, then each interval has the width of 0.5. That is, values in the range [-1 , -0.5) lie in interval 0, values in the range [-0.5, -0) lie in interval 1 , and so on. If e equals 3, then an embedding vector of {-0.42, -0.62, 0.67} is discretized to the interval-vector {1 ,0,3}, and assigned the multi-label of {0x4 + 1 , 1 x4 + 0, 2x4 + 3}, i.e., {1 , 4, 1 1 }.
[0128] The disclosed language model used by the pipeline architecture has greater than bxe columns with at least one column assigned to each label. As shown by Figure 23, in each time-step, the language model computes an output multi-label along with a confidence vector for each label in the multi-label. Here, each label corresponds to a specific interval in an embedding dimension. To generate the output embeddings, an output embedding generator 2304 first organizes all the labels from each dimension into a separate group. Then for each embedding dimension, the output value is generated based on the labels in that group, and their corresponding confidence score and interval range.
[0129] The output embeddings are presented to the output token generator 2308. An open-source pre-trained model can be used for output token generation. Such a model can include a perceptron network with a single hidden layer and can predict a token whose native embedding is the closest match to the generated outputembedding. Because the output token generator 2308 is pretrained, the output sequence generation from the output of the language model does not require any further training in the pipeline architecture. The use of the language model in the supervised mode also improves the prediction accuracy compared to existing implementations. Additionally, the number of columns in the language model can be significantly smaller than the dictionary size, and therefore Implementation III lends itself to more efficient and scalable computational systems.Remarks
[0130] The foregoing description of various embodiments of the claimed subject matter has been provided for the purposes of illustration and description. It is not intended to be exhaustive or to limit the claimed subject matter to the precise forms disclosed. Many modifications and variations will be apparent to one skilled in the art. Embodiments were chosen and described in order to best describe the principles of the invention and its practical applications, thereby enabling those skilled in the relevant art to understand the claimed subject matter, the various embodiments, and the various modifications that are suited to the particular uses contemplated.
[0131] Although the Detailed Description describes certain embodiments and the best mode contemplated, the technology can be practiced in many ways no matter how detailed the Detailed Description appears. Embodiments may vary considerably in their implementation details, while still being encompassed by the specification. Particular terminology used when describing certain features or aspects of various embodiments should not be taken to imply that the terminology is being redefined herein to be restricted to any specific characteristics, features, or aspects of the technology with which that terminology is associated. In general, the terms used in the following claims should not be construed to limit the technology to the specific embodiments disclosed in the specification, unless those terms are explicitly defined herein. Accordingly, the actual scope of the technology encompasses not only the disclosed embodiments, but also all equivalent ways of practicing or implementing the embodiments.
[0132] The language used in the specification has been principally selected for readability and instructional purposes. It may not have been selected to delineate or circumscribe the subject matter. It is therefore intended that the scope of thetechnology be limited not by this Detailed Description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of various embodiments is intended to be illustrative, but not limiting, of the scope of the technology as set forth in the following claims.
Claims
CLAIMSI / We Claim:1 . A method performed by a computational system that comprises a spatial pooler with multiple columns, each of which includes a neural computational unit that has multiple synaptic connections to an input Sparse Distributed Representation (iSDR), the method comprising: determining overlap between the multiple synaptic connections and the iSDR by establishing, for each of the multiple columns, a distance between the iSDR and the multiple synaptic connections; selecting one or more columns from among the multiple columns based on the distances established for the multiple columns; producing an output Sparse Distributed Representation (oSDR) comprising a bit-vector based on the one or more columns, wherein each bit in the bit-vector corresponds to a corresponding column of the multiple columns and indicates whether the corresponding column is one of the one or more columns; and updating strengths of the multiple synaptic connections based on the iSDR, so as to increase permanence of synaptic connections corresponding to the one or more columns.
2. The method of claim 1 , wherein each of the one or more columns is associated with a corresponding one of one or more column labels and the iSDR is associated with a ground truth label, and wherein said updating comprises: updating strengths of synaptic connections corresponding to the one or more columns to better align with the iSDR in response to a determination that the one or more column labels match the ground truth label; and updating strengths of synaptic connections corresponding to the one or more columns to diverge from the iSDR in response to a determination that the one or more column labels do not match the ground truth label.
3. The method of claim 1 , wherein said updating comprises:accumulating a batch of updates to the strengths of the multiple synaptic connections; and applying the accumulated batch of updates to the strengths of the multiple synaptic connections at a boundary of the batch.
4. The method of claim 1 , wherein the multiple columns are arranged among a plurality of ensembles, and wherein the spatial pooler is constructed across multiple iterations by training at least some of the plurality of ensembles in each iteration.
5. The method of claim 1 , wherein said selecting comprises: comparing the distances established for the multiple columns to thresholds associated with the multiple columns; and identifying the one or more columns in response to a determination that the distances established for the one or more columns are less than the thresholds associated with the one or more columns.
6. The method of claim 1 , wherein the iSDR is associated with a ground truth label, and wherein the spatial pooler operates in one of a supervised mode, a semisupervised mode, or an unsupervised mode based on the ground truth label.
7. The method of claim 1 , further comprising: identifying a candidate iSDR for each of the multiple columns, so as to identify multiple candidate iSDRs; and initializing the strengths and the multiple synaptic connections based on the multiple candidate iSDRs.
8. A computational system that, in operations, is representative of a temporal memory, the computational system comprising: a first spatial pooler with a first set of columns, each of which includes a neural computational unit, the first spatial pooler being configured to:generate an output Sparse Distributed Representation (oSDR) based on an input Sparse Distributed Representation (iSDR) and an accompanying label; and a second spatial pooler with a second set of columns, each of which includes one or more neural computational units, the second spatial pooler being configured to: receive the oSDR from the first spatial pooler; determine that a first one or more neural computational units of the second spatial pooler are active at a first time step based on an analysis of the oSDR; predict, based on the first one or more neural computational units, a second one or more neural computational units of the second spatial pooler that will be active at a second time step subsequent to the first time step; and establish a temporal pattern based on the second one or more neural computational units predicted to be active at the second time step.
9. The computational system of claim 8, wherein the first spatial pooler is further configured to select one or more columns from among the first set of columns based at least on the iSDR, and wherein the first one or more neural computational units of the second spatial pooler are determined to be active at the first time step based further on the one or more columns selected from among the first set of columns.
10. The computational system of claim 9, wherein the second spatial pooler is further configured to iteratively predict which neural computational units of the second spatial pooler will be active over a series of time steps.1 1 . The computational system of claim 8, wherein the second spatial pooler is further configured to: predict that no corresponding neural computational units of a column of the second set of columns will be active at the second time step; regenerate a predictive score for the column based on the first one or more neural computational units; and select a representative active neural computational unit for the column at the second time step based on the predictive score.
12. The computational system of claim 8, wherein the second spatial pooler is further configured to: update distal synaptic connections between the first one or more neural computational units and the second one or more neural computational units.
13. The computational system of claim 8, wherein the computational system is configured to: evaluate a performance of establishing the temporal pattern by the second spatial pooler; and add an ensemble of neural computational units to the second set of columns based on evaluating the performance.
14. The computational system of claim 8, wherein the first spatial pooler is configured to generate the oSDR by performing spatial pattern recognition on the iSDR, and wherein the second spatial pooler is configured to establish the temporal pattern by performing temporal sequence learning on the oSDR.
15. A non-transitory medium with instructions stored thereon that, when executed by a processing unit of a computational system, cause the computational system to perform operations comprising: acquiring a sequence of tokens that are collectively representative of input text that is provided to, or obtained by, the computational system as input; encoding the sequence of tokens as an internal state of an encoder-decoder architecture that includes (i) a first spatial pooler and (ii) a second spatial pooler that is synchronized with the first spatial pooler and that takes outputs produced by the first spatial pooler as input; andgenerating, in an autoregressive manner, at least one token with the encoderdecoder architecture that uses the internal state as a basis for predicting output text that follows the input text.
16. The non-transitory medium of claim 15, wherein said encoding is performed using the encoder-decoder architecture in an autoregressive encoder mode.
17. The non-transitory medium of claim 15, wherein said generating is performed using the encoder-decoder architecture in an autoregressive decoder mode.
18. The non-transitory medium of claim 15, wherein the operations further comprise: generating multiple representations of the sequence of tokens based on multiple contexts of the input text, and wherein the output text is predicted based on the multiple representations of the sequence of tokens.
19. The non-transitory medium of claim 15, wherein the operations further comprise: predicting the output text based on the generated at least one token; and generating an EOS token to end the output text based on a context of the input text and the output text, wherein the computational system is trained, using training text that includes EOS tokens, to predict occurrence of the EOS tokens in the training text.
20. The non-transitory medium of claim 15, wherein the computational system comprises multiple neural computational units connected by multiple distal synaptic connections, and wherein the operations further comprise: determining a temporal pattern in the input text; andupdating strengths of the multiple distal synaptic connections based on the temporal pattern.
21. A computational system for language processing, the computational system comprising: a tokenizer configured to convert an input text sequence into a sequence of input tokens that are arranged in temporal order; an embedder configured to compute embeddings for the sequence of input tokens; a Sparse Distributed Representation (SDR) generator configured to generate an input SDR (iSDR) for the sequence of input tokens based on an analysis of the embeddings; and a language model that comprises (i) a first spatial pooler and (ii) a second spatial pooler that is synchronized with the first spatial pooler, wherein the language model is configured to: determine a temporal pattern represented in the iSDR; and produce, based on the temporal pattern, an output text sequence that is predicted to follow the input text sequence.
22. The computational system of claim 21 , wherein the embeddings are arranged among discretized intervals, and wherein the computational system further comprises a label generator configured to: generate labels for the sequence of input tokens based on the discretized intervals, wherein the labels are used by the language model to determine the temporal pattern.
23. The computational system of claim 21 , wherein the language model is configured to: generate an output SDR (oSDR) based on the temporal pattern represented in the iSDR, and wherein the computational system further comprises a classifier configured to:generate, based on the oSDR, at least one token for producing the output text sequence; and produce, using the at least one token, a natural language response to the input text sequence.
24. The computational system of claim 21 , wherein the sequence of input tokens is associated with unknown labels, and wherein the language model operates in an unsupervised mode to determine the temporal pattern represented in the iSDR using the sequence of input tokens that is associated with the unknown labels.
25. The computational system of claim 21 , wherein the language model includes multiple columns, each of which includes one or more neural computational units, wherein a number of the multiple columns is greater than or equal to a size of a token dictionary that includes the sequence of input tokens, and wherein the language model is configured to operate in a supervised mode to determine the temporal pattern represented in the iSDR using the one or more neural computational units.
26. The computational system of claim 21 , wherein the embeddings have values that fall within value ranges, and wherein the computational system further comprises: a label generator configured to: generate multi-labels for the sequence of input tokens based on the value ranges, wherein the temporal pattern represented in the iSDR is determined based on the multilabels.
27. The computational system of claim 21 , wherein the embeddings are static embeddings obtained from another language model that is different from the language model.
28. The computational system of claim 21 , wherein the embeddings are generated during training of the language model.
Citation Information
Patent Citations
Storage compute device with tiered memory processing
US20150324125A1
Artificial Intelligence (AI) System for Learning Spatial Patterns in Sparse Distributed Representations (SDRs) and Associated Methods
US20230072539A1
Conceptual computation system using a hierarchical network of modules
US9524461B1