Answering cognitive queries from sensor input signals

By converting sensor input signals into hypervectors and storing them in associative memory, the sensor data encoding problem is solved, enabling efficient processing of unstructured data and finding the optimal solution.

CN113474795BActive Publication Date: 2025-11-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080016182.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-25
Filing Date
2020-02-14
Publication Date
2025-11-18
Estimated Expiration
2040-02-14

AI Technical Summary

Technical Problem

Existing technologies struggle to directly encode hyperdimensional (HD) vectors from sensor data, leading to challenges in finding analogies across different knowledge domains, and traditional analytical methods struggle to handle semi-structured or unstructured data.

Method used

By feeding sensor input signals into the input layer of an artificial neural network, a pseudo-random bit sequence group is generated and converted into a hypervector stored in an associative memory. The answer to the cognitive query is determined using measurement methods such as Hamming distance.

Benefits of technology

It enables direct encoding from sensor data to hyperdimensional vectors, improving learning speed and the reliability of the response process, and effectively processing unstructured data to find the best answer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113474795B_ABST
    Figure CN113474795B_ABST
Patent Text Reader

Abstract

A computer-implemented method for answering cognitive queries from sensor input signals can be provided. The method comprises feeding a sensor input signal to an input layer of an artificial neural network comprising a plurality of hidden neuron layers and an output neuron layer, determining a hidden layer output signal from each of the plurality of hidden neuron layers and an output signal from the output neuron layer, and applying a set of mapping functions by using the output signal of the output layer and the hidden layer output signal of one of the hidden neuron layers as input data for one of the mapping functions to generate a set of pseudo-random bit sequences. Further, the method comprises determining a hyper vector using the set of pseudo-random bit sequences and storing the hyper vector in an associative memory, wherein a distance between different hyper vectors is determinable.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] The present disclosure relates generally to machine learning, and more specifically, to a computer-implemented method for answering cognitive queries from sensor input signals. The present disclosure also relates to a related machine learning system for answering cognitive queries from sensor input signals, and a related computer program product.

[0002] Machine learning is one of the hottest topics in science and one of the topics for enterprise information technology (IT) organizations. In the past few years, the amount of data collected by enterprises has increased and more and more complex analysis tools are needed. Traditional business intelligence / business analytics tools have proven to be very useful for information technology (IT) and business users. However, it becomes more and more difficult to analyze semi-structured or unstructured data with traditional analysis methods.

[0003] Storage and computing power have increased significantly in the past few years, enabling relatively easy implementation of artificial intelligence (AI) systems, whether standalone systems or integrated into any type of application. These (AI) systems do not need to be programmed in a procedural way, but can be trained with example data in order to develop a model that, for example, recognizes and / or classifies unknown data.

[0004] Therefore, today, we are still in the era of narrow AI, which means that AI systems can learn and recognize patterns after being trained with training data, however, only within very tight boundaries. The industry is still far from general AI that is able to transfer abstract knowledge from one domain to another. Therefore, finding analogies between different knowledge domains is one of the challenges imposed on traditional von-Neumann computing architectures.

[0005] The ability of artificial neural systems to simulate and answer cognitive queries depends to a large extent on the representational power of such systems. Simulation requires a complex, relevant representation of cognitive structures that is not easily available through existing artificial neural networks. Vector Symbolic Architecture is a class of distributed representation schemes that have been shown to be able to represent and manipulate cognitive structures.

[0006] However, the missing step is the encoding of cognitive structures directly from sensor data, which is widely referred to in the existing literature as the "encoding problem".

[0007] The present disclosure can in one aspect increase the power of artificial neural networks, solve the information bottleneck and pass, for example, well-defined classical intelligence tests, such as the Ravens Progressive Matrices test. SUMMARY

[0008] According to an aspect of the present application, a computer-implemented method for answering cognitive queries from sensor input signals can be provided. The method can comprise feeding a sensor input signal to an input layer of an artificial neural network comprising a plurality of hidden neuron layers and one output neuron layer; determining a hidden layer output signal from each of the plurality of hidden neuron layers and an output signal from the output neuron layer, and applying a set of mapping functions by using the output signal of the output layer and the hidden layer output signal of one of the hidden neuron layers as input data for one mapping function, in particular at least one mapping function, thereby generating a set of pseudo-random bit sequences.

[0009] Additionally, the method can comprise determining hyper-vectors using the set of pseudo-random bit sequences and storing the set of hyper-vectors in an associative memory, wherein a distance between different hyper-vectors is determinable.

[0010] According to another aspect of the present application, a related machine learning system for answering cognitive queries from sensor input signals can be provided. The machine learning system can comprise a sensor adapted to feed a sensor input signal to an input layer of an artificial neural network comprising a plurality of hidden neuron layers and one output neuron layer; a first determining unit adapted to determine a hidden layer output signal from each of the plurality of hidden neuron layers and an output signal from the output neuron layer; and a generator unit adapted to apply a set of mapping functions by using the output signal of the output layer and the hidden layer output signal of one of the hidden neuron layers as input data for one mapping function, in particular at least one mapping function, thereby generating a set of pseudo-random bit sequences.

[0011] The machine learning system can further comprise a second determining unit adapted to determine hyper-vectors using the set of pseudo-random bit sequences; and a storage module adapted to store the hyper-vectors in an associative memory, wherein a distance between different hyper-vectors is determinable.

[0012] The proposed computer-implemented method for answering cognitive queries from sensor input signals can provide a number of advantages and technical effects:

[0013] This system offers an excellent way to solve the previously disclosed problem of directly encoding hyper-dimensional (HD) binary vectors from sensor data to form the components of a structure. This problem is widely referred to as the "encoding problem." The idea is to use information from sensor data from different layers of an artificial neural network to form HD vectors as input to an associative memory. Thus, it can be represented that sensor data from components of the same class can be transformed into information leading to similar binary hypervectors. Similarity here can be represented by relatively short distances between them in the associative memory. These distances can be measured by a metric between the query and potential answers. The closer the HD vectors are to each other, the better the potential answer chosen from the group including candidate answers, and the shorter the distance can be. The distance can be elegantly measured using the Hamming distance function. The shortest distance between the cognitive query and the candidate answers can be used as the best answer for a given query.

[0014] The proposed solution utilizes the implicit intrinsic states of a trained artificial neural network to derive HD vectors for addressing associative memory. This is achieved using two distinct timeframes. The first timeframe can be defined by the end of the neural network's training. This can be determined by gradient descent applied to the layers of the neural network to pinpoint the end of the training session. The generated HD role vectors are generated only once, at the end of the training cycle. They represent the ground state of the neural network associated with the combined representation elements of the cognitive query.

[0015] After training, HD padding vectors can be continuously generated as long as sensor input signals representing cognitive queries and candidate answers are fed into the system. Each candidate answer generates at least one HD padding vector to be fed into the associative memory. The response to the query can be derived based on the distance between the combined HD role vector representing the candidate answer and the padding vector.

[0016] Therefore, by addressing the problem of directly encoding hyperdimensional (HD) vectors from sensor data for composite structures, determining the answer to a cognitive query becomes elegantly possible.

[0017] The proposed non-von Neumann vector symbol architecture also allows for a global transformation, where information in the HD vector is moderately degraded relative to the number of failure bits, regardless of their position. Furthermore, it employs efficient HD computation, characterized by mapping constituent structures onto other constituent structures through simple binary operations, without having to decompose the sensor signal representation into their components.

[0018] Further embodiments of the inventive concept will now be described.

[0019] According to one embodiment, the method may include combining—i.e., using binding operators and / or bundle operators—different hypervectors stored in an associative memory to derive cognitive queries and candidate answers, each candidate answer obtained from sensor input signals. Therefore, query processing and candidate answer processing can be implemented in a similar manner, which implicitly makes comparison and the final decision process easier.

[0020] According to another embodiment, the method may include measuring the distance between a hypervector representing the cognitive query and a hypervector representing the candidate answer. One option for measuring the distance between binary hypervectors could be using the Hamming distance between two binary hypervectors. Other vector distance methods can also be successfully implemented. The key point is that by utilizing distance functions, good comparability can be achieved using controllable computational power. For completeness reasons, it may be noted that the terms hypervector and hyperdimensional vector are used interchangeably in this document.

[0021] According to one embodiment, the method may further include selecting a hypervector associated with the candidate answer, the selected hypervector being the closest to the hypervector representing the cognitive query. This relatively straightforward approach to selecting an answer for a given cognitive query can represent a very elegant way to derive the best possible answer from the option space of potential answers for a given cognitive query.

[0022] According to another embodiment of the method, feeding the sensor input data to the input layer of the artificial neural network may further include feeding the sensor input data to the input layers of multiple artificial neural networks. Utilizing this additional feature, the learning speed and the reliability of the response process can be significantly improved.

[0023] According to one embodiment of the method, generating a set of pseudo-random bit sequences may further include using the output signal of the output layer and the hidden layer output signals of multiple hidden neuron layers as input data to at least one mapping function. Using the hidden neuron layers of the neural network, rather than just the output layer, allows for the relatively easy generation of more complex structures, particularly one of the hypervectors, since the mapping function can actually be multiple separate mapping functions, and information can also be received from the intermediate states of the neural network, i.e., from the hidden layers. This can represent a much better underlying dataset for forming binary HD vectors for addressing associative memories.

[0024] According to a licensed embodiment of the method, the neural network can be a convolutional neural network. Convolutional neural networks have proven very successful in recognizing (i.e., predicting) images or processing sound (e.g., in the form of spoken words). One or more convolutional stages, typically located at the beginning of a sequence of hidden neuron layers, make it possible to very quickly reduce the number of neurons from one neural network layer to the next. Therefore, they help reduce the computational requirements of neural networks.

[0025] According to a further embodiment, the method may also include generating multiple random hypervectors (HD vectors) as character vectors at the end of training of the neural network, to be stored in associative memory. The term "at the end" can indicate that the training of the neural network is nearing completion. This completion can be determined using a gradient descent method that relies on the error determined by the difference between the results of the training data and the labels in a supervised training session. Using the SGD (Stochastic Gradient Descent) algorithm and monitoring gradient statistics can be an option to determine the end of the training session in order to generate character vectors. Those skilled in the art will know that several other methods besides the SDG algorithm exist to determine that the end of the learning session is approaching. Simply determining the difference in the variance of the stochastic gradients between two training epochs using two different training datasets and comparing it with a threshold can trigger the generation of character vectors.

[0026] According to an additional embodiment, the method may further include continuously generating multiple pseudo-random hypervectors as padding vectors for each new sensor input dataset based on the output data of multiple neural networks, and these pseudo-random hypervectors will be stored in an associated memory. In contrast to role vectors, these pseudo-random vectors can be continuously generated as long as the sensor input signals representing cognitive queries and candidate answers are fed into the relevant machine learning system.

[0027] According to one embodiment of the method, the combination may include binding different hypervectors via a vector-element-wise binary XOR operation. This operation can represent a low-cost and efficient digital function in terms of computational power to establish bindings between different hypervectors.

[0028] According to another related embodiment of the method, the combination may further include binding different hypervectors by binary averaging of vector elements, with a binding operator defined as an element-wise binary XOR operation. Conversely, the binding operator <·> can be defined as an element-wise binary average of the hypervectors, i.e., the majority rule. Therefore, simple binary computation can be used to define the binding and bounding of different hypervectors, thus requiring only limited computational power.

[0029] According to another embodiment of the method, if the stochastic gradient descent value determined between two sets of input vectors of the sensor input signal to the neural network remains below a predefined threshold, then the role vector can be stored in the associative memory at the end of the training of the artificial neural network. In other words, the difference between adjacent SGD values ​​can be so small that further training efforts may become increasingly less important. This method step can also be used to reduce the training dataset to a reasonable size in order to keep the training time as short as possible.

[0030] Furthermore, embodiments may take the form of an associated computer program product accessible from a computer-usable or computer-readable medium, which provides program code used by or in conjunction with a computer or any instruction execution system. For the purposes of this specification, the computer-usable or computer-readable medium may be any apparatus that may contain means for storing, transmitting, disseminating, or transporting a program used by or in conjunction with an instruction execution system, apparatus, or device. Attached Figure Description

[0031] It should be noted that embodiments of the present invention are described with reference to different subject matter. In particular, some embodiments are described with reference to method type claims, while others are described with reference to apparatus type claims. However, those skilled in the art will understand from the above and below description that, unless otherwise indicated, any combination of features related to different subject matter, in particular any combination of features between features of method type claims and features of apparatus type claims, is also considered to be disclosed herein, except for any combination of features belonging to one type of subject matter.

[0032] The above and other aspects of the invention will be apparent from the examples of embodiments described below, and will be explained with reference to the examples of embodiments, but the invention is not limited thereto.

[0033] Preferred embodiments of the invention will be described by way of example only and with reference to the following figures:

[0034] Figure 1 A block diagram illustrating an embodiment of a computer-implemented method of the present invention for answering cognitive queries from sensor input signals is shown.

[0035] Figure 2 Support and extensions are shown. Figure 1 A block diagram of an embodiment of the method with additional steps.

[0036] Figure 3 A block diagram illustrating a portion of an embodiment of a neural network is shown.

[0037] Figure 4 A block diagram illustrating an embodiment of the generally proposed basic concept is shown.

[0038] Figure 5 Block diagrams illustrating embodiments of the concepts presented herein and in further detail are shown.

[0039] Figure 6 A block diagram of a complete system embodiment of the proposed concept is shown.

[0040] Figures 7A-7C Examples of results are shown, including the bundle and bundle formulas, the potential input data, and the measures of the average of correct decisions and the average of incorrect decisions.

[0041] Figure 8 A block diagram of an embodiment of a machine learning system for answering cognitive queries from sensor input signals is shown.

[0042] Figure 9 The diagram shows portions of the proposed method used to perform the method and / or includes those based on... Figure 8 A block diagram of an embodiment of the computational system for a machine learning system.

[0043] Figure 10 An example form of the Raven progressive matrix test is shown. Detailed Implementation

[0044] In the context of this specification, the following conventions, terms and / or expressions may be used:

[0045] The term "cognitive query" can refer to a query to a machine learning system that, instead of being formulated as an SQL (Structured Query Language) statement or another formal query language, can be derived directly from sensor input data. A machine learning system (i.e., a relevance-based approach) is used to test a series of candidate answers against the cognitive query in order to determine the best possible response within the learning context, i.e., to obtain the desired information from a base trained neural network (or from a further trained neural network).

[0046] The term "sensor input signal" can refer to data derived directly from a sensor used for visual data (or data that can be recognized by a human), such as image data from an image sensor (like a camera). Alternatively, sound data can also be sensor data from a microphone.

[0047] The term "artificial neural network" (ANN), or in short neural networks (the two terms are used interchangeably in this document), can refer to a circuit or computational system vaguely inspired by the biological neural networks that make up the human brain. A neural network itself is not an algorithm, but rather a framework in which many different machine learning algorithms work together to process complex data inputs. Such a system can "learn" to perform tasks by considering examples, typically without being programmed with any task-specific rules, such as automatically generating identifying characteristics from the learning material being processed.

[0048] ANNs are based on a collection of connection units or nodes called artificial neurons, and connections similar to synapses in the biological brain, which are adapted to transmit signals from one artificial neuron to another. The receiving artificial neuron processes the signal and then signals additional artificial neurons connected to that signal.

[0049] In common ANN implementations, the signals at the connections between artificial neurons can be represented by real numbers, and the output of each artificial neuron can be calculated as a nonlinear function of the sum of its inputs (using an activation function). The connections between artificial neurons are often called "edges." Edges connecting artificial neurons typically have weights that adjust as learning progresses. These weights increase or decrease the signal strength at the connection. Artificial neurons may have thresholds, such that a signal is only sent when the aggregate signal crosses that threshold. Typically, artificial neurons are grouped into layers, such as input layers, hidden layers, and output layers. The signal may propagate from the first layer (input layer) to the last layer (output layer) after passing through the layers multiple times.

[0050] The term "hidden neuron layer(s)" can indicate that a layer in a neural network is neither an input layer nor an output layer.

[0051] The term "pseudo-random bit sequence" (PRBS) can refer to a binary sequence that, although generated using deterministic algorithms, may be difficult to predict and exhibit statistical behavior similar to that of a truly random sequence.

[0052] The term "mapping function" here refers to the algorithm used to generate PRBS from a given input. Typically, if the generated PRBS is viewed as a vector, it has a much higher dimension than the corresponding input vector.

[0053] The term "hyper-vector" can refer to a vector with a large number of dimensions. In the concept presented here, the dimension can be in the range of 10,000. Therefore, the term "hyper-dimensional vector" (HD vector) can also be used interchangeably in this paper.

[0054] The terms "associative memory," or Content Addressable Memory (CAM), or associative storage, can refer to a special type of computer memory used, for example, in some very high-speed search applications. Associative memory compares input data (such as an HD vector) with stored data and can return the address of the matching data.

[0055] Unlike standard computer memory (Random Access Memory or RAM) where the user provides a memory address and RAM returns the data word stored at that address, CAM is designed so that the user provides a data word and CAM searches its entire memory to see if that data word is stored anywhere within it. If the data word or a similar data word (i.e., a data vector) is found, CAM returns a list of one or more memory addresses where the word was found (and in some architectural implementations, it also returns the contents of that memory address or other associated data fragments). In other words, associative memory can be similar to ordinary computer memory; that is, when an HD vector X is stored using another HD vector A as an address, X can be retrieved later by addressing the associative memory using A. Furthermore, X can also be retrieved by addressing the memory using an HD vector A' similar to A.

[0056] The term "distance between different hyper-vectors" can be used to describe the result of passing a function that represents the value of a distance. If hyper-vectors are considered as strings, then the distance between two (equal-length) hyper-vectors can be equal to the number of positions where corresponding symbols in the strings differ (Hamming distance). In other words, it can measure the minimum number of substitutions required to transform one string into another, or alternatively, the minimum number of errors required to transform one string into another. Therefore, the term "minimum distance" can represent the shortest distance between a given vector and a set of other vectors.

[0057] The term "candidate answer" can refer to one of a set of response patterns that can be tested against input to a machine learning system represented as a query or cognitive query. The machine learning system can test each candidate answer against the query using a combination of at least one neural network via HD vector combinations and an associative memory, in order to determine the closest candidate answer (i.e., the one closest to the query) as the appropriate response.

[0058] The term "role vector" can refer to a set of HD vectors determined at the end of training of a system's neural network. Role vectors can be associated with combined representation elements of cognitive queries. They can be generated by using data from hidden and output layers as input to a random (or pseudo-random) bit sequence generator to produce random bit sequences used as role vectors.

[0059] The term "filler vector" can refer to an HD vector with the same dimensions as the character vector. Unlike the character vector, the filler vector can be generated continuously during the feeding of sensor input signals representing cognitive queries and candidate answers into at least one neural network of a machine learning system. The filler vector can be generated as a pseudo-random bit sequence.

[0060] The term "Modified National Institute of Standards and Technology (MNIST) database" refers to a large database of handwritten digits commonly used to train various image processing systems. This database is also widely used for training and testing in the field of machine learning. It was created by "remixing" samples from the original NIST dataset. The creators felt that because the NIST training dataset was taken from US Census Bureau employees, while the testing dataset was taken from US high school students, it was not ideally suited for machine learning experiments. Furthermore, the black and white images from NIST were normalized to fit 28×28 pixel bounding boxes and anti-aliasing was applied, which introduced grayscale levels.

[0061] Before providing a more detailed description of the embodiments by describing the various examples, a brief description of the theoretical background that partially reveals its implementation in the embodiments should be given.

[0062] The mutual information between random vectors X and T, which have individual entropies H(X), H(T) and joint entropy H(X,T), can be expressed as:

[0063]

[0064] In the context of the proposed method, it can be assumed that T is a quantized codebook of X, characterized by a conditional distribution p(t|x), causing a soft partition of X. As the cardinality of X increases, the complexity of representing X typically also increases. To determine the quality of the compressed representation, a distortion function d(X,T) can be introduced, which measures the “distance” between the random variable X and its new representation T. A standard measure of this quantity can be given by the rate of the code “transmitted” between X and T, i.e., the interactive information constrained by the maximum permissible average distortion D. Therefore, the rate-distortion function can be defined as follows:

[0065]

[0066] Finding the rate-distortion function may require minimizing a convex function on a convex set of all normalized conditional distributions p(t|x) that satisfy the distortion constraint. This problem can be solved by introducing Lagrange multipliers and solving the associated minimization task.

[0067] Therefore, the information bottleneck method (see above) that uses rate distortion to measure the quality of the compressed representation of X has three drawbacks:

[0068] (i) Distortion measurement is part of the established problem;

[0069] (ii) The cardinality of the compressed representation of X is chosen arbitrarily; and

[0070] (iii) The concept of relevance of the information contained in X is completely absent. This can represent the main elements in the defect list.

[0071] By employing the so-called information bottleneck method—that is, by introducing a "relevant" random vector Y, defining the information of interest to the user, and the joint statistics between X and Y, p(x,y)—a solution that eliminates all drawbacks can be obtained. Therefore, the problem is formulated as finding a compressed representation T of X that maximizes the mutual information between T and Y, I(T;Y). Then, the correlation compression function for a given joint distribution p(x,y) can be defined as…

[0072]

[0073] Here, given X (Markovian conditions), T is conditionally independent of Y, and minimization is performed on all normalized conditional distributions p(t|x) that satisfy the constraints.

[0074] We can now find the correlation-compression function ^R(^D) by again resorting to the Lagrange multiplier and then minimizing the correlation function that follows the normalization and assuming the Markovian condition p(t|x,y)=p(t|x).

[0075] Now turning to artificial neural networks, the random vectors associated with different layers of an artificial neural network can be viewed as Markov chains of continuous internal representations of the input X. Any representation of T can then be defined by the encoder P(T|X) and the correlated decoder P(^Y|T), and can be defined respectively by their information plane coordinates I. x =I(X;T) and I y =I(T;Y) is used for quantification.

[0076] The information bottleneck boundary represents the optimal representation T, which compresses the input X to the maximum extent to obtain the given mutual information about the desired output Y.

[0077] After training a fully connected artificial neural network with an actual output ^Y, the ratio I(Y;^Y) / I(X;Y) quantifies how much relevant information can be captured by the artificial neural network. Thus, in a fully connected neural network, each neuron in a layer within the network can be connected using relevant weighting factors, i.e., linked to each neuron in the next downstream layer of the network.

[0078] Furthermore, the term "encoding problem" can be addressed here: directly encoding hyperdimensional (HD) vectors from sensor data to form structures—that is, encoding cognitive elements as long binary sequences—has been an open problem and is widely referred to as the "encoding problem." The solution proposed here is to solve this encoding problem.

[0079] Essentially, the idea is to use information from sensor data from different layers of an artificial neural network to form hypervector inputs to an associative memory. Therefore, it is speculated that sensor data representing cognitive elements of the same class are converted into binary information that results in similar hypervectors. This can be explained later. Figure 4 We have come to realize this.

[0080] Now we turn to the topic of Vector Symbolic Architecture (VSA), which is based on a set of operators for a fixed-length high-dimensional vector (i.e., a mapping vector) to represent a simplified description of a complete concept. The fixed length of the vectors used for representation can mean that new constituent structures can be formed from simpler structures without increasing the size of the representation, however, at the cost of increased noise levels.

[0081] The simplified representation and cognitive model are essentially manipulated using two operators, performing "binding" and "bounding" functions, which for binary vectors are:

[0082] (i) Binding operators It is defined as an element-wise binary XOR operation, and

[0083] (ii) The binding operator <·>, which is defined as element-wise binary average.

[0084] Then, the concept of a mapping vector can be combined with a sparse distributed memory (SDM) model of associative memory to form an analog mapping unit (AMU), which enables the learning of mappings in a simple way. For example, the concept of "a circle on top of a square" can be derived from the relation The encoding is as follows: "a" is the relation name (above), a1 and a2 are relation roles (variables), and the geometry indicates what the associated filler (value) is. More on this later... Figure 6 The concepts of character vector and fill vector are described in the text.

[0085] Now let's turn to Sparse Distributed Memory (SDM): SDM is a suitable memory for storing analog mapping vectors, which automatically bundles similar mapping examples into well-organized mapping vectors. Basically, an SDM consists of two parts, a binary address matrix A and an integer content matrix C, which are initialized as follows:

[0086] (i) Matrix A is randomly filled with zeros and ones; rows of A represent address vectors;

[0087] (ii) Matrix C is initialized to zero; rows of C are counter vectors.

[0088] Using this setup, a one-to-one link exists between the address vector and the counter vector, ensuring that an active address vector is always accompanied by an active counter vector. Thus, the address and counter vectors typically have a dimension D of approximately 10,000. The number of address and counter vectors defines the required memory size S, and then the address is activated if the Hamming distance representing the correlation between the query (query vector) and the address is less than a threshold θ, where θ is approximately equal to 0.1.

[0089] It can be shown that a vector symbolic architecture including an AMU can be used to construct system modeling analogies representing key cognitive functions. An AMU including an SDM enables non-commutative assemblage of the constituent structures, making it possible to predict new patterns in a given sequence. Such a system can be used to solve commonly used intelligence tests, known as Raven asymptotic matrices (comparisons). Figure 10 ).

[0090] When bundling multiple mapping types

[0091]

[0092] The compositional structure of the role / filler relationship is integrated into the mapping vector to accurately summarize the new compositional structure. The experimental results clearly show that the symbol representing three triangles experiences a much higher correlation value than other candidate symbols.

[0093] Typically, an Analog Mapping Unit (AMU) consists of an SDM with additional input / output circuitry. Therefore, an AMU includes learning and mapping circuitry: it learns type x from an example. k →y k The mapping is performed, and the output vector y' is computed using the bundled mapping vector stored in the SDM. k .

[0094] The accompanying drawings will now be described in detail. All descriptions in the drawings are illustrative. First, a block diagram of an embodiment of a computer-implemented method of the present invention for answering cognitive queries from sensor input signals is given. Then, other embodiments and examples of a machine learning system for answering cognitive queries from sensor input signals will be described.

[0095] Figure 1 A block diagram of an embodiment of a method 100 for answering cognitive queries from sensor input signals is shown. Therefore, testing requires the use of simulation to determine the answer. Sensor input data or sensor input signals can generally be considered as image data, i.e., a rasterized image at a specific resolution. Alternatively, sound files (speech or other sound data) can also be used as sensor input data.

[0096] Method 100 includes feeding a sensor input signal 102 to an input layer of an artificial neural network (at least one) comprising one of the smallest hidden neuron layers among a plurality of hidden neuron layers and an output neuron layer, and determining 104 a hidden layer output signal from each of the plurality of hidden neuron layers and an output signal from the output neuron layer.

[0097] As a next step, method 100 includes generating 106 sets of pseudo-random bit sequences by applying a set of mapping functions, wherein the output signals of the output layer and the hidden layer output signals of at least one or all hidden neuron layers are used as input data for at least one mapping function. Thus, one signal from each artificial neuron can be used to represent a sub-signal of a layer in the neural network. Therefore, each layer of the neural network can represent a signal vector of different dimensions, depending on the number of artificial neurons in a particular layer.

[0098] Furthermore, method 100 includes determining 108 hypervectors using a pseudo-random bit sequence (e.g., generated or derived from a cascade of signal vector sources of an artificial neural network), and storing (one or more) hypervectors 110 in an associative memory, i.e., storing (one or more) hypervectors at locations given by their values, wherein the distance between different hypervectors is determinable (i.e., measurable), for example, Hamming distance values.

[0099] Figure 2 The representation method 100 is shown. Figure 1 The flowchart is a continuation of the embodiment of block diagram 200. Additional steps may include combining 202 different hypervectors stored in the associative memory to derive cognitive queries and candidate answers, and measuring 204 the distance between the cognitive query and the candidate answer, and finally selecting 206 the hypervector associated with the candidate answer, i.e., the associated hypervector, which has the minimum distance, such as the minimum Hamming distance.

[0100] Using this setup, given that a relevant neural network has already been trained using a dataset and candidate answers related to the query domain, it is possible to select the candidate answer with the highest probability of being the correct answer from a given set of candidate answers for a matching query, i.e., a cognitive query. Thus, both the query and the candidate answers are represented by hypervectors.

[0101] Figure 3 A block diagram illustrating a portion of an embodiment of a neural network 300 is shown. For example, 400 neurons x1 to x2 of the first layer 302 are shown. 400 This can be interpreted as an input layer. For example, if 400 input neurons are available, sensor data from an image with 20×20 pixels can be used as the input to neural network 300. The neurons in layer 302 showing the "+a" term can be interpreted as bias factors of layer 302. The same assumption can be made for the "+b" term in hidden layer 304. The final layer 306 does not include bias factors and can therefore also be interpreted as including 10 neurons H h. 2,1 to h 2,10 The output layer shows that the exemplary neural network 300 can be designed to distinguish between 10 different alternatives, namely, the classification of the 10 different input symbols shown in the above image.

[0102] It can also be noted that the exemplary portion of neural network 300 illustrates the concept of a fully connected neural network, where each node of a layer (e.g., layer 302) is connected to each neuron of the next downstream layer (here, layer 304). Each such link may have a weighting factor.

[0103] Figure 4A block diagram of an embodiment 400 of the generally proposed basic concept is shown. Sensor data 402 is fed into the input layer of a neural network 404. Binary information is fed from each hidden layer and output layer of the neural network 404 to a PRBS1 (pseudo-random binary sequence) generator 406 or a PRBS2 generator 408 to generate sequences S1 and S2. The input signals of the PRBSx generators represent their respective initial states. In the case that the neural network 404 is a convolutional neural network, a convolution step is typically performed between the input layer and the first hidden layer to significantly reduce the number of neurons required.

[0104] Sequences S1 and S2 can be combined, for example, cascaded in CAT unit 410, whose output is a hyperdimensional vector v used as input to associative memory 412. It can also be noted that the input signals from the hidden and output layers of neural network 404 are also vector data, typically having a single value, i.e., one dimension for each neuron in their associated layers.

[0105] Figure 5 A block diagram of embodiment 500 of the concept presented herein, along with further details, is shown. Sensor data 502 is fed into neural network 504, as already shown... Figure 4 As explained in the context, neural network 504 has an input layer 518, hidden layers 520, 522, 524, 526, and an output layer 528 for instructing the decision 506Y^ of neural network 504. It should also be noted that the number of nodes in each neural network layer of neural network 504 will be understood as an example; real neural networks can have more artificial neurons per layer, such as hundreds or thousands, and can also have more hidden layers.

[0106] As shown in the figure, the decision information 506Y^ is fed to mappers MAP1(Y^)510 and MAP2(.;Y^)512 via line 508, until MAP N (.;Y^)514. Additionally, the vector information derived from the hidden layers 520, 522, 524, and 526 of the neural network 504 is used as additional vector inputs to mappers 510, 512, and up to 514. Then, the generated bit sequences S1, S2, and up to S... N It can be combined by cascaded units CAT 516 so that it can be used as an input to the associative memory 518 in the form of a hypervector v.

[0107] The effect used here is that sensor data 502, which shows a similar pattern, can be transformed into hypervectors stored in an associative memory, which can be relatively close to each other, i.e., have only a short distance from each other, and are derived, for example, from a distance function between any two hypervectors using Hamming distance.

[0108] Figure 6A block diagram of a complete system embodiment 600 of the proposed concept is shown. Sensor data input 602 is fed into multiple neural networks, shown as classifiers 1 604, 2 606, up to classifier n 608. Essentially, these classifiers 604, 606, 608 pass the input signal to a unit 610 for monitoring gradient descent training error. If the gradient descent value (e.g., SGD stochastic gradient descent) falls below a predefined threshold, the end of training can be signaled, thereby generating multiple hyperdimensional random vectors 1 to n with dimensions in the thousands (e.g., 10,000) by units 612, 614, 616 that construct the role vectors role 1, role 2, up to role n, for storage in associative memory 618. This can represent a confirmation that the training of neural networks 604, 606, up to 608 is complete.

[0109] On the other hand, the sequence of padding vectors 1, 2 to n generated by the hyperdimensional pseudo-random vector generation units 620, 622, 624 is based on the sequence of input signals from sensor input data 602. Symbolically, query input 626 and input values ​​628 from the candidate answer input sequence are also shown. In any case, the response 630 is generated by taking query 626 as a hyperdimensional vector input and candidate answer sequence 628 as a hyperdimensional vector sequence input. Based on the Hamming distance in the association memory 618 between the hyperdimensional vector representing query 626 and one of the candidate answer sequences 628 (also represented as hyperdimensional vectors), the most likely correct HD vector from candidate answer sequence 628 constructs the response 630 from the candidate answer 628 with the shortest distance to query 626.

[0110] Using this general embodiment, more specific embodiments can be used to demonstrate the capabilities of the concepts presented herein. Three different types of classifiers are used: a first classifier (e.g., classifier 604) is trained using the MNIST database; a second classifier (e.g., classifier 606) can be trained to identify how many digits are seen in a specific field of the input data (e.g., reference). Figure 7A Examples of input data; as a third classifier (e.g., classifier 608), a position counter can be used and it can identify whether to detect one, two, or three digits. Here, as role vectors, r_shape, r_num, and r_pos HD vectors can be generated by units 612, 614, and 616 and stored in the associative memory. Furthermore, sequences of f_shape, f_num, and f_pos vectors can be generated from units 620, 622, and 624 and also fed into the associative memory. The results of the metric set are... Figure 7C As shown in the image.

[0111] Figures 7A-7CResults are presented for the binding and bundling formulas, the potential input data, and the average measures of correct and incorrect decisions. The results show a significant difference between the two alternatives, demonstrating the effectiveness of the concept presented here.

[0112] Figure 7A The program shows potential inputs of digits (i.e., integers) 1, 5, 9 in groups of one, two, or three digits, as well as an empty input for a random mode. Figure 7B The relevant binding and bundling combinations in a set of equations related to the embodiment described last are shown, as well as combinations that lead to candidate answers and potential responses. It can also be noted that the binding and bundling functional units can be integrated into the associative memory 618.

[0113] Figure 7C The results of this setup are shown in an xy plot, where the metric is displayed on the y-axis and the accuracy at 100% (over 100 trials) is shown on the x-axis. The results show that the average metric 702 for incorrect decisions is significantly different from the average metric 704 for correct decisions. Furthermore, the variance for correct decisions is much smaller.

[0114] For reasons of completeness, Figure 8 A block diagram of an embodiment of a machine learning system 800 for answering cognitive queries from sensor input signals is shown. The machine learning system 800 includes: a sensor 802 adapted to feed sensor input signals to an input layer of an artificial neural network including a plurality of hidden neuron layers and an output neuron layer; a first determining unit 804 adapted to determine hidden layer output signals from each of the plurality of hidden neuron layers and output signals from the output neuron layer; and a generator unit 806 adapted to generate a set of pseudo-random bit sequences by applying a mapping function set using the output signals of the output layers and the hidden layer output signals of one of the hidden neuron layers as input data for a mapping function.

[0115] Furthermore, the machine learning system 800 includes a second determining unit 808 adapted to determine hypervectors using the pseudo-random bit sequence set, and a storage module 810 adapted to store hypervectors in an associative memory, wherein the distance between different hypervectors is determinable. As mentioned above, Hamming distance can be a good tool for determining the distance between hyperdimensional vectors in the storage module (i.e., the associative memory), as described above.

[0116] Embodiments of the present invention can be implemented with any type of computer, regardless of whether the platform is suitable for storing and / or executing program code. Figure 9 As an example, a computing system 900 suitable for executing program code related to the proposed method is shown.

[0117] The computing system 900 is merely one example of a suitable computer system and is not intended to impose any limitation on the scope or functionality of the embodiments of the invention described herein, regardless of whether the computer system 900 can be implemented and / or perform any of the functions set forth above. Within the computer system 900, there are components that can operate with a multitude of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the computer system / server 900 include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computers, and distributed cloud computing environments that include any of the aforementioned systems or devices. The computer system / server 900 can be described in the general context of computer system executable instructions, such as program modules executed by the computer system 900. Typically, program modules may include routines, programs, objects, components, logic, data structures, etc., that perform a specific task or implement a specific abstract data type. The computer system / server 900 can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can reside on local and remote computer system storage media, including memory storage devices.

[0118] As shown in the figure, the computer system / server 900 is illustrated as a general-purpose computing device. Components of the computer system / server 900 may include, but are not limited to, one or more processors or processing units 902, system memory 904, and a bus 906 that couples various system components, including system memory 904, to the processor 902. Bus 906 represents one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of various bus architectures. By way of example and not limitation, these architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses. The computer system / server 900 typically includes various computer system readable media. Such media can be any available media accessible to the computer system / server 900, and it includes volatile and non-volatile media, removable and non-removable media.

[0119] System memory 904 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 908 and / or cache memory 910. Computer system / server 900 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 912 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, and generally referred to as a "hard disk drive"). Although not shown, a disk drive may be provided for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk"), and an optical disk drive may be provided for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media. In such instances, each may be connected to bus 906 via one or more data media interfaces. As will be further described and depicted below, memory 904 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the invention.

[0120] A program / utility having at least one set of program modules 916, along with an operating system, one or more applications, other program modules, and program data, may be stored in memory 904, as an example and not a limitation. Each of the operating system, one or more applications, other program modules, and program data, or some combination thereof, may include an implementation of a networking environment. Program modules 916 generally perform the functions and / or methods of embodiments of the invention as described herein.

[0121] The computer system / server 900 can also communicate with one or more external devices 918 (e.g., keyboard, pointing device, display 920, etc.), one or more devices that enable users to interact with the computer system / server 900, and / or any device that enables the computer system / server 900 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 914. Furthermore, the computer system / server 900 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 922. As shown, network adapter 922 communicates with other modules of the computer system / server 900 via bus 906. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the computer system / server 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0122] Additionally, a machine learning system 800 for answering cognitive queries from sensor input signals can be attached to a bus system 906. The computing system 900 can be used as an input / output device.

[0123] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies on the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0124] This invention can be implemented as a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.

[0125] The medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system used for propagation. Examples of computer-readable media include semiconductor or solid-state memory, magnetic tape, removable computer disks, random access memory (RAM), read-only memory (ROM), hard disks, and optical discs. Current examples of optical discs include optical disc read-only memory (CD-ROM), optical disc read / write (CD-R / W), DVDs, and Blu-ray discs.

[0126] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0127] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0128] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0129] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0130] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0131] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0133] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0134] All means or steps in the following claims, plus corresponding structures, materials, actions, and equivalents of the functional elements, are intended to include any structure, material, or action for performing a function in combination with other claimed elements as specifically claimed. The invention has been described for purposes of illustration and description, but this description is not exhaustive or intended to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments were chosen and described in order to best explain the principles and practical application of the invention and to enable others skilled in the art to understand various embodiments of the invention with various modifications, as suitable for the particular intended use.

Claims

1. A method for answering cognitive queries from sensor input signals, the method comprising: The sensor input signal is fed into the input layer of an artificial neural network that includes multiple hidden neuron layers and output neuron layers; Determine the hidden layer output signal from each of the plurality of hidden neuron layers and the output signal from the output neuron layer; A group of pseudo-random bit sequences is generated by applying multiple mapping functions, wherein at least the output signal of the output neuron layer and the hidden layer output signal of at least one hidden neuron layer directly from at least one hidden neuron layer in the hidden neuron layer are fed as input data to a mapping function, wherein the at least one hidden neuron layer includes at least one hidden neuron layer that is not immediately preceding the output neuron layer, and each of the multiple mapping functions is fed at least one item selected from the group as input, the group including: signals received directly from at least two hidden neuron layers in the multiple hidden neuron layers, and signals received directly from the output neuron layer and at least one hidden neuron layer in the hidden neuron layer, wherein each of the multiple hidden neuron layers is used to feed directly when generating the group of pseudo-random bit sequences; The pseudo-random bit sequence group generated by the corresponding plurality of mapping functions is used to determine the hypervector; as well as The hypervectors are stored in an associative memory, where the distance between different hypervectors is deterministic.

2. The method according to claim 1, further comprising: Different hypervectors stored in the associated memory are combined to derive the cognitive query and candidate answers, each candidate answer being obtained from the sensor input signal.

3. The method according to claim 2, further comprising: Measure the distance between the hypervector representing the cognitive query and the hypervector representing the candidate answer.

4. The method according to claim 3, further comprising: The supervector associated with the candidate answer is selected such that the selected supervector has the minimum distance to the supervector representing the cognitive query.

5. The method of claim 2, wherein feeding the sensor input data to the input layer of the artificial neural network further comprises: The sensor input data is fed into the input layer of multiple artificial neural networks.

6. The method according to claim 1, wherein generating the pseudo-random bit sequence group further comprises: The output signal of the output neuron layer and the hidden layer output signals of the plurality of hidden neuron layers are used as input data for a mapping function.

7. The method according to claim 1, wherein the neural network is a convolutional neural network.

8. The method according to claim 1, further comprising: At the end of the training of the neural network, multiple random hypervectors are generated as role vectors and stored in the associated memory.

9. The method according to claim 5, further comprising: Based on the output data of the multiple neural networks, multiple pseudo-random hypervectors are continuously generated as padding vectors for each new sensor input dataset, and the pseudo-random hypervectors are stored in the associated memory.

10. The method of claim 2, wherein the combination comprises: Different hypervectors are bound together by performing a binary XOR operation on each element of the vector.

11. The method of claim 2, wherein the combination comprises: Different hypervectors are bound together by the binary average of each vector element.

12. The method of claim 8, wherein if the stochastic gradient descent value determined between two sets of input vectors of the sensor input signal to the neural network remains below a predefined threshold, then at the end of the training of the artificial neural network, the role vector is stored in the associated memory.

13. A machine learning system for answering cognitive queries from sensor input signals, the machine learning system comprising: A sensor adapted to feed sensor input signals to the input layer of an artificial neural network comprising multiple layers of hidden neurons and output neurons; At least one computer processor is adapted to determine hidden layer output signals from each of the plurality of hidden neuron layers and output signals from the output neuron layers; The at least one computer processor is further adapted to generate a group of pseudo-random bit sequences by applying a plurality of mapping functions, wherein the output signal of the output neuron layer and the hidden layer output signal of the at least one hidden neuron layer of the hidden neuron layer directly from the at least one hidden neuron layer of the hidden neuron layer are fed as input data to a mapping function, wherein the at least one hidden neuron layer of the hidden neuron layer includes at least one hidden neuron layer that is not immediately preceding the output neuron layer, and each of the plurality of mapping functions is fed at least one item selected from the group as input, the group including: signals received directly from at least two hidden neuron layers of the plurality of hidden neuron layers, and signals received directly from the output neuron layer and the at least one hidden neuron layer of the hidden neuron layer, wherein each of the plurality of hidden neuron layers is used to feed directly when generating the group of pseudo-random bit sequences; The at least one computer processor is further adapted to determine the hypervector by cascading the pseudo-random bit sequence group generated by the respective plurality of mapping functions; as well as A storage module adapted to store the hypervectors in an associative memory, wherein the distance between different hypervectors is determinable.

14. The machine learning system of claim 13, wherein the at least one computer processor further comprises: It is suitable to combine different hypervectors stored in the associated memory to derive the cognitive query and candidate answers, each candidate answer being obtained from the sensor input signal.

15. The machine learning system of claim 14, wherein the at least one computer processor further comprises: Suitable for measuring the distance between the hypervector representing the cognitive query and the hypervector representing the candidate answer.

16. The machine learning system of claim 15, wherein the at least one computer processor further comprises: It is suitable to select the supervector associated with the candidate answer, wherein the selected supervector has the minimum distance to the supervector representing the cognitive query.

17. The machine learning system of claim 14, wherein the sensor for feeding sensor input data to the input layer of the artificial neural network is further adapted to: The sensor input data is fed into the input layer of multiple artificial neural networks.

18. The machine learning system of claim 13, wherein generating the pseudo-random bit sequence group further comprises: The output signal of the output neuron layer and the hidden layer output signals of the plurality of hidden neuron layers are used as input data for a mapping function.

19. The machine learning system of claim 13, wherein the neural network is a convolutional neural network.

20. The machine learning system of claim 13, wherein the at least one computer processor further comprises: It is suitable to generate multiple random hypervectors as role vectors at the end of the training of the neural network, and store them in the associated memory.

21. The machine learning system of claim 17, wherein the at least one computer processor further comprises: It is suitable to continuously generate multiple pseudo-random hypervectors as padding vectors for each new sensor input dataset based on the output data of the multiple neural networks, and the pseudo-random hypervectors are stored in the associated memory.

22. The machine learning system of claim 14, wherein the at least one computer processor further comprises: It is suitable for binding different hypervectors by performing binary XOR operations on each vector element.

23. The machine learning system of claim 14, wherein the at least one computer processor further: It is suitable for binding different hypervectors by binary averaging of vector elements.

24. The machine learning system of claim 20, wherein if a stochastic gradient descent value determined between two sets of input vectors of the sensor input signal to the neural network remains below a predefined threshold, then at the end of the training of the artificial neural network, the role vector is stored in the associated memory.

25. A computer program product for answering cognitive queries from sensor input signals, the computer program product having a computer-readable storage medium thereon containing program instructions executable by a computing system to cause the computing system to: The sensor input signal is fed into the input layer of an artificial neural network that includes multiple hidden neuron layers and output neuron layers; Determine the hidden layer output signal from each of the plurality of hidden neuron layers and the output signal from the output neuron layer; A group of pseudo-random bit sequences is generated by applying multiple mapping functions, wherein at least the output signal of the output neuron layer and the hidden layer output signal of at least one hidden neuron layer directly from at least one hidden neuron layer in the hidden neuron layer are fed as input data to a mapping function, wherein the at least one hidden neuron layer includes at least one hidden neuron layer that is not immediately preceding the output neuron layer, and each of the multiple mapping functions is fed at least one item selected from the group as input, the group including: signals received directly from at least two hidden neuron layers in the multiple hidden neuron layers, and signals received directly from the output neuron layer and at least one hidden neuron layer in the hidden neuron layer, wherein each of the multiple hidden neuron layers is used to feed directly when generating the group of pseudo-random bit sequences; The pseudo-random bit sequence group generated by the corresponding plurality of mapping functions is used to determine the hypervector; as well as The hypervectors are stored in an associative memory, where the distance between different hypervectors is deterministic.