Characterizing the activity in a recurrent artificial neural network and encoding and decoding information
By training the neural network to generate and process the topological representation of the active patterns in the source neural network, the problem of featureization and encoding of the activity patterns of the recurrent artificial neural network is solved, and the efficiency of identification of decision moments and signal processing is achieved.
Patent Information
- Application Number
- CN201980053141.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-11
- Filing Date
- 2019-06-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-06-06
AI Technical Summary
The prior art is difficult to effectively characterize the activity patterns in recurrent artificial neural networks, especially when identifying decision moments and encoding/decoding signals.
These representations are used to encode and decode by training the neural network in response to input to generate representations of topological structures in the active pattern that occurs in the source neural network. The device includes a processor to receive and process these representations and is used in different scenarios.
It realizes effective feature and encoding of the activity mode of the recurrent artificial neural network, supporting the identification of decision-making moments and efficient transmission and storage of signals.
Smart Images

Figure CN112567389B_ABST
Abstract
Description
Background Art
[0001] This specification relates to the characterization of activities in a recurrent artificial neural network. The characterization of activities can be used, for example, in the identification of decision moments, and in encoding / decoding signals in scenarios such as transmission, encryption, and data storage. It also relates to encoding and decoding information, and systems and techniques for using the encoded information in various scenarios. The encoded information can represent activities in a neural network (e.g., a recurrent neural network).
[0002] An artificial neural network is a device inspired by the structural and functional aspects of biological neural networks. In particular, an artificial neural network uses a system of interconnected structures called nodes to simulate the information encoding and other processing capabilities of biological neural networks. The arrangement and strength of the connections between nodes in an artificial neural network determine the result of information processing or information storage through the artificial neural network.
[0003] A neural network can be trained to produce a desired signal flow in the network and achieve a desired information processing or information storage result. Typically, training a neural network will change the arrangement and / or strength of the connections between nodes during a learning phase. When a neural network achieves a sufficiently appropriate processing result for a given set of inputs, the neural network can be considered trained.
[0004] Artificial neural networks can be used in a wide variety of different devices to perform non-linear data processing and analysis. Non-linear data processing does not satisfy the superposition principle, i.e., the variable to be determined cannot be written as a linear sum of independent components. Examples of scenarios where non-linear data processing is useful include pattern and sequence recognition, speech processing, novelty detection and sequential decision-making, complex system modeling, and systems and techniques in a wide variety of other scenarios.
[0005] Both encoding and decoding convert information from one form or representation to another. Different representations can provide different features that are more or less useful in different applications. For example, some forms or representations of information (e.g., natural language) may be more easily understood by humans. Other forms or representations may be smaller in size (e.g., "compressed") and easier to transmit or store. Still other forms or representations may deliberately obscure the information content (e.g., the information can be encrypted).
[0006] Regardless of the specific application, the encoding or decoding process will generally follow a predefined set of rules or algorithms that establish a correspondence between different forms or representations of information. For example, an encoding process that produces binary code can assign a role or meaning to each bit based on its position within a binary sequence or vector. Summary of the Invention
[0007] This document relates to encoding and decoding information, as well as systems and techniques for using encoded information in various scenarios. For example, in one implementation, a device includes a neural network that is trained to produce an approximation of a first representation of a topology in a pattern of activity that occurs in a source neural network in response to a first input, an approximation of a second representation of a topology in a pattern of activity that occurs in the source neural network in response to a second input, and an approximation of a third representation of a topology in a pattern of activity that occurs in the source neural network in response to a third input.
[0008] This implementation and other implementations may include one or more of the following features. The topologies may all include two or more nodes in the source neural network and one or more edges between the nodes. The topologies may include simplices. The topologies may enclose cavities. Each of the first representation, the second representation, and the third representation may represent a topology that occurs in the source neural network only at a time during which the pattern of activity has a complexity distinguishable from the complexity of other activity in response to the corresponding input among the inputs. The device may further include a processor that is coupled to receive the approximations of the representations produced by the neural network device and process the received approximations. The processor may include a second neural network that has been trained to process the representations produced by the neural network. Each of the first representation, the second representation, and the third representation may include multi-valued, non-binary digits. Each of the first representation, the second representation, and the third representation may represent the occurrence of the topology without specifying where in the source neural network the pattern of activity appears. The device may include a smart phone. The source neural network may be a recurrent neural network.
[0009] In another implementation, a device includes a neural network that is coupled to input a representation of a topology in a pattern of activity that occurs in a source neural network in response to a plurality of different inputs. The neural network is trained to process the representation and produce a response output.
[0010] This implementation and other implementations may include one or more of the following features. The topology may all include two or more nodes in the source neural network and one or more edges between the nodes. The topology may include a simplex. The representation of the topology may represent the topology that occurs in the source neural network only at a time when the pattern of activity during that time has a complexity distinguishable from the complexity of other activity in response to a corresponding input in the input. The device may include a neural network trained to produce a corresponding approximation of the representation of the topology in the pattern of activity that occurs in the source neural network in response to a plurality of different inputs. The representation of the topology may include multi-valued, non-binary numbers. The representation of the topology may represent the occurrence of the topology without specifying where in the source neural network the pattern of activity appears. The source neural network may be a recurrent neural network.
[0011] In another implementation, a method is implemented by a neural network device and includes: inputting a representation of a topology in a pattern of activity in a source neural network, where the activity is in response to an input into the source neural network; processing the representation; and outputting a result of the processing of the representation. The processing is consistent with training the neural network to process different such representations of the topology in the pattern of activity in the source neural network.
[0012] This implementation and other implementations may include one or more of the following features. The topology may all include two or more nodes in the source neural network and one or more edges between the nodes. The topology may include a simplex. The topology may enclose a cavity. The representation of the topology may represent the topology that occurs in the source neural network only at a time when the pattern of activity during that time has a complexity distinguishable from the complexity of other activity in response to a corresponding input in the input. The representation of the topology may include multi-valued, non-binary numbers. The representation of the topology may represent the occurrence of the topology without specifying where in the source neural network the pattern of activity appears. The source neural network may be a recurrent neural network.
[0013] Details of one or more implementations described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic illustration of the structure of a recurrent artificial neural network device.
[0015] Figure 2 and Figure 3 are schematic illustrations of cycling the functionality of a recurrent artificial neural network device within different time windows.
[0016] Figure 4 is a flowchart of a process for identifying decision moments in the network based on the characterization of activity in a recurrent artificial neural network.
[0017] Figure 5 is a schematic illustration of a pattern of activity that can be identified and used to identify decision moments in a recurrent artificial neural network.
[0018] Figure 6 is a schematic illustration of a pattern of activity that can be identified and used to identify decision moments in a recurrent artificial neural network.
[0019] Figure 7 is a schematic illustration of a pattern of activity that can be identified and used to identify decision moments in a recurrent artificial neural network.
[0020] Figure 8 is a schematic illustration of a data table for determining the complexity or degree of ordering in a pattern of activity that can be used in a recurrent artificial neural network device.
[0021] Figure 9 is a schematic illustration of the determination of a specific timing of a pattern of activity having distinguishable complexity.
[0022] Figure 10 is a flowchart of a process for encoding a signal using the network based on the characterization of activity in a recurrent artificial neural network.
[0023] Figure 11 is a flowchart of a process for decoding a signal using the network based on the characterization of activity in a recurrent artificial neural network.
[0024] Figure 12 、 Figure 13 and Figure 14 are schematic illustrations of a binary form or representation of a topology.
[0025] Figure 15 and Figure 16 schematically illustrate an example of how the presence or absence of features corresponding to different bits are not independent of each other.
[0026] Figure 17 、 Figure 18 、 Figure 19 、 Figure 20 are schematic illustrations of the use of representations of the occurrence of a topology in the activity in a neural network in four different classification systems.
[0027] Figure 21 and Figure 22 is a schematic illustration of an edge device including a local artificial neural network that can be trained using an occurrence representation corresponding to the topology of the activities in the source neural network.
[0028] Figure 23 is a schematic illustration of a system in which a local neural network can be trained using an occurrence representation corresponding to the topology of the activities in the source neural network.
[0029] Figure 24 and Figure 25 and Figure 26 and Figure 27 are schematic illustrations of using an occurrence representation of the topology in the activities of a neural network in four different systems.
[0030] Figure 28 is a schematic illustration of a system 0 that includes an artificial neural network that can be trained using an occurrence representation corresponding to the topology of the activities in the source neural network.
[0031] Like reference numerals in the various figures indicate like elements. DETAILED DESCRIPTION
[0032] Figure 1 is a schematic illustration of the structure of a recurrent artificial neural network device 100. The recurrent artificial neural network device 100 is a device that uses a system of interconnected nodes to simulate the information encoding and other processing capabilities of a biological neuron network. The recurrent artificial neural network device 100 can be implemented in hardware, software, or a combination thereof.
[0033] An illustration of the recurrent artificial neural network device 100 includes a plurality of nodes 101, 102, ……, 107 interconnected by a plurality of structural links 110. The nodes 101, 102, ……, 107 are discrete information processing constructs similar to neurons in a biological network. The nodes 101, 102, ……, 107 typically process one or more input signals received through one or more of the links 110 to produce one or more output signals output through one or more of the links 110. For example, in some implementations, the nodes 101, 102, ……, 107 can be artificial neurons that weight and sum multiple input signals, pass the sum through one or more non-linear activation functions, and output one or more output signals.
[0034] Nodes 101, 102, ……, 107 can operate as accumulators. For example, nodes 101, 102, ……, 107 can operate according to an integrate-and-fire model, in which one or more signals accumulate in a first node until a threshold is reached. After reaching the threshold, the first node fires by transmitting an output signal to a connected second node along one or more of links 110. In turn, the second nodes 101, 102, ……, 107 accumulate the received signals, and if the threshold is reached, the second nodes 101, 102, ……, 107 transmit another output signal to another connected node.
[0035] Structural links 110 are connections capable of transmitting signals between nodes 101, 102, ……, 107. For convenience, all structural links 110 are considered to be the same bi-directional links herein, which transmit signals from each first node among nodes 101, 102, ……, 107 to each second node among nodes 101, 102, ……, 107 in the same way as signals are transmitted from the second node to the first node. However, this may not necessarily be the case. For example, some or all of the structural links 110 can be unidirectional links that transmit signals from a first node among nodes 101, 102, ……, 107 to a second node among nodes 101, 102, ……, 107 without transmitting signals from the second node to the first node.
[0036] As another example, in some implementations, the structural links 110 can have various characteristics different from or in addition to directionality. For example, in some implementations, different structural links 110 can carry signals of different amplitudes - resulting in different interconnection strengths between the corresponding nodes among nodes 101, 102, ……, 107. As another example, different structural links 110 can carry different types of signals (e.g., inhibitory and / or excitatory signals). In fact, in some implementations, the structural links 110 can mimic the links between somatic cells in a biological system and reflect at least a portion of the great morphological, chemical, and other diversity of such links.
[0037] In the illustrated implementation, the recurrent artificial neural network device 100 is a clique network (or sub-network) because each node 101, 102, ……, 107 is connected to every other node 101, 102, ……, 107. This need not be the case. Instead, in some implementations, each node 101, 102, ……, 107 may be connected to a proper subset of nodes 101, 102, ……, 107 (by the same link or various links, as the case may be).
[0038] For the sake of clear illustration, the recurrent artificial neural network device 100 is illustrated as having only seven nodes. Typically, real-world neural network devices will include a significantly larger number of nodes. For example, in some implementations, a neural network device may include hundreds of thousands, millions, or even billions of nodes. Thus, the recurrent neural network device 100 can be a small part (i.e., a sub-network) of a larger recurrent artificial neural network.
[0039] In a biological neural network device, the process of summing and signal transmission requires the passage of time in the real world. For example, the soma of a neuron integrates the inputs received over time, and the signal transmission from neuron to neuron takes time, which is determined by, for example, the signal transmission speed and the nature and length of the links between neurons. Thus, the state of a biological neural network device is dynamic and changes over time.
[0040] In an artificial recurrent neural network device, time is artificial and represented using a mathematical construct. For example, the transmission of a signal from node to node does not require the passage of real-world time and can be represented in artificial units that are typically independent of the passage of real-world time, such as measured by a computer clock cycle or otherwise. However, the state of an artificial recurrent neural network device can be described as “dynamic” because it changes with respect to these artificial units.
[0041] Note that, for convenience, these artificial units are referred to as “time” units in this document. However, it should be understood that these units are artificial and generally do not correspond to the passage of real-world time.
[0042] Figure 2 and Figure 3is a schematic illustration of cycling the functionality of the recurrent artificial neural network device 100 over different time windows. Since the state of the device 100 is dynamic, the signal transmission activities occurring within one window can be used to represent the functionality of the device 100. Such functional illustrations typically show the activities in only a small portion of the links 110. In particular, since not every link 110 typically transmits signals within a particular window, not every link 110 is illustrated as actively contributing to the functionality of the device 100 in these illustrations.
[0043] In Figure 2 and Figure 3 illustrations, the active links 110 are illustrated as relatively thick solid lines connecting a pair of nodes 101, 102, …, 107. In contrast, the inactive links 110 are illustrated as dashed lines. This is for illustrative purposes only. In other words, the structural connections formed by the links 110 exist regardless of whether the link 110 is active. However, this formalism highlights the activities and functionality of the device 100.
[0044] In addition to schematically illustrating the presence of activities along the links, the direction of the activities is also schematically illustrated. In particular, the relatively thick solid lines illustrating the active links 110 also include arrows indicating the direction of signal transmission along the link during the relevant window. Typically, the direction of signal transmission within a single window does not deterministically limit the link to a unidirectional link with the indicated directionality. Instead, in a first functional illustration for a first time window, the link can be active in a first direction. In a second functional illustration for a second window, the link can be active in the opposite direction. However, in some cases, such as, for example, in a recurrent artificial neural network device 100 that includes only unidirectional links, the directionality of signal transmission will deterministically indicate the directionality of the link.
[0045] In a feedforward neural network device, information moves only in a single direction (i.e., forward) to the node output layer at the end of the network. The feedforward neural network device indicates that a “decision” has been reached and information processing is complete by the propagation of signals through the network to the output layer.
[0046] In contrast, in a recurrent neural network, the connections between nodes form loops, and the activities of the network progress dynamically without an easily identifiable decision. For example, even in a three-node recurrent neural network, the first node can transmit a signal to the second node, and in response, the second node can transmit a signal to the third node. In response, the third node can transmit a signal back to the first node. The signal received by the first node can be—at least in part—in response to the signal transmitted from that same node.
[0047] Schematic functional illustration Figure 2 and Figure 3 This is illustrated in a network that is only slightly larger than a three-node recurrent neural network. Figure 2 The functional illustration shown in can illustrate the activity within the first window, and Figure 3 can illustrate the activity within the immediately following second window. As shown, the set of signal transmission activities appears to originate in node 104 and progresses in a generally clockwise direction through device 100 during the first window. Within the second window, at least some of the signal transmission activities generally appear to return to node 104. Even in such a simplistic illustration, the signal transmission does not occur in a manner that produces a clearly identifiable output or ending.
[0048] When considering a recurrent neural network with, for example, thousands of nodes or more, it can be recognized that signal propagation can occur through a large number of paths and that these signals lack a clearly identifiable "output" location or time. Although the network can be designed to return to a quiescent state where only background activity occurs or even no signal transmission activity occurs, the quiescent state itself does not indicate the result of information processing. Regardless of the input, the recurrent neural network always returns to the quiescent state. Thus, the "output" or result of information processing is encoded in the activity that occurs in the recurrent neural network in response to a specific input.
[0049] Figure 4 is a flowchart of a process 400 for characterizing decision moments in a network based on the characteristics of the activity in a recurrent artificial neural network. A decision moment is a time point at which the activity in a recurrent artificial neural network indicates the result of the network's information processing in response to an input. Process 400 can be executed by a system of one or more data processing devices that perform operations based on a set of one or more machine-readable instructions. For example, process 400 can be executed by the same system of one or more computers that execute software for implementing the recurrent artificial neural network used in process 400.
[0050] At 405, the system executing process 400 receives notice that a signal has been input into the recurrent artificial neural network. In some cases, the input of the signal is a discrete injection event, where, for example, information is injected into one or more nodes and / or one or more links of the neural network. In other cases, the input of the signal is an information stream that is injected into one or more nodes and / or links of the neural network over a period of time. This notice indicates that the artificial neural network is actively processing information and is not, for example, in a quiescent state. In some cases, this notice is received from the neural network itself, for example, such as when the neural network exits a recognizable quiescent state.
[0051] At 410, the system executing process 400 divides the response activities in the network into a set of windows. In cases where the injection is a discrete event, the window can subdivide the time between the injection and the return to the quiescent state into multiple periods during which the activity exhibits variable complexity. In cases where the injection is an information flow, the duration of the injection (and optionally the time to return to the quiescent state after the injection is complete) can be subdivided into windows during which the activity exhibits variable complexity. Various methods for determining the complexity of the activity are discussed further below.
[0052] In some implementations, all the windows have the same duration, but this need not be the case. Instead, in some implementations, the windows can have different durations. For example, in some implementations, the duration can increase as the time since the discrete injection event has occurred increases.
[0053] In some implementations, the windows can be a continuous series of individual windows. In other implementations, the windows overlap in time such that one window starts before the previous window ends. In some cases, the windows can be moving windows that move in time.
[0054] In some implementations, different window durations are defined for different determinations of the complexity of the activity. For example, for an activity pattern defined as occurring between a relatively large number of nodes, the window can have a relatively longer duration compared to the window defined for an activity pattern defined as occurring between a relatively small number of nodes. For example, in the scenario of activity pattern 500 ( Figure 5 ), the window defined for identifying the activity commensurate with pattern 530 can be longer than the window defined for identifying the activity commensurate with pattern 505.
[0055] At 415, the system executing process 400 identifies patterns in the activities within different windows in the network. As discussed further below, patterns in the activities can be identified by treating the functional graph as a topological space with nodes as points. In some implementations, the identified activity pattern is a clique in the functional graph of the network, for example, a directed clique.
[0056] At 420, the system executing process 400 determines the complexity of the activity patterns within different windows. The complexity can be a measure of the likelihood that an ordered pattern of the activity appears within a window. Thus, an activity pattern that appears randomly will be relatively simple. On the other hand, an activity pattern that shows a non-random order is relatively complex. For example, in some implementations, the complexity of the activity pattern can be measured using, for example, the simplex count or the Betti number of the activity pattern.
[0057] At 425, the system of process 400 determines a particular time of an activity pattern having a distinguishable complexity. The particular activity pattern may be distinguishable in complexity based on, for example, a deviation upward or downward from a fixed or variable baseline. In other words, a particular time may be determined of a particularly high or particularly low activity pattern indicating a non-random order in the activity.
[0058] For example, in the case where the signal input is a discrete injection event, a deviation, e.g., a deviation relative to a stable baseline or relative to a curve characteristic of the average response of a neural network to a variety of different discrete injection events, may be used to determine the particular time of a distinguishable complex activity pattern. As another example, in the case where information is input in a stream form, a large change in complexity during streaming may be used to determine the particular time of a distinguishable complex activity pattern.
[0059] At 430, the system of process 400 schedules the reading of the output from the neural network based on the particular time of the distinguishable complex activity pattern. For example, in some implementations, the output of the neural network may be read at the same time the distinguishable complex activity pattern occurs. In an implementation where the complexity deviation indicates a relatively high non-random order in the activity, the observed activity pattern itself may also be treated as the output of a recurrent artificial neural network.
[0060] Figure 5 is an illustration of an activity pattern 500 that can be recognized and used to identify decision moments in a recurrent artificial neural network. For example, pattern 500 may be recognized at 415 in process 400 ( Figure 4 ).
[0061] Pattern 500 is an illustration of an activity in a recurrent artificial neural network. During the application of pattern 500, the functional graph is regarded as a topological space with nodes as points. The activity in the nodes and links commensurate with pattern 500 can be recognized as ordered, regardless of the identity of the particular nodes and / or links participating in the activity. For example, a first pattern 505 may represent Figure 2 the activity among nodes 101, 104, 105 in Figure 3 with point 0 in pattern 505 as node 104, point 1 as node 105, and point 2 as node 101. As another example, the first pattern 505 may also represent Figure 3 the activity among nodes 104, 105, 106 in Figure 3 with point 0 in pattern 505 as node 106, point 1 as node 104, and point 2 as node 105. The order of the activity in the directed clique is also specified. For example, in pattern 505, the activity between point 1 and point 2 occurs after the activity between point 0 and point 1.
[0062] In the illustrated implementations, the patterns 500 are all directed cliques or directed simplices. In such patterns, the activity originates from a source node that transmits signals to every other node in the pattern. In pattern 500, such a source node is designated as point 0, and the other nodes are designated as points 1, 2, …. Further, in a directed clique or simplex, one of the nodes acts as a sink and receives signals transmitted from every other node in the pattern. In pattern 500, such a sink node is designated as the highest numbered point in the pattern. For example, in pattern 505, the sink node is designated as point 2. In pattern 510, the sink node is designated as point 3. In pattern 515, the sink node is designated as point 3, and so on. Thus, the activity represented by pattern 500 is ordered in a distinguishable manner.
[0063] Each of the patterns 500 has a different number of points and reflects the ordered activity among a different number of nodes. For example, pattern 505 is a two-dimensional simplex and reflects the activity among three nodes, pattern 510 is a three-dimensional simplex and reflects the activity among four nodes, and so on. As the number of points in the pattern increases, the degree of ordering and complexity of the activity also increases. For example, for a large set of nodes with a certain degree of random activity within a window, some of the activity may by chance match pattern 505. However, the random activity will increasingly less likely match the corresponding patterns in patterns 510, 515, 520, …. The presence of activity matching pattern 530 indicates a relatively higher degree of ordering and complexity in the activity compared to the presence of activity matching pattern 505.
[0064] As previously discussed, in some implementations, windows of different durations can be defined for different determinations of the complexity of the activity. For example, when identifying activity matching pattern 530, a window of longer duration can be used compared to when identifying activity matching pattern 505.
[0065] Figure 6 is an illustration of a pattern 600 of activity that can be recognized and used to identify decision moments in a recurrent artificial neural network. For example, pattern 600 can be recognized at 415 in process 400( Figure 4 ).
[0066] Pattern 600 is an illustration of the activity in a recurrent artificial neural network, like pattern 500. However, pattern 600 deviates from the strict ordering of pattern 500 because pattern 600 is not all directed cliques or directed simplices. In particular, patterns 605, 610 have lower directionality than pattern 515. In fact, pattern 605 completely lacks convergent nodes. However, patterns 605, 610 indicate an ordered activity level that exceeds the level of ordered activity expected by random chance and can be used to determine the complexity of the activity in a recurrent artificial neural network.
[0067] Figure 7 is an illustration of pattern 700 of the activity that can be recognized and used to recognize decision-making moments in a recurrent artificial neural network. For example, pattern 700 can be recognized at 415 in process 400( Figure 4 ).
[0068] Pattern 700 is a group of directed cliques or directed simplices with the same dimension (i.e., having the same number of points), which defines a pattern involving more points than a single clique or simplex and encloses a cavity within the group of directed simplices.
[0069] By way of example, pattern 705 includes six different three-point, two-dimensional patterns 505 that together define a homology class of level 2, while pattern 710 includes eight different three-point, two-dimensional patterns 505 that together define a second homology class of level 2. Each of the three-point, two-dimensional patterns 505 in patterns 705, 710 can be considered to enclose a corresponding cavity. The nth Betti number associated with a directed graph provides a count of such homology classes in a topological representation.
[0070] The activity illustrated by patterns such as pattern 700 illustrates a relatively high level of ordering of the activity in the network that is unlikely to occur by random chance. Pattern 700 can be used to characterize the complexity of this activity.
[0071] In some implementations, only some patterns of the activity are recognized and / or a portion of the recognized patterns of the activity are discarded or otherwise ignored during the recognition of decision-making moments. For example, referring to Figure 5 , the activity commensurate with a five-point, four-dimensional simplex pattern 515 inherently includes the activity commensurate with a four-point, three-dimensional and a three-point, two-dimensional simplex patterns 510, 505. For example, Figure 5 points 0, 2, 3, 4 and points 1, 2, 3, 4 in the four-dimensional simplex pattern 515 of
[0072] As another example, only some patterns of activity need to be recognized. For example, in some implementations, only patterns with odd points (3, 5, 7, ……) or even dimensions (2, 4, 6, ……) are used in the recognition at the decision moment.
[0073] The complexity or degree of ordering in the activity patterns in different windows of the recurrent artificial neural network device can be determined in a variety of different ways. Figure 8 is a schematic illustration of a data table 800 that can be used in such a determination. The data table 800 can be used to determine the complexity of the activity pattern either in isolation or in combination with other activities. For example, the data table 800 can be used at 420 in process 400( Figure 4 ).
[0074] More specifically, table 800 includes a count of the number of occurrences of patterns during window “N”, where the counts of the activities of patterns matching different dimensions are presented in different rows. For example, in the illustrated example, row 805 includes the count of the number of occurrences of activities matching one or more three-point, two-dimensional patterns (i.e., “2032”), and row 810 includes the count of the number of occurrences of activities matching one or more four-point, three-dimensional patterns (i.e., “877”). Since the occurrence of a pattern indicates that the activity has a non-random order, the count also provides a generalized characterization of the overall complexity of the activity pattern. A table similar to table 800 can be formed for each window defined, for example, at 410 in process 400( Figure 4 ).
[0075] Although table 800 includes separate rows and separate entries for each type of activity pattern, this may not necessarily be the case. For example, one of the counts (e.g., the count of simpler patterns) can be omitted from table 800 and from the determination of complexity. As another example, in some implementations, a single row or entry can include the counts of the occurrences of multiple activity patterns.
[0076] Although Figure 8 the counts are presented in table 800, this may not necessarily be the case. For example, the counts can be presented as a vector (e.g., <2032, 877, 133, 66, 48, ……>). Regardless of how the counts are presented, in some implementations, the counts can be expressed in binary and can be compatible with the digital data processing infrastructure.
[0077] In some implementations, the counts of the occurrences of patterns can be weighted or combined to determine the degree of ordering or complexity, for example, at 420 in process 400( Figure 4 ). For example, the Euler characteristic can provide an approximation of the complexity of the activity and is given by the following equation:
[0078] The equation S0 - S1 + S2 - S3 +... Equation 1
[0079] where S n is the number of occurrences of a pattern of n points (i.e., a pattern of dimension n - 1). The pattern can be, for example, a directed clique pattern 500( Figure 5 ).
[0080] As another example of how the number of occurrences of a pattern can be weighted to determine the degree of ordering or complexity, in some implementations, the occurrences of a pattern can be weighted based on the weights of the active links. More specifically, as previously discussed, the strength of the connections between nodes in an artificial neural network can vary, e.g., due to the degree of activity of the connections during training. The occurrence of a pattern of activity along a set of relatively strong links can be weighted differently from the occurrence of the same pattern of activity along a set of relatively weak links. For example, in some implementations, the sum of the weights of the active links can be used to weight the occurrences.
[0081] In some implementations, the Euler characteristic or other measures of complexity can be normalized by the total number of patterns matched within a particular window and / or the total number of patterns that a given network might form considering the structure of the given network. An example of normalizing by the total number of patterns that a network might form is given below in Equation 2 and Equation 3.
[0082] In some implementations, the occurrences of higher-dimensional patterns involving a larger number of nodes can be weighted more heavily compared to the occurrences of lower-dimensional patterns involving a smaller number of nodes. For example, the probability of forming a directed clique decreases rapidly with increasing dimension. In particular, to form an n - clique from n + 1 nodes, (n + 1)n / 2 edges all need to be correctly oriented. This probability can be reflected in the weighting.
[0083] In some implementations, both the dimension and the directionality of a pattern can be used to weight the occurrences of the pattern and determine the complexity of the activity. For example, referring to Figure 6 , based on the difference in the directionality of the five - point, four - dimensional pattern 515 and the five - point, four - dimensional patterns 605, 610, the occurrences of the five - point, four - dimensional pattern 515 can be weighted more heavily compared to the occurrences of the five - point, four - dimensional patterns 605, 610.
[0084] An example of using both the directionality and dimension of a pattern to determine the degree of ordering or complexity of the activity can be given by the following equation:
[0085]
[0086] where Sx active Indicates the number of active occurrences of the pattern of n points, and ERN is calculated for an equivalent random network (i.e., a network with the same number of nodes with random connections). Additionally, SC is given by the following equation:
[0087]
[0088] where S x silent Indicates the number of occurrences of the pattern of n points when the recurrent artificial neural network is silent and can be considered to embody the total number of patterns that the network may form. In Equations 2 and 3, the pattern can be, for example, a directed clique pattern 500( Figure 5 ).
[0089] Figure 9 is a schematic illustration of the determination of a specific time of an activity pattern having distinguishable complexity. The determination illustrated in Figure 9 can be performed in isolation or in combination with other activities. For example, the determination can be performed at 425 in process 400( Figure 4 ).
[0090] Figure 9 Includes graph 905 and graph 910. Graph 905 illustrates the occurrence of the pattern as a function of time along the x-axis. In particular, each occurrence is schematically illustrated as vertical lines 906, 907, 908, 909. The occurrences in each row can be instances of activities that match the corresponding pattern or class of patterns. For example, the occurrences in the top row can be instances of activities that match pattern 505( Figure 5 ), the occurrences in the second row can be instances of activities that match pattern 510( Figure 5 ), the occurrences in the third row can be instances of activities that match pattern 515( Figure 5 ), and so on.
[0091] Graph 905 also includes dashed rectangles 915, 920, 925 that schematically depict different time windows when the activity pattern has distinguishable complexity. As shown, during the windows depicted by dashed rectangles 915, 920, 925, the likelihood that the activity in the recurrent artificial neural network matches a pattern indicating complexity is higher than outside those windows.
[0092] Chart 910 illustrates the complexity associated with these occurrences as a function of time along the x-axis. Chart 910 includes: a first peak 930 of complexity, which coincides with the window depicted by the dashed rectangle 915; and a second peak 935 of complexity, which coincides with the windows depicted by the dashed rectangles 920, 925. As shown, the complexity exemplified by peaks 930, 925 is distinguishable from the complexity that can be considered a baseline level 940 of complexity.
[0093] In some implementations, the time at which the output of the recurrent artificial neural network will be read coincides with the occurrence of an activity pattern having distinguishable complexity. For example, in Figure 9 the illustrative scenario, the output of the recurrent artificial neural network can be read at peaks 930, 925, i.e., during the windows depicted by the dashed rectangles 915, 920, 925.
[0094] When the input is a data stream, the identification of distinguishable levels of complexity in a recurrent artificial neural network is particularly beneficial. Examples of data streams include, for example, video or audio data. Although a data stream has a beginning, it is generally desirable to process information in the data stream that does not have a predefined relationship to the beginning of the data stream. By way of example, a neural network can perform object recognition, such as, for example, identifying a cyclist near a car. Such a neural network should be able to identify the cyclist regardless of when those cyclists appear in the video stream, i.e., regardless of the time since the beginning of the video. Continuing this example, when the data stream is input into an object recognition neural network, any pattern of activity in the neural network will generally exhibit a low or stationary level of complexity. These low or stationary levels of complexity are exhibited regardless of the continuous (or near continuous) input of streaming data into the neural network device. However, when the object of interest appears in the video stream, the complexity of the activity will become distinguishable and indicate the time at which the object is recognized in the video stream. Thus, the specific time of the distinguishable level of complexity of the activity can also serve as a yes / no output regarding whether the data in the data stream meets certain criteria.
[0095] In some implementations, the activity pattern having distinguishable complexity gives not only the specific time at which the output of the recurrent artificial neural network is given, but also the content of the output of the recurrent artificial neural network. In particular, the identity and activity of the nodes participating in the activity commensurate with the activity pattern can be considered the output of the recurrent artificial neural network. Thus, the identified activity pattern can exemplify the result of processing by the neural network, as well as the specific time at which this decision will be read.
[0096] The content of a decision can be expressed in a variety of different forms. For example, in some implementations and as discussed further below, the content of a decision can be expressed as a binary vector or matrix of 1s and 0s. Each digit can indicate, for example, whether a pattern of activity exists for a predefined group of nodes and / or a predefined duration. In such an implementation, the content of the decision is in binary representation and can be compatible with traditional digital data processing infrastructure.
[0097] Figure 10 FIG. 1000 is a flow chart of a process 1000 for encoding a signal using a recurrent artificial neural network based on the characterization of activity in the network. The signal can be encoded in a variety of different scenarios (such as, for example, transmission, encryption, and data storage). Process 1000 can be executed by a system of one or more data processing devices that perform operations according to one or more sets of machine-readable instructions. For example, process 1000 can be executed by the same system of one or more computers that execute software for implementing the recurrent artificial neural network used in process 1000. In some instances, process 1000 can be executed by the same data processing device that executes process 400. In some instances, process 1000 can be executed by, for example, an encoder in a signal transmission system or an encoder in a data storage system.
[0098] At 1005, the system executing process 1000 inputs the signal into the recurrent artificial neural network. In some cases, the input of the signal is a discrete injection event. In other cases, the input signal is streamed into the recurrent artificial neural network.
[0099] At 1010, the system executing process 1000 identifies one or more decision moments in the recurrent artificial neural network. For example, the system can identify one or more decision moments by executing process 400 ( Figure 4 ).
[0100] At 1015, the system executing process 1000 reads the output of the recurrent artificial neural network. As discussed above, in some implementations, the content of the output of the recurrent artificial neural network is the activity in the neural network that matches the pattern used to identify the decision point.
[0101] In some implementations, a separate "reader node" can be added to the neural network to identify the occurrence of a particular pattern of activity at a particular set of nodes and thus to read the output of the recurrent artificial neural network at 1015. The reader node can fire if and only if the activity at the particular set of nodes meets a particular time (and also possibly amplitude) criterion. For example, in order to read pattern 505 at nodes 104, 105, 106 ( Figure 2 , Figure 3 ), Figure 5) can occur by connecting a reader node to nodes 104, 105, 106 (or the link 110 between them). The reader node itself becomes active only when a pattern of activity involving nodes 104, 105, 106 (or their links) occurs.
[0102] The use of such reader nodes eliminates the need to define a time window for the recurrent artificial neural network as a whole. In particular, individual reader nodes can be connected to different nodes and / or multiple nodes (or the links between them). Individual reader nodes can be set to have a customized response (e.g., incorporating different decay times in the firing model) to identify different activity patterns. At 1020, the system executing process 1000 transmits or stores the output of the recurrent artificial neural network. The specific action performed at 1020 can reflect the scenario in which process 1000 is being used. For example, in a scenario where secure or compressed communication is desired, the system executing process 1000 can transmit the output of the recurrent neural network to a receiver that can access the same or a similar recurrent neural network. As another example, in a scenario where secure or compressed data storage is desired, the system executing process 1000 can record the output of the recurrent neural network in one or more machine-readable data storage devices for later access.
[0103] In some implementations, the complete output of the recurrent neural network is not transmitted or stored. For example, in an implementation where the content of the output of the recurrent neural network is the activity in a neural network that matches a pattern indicating the complexity in the activity, only the activity that matches relatively complex or higher-dimensional activity can be transmitted or stored. By way of example, referring to pattern 500( Figure 5 ), in some implementations, only the activities that match patterns 515, 520, 525, and 530 are transmitted or stored, while the activities that match patterns 505, 510 are ignored or discarded. In this way, lossy processing allows reducing the amount of data transmitted or stored at the expense of the integrity of the encoded information.
[0104] Figure 11FIG. 1100 is a flow diagram of a process 1100 for decoding a signal using a recurrent artificial neural network based on the characterization of the activity in the network. The signal can be decoded in a variety of different scenarios such as, for example, signal reception, decryption, and reading data from storage. Process 1100 can be executed by a system of one or more data processing devices that perform operations according to one or more sets of machine-readable instructions. For example, process 1100 can be executed by the same system of one or more computers that execute software for implementing the recurrent artificial neural network used in process 1100. In some instances, process 1100 can be executed by the same data processing devices that execute process 400 and / or process 1000. In some instances, process 1100 can be executed by, for example, a decoder in a signal reception system or a decoder in a data storage system.
[0105] At 1105, the system executing process 1100 receives at least a portion of the output of the recurrent artificial neural network. The particular action performed at 1105 can reflect the scenario in which process 1100 is being used. For example, the system executing process 1000 can receive a transmission signal that includes the output of the recurrent artificial neural network or read a machine-readable data storage device that stores the output of the recurrent artificial neural network.
[0106] At 1110, the system executing process 1100 reconstructs the input to the recurrent artificial neural network from the received output. The reconstruction can be performed in a variety of different ways. For example, in some implementations, a second artificial neural network (recurrent or non-recurrent) can be trained to reconstruct the input to the recurrent neural network from the output received at 1105.
[0107] As another example, in some implementations, a decoder that has been trained using machine learning (including but not limited to deep learning) can be trained to reconstruct the input to the recurrent neural network from the output received at 1105.
[0108] As yet another example, in some implementations, the input to the same recurrent artificial neural network or to a similar recurrent artificial neural network can be iteratively permuted until the output of the recurrent artificial neural network matches the output received at 1105 to some extent.
[0109] In some implementations, process 1100 may include receiving user input specifying the extent to which an input is to be reconstructed, and in response, adjusting the reconstruction accordingly at 1110. For example, the user input may specify that a full reconstruction is not required. In response, the system performing process 1100 adjusts the reconstruction. For example, in an implementation where the content of the output of a recurrent neural network is the activity in a neural network that matches a pattern indicating complexity in an activity, only the output that characterizes the activity matching relatively complex or higher-dimensional activities will be used to reconstruct the input. By way of example, referring to pattern 500( Figure 5 ), in some implementations, only the activities matching patterns 515, 520, 525, and 530 may be used to reconstruct the input, while the activities matching patterns 505, 510 may be ignored or discarded. In this way, lossy reconstruction can be performed in selected cases.
[0110] In some implementations, processes 1000, 1100 may be used for peer-to-peer encrypted communication. In particular, both the transmitter (i.e., encoder) and the receiver (i.e., decoder) may be provided with the same recurrent artificial neural network. There are several ways to customize the shared recurrent artificial neural network to ensure that a third party cannot reverse engineer it and decrypt the signal, including:
[0111] — the structure of the recurrent artificial neural network
[0112] — the functional settings of the recurrent artificial neural network, including node states and edge weights,
[0113] — the size (or dimension) of the pattern, and
[0114] — a subset of the pattern in each dimension.
[0115] These parameters can be considered to together ensure multiple layers of transmission security. Additionally, in some implementations, the decision-making moment time point may be used as the key to decrypt the signal.
[0116] Although processes 1000, 1100 are presented in terms of encoding and decoding a single recurrent artificial neural network, processes 1000, 1100 may also be applied in systems and processes that rely on multiple recurrent artificial neural networks. These recurrent artificial neural networks may operate in parallel or in series.
[0117] As an example of serial operation, the output of the first recurrent artificial neural network can be used as the input to the second recurrent artificial neural network. The resulting output of the second recurrent artificial neural network is a twice-encoded (or twice-encrypted) version of the input into the first recurrent artificial neural network. Such a serial arrangement of recurrent artificial neural networks can be useful in situations where different parties have different levels of access to information. For example, in a medical record system, where patient identity information may be inaccessible to parties that will use and can access the remainder of the medical record.
[0118] As an example of parallel operation, the same information can be input into multiple different recurrent artificial neural networks. The different outputs of these neural networks can be used, for example, to ensure that the input can be reconstructed with high fidelity.
[0119] Although multiple implementations have been described, various modifications can be made. For example, although an application typically implies that the activity in a recurrent artificial neural network should match a pattern indicating an order, this may not be the case. Instead, in some implementations, the activity in a recurrent artificial neural network can be commensurate with a pattern without necessarily exhibiting activity that matches the pattern. For example, an increase in the likelihood that a recurrent neural network will exhibit activity that matches a pattern can be considered a non-random ordering of the activity.
[0120] As yet another example, in some implementations, different groups of patterns can be customized for use in characterizing the activity in different recurrent artificial neural networks. The patterns can be customized, for example, according to the effectiveness of the pattern in characterizing the activity of different recurrent artificial neural networks. The effectiveness can be quantified, for example, based on the magnitude of a table or vector representing the occurrence counts of different patterns.
[0121] As yet another example, in some implementations, the patterns used to characterize the activity in a recurrent artificial neural network can take into account the strength of the connections between nodes. In other words, the patterns described previously herein treat all signal transmission activities between two nodes in a binary manner (i.e., the activity is either present or absent). This may not be the case. Instead, in some implementations, being commensurate with a pattern may require treating the activity of connections having a certain level or strength as indicating an ordered complexity in the activity of the recurrent artificial neural network.
[0122] As yet another example, the content of the output of a recurrent artificial neural network can include activity patterns that occur outside of a time window within which the activity in the neural network has a distinguishable level of complexity. For example, the output of a recurrent artificial neural network read at 1015 and transmitted or stored at 1020 ( Figure 10 ) can include, for example, an indication of what occurred in graph 905 ( Figure 9) the information encoding the active patterns outside the dashed rectangles 915, 920, 925. By way of example, the output of a recurrent artificial neural network can characterize only the highest-dimensional patterns of the activity, regardless of when those patterns of the activity occur. As another example, the output of a recurrent artificial neural network can characterize only the patterns of the activity surrounding a cavity, regardless of when those patterns of the activity occur.
[0123] Figure 12 , Figure 13 and Figure 14 are schematic illustrations of the binary form or representation 1200 of a topology (such as, for example, a pattern of activity in a neural network). Figure 12 , Figure 13 and Figure 14 The topologies illustrated in all include the same information, that is, an indication of the presence or absence of features in the diagram. The features can be, for example, the activity in a neural network device. In some implementations, the activity is identified based on or during a time period within which the activity in the neural network has a complexity distinguishable from other activities in response to an input.
[0124] As shown, the binary representation 1200 includes bits 1205, 1207, 1211, 1293, 1294, 1297 and additional, any number of bits (represented by the ellipsis “…”). For pedagogical purposes, bits 1205, 1207, 1211, 1293, 1294, 1297… are illustrated as discrete rectangular shapes that are filled or unfilled to indicate the binary value of the bit. In the schematic illustration, the representation 1200 appears on the surface to be a one-dimensional vector of bits ( Figure 12 , Figure 13 ) or a two-dimensional matrix of bits ( Figure 14 ). However, the representation 1200 differs from a vector, matrix, or other ordered set of bits in that the same information can be encoded regardless of the order of the bits—that is, regardless of the position of the individual bits within the set.
[0125] For example, in some implementations, each individual bit 1205, 1207, 1211, 1293, 1294, 1297… can represent the presence or absence of a topological feature—regardless of the position of the feature in the diagram. By way of example, referring to Figure 2 , a bit such as 1207 can indicate a pattern 505 ( Figure 5) the presence of corresponding topological features, regardless of whether the activity occurs between nodes 104, 105, 101 or between nodes 105, 101, 102. Thus, although each individual bit 1205, 1207, 1211, 1293, 1294, 1297... can be associated with a specific feature, the position of the feature in the diagram need not be encoded, e.g., by the corresponding position of the bit in the representation 1200. In other words, in some implementations, the representation 1200 can provide only an isomorphic topological reconstruction of the diagram.
[0126] In addition, in other implementations, it may be the case that the positions of the individual bits 1205, 1207, 1211, 1293, 1294, 1297... do encode information such as, for example, the position of the feature in the diagram. In these implementations, the source graph can be reconstructed using the representation 1200. However, such encoding does not necessarily exist.
[0127] Given that bits are able to represent the presence or absence of a topological feature regardless of the position of the topological feature in the diagram, in Figure 1 , at the start of the representation 1200, bit 1205 appears before bit 1207, and bit 1207 appears before bit 1211. Conversely, in Figure 2 and Figure 3 , the order of bits 1205, 1207, and 1211 in the representation 1200 - as well as the positions of bits 1205, 1207, and 1211 relative to the other bits in the representation 1200 - has changed. However, the binary representation 1200 remains the same - and the set of rules or algorithm defining the process for encoding the information in the binary representation 1200 also remains unchanged. As long as the correspondence between bits and features is known, the position of the bits in the representation 1200 is irrelevant.
[0128] More specifically, each of the bits 1205, 1207, 1211, 1293, 1294, 1297... individually represents the presence or absence of a feature in a graph. A graph is a set of nodes and a set of edges between these nodes. The nodes can correspond to objects. Examples of objects can include, for example, artificial neurons in a neural network, individuals in a social network, etc. The edges can correspond to some relationships between the objects. Examples of relationships include, for example, structural connections or the activity along such connections. In the scenario of a neural network, artificial neurons can be associated through structural connections between neurons or through the transmission of information along the structural connections. In the scenario of a social network, individuals can be connected through "friends" or other relationships or through the transmission of information (e.g., posts) along such connections. Thus, an edge can characterize a relatively long-term existing structural feature of a set of nodes or a relatively short-term occurring activity within a defined time frame. In addition, an edge can be directed or bidirectional. A directed edge indicates the directionality of the relationship between objects. For example, the transmission of information from a first neuron to a second neuron can be represented by a directed edge indicating the direction of the transmission. As another example, in a social network, a relationship connection can indicate that a second user will receive information from a first user, rather than the first user receiving information from the second user. In topological terms, a graph can be expressed as a set of unit intervals [0, 1], where 0 and 1 are identified with the corresponding nodes connected by edges.
[0129] The features whose presence or absence is indicated by bits 1205, 1207, 1211, 1293, 1294, 1297 can be, for example, a node, a set of nodes, a set of nodes within a set of multiple sets of nodes, a set of edges, a set of edges within a set of multiple sets of edges, and / or additional hierarchically more complex features (e.g., a set of sets of nodes within a set of multiple sets of nodes). Bits 1205, 1207, 1211, 1293, 1294, 1297 generally represent the presence or absence of features at different hierarchical levels. For example, bit 1205 can represent the presence or absence of a single node, while bit 1205 can represent the presence or absence of a set of nodes.
[0130] In some implementations, bits 1205, 1207, 1211, 1293, 1294, 1297 can represent features of some properties in a graph with a threshold level. For example, bits 1205, 1207, 1211, 1293, 1294, 1297 can not only represent the presence of activity in a set of edges, but also represent weighting this activity above or below the threshold level. The weight can, for example, embody the training of a neural network device for a specific purpose or can be an inherent feature of the edge.
[0131] The above Figure 5 、 Figure 6 and Figure 8Illustrates features whose presence or absence can be represented by bits 1205, 1207, 1211, 1293, 1294, 1297, ….
[0132] The directed simplices in sets 500, 600, 700 view a functional diagram or a structural diagram as a topological space with nodes as points. Structures or activities involving one or more nodes and links commensurate with the simplices in sets 500, 600, 700 can be represented by bits, regardless of the identity of the specific nodes and / or links participating in the activity.
[0133] In some implementations, only some patterns of identifying a structure or activity and / or discarding or otherwise ignoring a portion of the pattern of the identified structure or activity. For example, referring to Figure 5 , a structure or activity commensurate with the five-point, four-dimensional simplex pattern 515 inherently includes structures or activities commensurate with the four-point, three-dimensional and three-point, two-dimensional simplex patterns 510, 505. For example, Figure 5 points 0, 2, 3, 4 and points 1, 2, 3, 4 in the four-dimensional simplex pattern 515 of
[0134] are both commensurate with the three-dimensional simplex pattern 510. In some implementations, simplex patterns containing fewer points - and thus having a lower dimension - can be discarded or otherwise ignored.
[0135] Returning to Figure 12 , Figure 13 , Figure 14 , the features whose presence or absence is represented by bits 1205, 1207, 1211, 1293, 1294, 1297, … may not be independent of each other. By way of explanation, if bits 1205, 1207, 1211, 1293, 1294, 1297 represent the presence or absence of zero-dimensional simplices - each of which reflects the presence or activity of a single node, then bits 1205, 1207, 1211, 1293, 1294, 1297 are independent of each other. However, if bits 1205, 1207, 1211, 1293, 1294, 1297 represent the presence or absence of higher-dimensional simplices - each of which reflects the presence or activity of multiple nodes, then the information encoded by the presence or absence of each individual feature may not be independent of the presence or absence of other features.
[0136] Figure 15Schematically illustrates an example of how the presence or absence of features corresponding to different bits are not independent of each other. In particular, it illustrates a sub-diagram 1500 including four nodes 1505, 1510, 1515, 1520 and six directed edges 1525, 1530, 1535, 1540, 1545, 1550. In particular, edge 1525 points from node 1525 to node 1510, edge 1530 points from node 1515 to node 1505, edge 1535 points from node 1520 to node 1505, edge 1540 points from node 1520 to node 1510, edge 1545 points from node 1515 to node 1510, and edge 1550 points from node 1515 to node 1520.
[0137] A single bit in representation 1200 (e.g., Figure 12 , Figure 13 , Figure 14 the filled bit 1207 in ) can indicate the presence of a directed 3-simplex. For example, such a bit can indicate the presence of a 3-simplex formed by nodes 1505, 1510, 1515, 1520 and edges 1525, 1530, 1535, 1540, 1545, 1550. A second bit in representation 1200 (e.g., Figure 12 , Figure 13 , Figure 14 the filled bit 1293 in ) can indicate the presence of a directed 2-simplex. For example, such a bit can indicate the presence of a 2-simplex formed by nodes 1515, 1505, 1510 and edges 1525, 1530, 1545. In this simple example, the information encoded by bit 1293 is completely redundant with the information encoded by bit 1207.
[0138] Note that the information encoded by bit 1293 can also be redundant with the information encoded by yet another bit. For example, the information encoded by bit 1293 will be redundant with both a third bit and a fourth bit indicating the presence of additional directed 2-simplices. Examples of these simplices are formed by nodes 1515, 1520, 1510 and edges 1540, 1545, 1550 and nodes 1520, 1505, 1510 and edges 1525, 1535, 1540.
[0139] Figure 16Schematically illustrates another example of how the presence or absence of features corresponding to different bits are not independent of each other. In particular, sub-diagram 1600 including four nodes 1605, 1610, 1615, 1620 and five directed edges 1625, 1630, 1635, 1640, 1645 is illustrated. Nodes 1505, 1510, 1515, 1520 and edges 1625, 1630, 1635, 1640, 1645 generally correspond to nodes 1505, 1510, 1515, 1520 and edges 1525, 1530, 1535, 1540, 1545 in sub-diagram 1500( Figure 15 ) However, in contrast to sub-diagram 1500 in which nodes 1515, 1520 are connected by edge 1550, nodes 1615, 1620 are not connected by an edge.
[0140] A single bit in representation 1200 (e.g., Figure 12 , Figure 13 , Figure 14 the unfilled bit 1205 in) can indicate the absence of a directed three-dimensional simplex (such as, for example, the directed three-dimensional simplex including nodes 1605, 1610, 1615, 1620). The second bit in representation 1200 (e.g., Figure 12 , Figure 13 , Figure 14 the filled bit 1293 in) can indicate the presence of a two-dimensional simplex. An exemplary directed two-dimensional simplex is formed by nodes 1615, 1605, 1610 and edges 1625, 1630, 1645. This combination of the filled bit 1293 and the unfilled bit 1205 provides information indicating the presence or absence of other features (and the states of other bits) that may be present or absent in representation 1200. In particular, the combination of the absence of the directed three-dimensional simplex and the presence of the directed two-dimensional simplex indicates that at least one edge is absent from:
[0141] a) a possible directed two-dimensional simplex formed by nodes 1615, 1620, 1610 or
[0142] b) a possible directed two-dimensional simplex formed by nodes 1620, 1605, 1610.
[0143] Thus, the state of the bits indicating the presence or absence of any of these possible simplices is not independent of the states of bits 1205, 1293.
[0144] Although these examples have been discussed in terms of features having different numbers of nodes and hierarchical relationships, this need not be the case. For example, it is possible to have a representation 1200 including only a set of bits corresponding to, for example, the presence or absence of three-dimensional simplices.
[0145] The use of a single bit to indicate the presence or absence of a feature in a graph yields certain properties. For example, the encoding of the information is error tolerant and provides "graceful degradation" of the encoded information. In particular, the loss of a particular bit (or group of bits) may increase the uncertainty about the presence or absence of a feature. However, the likelihood of the presence or absence of a feature can still be assessed based on other bits indicating the presence or absence of neighboring features.
[0146] Likewise, as the number of bits increases, the certainty regarding the presence or absence of a feature increases.
[0147] As another example, as discussed above, the ordering or arrangement of the bits is irrelevant to the isomorphic reconstruction of the graph represented by the bits. All that is required is a known correspondence between the bits and specific nodes / structures in the graph.
[0148] In some implementations, the pattern of activity in the neural network can be represented in 1200 ( Figure 12 , Figure 13 and Figure 14 ). Typically, the pattern of activity in a neural network is the result of many features of the neural network, such as, for example, the structural connections between nodes of the neural network, the weights between nodes, and a large number of possible other parameters. For example, in some implementations, the neural network may have been trained prior to encoding the pattern of activity in representation 1200.
[0149] However, regardless of whether the neural network is untrained or trained, for a given input, the response pattern of activity can be considered a "representation" or "abstract" of that input in the neural network. Thus, while representation 1200 may appear to be a straightforward-appearing collection of (in some cases, binary) numbers, each of the numbers may encode a relationship or correspondence between a particular input and the associated activity in the neural network.
[0150] Figure 17 , Figure 18 , Figure 19 , Figure 20It is a schematic illustration of the occurrence of a topology in the activities in a neural network being used in four different classification systems 1700, 1800, 1900, 2000. Each of the classification systems 1700, 1800 classifies a representation of a pattern of activities in the neural network as part of an input classification. Each of the classification systems 1900, 2000 approximates a representation of a pattern of activities in the neural network as part of an input classification. In the classification systems 1700, 1800, the pattern of activities being represented occurs in a source neural network device 1705 that is part of the classification systems 1700, 1800 and is read from the source neural network device 1705. In contrast, in the classification systems 1900, 2000, the pattern of activities being approximately represented occurs in a source neural network device that is not part of the classification systems 1700, 1800. However, the approximation of the representation of those patterns of activities is read from an approximator 1905 that is part of the classification systems 1900, 2000.
[0151] More specifically, turning to Figure 17 , the classification system 1700 includes a source neural network 1705 and a linear classifier 1710. The source neural network 1705 is a neural network device that is configured to receive inputs and present a representation of the occurrence of a topology in the activities in the source neural network 1705. In the illustrated implementation, the source neural network 1705 includes an input layer 1715 that receives inputs. However, this need not be the case. For example, in some implementations, some or all of the inputs can be injected into different layers and / or edges or nodes throughout the source neural network 1705.
[0152] The source neural network 1705 can be any of a variety of different types of neural networks. Generally, the source neural network 1705 is a recurrent neural network, such as, for example, a recurrent neural network that mimics a biological system. In some cases, the source neural network 1705 can simulate the morphology, chemistry, and other characteristics of a biological system to some extent. Generally, the source neural network 1705 is implemented on one or more computing devices (e.g., supercomputers) having a relatively high level of computing performance. In such a case, the classification system 1700 will generally be a distributed system in which a remote classifier 1710 communicates with the source neural network 1705, e.g., via a data communication network.
[0153] In some implementations, the source neural network 1705 can be untrained and the activities being represented can be the intrinsic activities of the source neural network 1705. In other implementations, the source neural network 1705 can be trained and the activities being represented can embody this training.
[0154] The representation read from the source neural network 1705 can be, such as representation 1200(Figure 12 , Figure 13 , Figure 14 ) representation. The representation can be read from the source neural network 1705 in a variety of ways. For example, in the illustrated example, the source neural network 1705 includes a "reader node" that reads patterns of activity among other nodes in the source neural network 1705. In other implementations, the activity in the source neural network 1705 is read by a data processing component that is programmed to monitor relatively highly ordered patterns of activity in the source neural network 1705. In other implementations, the source neural network 1705 can include an output layer, and for example, when the source neural network 1705 is implemented as a feedforward neural network, the representation 1200 can be read from that output layer.
[0155] The linear classifier 1710 is a device that classifies an object - that is, a representation of a pattern of activity in the source neural network 1705 - based on a linear combination of the features of the object. The linear classifier 1710 includes an input 1720 and an output 1725. The input 1720 is coupled to receive a representation of a pattern of activity in the source neural network 1705. In other words, the representation of a pattern of activity in the source neural network 1705 is a feature vector of the features of the input into the source neural network 1705, and the features are used by the linear classifier 1710 to classify the input. The linear classifier 1710 can receive the representation of the pattern of activity in the source neural network 1705 in a variety of ways. For example, the representation of the pattern of activity can be received as a discrete event or as a continuous stream over a real-time or non-real-time communication channel.
[0156] The output 1725 is coupled to output the classification result from the linear classifier 1710. In the illustrated implementation, the output 1725 is schematically illustrated as a parallel port with multiple channels. This need not be the case. For example, the output 1725 can output the classification result through a serial port or a port with combined parallel and serial capabilities.
[0157] In some implementations, the linear classifier 1710 can be implemented on one or more computing devices with relatively limited computing performance. For example, the linear classifier 1710 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer.
[0158] In Figure 18In this case, classification system 1800 includes source neural network 1705 and neural network classifier 1810. Neural network classifier 1810 is a neural network device that classifies an object - that is, a representation of the pattern of activity in source neural network 1705 - based on a non-linear combination of the features of the object. In the illustrated implementation, neural network classifier 1810 is a feed-forward network that includes input layer 1820 and output layer 1825. Like linear classifier 1710, neural network classifier 1810 can receive a representation of the pattern of activity in source neural network 1705 in a variety of ways. For example, the representation of the pattern of activity can be received as a discrete event or as a continuous stream over a real-time or non-real-time communication channel.
[0159] In some implementations, neural network classifier 1810 can perform inference on one or more computing devices with relatively limited computational performance. For example, neural network classifier 1810 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer, e.g., in the neural processing unit of such a device. Like classification system 1700, classification system 1800 will typically be a distributed system in which remote neural network classifier 1810 communicates with source neural network 1705, e.g., via a data communication network.
[0160] In some implementations, neural network classifier 1810 can be, for example, a deep neural network, such as a convolutional neural network that includes convolutional layers, pooling layers, and fully connected layers. The convolutional layers can generate feature maps, e.g., using linear convolutional filters and / or non-linear activation functions. The pooling layers reduce the number of parameters and control overfitting. The computations performed by the different layers in image classifier 1820 can be defined in different ways in different implementations of image classifier 1820.
[0161] In Figure 19 this case, classification system 1900 includes source approximator 1905 and linear classifier 1710. As discussed further below, source approximator 1905 is a relatively simple neural network that is trained to - at input layer 1915 or elsewhere - receive an input and output a vector that approximates a representation of the topology that occurs in the pattern of activity in a relatively more complex neural network. For example, source approximator 1905 can be trained to approximate a recurrent source neural network, such as, e.g., a recurrent neural network that mimics a biological system and includes aspects of the morphology, chemistry, and other characteristics of the biological system. In the illustrated implementation, source approximator 1905 includes input layer 1915 and output layer 1920. Input layer 1915 is coupled to receive input data. Output layer 1920 is coupled to output an approximation of the representation of the activity within the neural network device for receipt by input 1720 of the linear classifier. For example, output layer 1920 can output representation 1200(Figure 12 , Figure 13 , Figure 14 ) is approximately 1200'. Additionally, Figure 17 , Figure 18 The representation 1200 schematically illustrated in Figure 19 , Figure 20 and the approximation 1200' of the representation 1200 schematically illustrated in
[0162] are the same. This is merely for convenience. Generally, the approximation 1200' will differ from the representation 1200 in at least some respects. Despite these differences, the linear classifier 1710 can still classify the approximation 1200'.
[0163] In Figure 20 , the classification system 2000 includes a source approximator 1905 and a neural network classifier 1810. The output layer 1920 of the source approximator 1905 is coupled to output an approximation 1200' of the representation of the activities within the neural network device for receipt by the input 1820 of the neural network classifier 1810. Despite any differences between the approximation 1200' and the representation 1200, the neural network classifier 1810 can still classify the approximation 1200'. Generally and like the classification system 1900, the classification system 1900 will typically be housed within a single housing, e.g., where the source approximator 1905 and the neural network classifier 1810 are implemented on the same data processing device or on data processing devices coupled by a hardwired connection.
[0164] Figure 21is a schematic illustration of an edge device 2100 that includes a local artificial neural network that can be trained using an occurrence representation corresponding to the topology of the activity in a source neural network. In this scenario, the local artificial neural network can be, for example, an artificial neural network that executes entirely on one or more local processors that do not require a communication network to exchange data. Typically, the local processors will be connected by hardwiring. In some instances, the local processors can be housed within a single housing, such as a single personal computer or a single handheld, mobile device. In some instances, the local processors can be controlled and accessed by a single individual or a limited number of individuals. In fact, by training a simpler and / or less highly trained but more unique second neural network using an occurrence representation of the topology in a more complex source neural network (e.g., using supervised learning or reinforcement learning techniques), an individual with limited computing resources and a limited number of training samples can train a neural network as needed. The storage requirements and computational complexity during training are reduced, and resources such as battery life are conserved.
[0165] In the illustrated implementation, the edge device 2100 is schematically illustrated as a security camera device that includes an optical imaging system 2110, image processing electronics 2115, a source approximator 2120, a representation classifier 2125, and a communication controller and interface 2130.
[0166] The optical imaging system 2110 can include, for example, one or more lenses (or even a pinhole) and a CCD device. The image processing electronics 2115 can read the output of the optical imaging system 2110 and can generally perform basic image processing functions. The communication controller and interface 2130 is a device configured to control the flow of information to and from the device 2100. Among the operations that the communication controller and interface 2130 can perform, as discussed further below, are transmitting an image of interest to other devices and receiving training information from other devices. Thus, the communication controller and interface 2130 can include both a data transmitter and a receiver that can communicate through, for example, a data port 2135. The data port 2135 can be a wired port, a wireless port, an optical port, etc.
[0167] The source approximator 2120 is a relatively simple neural network that is trained to output a vector that approximates a representation of the topology that occurs in the pattern of activity in a relatively more complex neural network. For example, the source approximator 2120 can be trained to approximate a recurrent source neural network, such as, for example, a recurrent neural network that mimics a biological system and includes aspects of the morphology, chemistry, and other characteristics of the biological system.
[0168] Indicates that classifier 2125 is a linear classifier or a neural network classifier, which is coupled to receive an approximation of the representation of the pattern of activities in the source neural network from source approximator 2120 and output a classification result. The representation classifier 2125 can be, for example, a deep neural network, such as a convolutional neural network including convolutional layers, pooling layers, and fully connected layers. The convolutional layer can generate feature maps, for example, using linear convolutional filters and / or non-linear activation functions. The pooling layer reduces the number of parameters and controls overfitting. The computations performed by the different layers in the representation classifier 2125 can be defined in different ways in different implementations of the representation classifier 2125.
[0169] In some implementations, in operation, the optical imaging system 2110 can generate a raw digital image. The image processing electronic device 2115 can read the raw image and generally perform at least some basic image processing functions. The source approximator 2120 can receive the image from the image processing electronic device 2115 and perform an inference operation to output a vector that approximates the representation of the topology that appears in the pattern of activities in a relatively complex neural network. This approximate vector is input into the representation classifier 2125, which determines whether the approximate vector satisfies one or more sets of classification criteria. Examples include face recognition and other machine vision operations. In the case where the representation classifier 2125 determines that the approximate vector satisfies a set of classification criteria, the representation classifier 2125 can instruct the communication controller and interface 2130 to transmit information about the image. For example, the communication controller and interface 2130 can transmit the image itself, the classification, and / or other information about the image.
[0170] Sometimes, it may be desirable to change the classification process. In these cases, the communication controller and interface 2130 can receive a training set. In some implementations, the training set can include raw or processed image data and the representation of the topology that appears in the pattern of activities in a relatively complex neural network. Such a training set can be used to retrain the source approximator 2120, for example, using supervised learning or reinforcement learning techniques. In particular, the representation is used as the target answer vector, and represents the desired result of the source approximator 2120 processing the raw or processed image data.
[0171] In other implementations, the training set can include the representation of the topology that appears in the pattern of activities in a relatively complex neural network and the desired classification of those representations of the topology. Such a training set can be used to retrain the neural network representation classifier 2125, for example, using supervised learning or reinforcement learning techniques. In particular, the desired classification is used as the target answer vector, and represents the desired result of the representation classifier 2125 processing the representation of the topology.
[0172] Regardless of whether the source approximator 2120 or the representation classifier 2125 is retrained, the inference operations at device 2100 can be easily adapted to changing situations and objectives without large training data sets and time - intensive and computationally intensive iterative training.
[0173] Figure 22 is a schematic illustration of a second - edge device 2200 that includes a local artificial neural network that can be trained using occurrences of representations corresponding to the topology of the activity in the source neural network. In the illustrated implementation, the second - edge device 2200 is schematically illustrated as a mobile computing device such as a smart phone or a tablet computer. Device 2200 includes an optical imaging system (e.g., on the back of device 2200, not shown), image - processing electronics 2215, a representation classifier 2225, a communication controller and interface 2230, and a data port 2235. These components can have features and perform actions corresponding to the actions of the optical imaging system 2110, image - processing electronics 2115, representation classifier 2125, communication controller and interface 2130, and data port 2135 in device 2100 ( Figure 21 )).
[0174] The illustrated implementation of device 2200 additionally includes one or more additional sensors 2240 and a multi - input source approximator 2245. The sensors 2240 can sense one of a plurality of characteristics of the environment around device 2200 or of device 2200 itself. For example, in some implementations, the sensor 2240 can be an accelerometer that senses the acceleration that device 2200 is subjected to. As another example, in some implementations, the sensor 2240 can be an acoustic sensor, such as a microphone, that senses noise in the environment of device 2200. Yet another example of the sensor 2240 includes chemical sensors (e.g., "artificial nose", etc.), humidity sensors, radiation sensors, etc. In some cases, the sensor 2240 is coupled to processing electronics that can read the output of the sensor 2240 (or other information, such as, for example, a contact list or a map) and perform basic processing functions. Thus, different implementations of the sensor 2240 can have different "modalities" because the physical parameters sensed physically change from sensor to sensor.
[0175] The multi - input source approximator 2245 is a relatively simple neural network that is trained to output a vector that approximates a representation of the topology that occurs in the pattern of activity in a relatively more complex neural network. For example, the multi - input source approximator 2245 can be trained to approximate a recurrent source neural network, such as, for example, a recurrent neural network that mimics biological systems and includes aspects of the morphology, chemistry, and other characteristics of biological systems.
[0176] Unlike the source approximator 2120, the multi-input source approximator 2245 is coupled to receive raw or processed sensor data from multiple sensors and returns an approximation of a representation of a topology that occurs in the pattern of activity in a relatively complex neural network based on that data. For example, the multi-input source approximator 2245 can receive processed image data from the image processing electronics 2215 and, for example, acoustic, acceleration, chemical, or other data from one or more sensors 2240. The multi-input source approximator 2245 can be, for example, a deep neural network, such as a convolutional neural network including convolutional layers, pooling layers, and fully connected layers. The computations performed by the different layers in the multi-input source approximator 2245 can be dedicated to a single type of sensor data or multiple forms of sensor data.
[0177] Regardless of the specific organization of the multi-input source approximator 2245, the multi-input source approximator 2245 will return an approximation based on raw or processed sensor data from multiple sensors. For example, processed image data from the image processing electronics 2215 and acoustic data from a microphone sensor 2240 can be used by the multi-input source approximator 2245 to approximate a representation of a topology that will occur in the pattern of activity in a relatively complex neural network receiving the same data.
[0178] Sometimes, it may be desirable to change the classification process at the device 2200. In these cases, the communication controller and interface 2230 can receive a training set. In some implementations, the training set can include raw or processed images, sounds, chemicals, or other data and a representation of a topology that occurs in the pattern of activity in a relatively complex neural network. Such a training set can be used to retrain the multi-input source approximator 2245, for example, using supervised learning or reinforcement learning techniques. In particular, the representation is used as a target answer vector and represents the desired result of the multi-input source approximator 2245 processing the raw or processed image or sensor data.
[0179] In other implementations, the training set can include a representation of a topology that occurs in the pattern of activity in a relatively complex neural network and the desired classification of those representations of the topology. Such a training set can be used to retrain the neural network representation classifier 2225, for example, using supervised learning or reinforcement learning techniques. In particular, the desired classification is used as a target answer vector and represents the desired result of the representation classifier 2225 processing the representation of the topology.
[0180] Whether the multi-input source approximator 2245 or the representation classifier 2225 is retrained, the inference operation at device 2200 can easily adapt to changing situations and goals without a large training dataset and time-intensive and computationally intensive iterative training.
[0181] Figure 23 FIG. 23 is a schematic illustration of a system 2300 in which a local neural network can be trained using an occurrence representation corresponding to the topology of the activity in a source neural network. The target neural network is implemented on a relatively simple and less expensive data processing system, while the source neural network can be implemented on a relatively complex and more expensive data processing system.
[0182] System 2300 includes a variety of devices 2305 with local neural networks, telephone base stations 2310, wireless access points 2315, server systems 2320, and one or more data communication networks 2325.
[0183] The local neural network device 2305 is a device configured to process data using a computationally - lower - intensive target neural network. As illustrated, the local neural network device 2305 can be implemented as a mobile computing device, a camera, a vehicle, or any one of a large number of other appliances, fixtures, and mobile components, as well as devices of different brands and models within each category. Different local neural network devices 2305 can belong to different owners. In some implementations, access to the data processing functions of the local neural network device 2305 will generally be restricted to these owners and / or the owners' designees.
[0184] Each local neural network device 2305 can include one or more source approximators that are trained to output a vector that approximates the representation of the topology that occurs in the pattern of activity in a relatively complex neural network. For example, the relatively complex neural network can be a recurrent source neural network, such as, for example, a recurrent neural network that mimics a biological system and includes aspects of the morphology, chemistry, and other characteristics of the biological system.
[0185] In some implementations, in addition to processing data using the source approximator, the local neural network device 2305 can also be programmed to use a representation of a topology that occurs in the pattern of activity in a relatively complex neural network as a target answer vector to retrain the source approximator. For example, the local neural network device 2305 can be programmed to perform one or more iterative training techniques (e.g., gradient descent or stochastic gradient descent). In other implementations, the source approximator in the local neural network device 2305 is trainable by, for example, a dedicated training system or by a training system installed on a personal computer that can interact with the local neural network device 2305 to train the source approximator.
[0186] Each local neural network device 2305 includes one or more wireless or wired data communication components. In the illustrated implementation, each local neural network device 2305 includes at least one wireless data communication component, such as a mobile phone transceiver, a wireless transceiver, or both. The mobile phone transceiver is capable of exchanging data with a phone base station 2310. The wireless transceiver is capable of exchanging data with a wireless access point 2315. Each local neural network device 2305 may also be capable of exchanging data with peer mobile computing devices.
[0187] The phone base station 2310 and the wireless access point 2315 are connected for data communication with one or more data communication networks 2325 and can exchange information with a server system 2320 via the network. Thus, the local neural network device 2305 is generally also in data communication with the server system 2320. However, this may not necessarily be the case. For example, in implementations where the local neural network device 2305 is trained by other data processing devices, the local neural network device 2305 only needs to communicate with these other data processing devices at least once.
[0188] The server system 2320 is a system of one or more data processing devices programmed to perform data processing activities according to one or more machine-readable instruction sets. The activities can include providing a training set to a training system for the mobile computing device 2305. As discussed above, the training system can be within the mobile local neural network device 2305 itself or on one or more other data processing devices. The training set can include representations of the occurrences of topologies corresponding to the activities in the source neural network and corresponding input data.
[0189] In some implementations, the server system 2320 also includes the source neural network. However, this may not necessarily be the case, and the server system 2320 can receive a training set from another system of data processing devices implementing the source neural network.
[0190] In operation, after the server system 2320 receives a training set (from a source neural network found either at the server system 2320 itself or elsewhere), the server system 2320 can provide the training set to a trainer of the training mobile computing device 2305. The training set can be used to train a source approximator in the target local neural network device 2305 such that the target neural network approximates the operation of the source neural network.
[0191] Figure 24 , Figure 25 , Figure 26 , Figure 27 are schematic illustrations of the occurrence of representations of topologies in the activity of neural networks used in four different systems 2400, 2500, 2600, 2700. The systems 2400, 2500, 2600, 2700 can be configured to perform any one of a number of different operations. For example, the systems 2400, 2500, 2600, 2700 can perform object localization operations, object detection operations, object segmentation operations, object detection operations, prediction operations, action selection operations, etc.
[0192] An object localization operation locates an object within an image. For example, a bounding box can be constructed around the object. In some cases, object localization can be combined with object recognition, in which the located object is labeled with an appropriate designation.
[0193] An object detection operation classifies image pixels as belonging to a particular class (e.g., belonging to an object of interest) or not belonging to a particular class. Typically, object detection is performed by grouping pixels and forming a bounding box around the pixel group. The bounding box should fit tightly around the object.
[0194] Object segmentation typically assigns a class label to each image pixel. Thus, object segmentation is performed on a pixel-by-pixel basis and typically requires assigning only a single label to each pixel, rather than a bounding box.
[0195] A prediction operation seeks to draw conclusions outside the scope of observed data. Although a prediction operation can seek to predict future occurrences (e.g., based on information about past and current states), a prediction operation can also seek to draw conclusions about past and current states based on incomplete information about those states.
[0196] An action selection operation seeks to select an action based on a set of conditions. Action selection operations have traditionally been decomposed into different methods, such as symbol-based systems (classical planning), distributed solutions, and reactive or dynamic programming.
[0197] Classification systems 2400, 2500 each perform desired operations on the representation of patterns of activity in a neural network. Systems 2600, 2700 each perform desired operations on an approximation of the representation of patterns of activity in a neural network. In systems 2400, 2500, the patterns of activity being represented occur in source neural network device 1705 that is part of systems 2400, 2500 and are read from that source neural network device 1705. In contrast, in systems 2400, 2500, the patterns of activity being approximately represented occur in a source neural network device that is not part of systems 2400, 2500. However, the approximation of the representation of those patterns of activity is read from approximator 1905 that is part of systems 2400, 2500.
[0198] More specifically, turning to Figure 24 , system 2400 includes source neural network 1705 and linear processor 2410. Linear processor 2410 is a device that performs operations based on a linear combination of features of a representation (or an approximation of such a representation) of patterns of activity in a neural network. The operations can be, for example, object localization operations, object detection operations, object segmentation operations, prediction operations, action selection operations, etc.
[0199] Linear processor 2410 includes input 2420 and output 2425. Input 2420 is coupled to receive the representation of patterns of activity in source neural network 1705. Linear processor 2410 can receive the representation of patterns of activity in source neural network 1705 in a variety of ways. For example, the representation of patterns of activity can be received as discrete events or as a continuous stream over a real-time or non-real-time communication channel. Output 2525 is coupled to output the processing result from linear processor 2410. In some implementations, linear processor 2410 can be implemented on one or more computing devices with relatively limited computational performance. For example, linear processor 2410 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer.
[0200] Turning to Figure 24 , system 2400 includes source neural network 1705 and linear processor 2410. Linear processor 2410 is a device that performs operations based on a linear combination of features of a representation (or an approximation of such a representation) of patterns of activity in a neural network. The operations can be, for example, object localization operations, object detection operations, object segmentation operations, prediction operations, action selection operations, etc.
[0201] The linear processor 2410 includes an input 2420 and an output 2425. The input 2420 is coupled to receive a representation of the pattern of activity in the source neural network 1705. The linear processor 2410 can receive the representation of the pattern of activity in the source neural network 1705 in a variety of ways. For example, the representation of the pattern of activity can be received as discrete events or as a continuous stream over a real-time or non-real-time communication channel. The output 2525 is coupled to output the processing result from the linear processor 2410. In some implementations, the linear processor 2410 can be implemented on one or more computing devices having relatively limited computational performance. For example, the linear processor 2410 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer.
[0202] In Figure 25 , the classification system 2500 includes a source neural network 1705 and a neural network 2510. The neural network 2510 is a neural network device that is configured to perform operations based on a non-linear combination of features of a representation (or an approximation of such a representation) of the pattern of activity in the neural network. The operations can be, for example, object localization operations, object detection operations, object segmentation operations, prediction operations, action selection operations, etc. In the illustrated implementation, the neural network 2510 is a feed-forward network that includes an input layer 2520 and an output layer 2525. Like the linear processor 2410, the neural network 2510 can receive the representation of the pattern of activity in the source neural network 1705 in a variety of ways.
[0203] In some implementations, the neural network 2510 can perform inference on one or more computing devices having relatively limited computational performance. For example, the neural network 2510 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer, for example, in the neural processing unit of such a device. Like system 2400, system 2500 will typically be a distributed system in which the remote neural network 2510 communicates with the source neural network 1705, for example, via a data communication network. In some implementations, the neural network 2510 can be, for example, a deep neural network, such as a convolutional neural network.
[0204] In Figure 26 , the system 2600 includes a source approximator 1905 and a linear processor 2410. Despite any differences between the approximation 1200' and the representation 1200, the processor 2410 can still perform operations on the approximation 1200'.
[0205] In Figure 27 , the system 2700 includes a source approximator 1905 and a neural network 2510. Despite any differences between the approximation 1200' and the representation 1200, the neural network 2510 can still perform operations on the approximation 1200'.
[0206] In some implementations, systems 2600, 2700 can be implemented on edge devices (such as, for example, edge devices 2100, 2200( Figure 21 , Figure 22 ))). In some implementations, systems 2600, 2700 can be implemented as part of a system (such as system 2300( Figure 23 )) in which a local neural network can be trained using an occurrence representation corresponding to the topology of the activity in the source neural network.
[0207] Figure 28 FIG. is a schematic illustration of a reinforcement learning system 2800 that includes an artificial neural network that can be trained using an occurrence representation corresponding to the topology of the activity in the source neural network. Reinforcement learning is a type of machine learning in which an artificial neural network learns from feedback on the results of actions taken in response to decisions of the artificial neural network. A reinforcement learning system moves from one state in an environment to another by performing actions and receiving information that characterizes the new state and rewards and / or regrets that characterize the success (or lack of success) of the actions. Reinforcement learning seeks to maximize the total reward (or minimize the regret) through a learning process.
[0208] In the illustrated implementation, the artificial neural network in reinforcement learning system 2800 is a deep neural network 2805 (or other deep learning architecture) trained using reinforcement learning methods. In some implementations, deep neural network 2805 can be a local artificial neural network (such as neural network 2510( Figure 25 , Figure 27 )) and is implemented locally on, for example, an automobile, an airplane, a robot, or other device. However, this need not be the case, and in other implementations, deep neural network 2805 can be implemented on a system of networked devices.
[0209] In addition to source approximator 1905 and deep neural network 2805, reinforcement learning system 2800 also includes actuator 2810, one or more sensors 2815, and a teacher module 2820. In some implementations, reinforcement learning system 2800 also includes one or more additional data sources 2825.
[0210] Actuator 2810 is a device that controls an agency or system that interacts with environment 2830. In some implementations, actuator 2810 controls a physical agency or system (e.g., the steering of an automobile or the positioning of a robot). In other implementations, actuator 2810 may control a virtual agency or system (e.g., a virtual game board or an investment portfolio). Thus, environment 2830 may also be physical or virtual.
[0211] Sensor 2815 is a device that measures characteristics of environment 2830. At least some of the measurements characterize the interaction between the agency or system being controlled and other aspects of environment 2830. For example, when actuator 2810 maneuvers an automobile, sensor 2815 may measure one or more of the speed, direction, and acceleration of the automobile, the proximity of the automobile to other features, and the response of other features to the automobile. As another example, when actuator 2810 controls an investment portfolio, sensor 2815 may measure the value and risk associated with the portfolio.
[0212] Typically, both source approximator 1905 and teacher module 2820 are coupled to receive at least some of the measurements obtained by sensor 2815. For example, source approximator 1905 may receive measurement result data at input layer 1915 and output an approximation 1200’ of the representation of the topology that appears in the pattern of activity in the source neural network.
[0213] Teacher module 2820 is a device configured to interpret the measurements received from sensor 2815 and provide rewards and / or regrets to deep neural network 2805. A reward is positive and indicates successful control of the agency or system. A regret is negative and indicates unsuccessful or suboptimal control. Typically, teacher module 2820 also provides a characterization of the measurements and the rewards / regrets for reinforcement learning. Typically, the characterization of the measurements is an approximation of the representation of the topology that appears in the pattern of activity in the source neural network (such as approximation 1200’). For example, teacher module 2820 may read approximation 1200’ output from source approximator 1905 and pair the read approximation 1200’ with corresponding reward / regret values.
[0214] In multiple implementations, reinforcement learning does not occur in real time in system 2800 or during the active control of actuator 2810 by deep neural network 2805. Instead, training feedback can be collected by teacher module 2820 and used for reinforcement training when deep neural network 2805 is not actively instructing actuator 2810. For example, in some implementations, teacher module 2820 can be remote from deep neural network 2805 and only communicate with deep neural network 2805 intermittently for data. Regardless of whether the reinforcement learning is intermittent or continuous, deep neural network 2805 can be evolved, e.g., to optimize rewards and / or reduce regret using the information received from teacher module 2820.
[0215] In some implementations, system 2800 also includes one or more additional data sources 2825. Source approximator 1905 can also receive data from data sources 2825 at input layer 1915. In these instances, approximation 1200’ will be caused by processing both sensor data and data from data sources 2825.
[0216] In some implementations, data collected by a reinforcement learning system 2800 can be used for the training or reinforcement learning of other systems (including other reinforcement learning systems). For example, the characterization of measurements and the reward / regret values can be provided by teacher module 2820 to a data exchange system that collects such data from a variety of reinforcement learning systems and redistributes the data among them. Additionally, as discussed above, the characterization of measurements can be an approximation of the representation of topologies that occur in the pattern of activities in the source neural network, such as approximation 1200’.
[0217] The specific operations performed by reinforcement learning system 2800 will of course depend on the specific operating scenario. For example, in a scenario where source approximator 1905, deep neural network 2805, actuator 2810, and sensor 2815 are part of an automobile, deep neural network 2805 can perform object localization and / or detection operations while maneuvering the automobile.
[0218] In implementations where data collected by reinforcement learning system 2800 is used for the training or reinforcement learning of other systems, the reward / regret values and approximation 1200’ that characterize the state of the environment when performing object localization and / or detection operations can be provided to the data exchange system. The data exchange system can then distribute the reward / regret values and approximation 1200’ to other reinforcement learning systems 2800 associated with other vehicles for reinforcement learning at these other vehicles. For example, reinforcement learning can be used to improve object localization and / or detection operations at a second vehicle using the reward / regret values and approximation 1200’.
[0219] However, the operations learned at other vehicles need not be the same as those performed by the deep neural network 2805. For example, a travel-time based reward / regret value and an approximation of 1200’ resulting from an input of sensor data characterizing an unexpectedly wet road at a location identified, e.g., via the GPS data source 2825, can be used for route planning operations at another vehicle.
[0220] Embodiments of the operations and subject matter described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware—including the structures disclosed in this specification and their structural equivalents—or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Optionally or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination or subset of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be one or more separate physical components or media (such as multiple CDs, disks, or other storage devices) or a combination or subset of one or more of them included therein.
[0221] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0222] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, a system-on-a-chip, or multiple or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination or subset of one or more of them. The apparatus and the execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0223] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language - including compiled or interpreted languages, declarative or procedural languages - and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0224] The processes and logical flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit).
[0225] Processors suitable for executing computer programs include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions in accordance with the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to receive data from one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or to transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks) or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game controller, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0226] To provide for interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including acoustic, speech, or tactile input. Additionally, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page in response to a request received from a web browser to the web browser on a client device of the user.
[0227] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination within a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. In addition, although features may be described above as acting in certain combinations and even as initially claimed in such combinations, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.
[0228] Similarly, while operations are depicted in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components described above in the embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0229] Accordingly, specific implementations of a subject matter have been described. Other implementations are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order shown or sequential order to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0230] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, although 1200 is shown as a binary representation where each bit individually represents the presence or absence of a feature in a diagram, other representations of information are possible. For example, a vector or matrix of multi-valued, non-binary numbers can be used to represent, for example, the presence or absence of features and perhaps other features of these features. An example of such a feature is the weight of an active edge that constitutes a feature.
[0231] Accordingly, other embodiments are within the scope of the appended claims.
Claims
1. A classification device, comprising: A neural network, the neural network being trained to: Produce an approximation of a first representation of the presence or absence of a plurality of individual topologies in a pattern that is only a portion of the signal transmission activity that occurs in a source recurrent artificial neural network in response to a first input, wherein the nodes in the source recurrent artificial neural network operate as accumulators, and wherein the topologies all include more than two nodes in the source recurrent artificial neural network and more than one edge between the nodes, Produce an approximation of a second representation of the presence or absence of the plurality of individual topologies in a pattern that is only a portion of the signal transmission activity that occurs in the source recurrent artificial neural network in response to a second input, and Produce an approximation of a third representation of the presence or absence of the plurality of individual topologies in a pattern that is only a portion of the signal transmission activity that occurs in the source recurrent artificial neural network in response to a third input; And A classifier, the classifier being coupled to receive the approximations of the representations produced by the neural network and classify the received approximations, Wherein each of the first input, the second input, and the third input is a data stream, and the data stream includes video or audio data.
2. The device according to claim 1, wherein the topology comprises a directed simplex.
3. The device according to claim 1, wherein the topology encloses a cavity.
4. The device according to claim 1, wherein each of the first representation, the second representation, and the third representation represents a separate topology that occurs in the source recurrent artificial neural network only at a time when the pattern that is active during that time has a complexity distinguishable from the complexity of other activities in response to a corresponding input in the input.
5. The device according to claim 1, wherein the classifier comprises a second neural network that has been trained to process the approximation of the representation produced by the neural network.
6. The device according to claim 1, wherein each of the first representation, the second representation, and the third representation comprises multi-valued, non-binary digits, where each digit indicates the presence or absence of a corresponding topology in the separate topology.
7. The device according to claim 1, wherein each of the first representation, the second representation, and the third representation represents the occurrence of the separate topology without specifying where in the source recurrent artificial neural network the topology occurs.
8. The device according to claim 1, wherein the device comprises a smart phone.
9. The device according to claim 1, wherein the device is a security camera device.
10. A device, comprising: A neural network, the neural network being coupled to input a representation of the presence or absence of a plurality of individual topologies in a pattern of activity that occurs in a source neural network in response to a plurality of different inputs, wherein the nodes in the source neural network operate as accumulators, and wherein the topologies all include three or more nodes in the source neural network and three or more edges between the nodes, and wherein the neural network is trained to process the representation and produce a response output, Wherein the input to the source neural network is a data stream, and the data stream includes video or audio data.
11. The device according to claim 10, wherein the topology comprises a directed simplex.
12. The device according to claim 10, wherein the representation of the topology represents a topology that occurs in the source neural network only at a time when the pattern that is active during that time has a complexity distinguishable from the complexity of other activities in response to the corresponding input in the input.
13. The apparatus according to claim 10, wherein the apparatus further comprises a neural network trained to produce a respective approximation of the presence or absence of the plurality of individual topologies in the pattern of activity that occurs in the source neural network in response to the plurality of different inputs in response to the plurality of different inputs.
14. The apparatus according to claim 10, wherein the representation of the presence or absence of the plurality of individual topologies comprises multi-valued, non-binary digits, wherein each digit indicates the presence or absence of a respective topology in the individual topologies.
15. The apparatus according to claim 10, wherein the representation of the presence or absence of the plurality of individual topologies represents the presence or absence of the individual topologies without specifying where in the source neural network the pattern of activity occurs.
16. The apparatus according to claim 10, wherein the source neural network is a recurrent neural network.
17. A method implemented by a neural network apparatus, the method comprising: Input a representation of the presence or absence of a plurality of individual topologies in a pattern of activity in a source neural network, wherein the activity is in response to an input to the source neural network and the topologies all include three or more nodes in the source neural network and three or more edges between the nodes; Process the representation, wherein the processing is consistent with training the neural network to process different such representations of the presence or absence of the plurality of individual topologies in a pattern of activity in the source neural network; And Output the result of the processing of the representation, Wherein the input to the source neural network is a data stream, and the data stream includes video or audio data.
18. The method according to claim 17, wherein the topology comprises a directed simplex.
19. The method according to claim 17, wherein the topology encloses a cavity.
20. The method according to claim 17, wherein the representation of the presence or absence of the plurality of individual topologies represents the presence or absence of the plurality of individual topologies only at times when the pattern of activity has a complexity distinguishable from the complexity of other activity in response to a corresponding input in the input.
21. The method according to claim 17, wherein the representation of the presence or absence of the plurality of individual topologies comprises multi-valued, non-binary digits, wherein each digit indicates the presence or absence of a respective topology in the individual topologies.
22. The method according to claim 17, wherein the representation of the presence or absence of the plurality of individual topologies represents the occurrence of the topology without specifying where in the source neural network the topology occurs.
23. The method according to claim 17, wherein the source neural network is a recurrent neural network.
Citation Information
Patent Citations
Computational nodes and computational-node networks that include dynamical-nanodevice connections
CN101669130A