Characterizing and encoding and decoding information in recurrent artificial neural networks
By characterizing activity patterns in recurrent artificial neural networks and leveraging topological structure and complexity analysis, we address the recognition and information encoding issues at the decision moment, improving the clarity and efficiency of information processing.
Patent Information
- Application Number
- CN202510727328.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-11
- Filing Date
- 2019-06-06
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies have difficulty effectively identifying and encoding decision moments in recurrent artificial neural networks, resulting in unclear information processing results.
By characterizing the activity patterns in recurrent artificial neural networks, we use topological representation and complexity analysis to identify decision moments and encode information.
It achieves accurate identification of decision moments and effective encoding of information in recurrent artificial neural networks, improving the clarity and efficiency of information processing.
Smart Images

Figure CN120805978A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application with the application date of 06 June 2019, the application number of 201980053141.0 (international application number of PCT / EP2019 / 064773), and the invention name of “Characterizing activity in recurrent artificial neural networks and encoding and decoding information”. TECHNICAL FIELD
[0002] The present application relates to characterizing activity in recurrent artificial neural networks and encoding and decoding information. BACKGROUND
[0003] The present specification relates to the characterization of activity in recurrent artificial neural networks. The characterization of activity can be used, for example, in the identification of decision moments, and in encoding / decoding signals in scenarios such as transmission, encryption, and data storage. It also relates to encoding and decoding information, and systems and techniques for using the encoded information in various scenarios. The encoded information can represent activity in a neural network (e.g., a recurrent neural network).
[0004] Artificial neural networks are devices inspired by structural and functional aspects of biological neuron networks. In particular, artificial neural networks use a system of interconnected constructs called nodes to mimic the information encoding and other processing capabilities of biological neuron networks. The arrangement and strength of connections between nodes in an artificial neural network determine the outcome of information processing or information storage by the artificial neural network.
[0005] Neural networks can be trained to produce a desired signal flow in the network and achieve a desired information processing or information storage outcome. Typically, training a neural network will change the arrangement and / or strength of connections between nodes during a learning phase. A neural network can be considered trained when it achieves a sufficiently appropriate processing outcome for a given set of inputs.
[0006] Artificial neural networks can be used in a wide variety of different devices to perform nonlinear data processing and analysis. Nonlinear data processing does not satisfy the superposition principle, i.e., the variable to be determined cannot be written as a linear sum of independent components. Examples of scenarios in which nonlinear data processing is useful include pattern and sequence recognition, speech processing, novelty detection and sequential decision making, complex system modeling, and systems and techniques in a wide variety of other scenarios.
[0007] Both encoding and decoding transform information from one form or representation to another. Different representations can provide different features that are more or less useful in different applications. For example, some forms or representations of information (e.g., natural languages) can be more easily understood by humans. Other forms or representations can be smaller in size (e.g., "compressed") and easier to transmit or store. Still other forms or representations can intentionally obscure the content of the information (e.g., the information can be encoded cryptographically).
[0008] Regardless of the particular application, the encoding or decoding process will generally follow a predefined set of rules or algorithm that establishes a correspondence between the information in the different forms or representations. For example, an encoding process that produces binary code can assign a role or meaning to individual bits according to their position in a binary sequence or vector. SUMMARY
[0009] This document relates to encoding and decoding information, and systems and techniques for using encoded information in various scenarios. For example, in one implementation, a device includes a neural network trained to produce, in response to a first input, an approximation of a first representation of a topology in a pattern of activity that occurs in a source neural network in response to the first input, to produce, in response to a second input, an approximation of a second representation of the topology in the pattern of activity that occurs in the source neural network in response to the second input, and to produce, in response to a third input, an approximation of a third representation of the topology in the pattern of activity that occurs in the source neural network in response to the third input.
[0010] This and other implementations can include one or more of the following features. The topologies can all include two or more nodes in the source neural network and one or more edges between the nodes. The topologies can include simplices. The topologies can enclose cavities. Each of the first representation, the second representation, and the third representation can represent topologies that occur in the source neural network only at times during which the pattern of activity has a complexity that is distinguishable from the complexity of other activity in response to respective ones of the inputs. The device can also include a processor coupled to receive the approximations of the representations produced by the neural network device and to process the received approximations. The processor can include a second neural network that has been trained to process representations produced by the neural network. Each of the first representation, the second representation, and the third representation can include multi-valued, non-binary digits. Each of the first representation, the second representation, and the third representation can represent occurrences of the topologies without specifying where in the source neural network the pattern of activity occurs. The device can include a smartphone. The source neural network can be a recurrent neural network.
[0011] In another implementation, a device includes a neural network coupled to input representations of topologies in patterns of activity that occur in a source neural network in response to a plurality of different inputs. The neural network is trained to process the representations and to produce a responsive output.
[0012] This and other implementations can include one or more of the following features. The topologies can all include two or more nodes in the source neural network and one or more edges between the nodes. The topologies can include simplices. The representations of topologies can represent topologies that occur in the source neural network only at times during which the pattern of activity has a complexity that is distinguishable from the complexity of other activity in response to respective ones of the inputs. The device can include a neural network trained to produce, in response to a plurality of different inputs, respective approximations of representations of topologies in patterns of activity that occur in the source neural network in response to the different inputs. The representations of topologies can include multi-valued, non-binary digits. The representations of topologies can represent occurrences of the topologies without specifying where in the source neural network the pattern of activity occurs. The source neural network can be a recurrent neural network.
[0013] In another implementation, a method is implemented by a neural network device and includes inputting a representation of a topology in a pattern of activity in a source neural network, where the activity is responsive to input into the source neural network; processing the representation; and outputting a result of the processing of the representation. The processing is consistent with training the neural network to process different such representations of topologies in patterns of activity in the source neural network.
[0014] This and other implementations can include one or more of the following features. The topologies can all include two or more nodes in the source neural network and one or more edges between the nodes. The topologies can include simplices. The topologies can enclose cavities. The representation of a topology can represent topologies that occur in the source neural network only at times during which the pattern of activity has a complexity that is distinguishable from complexities of other activity responsive to respective ones of the input. The representation of a topology can include multivalued, non-binary digits. The representation of a topology can represent an occurrence of the topology without specifying where in the source neural network the pattern of activity occurs. The source neural network can be a recurrent neural network.
[0015] The details of one or more implementations described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a schematic illustration of a structure of a recurrent artificial neural network device.
[0017] Figure 2 and Figure 3 is a schematic illustration of a function of a recurrent artificial neural network device over different time windows.
[0018] Figure 4 is a flowchart of a process for identifying decision instants in a recurrent artificial neural network based on characterization of activity in the network.
[0019] Figure 5 is a schematic illustration of a pattern of activity that can be identified and used to identify decision instants in a recurrent artificial neural network.
[0020] Figure 6 is a schematic illustration of a pattern of activity that can be identified and used to identify decision instants in a recurrent artificial neural network.
[0021] Figure 7is a schematic illustration of a pattern of activity that can be identified and used to identify decision instants in a recurrent artificial neural network.
[0022] Figure 8 is a schematic illustration of a data table that can be used in determining the complexity or degree of ordering in a pattern of activity in a recurrent artificial neural network device.
[0023] Figure 9 is a schematic illustration of the determination of the specific timing of a pattern of activity having a distinguishable complexity.
[0024] Figure 10 is a flowchart of a process for encoding a signal using a recurrent artificial neural network based on the characterization of activity in that network.
[0025] Figure 11 is a flowchart of a process for decoding a signal using a recurrent artificial neural network based on the characterization of activity in that network.
[0026] Figure 12 , Figure 13 and Figure 14 are schematic illustrations of binary forms or representations of topologies.
[0027] Figure 15 and Figure 16 schematically illustrate one example of how the presence or absence of features corresponding to different bits are not independent of one another.
[0028] Figure 17 , Figure 18 , Figure 19 , Figure 20 are schematic illustrations of representations of the occurrence of topologies in activity in a neural network used in four different classification systems.
[0029] Figure 21 , Figure 22 are schematic illustrations of edge devices including local artificial neural networks that can be trained using representations of the occurrence of topologies corresponding to activity in a source neural network.
[0030] Figure 23 is a schematic illustration of a system in which a local neural network can be trained using representations of the occurrence of topologies corresponding to activity in a source neural network.
[0031] Figure 24 , Figure 25 , Figure 26 , Figure 27is a schematic illustration of the use of representations of occurrences of topology in activity in neural networks in four different systems.
[0032] Figure 28 is a schematic illustration of a system 0 that includes an artificial neural network that can be trained using representations of occurrences of topology corresponding to activity in a source neural network.
[0033] Like reference symbols in the various drawings indicate like elements. DETAILED DESCRIPTION
[0034] Figure 1 is a schematic illustration of the structure of a recurrent artificial neural network device 100. Recurrent artificial neural network devices 100 are devices that use systems of interconnected nodes to simulate information encoding and other processing capabilities of biological neural networks. Recurrent artificial neural network devices 100 can be implemented in hardware, software, or a combination thereof.
[0035] An illustration of a recurrent artificial neural network device 100 includes a plurality of nodes 101, 102, 103, 104, 105, 106, 107 interconnected by a plurality of structural links 110. Nodes 101, 102, 103, 104, 105, 106, 107 are discrete information processing constructs analogous to neurons in biological networks. Nodes 101, 102, 103, 104, 105, 106, 107 typically process one or more input signals received through one or more of links 110 to produce one or more output signals output through one or more of links 110. For example, in some implementations, nodes 101, 102, 103, 104, 105, 106, 107 can be artificial neurons that weight and sum a plurality of input signals, pass the sum through one or more non-linear activation functions, and output one or more output signals.
[0036] Nodes 101, 102, 103, 104, 105, 106, 107 can operate as accumulators. For example, nodes 101, 102, 103, 104, 105, 106, 107 can operate according to an integrate-and-fire model in which one or more signals are accumulated in a first node until a threshold is reached. Upon reaching the threshold, the first node fires by transmitting an output signal along one or more of links 110 to a connected second node. In turn, the second node 101, 102, 103, 104, 105, 106, 107 accumulates received signals and, if a threshold is reached, the second node 101, 102, 103, 104, 105, 106, 107 transmits yet another output signal to one further connected node.
[0037] Structural links 110 are connections that enable the transmission of signals between nodes 101, 102,..., 107. For convenience, all structural links 110 are treated as identical bidirectional links in this document that carry signals from every first node of nodes 101, 102,..., 107 to every second node of nodes 101, 102,..., 107 in the same way that they carry signals from the second node to the first node. However, this need not be the case. For example, some or all of structural links 110 can be unidirectional links that carry signals from a first node of nodes 101, 102,..., 107 to a second node of nodes 101, 102,..., 107 without carrying signals from the second node to the first node.
[0038] As another example, in some implementations, structural links 110 can have a variety of characteristics other than, or in addition to, directionality. For example, in some implementations, different structural links 110 can carry signals of different amplitudes - resulting in different strengths of interconnection between respective ones of nodes 101, 102,..., 107. As another example, different structural links 110 can carry different types of signals (e.g., inhibitory and / or excitatory signals). Indeed, in some implementations, structural links 110 can mimic links between somatic cells in biological systems, and reflect at least a portion of the vast morphological, chemical, and other diversity of such links.
[0039] In the illustrated implementation, recurrent artificial neural network device 100 is a clique network (or subnetwork) because each node 101, 102,..., 107 is connected to every other node 101, 102,..., 107. This need not be the case. Rather, in some implementations, each node 101, 102,..., 107 can be connected to a proper subset of nodes 101, 102,..., 107 (via the same links or a variety of links, as the case can be).
[0040] For the sake of clarity of illustration, recurrent artificial neural network device 100 is illustrated as having only seven nodes. In general, real-world neural network devices will include a significantly larger number of nodes. For example, in some implementations, a neural network device can include hundreds of thousands, millions, or even billions of nodes. Thus, recurrent neural network device 100 can be a small portion (i.e., a subnetwork) of a larger recurrent artificial neural network.
[0041] In biological neural network devices, the accumulation and signal transmission processes require the passage of time in the real world. For example, a soma of a neuron integrates inputs received over time, and signal transmission from neuron to neuron takes time, determined by, for example, the speed of signal transmission and the nature and length of the links between neurons. Thus, the state of a biological neural network device is dynamic and changes over time.
[0042] In artificial recurrent neural network devices, time is artificial and represented using mathematical constructs. For example, signal transmission from node to node does not require the passage of time in the real world, and such signals can be represented in artificial units that are typically independent of the passage of time in the real world, as measured, for example, with a computer clock cycle or other means. However, the state of an artificial recurrent neural network device can be described as “dynamic” in that it changes with respect to these artificial units.
[0043] Note that these artificial units are referred to herein as “time” units for convenience. However, it is understood that these units are artificial and typically do not correspond to the passage of time in the real world.
[0044] Figure 2 and Figure 3 are schematic illustrations of the functioning of artificial neural network device 100 over different time windows. Because the state of device 100 is dynamic, the functioning of device 100 can be represented using signal transmission activity that occurs within one window. Such functional illustrations typically show activity in only a small fraction of links 110. In particular, since typically not every link 110 transmits a signal within a particular window, not every link 110 is illustrated as actively contributing to the functioning of device 100 in these illustrations.
[0045] In Figure 2 and Figure 3 illustrations, active links 110 are illustrated as relatively thick solid lines connecting pairs of nodes 101, 102,..., 107. In contrast, inactive links 110 are illustrated as dashed lines. This is for illustration purposes only. In other words, whether or not a link 110 is active, the structural connectivity formed by links 110 exists. However, this formalism highlights the activity and functioning of device 100.
[0046] In addition to illustratively showing the presence of activity along the links, the direction of the activity is also illustratively shown. In particular, the relatively thick solid lines illustrating active links 110 also include arrows indicating the direction of signal transmission along the links during the relevant window. Generally, the direction of signal transmission within a single window does not deterministically constrain the links to be unidirectional with the indicated directionality. Rather, in the first functional illustration for the first window of time, the links can be active in a first direction. In the second functional illustration for the second window, the links can be active in an opposite direction. In some cases, however, such as, for example, in a recurrent artificial neural network device 100 that includes only unidirectional links, the directionality of signal transmission will deterministically indicate the directionality of the links.
[0047] In feedforward neural network devices, information moves only in a single direction (i.e., forward) to a node output layer at the end of the network. Feedforward neural network devices indicate that a "decision" has been reached and information processing is complete by the propagation of signals through the network to the output layer.
[0048] In contrast, in recurrent neural networks, the connections between nodes form a loop, and the activity of the network dynamically progresses without an easily identifiable decision. For example, even in a three-node recurrent neural network, a first node can transmit a signal to a second node, which in response can transmit a signal to a third node. In response, the third node can transmit a signal back to the first node. The signal received by the first node can be, at least in part, responsive to the signal transmitted from that same node.
[0049] Illustrative functional illustration Figure 2 and Figure 3 This is illustrated in a network only slightly larger than a three-node recurrent neural network. Figure 2 The functional illustration shown in FIG. 1 can illustrate activity within a first window, and Figure 3 activity within a second window immediately following. As shown, the set of signal transmission activities appears to originate in node 104 and progress through the device 100 in a generally clockwise direction during the first window. Within the second window, at least some of the signal transmission activities generally appear to return to node 104. Even in such an oversimplified illustration, the signal transmission does not proceed in a manner that produces a clearly identifiable output or end.
[0050] When considering recurrent neural networks of, for example, thousands of nodes or more, it can be recognized that signal propagation can occur through a large number of paths, and that these signals lack a clearly identifiable "output" location or time. While by design the network can return to a quiescent state in which only background activity or even no signal transmission activity occurs, the quiescent state itself does not indicate the result of the information processing. Regardless of the input, the recurrent neural network always returns to the quiescent state. Thus, the "output" or result of the information processing is encoded in the activity that occurs in the recurrent neural network in response to a particular input.
[0051] Figure 4 is a flowchart of a process 400 for identifying a decision instant in a recurrent artificial neural network based on characterization of activity in the network. A decision instant is a point in time at which activity in a recurrent artificial neural network indicates a result of information processing by the network in response to an input. The process 400 can be performed by a system of one or more data processing devices executing operations in accordance with one or more machine-readable instruction sets. For example, the process 400 can be performed by the same system of one or more computers executing software for implementing a recurrent artificial neural network used in the process 400.
[0052] At 405, the system performing the process 400 receives a notification that a signal has been input into a recurrent artificial neural network. In some cases, the input of the signal is a discrete injection event in which, for example, information is injected into one or more nodes and / or one or more links of the neural network. In other cases, the input of the signal is a stream of information injected into one or more nodes and / or links of the neural network over a period of time. The notification indicates that the artificial neural network is actively processing information and is not, for example, in a quiescent state. In some cases, the notification is received from the neural network itself, for example, such as when the neural network exits an identifiable quiescent state.
[0053] At 410, the system performing the process 400 partitions the responsive activity in the network into a set of windows. In cases where the injection is a discrete event, the windows can subdivide the time between the injection and the return to the quiescent state into periods during which the activity displays variable complexity. In cases where the injection is a stream of information, the duration of the injection (and optionally the time to return to the quiescent state after the injection is complete) can be subdivided into windows during which the activity displays variable complexity. Various methods of determining the complexity of the activity are discussed further below.
[0054] In some implementations, the windows all have the same duration, although this need not be the case. Rather, in some implementations, the windows can have different durations. For example, in some implementations, the duration can increase as time since the discrete injection event has occurred increases.
[0055] In some implementations, the windows can be a continuous series of individual windows. In other implementations, the windows overlap in time, such that one window begins before a previous window ends. In some cases, the windows can be moving windows that move in time.
[0056] In some implementations, different durations of windows are defined for different complexities of patterns of activity. For example, a window can have a relatively long duration for defining a pattern of activity that occurs between a relatively large number of nodes than a window defined for a pattern of activity that occurs between a relatively small number of nodes. For example, in the scenario of the patterns of activity 500 ( Figure 5 ), a window defined for identifying activity commensurate with the pattern 530 can be longer than a window defined for identifying activity commensurate with the pattern 505.
[0057] At 415, the system executing the process 400 identifies patterns in activity within different windows in the network. As discussed further below, the patterns in activity can be identified by treating a functional graph as a topological space with nodes as points. In some implementations, the identified patterns of activity are cliques in the functional graph of the network, e.g., directed cliques.
[0058] At 420, the system executing the process 400 determines a complexity of the patterns of activity within the different windows. The complexity can be a measure of the likelihood of an ordered pattern of activity occurring within a window. Thus, a randomly occurring pattern of activity would be relatively simple. On the other hand, a pattern of activity that shows a non-random order would be relatively complex. For example, in some implementations, the complexity of a pattern of activity can be measured using, e.g., a simplex count or a Betti number of the pattern of activity.
[0059] At 425, the system executing the process 400 determines a particular time of a pattern of activity having a distinguishable complexity. The particular pattern of activity can be distinguishable based on a complexity that deviates upward or downward (e.g., from a fixed or variable baseline). In other words, a particular time of a pattern of activity can be determined that indicates a particularly high level or a particularly low level of non-random order in the activity.
[0060] For example, in cases where the signal input is a discrete injection event, a deviation, e.g., a deviation from a stable baseline or a deviation from a curve that is a characteristic of the average response of the neural network to a variety of different discrete injection events, can be used to determine the specific time of a distinguishably complex activity pattern. As another example, in cases where information is input in a stream, a large change in complexity during the streaming can be used to determine the specific time of a distinguishably complex activity pattern.
[0061] At 430, the system performing process 400 schedules a reading of the output from the neural network based on the specific time of the distinguishably complex activity pattern. For example, in some implementations, the output of the neural network can be read at the same time that the distinguishably complex activity pattern occurs. In implementations where the complexity deviation indicates a relatively high non-random order in the activity, the observed activity pattern itself can also be taken as the output of the recurrent artificial neural network.
[0062] Figure 5 is an example of a pattern 500 of activity that can be identified and used to identify a decision moment in a recurrent artificial neural network. For example, pattern 500 can be identified at 415 in process 400 Figure 4 ) of FIG. 4.
[0063] Pattern 500 is an example of activity in a recurrent artificial neural network. During the application of pattern 500, the functional graph is considered as a topological space with nodes as points. Activity in the nodes and links commensurate with pattern 500 can be identified as ordered, regardless of the identity of the particular nodes and / or links that participate in the activity. For example, first pattern 505 can represent activity between nodes 101, 104, 105 in Figure 2 , with point 0 in pattern 505 as node 104, point 1 as node 105, and point 2 as node 101. As another example, first pattern 505 can also represent activity between nodes 104, 105, 106 in Figure 3 , with point 0 in pattern 505 as node 106, point 1 as node 104, and point 2 as node 105. The order of activity in the directed clique is also specified. For example, in pattern 505, activity between point 1 and point 2 occurs after activity between point 0 and point 1.
[0064] In the illustrated implementations, the patterns 500 are all directed cliques or directed simplices. In such patterns, an activity originates from a source node that transmits a signal to every other node in the pattern. In the patterns 500, such a source node is designated as point 0, while the other nodes are designated as points 1, 2,.... Furthermore, in a directed clique or simplex, one of the nodes acts as a sink and receives the transmitted signals from every other node in the pattern. In the patterns 500, such a sink node is designated as the highest numbered point in the pattern. For example, in the pattern 505, the sink node is designated as point 2. In the pattern 510, the sink node is designated as point 3. In the pattern 515, the sink node is designated as point 3, and so on. Thus, the activities represented by the patterns 500 are ordered in a distinguishable manner.
[0065] Each of the patterns 500 has a different number of points and reflects ordered activities in a different number of nodes. For example, the pattern 505 is a two-dimensional simplex and reflects activities in three nodes, the pattern 510 is a three-dimensional simplex and reflects activities in four nodes, and so on. As the number of points in the pattern increases, the degree and complexity of the ordering of the activities also increases. For example, for a large set of nodes with a certain degree of random activity within a window, some of the activities can coincidentally be commensurate with the pattern 505. However, the random activities will increasingly and progressively be less likely to be commensurate with the respective patterns in the patterns 510, 515, 520,.... The existence of activities commensurate with the pattern 530 indicates a relatively high degree and complexity of ordering in the activities compared to the existence of activities commensurate with the pattern 505.
[0066] As previously discussed, in some implementations, different durations of windows can be defined for different determinations of complexity of activities. For example, when activities commensurate with the pattern 530 are to be identified, a longer duration of window can be used compared to when activities commensurate with the pattern 505 are to be identified.
[0067] Figure 6 is an illustration of a pattern 600 that can be identified and used to identify decision instants in activities in a recurrent artificial neural network. For example, the pattern 600 can be identified at 415 in the process 400 ( Figure 4 ) of FIG. 4.
[0068] Like pattern 500, pattern 600 is an illustration of activity in a recurrent artificial neural network. However, pattern 600 deviates from the strict ordering of pattern 500, as pattern 600 is not entirely a directed clique or a directed simplex. In particular, patterns 605, 610 have less directionality than pattern 515. Indeed, pattern 605 is entirely devoid of convergent nodes. However, patterns 605, 610 indicate a degree of ordered activity that exceeds that expected by random chance, and can be used to determine a complexity of activity in a recurrent artificial neural network.
[0069] Figure 7 Pattern 700 is an illustration of a pattern that can be identified and used to identify activity at a decision instant in a recurrent artificial neural network. For example, pattern 700 can be identified at 415 in process 400 Figure 4 ) of FIG. 4.
[0070] Pattern 700 is a group of directed cliques or directed simplices of the same dimensionality (i.e., having the same number of points) that define a pattern involving more points than an individual clique or simplex and enclose a cavity within the group of directed simplices.
[0071] By way of example, pattern 705 includes six distinct three-point, two-dimensional patterns 505 that together define a homology class of level 2, while pattern 710 includes eight distinct three-point, two-dimensional patterns 505 that together define a second homology class of level 2. Each of the three-point, two-dimensional patterns 505 in patterns 705, 710 can be considered to enclose a respective cavity. The nth Betti number associated with a directed graph provides a count of such homology classes in a topological representation.
[0072] Activity exemplified by patterns such as pattern 700 illustrates a relatively high degree of ordering of activity in a network that is unlikely to occur by random chance. Pattern 700 can be used to characterize a complexity of this activity.
[0073] In some implementations, only some patterns of activity are identified and / or some portion of the identified activity is discarded or otherwise ignored during identification at a decision instant. For example, with reference to Figure 5 Activity commensurate with five-point, four-dimensional simplex pattern 515 inherently includes activity commensurate with four-point, three-dimensional and three-point, two-dimensional simplex patterns 510, 505. For example, Figure 5 Points 0, 2, 3, 4 and points 1, 2, 3, 4 of four-dimensional simplex pattern 515 of FIG. 5 are commensurate with three-dimensional simplex pattern 510. In some implementations, patterns containing fewer points - and thus having lower dimensionality - can be discarded or otherwise ignored during identification at a decision instant.
[0074] As another example, only some patterns of activity need to be identified. For example, in some implementations, only patterns with an odd number of points (3, 5, 7, ...) or an even number of dimensions (2, 4, 6, ...) are used in identifying decision moments.
[0075] The degree of complexity or order in the activity patterns in a recurrent artificial neural network device within different windows can be determined in a variety of different ways. Figure 8 is a schematic illustration of a data table 800 that can be used in such a determination. Data table 800 can be used to determine the complexity of an activity pattern in isolation or in combination with other activities. For example, in process 400 ( Figure 4 ) uses data table 800 at 420.
[0076] In more detail, table 800 includes a count of the number of occurrences of the pattern during window "N," with counts of activities matching patterns of different dimensions presented in different rows. For example, in the illustrated example, row 805 includes a count of the number of occurrences of activities matching one or more three-point, two-dimensional patterns (i.e., "2032"), while row 810 includes a count of the number of occurrences of activities matching one or more four-point, three-dimensional patterns (i.e., "877"). Because the occurrence of a pattern indicates that the activity has a non-random order, the counts also provide a generalized characterization of the overall complexity of the activity pattern. This can be done, for example, in process 400 ( Figure 4 Each window defined at 410 in ) forms a table similar to table 800.
[0077] Although table 800 includes a separate row and separate entry for each type of activity pattern, this is not necessarily the case. For example, one of the multiple counts (e.g., the count of a simpler pattern) can be omitted from table 800 and from the complexity determination. As another example, some implementations may include a single row or entry that includes counts of occurrences of multiple activity patterns.
[0078] although Figure 8 While the number counts are presented in table 800, this need not be the case. For example, the number counts may be presented as a vector (e.g., <2032, 877, 133, 66, 48, ... >). Regardless of how the counts are presented, in some implementations, the counts may be expressed in binary and may be compatible with digital data processing infrastructure.
[0079] In some implementations, the counts of the number of occurrences of a pattern can be weighted or combined to determine the degree or complexity of the ranking, e.g., in process 400 ( Figure 4 ) at 420. For example, the Eulerian characteristic can provide an approximation of the complexity of the activity and is given by the following equation:
[0080] S0-S1+S2-S3+... Equation 1
[0081] where S n is the number of occurrences of a pattern of n points (i.e., a pattern of dimension n - 1). The pattern can be, for example, a directed clique pattern 500 Figure 5 ).
[0082] As another example of how the number of occurrences of a pattern can be weighted to determine the degree or complexity of an ordering, in some implementations, the occurrences of a pattern can be weighted based on the weights of the links that are active. In more detail, as previously discussed, the strength of connections between nodes in an artificial neural network can vary, for example, due to the degree of activity of the connections during training. The occurrences of a pattern of activity along a set of relatively strong links can be weighted differently than the occurrences of that same pattern of activity along a set of relatively weak links. For example, in some implementations, the sum of the weights of the links that are active can be used to weight the occurrences.
[0083] In some implementations, the Euler characteristic or other measures of complexity can be normalized by the total number of patterns that can be matched within a particular window and / or the total number of patterns that a given network can form given the structure of the network. One example of normalizing by the total number of patterns that a network can form is given below in Equations 2, 3.
[0084] In some implementations, occurrences of higher-dimensional patterns involving a larger number of nodes can be weighted more heavily than occurrences of lower-dimensional patterns involving a smaller number of nodes. For example, the probability of forming a directed clique decreases rapidly with increasing dimension. In particular, to form an n-clique from n + 1 nodes, it is required that (n + 1)n / 2 edges all be correctly oriented. This probability can be reflected in the weighting.
[0085] In some implementations, both the dimension and the directionality of a pattern can be used to weight the occurrences of the pattern and determine the degree of complexity of activity. For example, referring to Figure 6 , the occurrences of the five-point, four-dimensional pattern 515 can be weighted more heavily than the occurrences of the five-point, four-dimensional patterns 605, 610, in accordance with the difference in directionality of these patterns.
[0086] One example of using both the directionality and the dimension of a pattern to determine the degree of ordering or complexity of activity can be given by the following equation:
[0087]
[0088] where Sx active the number of occurrences of patterns of n points and ERN is a calculation for an equivalent random network (i.e., a network with the same number of nodes with random connections). Further, SC is given by the following equation:
[0089]
[0090] where S x silent the number of occurrences of patterns of n points when the recurrent artificial neural network is silent and can be considered to embody the total number of patterns that the network can form. In Equations 2, 3, the patterns can be, for example, directed clique patterns 500 Figure 5 ).
[0091] Figure 9 is a schematic illustration of a determination of a particular time of active patterns with distinguishable complexity. The determination illustrated in Equations 1-4 can be performed in isolation or in conjunction with other activities. For example, the determination can be performed at 425 in process 400 Figure 9 ). Figure 4
[0092] Figure 9 includes graphs 905 and 910. Graph 905 illustrates occurrences of patterns as a function of time along the x-axis. In particular, individual occurrences are schematically illustrated as vertical lines 906, 907, 908, 909. The occurrences of each row can be instances of activities matching the corresponding pattern or class of patterns. For example, the occurrences of the top row can be instances of activities matching pattern 505 Figure 5 ), the occurrences of the second row can be instances of activities matching pattern 510 Figure 5 ), the occurrences of the third row can be instances of activities matching pattern 515 Figure 5 ), and so on.
[0093] Graph 905 also includes dashed rectangles 915, 920, 925 that schematically depict different, time windows when the active patterns have distinguishable complexity. As shown, during the windows depicted by dashed rectangles 915, 920, 925, the activities in the recurrent artificial neural network match patterns of likelihood of complexity higher than outside those windows.
[0094] Graph 910 illustrates the complexity associated with the occurrences as a function of time along the x-axis. Graph 910 includes a first peak 930 of complexity coinciding with the window depicted by dashed rectangle 915, and a second peak 935 of complexity coinciding with the windows depicted by dashed rectangles 920, 925. As shown, the complexity illustrated by peaks 930, 925 is distinguishable from a baseline level 940 of complexity that can be considered to be the baseline level of complexity.
[0095] In some implementations, the time at which the output of the recurrent artificial neural network will be read coincides with the occurrence of an activity pattern having distinguishable complexity. For example, in the illustrative scenario of FIG. 9, the output of the recurrent artificial neural network can be read at peaks 930, 925, i.e., during the windows depicted by dashed rectangles 915, 920, 925. Figure 9
[0096] The identification of distinguishable levels of complexity in a recurrent artificial neural network is particularly beneficial when the input is a data stream. Examples of data streams include, for example, video or audio data. While a data stream has a beginning, it is often desirable to process information in the data stream that does not have a predefined relationship to the beginning of the data stream. By way of example, a neural network can perform object recognition, such as, for example, recognizing a bicyclist near a car. Such a neural network should be able to recognize the bicyclist regardless of when those bicyclists appear in the video stream, i.e., regardless of the time after the beginning of the video. Continuing this example, when a data stream is input into an object recognition neural network, any pattern of activity in the neural network will typically show a low or resting level of complexity. These low or resting levels of complexity are shown regardless of the continuous (or near-continuous) input of the stream of data into the neural network device. However, when the object of interest appears in the video stream, the complexity of the activity will become distinguishable and indicate the time at which the object was recognized in the video stream. Thus, the particular time of the distinguishable level of complexity of the activity can also act as a yes / no output as to whether the data in the data stream meets certain criteria.
[0097] In some implementations, not only is the particular time of the output of the recurrent artificial neural network given by the activity pattern having distinguishable complexity, but also the content of the output of the recurrent artificial neural network is given. In particular, the identity and activity of the nodes that participate in the activity commensurate with the activity pattern can be considered to be the output of the recurrent artificial neural network. Thus, the activity pattern that is identified can illustrate the result of the processing by the neural network, as well as the particular time at which this decision will be read.
[0098] The content of the decision can be expressed in a variety of different forms. For example, in some implementations and as discussed in further detail below, the content of the decision can be expressed as a binary vector or matrix of 1s and Os. Each number can indicate, for example, whether a pattern of activity exists for a predefined set of nodes and / or for a predefined duration of time. In such implementations, the content of the decision is expressed in binary and can be compatible with traditional digital data processing infrastructure.
[0099] Figure 10 is a flowchart of a process 1000 for encoding a signal using a recurrent artificial neural network based on the characterization of activity in the network. The signal can be encoded in a variety of different scenarios, such as, for example, transmission, encryption, and data storage. The process 1000 can be performed by a system of one or more data processing devices operating according to one or more machine-readable instruction sets. For example, the process 1000 can be performed by the same system of one or more computers executing software for implementing the recurrent artificial neural network used in the process 1000. In some instances, the process 1000 can be performed by the same data processing devices that perform the process 400. In some instances, the process 1000 can be performed by an encoder, for example, in a signal transmission system or a data storage system.
[0100] At 1005, the system performing the process 1000 inputs a signal into the recurrent artificial neural network. In some cases, the input of the signal is a discrete injection event. In other cases, the input signal is streamed into the recurrent artificial neural network.
[0101] At 1010, the system performing the process 1000 identifies one or more decision moments in the recurrent artificial neural network. For example, the system can identify the one or more decision moments by performing the process 400 Figure 4 ) to identify the one or more decision moments.
[0102] At 1015, the system performing the process 1000 reads the output of the recurrent artificial neural network. As discussed above, in some implementations, the content of the output of the recurrent artificial neural network is activity in the neural network that matches a pattern used to identify a decision point.
[0103] In some implementations, a separate "reader node" can be added to the neural network to identify the occurrence of a particular pattern of activity at a particular set of nodes and, thus, to read the output of the recurrent artificial neural network at 1015. The reader node can fire only when the activity at the particular set of nodes satisfies the particular temporal (and possibly also amplitude) criteria. For example, to read the pattern 505 Figure 2 、 Figure 3 ) at the nodes 104, 105, 106 Figure 5) occurs, the reader node may be connected to nodes 104, 105, 106 (or links 110 therebetween). The reader node itself may become active only when a pattern of activity involving nodes 104, 105, 106 (or their links) occurs.
[0104] The use of such reader nodes would eliminate the need to define time windows for the recurrent artificial neural network as a whole. In particular, separate reader nodes can be connected to different nodes and / or multiple nodes (or links between them). Separate reader nodes can be set to have customized responses (e.g., integrating different decay times in the excitation model) to identify different activity patterns. At 1020, the system performing process 1000 transmits or stores the output of the recurrent artificial neural network. The specific actions performed at 1020 can reflect the scenario in which process 1000 is being used. For example, in a scenario where secure or compressed communication is desired, the system performing process 1000 can transmit the output of the recurrent neural network to a receiver that can access the same or similar recurrent neural network. As another example, in a scenario where secure or compressed data storage is desired, the system performing process 1000 can record the output of the recurrent neural network in one or more machine-readable data storage devices for later access.
[0105] In some implementations, the complete output of the recurrent neural network is not transmitted or stored. For example, in implementations where the content of the output of the recurrent neural network is activity in the neural network matching a pattern indicating complexity in the activity, only activity matching a relatively complex or higher dimensional activity may be transmitted or stored. By way of example, referring to pattern 500 ( Figure 5 ), in some implementations, only activities that match patterns 515, 520, 525, and 530 are transmitted or stored, while activities that match patterns 505, 510 are ignored or discarded. In this way, lossy processing allows the amount of data transmitted or stored to be reduced at the expense of the integrity of the encoded information.
[0106] Figure 11is a flowchart of a process 1100 for decoding a signal using a recurrent artificial neural network based on a characterization of activity in the network. Signals can be decoded in a variety of different scenarios, such as, for example, signal reception, decryption, and reading data from storage. The process 1100 can be performed by a system of one or more data processing apparatus executing operations in accordance with one or more machine-readable instruction sets. For example, the process 1100 can be performed by the same system of one or more computers executing software for implementing a recurrent artificial neural network used in the process 1100. In some instances, the process 1100 can be performed by the same data processing apparatus that performs the process 400 and / or the process 1000. In some instances, the process 1100 can be performed by a decoder in, for example, a signal reception system or a decoder of a data storage system.
[0107] At 1105, the system performing the process 1100 receives at least a portion of an output of a recurrent artificial neural network. The particular action performed at 1105 can reflect the scenario in which the process 1100 is being used. For example, a system performing the process 1000 can receive a transmitted signal including an output of a recurrent artificial neural network or read a machine-readable data storage device storing an output of a recurrent artificial neural network.
[0108] At 1110, the system performing the process 1100 reconstructs an input to the recurrent artificial neural network from the received output. The reconstruction can be performed in a variety of different ways. For example, in some implementations, a second artificial neural network (recurrent or non-recurrent) can be trained to reconstruct an input into a recurrent neural network from the output received at 1105.
[0109] As another example, in some implementations, a decoder that has been trained using machine learning (including but not limited to deep learning) can be trained to reconstruct an input into a recurrent neural network from the output received at 1105.
[0110] As yet another example, in some implementations, an input into the same recurrent artificial neural network or into a similar recurrent artificial neural network can be permuted iteratively until an output of the recurrent artificial neural network matches the output received at 1105 to some degree.
[0111] In some implementations, the process 1100 can include receiving user input specifying a degree to which the input is to be reconstructed, and in response, adjusting the reconstruction accordingly at 1110. For example, the user input can specify that a full reconstruction is not needed. In response, the system executing the process 1100 adjusts the reconstruction. For example, in implementations in which the content of the output of the recurrent neural network is an activity in a neural network that matches a pattern indicative of a complexity in the activity, only the output that characterizes the activity matching a relatively more complex or higher dimensional activity will be used to reconstruct the input. By way of example, with reference to the pattern 500 Figure 5 ), in some implementations, only the activities matching the patterns 515, 520, 525, and 530 can be used to reconstruct the input, while the activities matching the patterns 505, 510 can be ignored or discarded. In this way, lossy reconstruction can be performed on a selective basis.
[0112] In some implementations, the processes 1000, 1100 can be used for peer-to-peer encrypted communications. In particular, both the sender (i.e., the encoder) and the receiver (i.e., the decoder) can be provided with the same recurrent artificial neural network. There are several ways in which the shared recurrent artificial neural network can be customized to ensure that third parties are unable to reverse engineer it and decrypt the signal, including:
[0113] — the structure of the recurrent artificial neural network
[0114] — the functional settings of the recurrent artificial neural network, including node states and edge weights,
[0115] — the size (or dimensionality) of the patterns, and
[0116] — a fraction of the patterns in each dimension.
[0117] These parameters can be thought of as multiple layers that together ensure the security of the transmission. Furthermore, in some implementations, the decision time instant can be used as a key to decrypt the signal.
[0118] Although the processes 1000, 1100 are presented in terms of encoding and decoding a single recurrent artificial neural network, the processes 1000, 1100 can also be applied in systems and processes that rely on multiple recurrent artificial neural networks. These recurrent artificial neural networks can operate in parallel or in series.
[0119] As one example of running in series, the output of a first recurrent artificial neural network can be used as input to a second recurrent artificial neural network. The resulting output of the second recurrent artificial neural network is a twice- encoded (or twice-encrypted) version of the input into the first recurrent artificial neural network. Such a series arrangement of recurrent artificial neural networks can be useful in situations where different parties have different levels of access to information, e.g., in a medical records system where patient identity information can be inaccessible to parties that will use and can access the remainder of the medical record.
[0120] As one example of running in parallel, the same information can be input into multiple different recurrent artificial neural networks. The different outputs of these neural networks can be used, e.g., to ensure that the input can be reconstructed with high fidelity.
[0121] While a number of implementations have been described, various modifications can be made. For example, while an application generally means that the activity in a recurrent artificial neural network should match an indicated ordering, this need not be the case. Rather, in some implementations, the activity in a recurrent artificial neural network can be commensurate with a pattern without necessarily showing activity that matches the pattern. For example, a recurrent neural network will show an increase in the likelihood of activity that will match a pattern can be considered a non-random ordering of activity.
[0122] As yet another example, in some implementations, different pattern sets can be tailored for use in characterizing activity in different recurrent artificial neural networks. The patterns can be tailored, e.g., according to the efficacy of the patterns in characterizing activity in different recurrent artificial neural networks. The efficacy can be quantified, e.g., based on the size of a table or vector of occurrence counts representing different patterns.
[0123] As yet another example, in some implementations, the patterns used to characterize activity in a recurrent artificial neural network can take into account the strength of connections between nodes. In other words, the patterns previously described herein treat all signaling activity between two nodes in a binary fashion (i.e., activity is present or not present). This need not be the case. Rather, in some implementations, it can be necessary to consider activity with a certain level or strength of connection as indicative of an ordered complexity in the activity of a recurrent artificial neural network to be commensurate with a pattern.
[0124] As yet another example, the content of the output of a recurrent artificial neural network can include patterns of activity that occur outside of a time window in which activity in the neural network has a distinguishable level of complexity. For example, the output of a recurrent artificial neural network read at 1015 and transmitted or stored at 1020 ( Figure 10 ) can include patterns of activity that occur outside of a time window in which activity in the neural network has a distinguishable level of complexity, e.g., the graph 905 ( Figure 9information that encodes the patterns of activity outside the dashed rectangles 915, 920, 925 in FIG. 9. By way of example, the output of a recurrent artificial neural network can characterize only the highest dimensional patterns of activity, regardless of when those patterns of activity occur. As another example, the output of a recurrent artificial neural network can characterize only the patterns of activity that surround a cavity, regardless of when those patterns of activity occur.
[0125] Figure 12 , Figure 13 and Figure 14 is a schematic illustration of a binary form or representation 1200 of a topology, such as, for example, a pattern of activity in a neural network. Figure 12 , Figure 13 and Figure 14 The topologies illustrated in
[0126] As illustrated, the binary representation 1200 includes bits 1205, 1207, 1211, 1293, 1294, 1297, and an additional, arbitrary number of bits (represented by ellipses “...”). For pedagogical purposes, the bits 1205, 1207, 1211, 1293, 1294, 1297... are illustrated as discrete rectangular shapes that are either filled or unfilled to indicate the binary value of the bit. In the schematic illustration, the representation 1200 superficially appears to be a one-dimensional vector of bits ( Figure 12 , Figure 13 ) or a two-dimensional matrix of bits ( Figure 14 ). However, the representation 1200 differs from a vector, matrix, or other ordered collection of bits in that the same information can be encoded regardless of the order of the bits— i.e., regardless of the position of the individual bits within the collection.
[0127] For example, in some implementations, each individual bit 1205, 1207, 1211, 1293, 1294, 1297... can represent the presence or absence of a topological feature— regardless of the position of that feature in the graph. By way of example, referring to Figure 2 , a bit such as the bit 1207 can indicate the presence or absence of the pattern 505 ( Figure 5the presence of a commensurate topological feature, regardless of whether the activity occurred between nodes 104, 105, 101 or between nodes 105, 101, 102. Thus, while each individual bit 1205, 1207, 1211, 1293, 1294, 1297... can be associated with a particular feature, the position of that feature in the graph does not need to be encoded, e.g., by the corresponding position of the bit in representation 1200. In other words, in some implementations, representation 1200 can provide only a homotopic reconstruction of the graph.
[0128] Furthermore, in other implementations, it can be that the position of individual bits 1205, 1207, 1211, 1293, 1294, 1297... does encode information such as, for example, the position of the feature in the graph. In these implementations, representation 1200 can be used to reconstruct the source graph. However, such encoding is not necessarily present.
[0129] In view of the fact that the presence or absence of a topological feature can be represented by a bit regardless of the position of the topological feature in the graph, Figure 1 At the beginning of representation 1200, bit 1205 occurs before bit 1207, which occurs before bit 1211. In contrast, in Figure 2 and Figure 3 representations 1200 have changed. However, the binary representation 1200 remains the same - as does the set of rules or algorithm that defines the process used to encode the information in binary representation 1200. As long as the correspondence between bits and features is known, the position of the bits in representation 1200 is irrelevant.
[0130] In more detail, each bit 1205, 1207, 1211, 1293, 1294, 1297... individually represents the presence or absence of a feature in the graph. The graph is a set of nodes and a set of edges between these nodes. The nodes can correspond to objects. Examples of objects can include, for example, artificial neurons in a neural network, individuals in a social network, etc. The edges can correspond to some relationship between the objects. Examples of relationships include, for example, a structural connection or activity along the connection. In the context of a neural network, artificial neurons can be related by structural connections between the neurons or by transmission of information along the structural connections. In the context of a social network, individuals can be related by a "friend" or other relationship connection or by transmission of information (e.g., posts) along such a connection. Thus, an edge can characterize a relatively long-lasting structural feature of a set of nodes or a relatively short-lived activity feature that occurs within a defined timeframe. Furthermore, an edge can be directed or undirected. A directed edge indicates directionality of a relationship between objects. For example, transmission of information from a first neuron to a second neuron can be represented by a directed edge that represents the direction of the transmission. As another example, in a social network, a relationship connection can indicate that a second user will receive information from a first user, but not that a first user will receive information from a second user. In topological terms, the graph can be expressed as a set of unit intervals [0, 1] where 0 and 1 are identified with respective nodes connected by an edge.
[0131] The features whose presence or absence is indicated by bits 1205, 1207, 1211, 1293, 1294, 1297 can be, for example, one node, a set of nodes, a set of nodes of a set of nodes, a set of edges, a set of edges of a set of edges, and / or additional hierarchically more complex features (e.g., a set of sets of nodes of a set of nodes). Bits 1205, 1207, 1211, 1293, 1294, 1297 generally represent the presence or absence of a feature at different hierarchical levels. For example, bit 1205 can represent the presence or absence of one node, while bit 1205 can represent the presence or absence of a set of nodes.
[0132] In some implementations, bits 1205, 1207, 1211, 1293, 1294, 1297 can represent features in the graph that have some characteristic above a threshold level. For example, bits 1205, 1207, 1211, 1293, 1294, 1297 can not only represent the presence of activity in a set of edges, but also that this activity is weighted above or below a threshold level. The weight can, for example, embody training of the neural network device for a particular purpose or can be an inherent feature of the edge.
[0133] The above Figure 5 , Figure 6 and Figure 8Features whose presence or absence can be represented by bits 1205, 1207, 1211, 1293, 1294, 1297... are exemplified.
[0134] Directed simplices in the collections 500, 600, 700 view a functional or structural graph as a topological space with nodes as points. Structures or activities commensurate with simplices in the collections 500, 600, 700 involve one or more nodes and links can be represented by bits regardless of the identity of the particular nodes and / or links participating in the activity.
[0135] In some implementations, only some patterns of structures or activities are identified and / or some portion of the identified structures or activities are discarded or otherwise ignored. For example, with reference to Figure 5 Structures or activities commensurate with the five-point, four-dimensional simplex pattern 515 inherently include structures or activities commensurate with the four-point, three-dimensional and three-point, two-dimensional simplex patterns 510, 505. For example, Figure 5 Points 0, 2, 3, 4 and points 1, 2, 3, 4 in the four-dimensional simplex pattern 515 of FIG. 5B are commensurate with the three-dimensional simplex pattern 510. In some implementations, simplex patterns containing fewer points - and thus having lower dimensionality - can be discarded or otherwise ignored.
[0136] As another example, only some patterns of structures or activities need to be identified. For example, in some implementations, only patterns having an odd number of points (3, 5, 7,...) or even number of dimensions (2, 4, 6,...) are used.
[0137] Returning to Figure 12 , Figure 13 , Figure 14 Features whose presence or absence are represented by bits 1205, 1207, 1211, 1293, 1294, 1297... can not be independent of one another. By way of explanation, if bits 1205, 1207, 1211, 1293, 1294, 1297 represent the presence or absence of zero-dimensional simplices - each reflecting the presence or activity of a single node - then bits 1205, 1207, 1211, 1293, 1294, 1297 are independent of one another. However, if bits 1205, 1207, 1211, 1293, 1294, 1297 represent the presence or absence of higher-dimensional simplices - each reflecting the presence or activity of multiple nodes - then the information encoded by the presence or absence of each individual feature can not be independent of the presence or absence of the other features.
[0138] Figure 15One example of how the presence or absence of features corresponding to different bits are not independent of one another is illustrated schematically. In particular, a sub-diagram 1500 is illustrated that includes four nodes 1505, 1510, 1515, 1520 and six directed edges 1525, 1530, 1535, 1540, 1545, 1550. In particular, edge 1525 points from node 1525 to node 1510, edge 1530 points from node 1515 to node 1505, edge 1535 points from node 1520 to node 1505, edge 1540 points from node 1520 to node 1510, edge 1545 points from node 1515 to node 1510, and edge 1550 points from node 1515 to node 1520.
[0139] A single bit in representation 1200 (e.g., filled bit 1207 in Figure 12 , Figure 13 , Figure 14 may indicate the presence of a directed three-dimensional simplex. For example, such a bit can indicate the presence of a three-dimensional simplex formed by nodes 1505, 1510, 1515, 1520 and edges 1525, 1530, 1535, 1540, 1545, 1550. A second bit in representation 1200 (e.g., filled bit 1293 in Figure 12 , Figure 13 , Figure 14 may indicate the presence of a directed two-dimensional simplex. For example, such a bit can indicate the presence of a two-dimensional simplex formed by nodes 1515, 1505, 1510 and edges 1525, 1530, 1545. In this simple example, the information encoded by bit 1293 is completely redundant together with the information encoded by bit 1207.
[0140] Note that the information encoded by bit 1293 can also be redundant together with the information encoded by yet another bit. For example, the information encoded by bit 1293 would be redundant together with both a third bit and a fourth bit indicating the presence of additional directed two-dimensional simplices. Examples of these simplices are formed by nodes 1515, 1520, 1510 and edges 1540, 1545, 1550 and by nodes 1520, 1505, 1510 and edges 1525, 1535, 1540.
[0141] Figure 16Schematically illustrates another example of how the presence or absence of features corresponding to different bits are not independent of each other. In particular, a subgraph 1600 is illustrated that includes four nodes 1605, 1610, 1615, 1620 and five directed edges 1625, 1630, 1635, 1640, 1645. The nodes 1505, 1510, 1515, 1520 and the edges 1625, 1630, 1635, 1640, 1645 generally correspond to the subgraph 1500 ( Figure 15 ) in the subgraph 1500 and the edges 1525, 1530, 1535, 1540, 1545. However, in contrast to the subgraph 1500 in which the nodes 1515, 1520 are connected by the edge 1550, the nodes 1615, 1620 are not connected by an edge.
[0142] Represents a single bit in 1200 (e.g., Figure 12 、 Figure 13 、 Figure 14 The unfilled bit 1205 in 1200 may indicate the absence of a directed three-dimensional simplex (such as, for example, a directed three-dimensional simplex containing nodes 1605, 1610, 1615, 1620). Figure 12 、 Figure 13 、 Figure 14 The presence of a two-dimensional simplex can be indicated by the filled bits 1293 in the representation 1200. An exemplary directed two-dimensional simplex is formed by nodes 1615, 1605, 1610 and edges 1625, 1630, 1645. This combination of filled bits 1293 and unfilled bits 1205 provides information indicating the presence or absence of other features (and the state of other bits) that may or may not be present in the representation 1200. In particular, the combination of the absence of a directed three-dimensional simplex and the presence of a directed two-dimensional simplex indicates that at least one edge is not present in:
[0143] a) A possible directed two-dimensional simplex formed by nodes 1615, 1620, 1610 or
[0144] b) A possible directed two-dimensional simplex formed by nodes 1620, 1605, 1610.
[0145] Therefore, the state of the bit representing the presence or absence of any of these possible simplexes is not independent of the state of bits 1205, 1293.
[0146] While these examples have been discussed in terms of features having different numbers of nodes and hierarchical relationships, this is not necessarily the case. For example, a representation 1200 comprising a set of bits corresponding only to the presence or absence of, for example, a three-dimensional simplex is possible.
[0147] Using separate bits to represent the presence or absence of a feature in a graph yields certain properties. For example, the encoding of information is fault-tolerant, and provides for "graceful degradation" of the encoded information. In particular, loss of a particular bit (or group of bits) can increase the uncertainty about the presence or absence of a feature. However, it will still be possible to evaluate the likelihood of the presence or absence of a feature from other bits indicating the presence or absence of neighboring features.
[0148] Likewise, as the number of bits increases, the certainty about the presence or absence of a feature increases.
[0149] As another example, as discussed above, the ordering or arrangement of bits is irrelevant to the isomorphic reconstruction of the graph represented by the bits. All that is required is a known correspondence between the bits and particular nodes / structures in the graph.
[0150] In some implementations, the pattern of activity in a neural network can be encoded in a representation 1200 Figure 12 , Figure 13 and Figure 14 Typically, the pattern of activity in a neural network is a result of many features of the neural network, such as, for example, structural connections between nodes of the neural network, weights between nodes, and a large number of other possible parameters. For example, in some implementations, the neural network can have been trained prior to the encoding of the pattern of activity in the representation 1200.
[0151] However, regardless of whether the neural network is untrained or trained, the response pattern of activity for a given input can be considered a "representation" or "abstraction" of that input in the neural network. Thus, although the representation 1200 can appear to be a straightforward-appearing set of (in some cases, binary) numbers, each of the numbers can encode a relationship or correspondence between a particular input and the related activity in the neural network.
[0152] Figure 17 , Figure 18 , Figure 19 , Figure 20is a schematic illustration of the use of representations of occurrences of topologies in activity in a neural network in four different classification systems 1700, 1800, 1900, 2000. Classification systems 1700, 1800 each classify representations of patterns of activity in a neural network as part of the classification of the input. Classification systems 1900, 2000 each approximate the classification of representations of patterns of activity in a neural network as part of the classification of the input. In classification systems 1700, 1800, the represented patterns of activity occur in a source neural network device 1705 that is part of classification systems 1700, 1800 and are read from source neural network device 1705. In contrast, in classification systems 1900, 2000, the approximated represented patterns of activity occur in a source neural network device that is not part of classification systems 1700, 1800. However, the approximation of those representations of patterns of activity is read from an approximator 1905 that is part of classification systems 1900, 2000.
[0153] In more detail, turning to Figure 17 Classification system 1700 includes a source neural network 1705 and a linear classifier 1710. Source neural network 1705 is a neural network device configured to receive input and to present representations of occurrences of topologies in activity in source neural network 1705. In the illustrated implementation, source neural network 1705 includes an input layer 1715 that receives input. However, this need not be the case. For example, in some implementations, some or all of the input can be injected into different layers and / or edges or nodes throughout source neural network 1705.
[0154] Source neural network 1705 can be any of a variety of different types of neural networks. Generally, source neural network 1705 is a recurrent neural network, such as, for example, a recurrent neural network that mimics a biological system. In some cases, source neural network 1705 can mimic biological systems to a degree in terms of morphology, chemistry, and other features. Generally, source neural network 1705 is implemented on one or more computing devices (e.g., a supercomputer) with a relatively high level of computing performance. In such cases, classification system 1700 will generally be a distributed system in which remote classifier 1710 communicates with source neural network 1705, e.g., via a data communication network.
[0155] In some implementations, source neural network 1705 can be untrained and the represented activity can be the inherent activity of source neural network 1705. In other implementations, source neural network 1705 can be trained and the represented activity can embody this training.
[0156] The representations read from source neural network 1705 can be representations 1200 such asFigure 12 、 Figure 13 、 Figure 14 The representation of the pattern of activity in the source neural network 1705 can be read from the source neural network 1705 in a variety of ways. For example, in the illustrated example, the source neural network 1705 includes a "reader node" that reads the pattern of activity between other nodes in the source neural network 1705. In other implementations, the activity in the source neural network 1705 is read by a data processing component that is programmed to monitor the relatively highly ordered pattern of activity in the source neural network 1705. In other implementations, the source neural network 1705 can include an output layer from which the representation 1200 can be read, for example, when the source neural network 1705 is implemented as a feedforward neural network.
[0157] The linear classifier 1710 is a device that classifies objects, that is, the representation of the pattern of activity in the source neural network 1705, based on a linear combination of features of the objects. The linear classifier 1710 includes an input 1720 and an output 1725. The input 1720 is coupled to receive the representation of the pattern of activity in the source neural network 1705. In other words, the representation of the pattern of activity in the source neural network 1705 is a feature vector that represents features of an input to the source neural network 1705 that are used by the linear classifier 1710 to classify the input. The linear classifier 1710 can receive the representation of the pattern of activity in the source neural network 1705 in a variety of ways. For example, the representation of the pattern of activity can be received as a discrete event or as a continuous stream over a real-time or non-real-time communication channel.
[0158] The output 1725 is coupled to output the results of the classification from the linear classifier 1710. In the illustrated implementation, the output 1725 is schematically illustrated as a parallel port having multiple channels. This need not be the case. For example, the output 1725 can output the results of the classification through a serial port or a port having combined parallel and serial capabilities.
[0159] In some implementations, the linear classifier 1710 can be implemented on one or more computing devices having relatively limited computing performance. For example, the linear classifier 1710 can be implemented on a personal computer or a mobile computing device such as a smart phone or a tablet computer.
[0160] In Figure 18 In the illustrated implementation, the classification system 1800 includes a source neural network 1705 and a neural network classifier 1810. The neural network classifier 1810 is a neural network device that classifies objects, that is, representations of patterns of activity in the source neural network 1705, based on a nonlinear combination of features of the objects. In the illustrated implementation, the neural network classifier 1810 is a feedforward network that includes an input layer 1820 and an output layer 1825. As with the linear classifier 1710, the neural network classifier 1810 can receive representations of patterns of activity in the source neural network 1705 in a wide variety of ways. For example, the representations of patterns of activity can be received as discrete events or as a continuous stream over a real-time or non-real-time communication channel.
[0161] In some implementations, the neural network classifier 1810 can perform inference on one or more computing devices with relatively limited computing performance. For example, the neural network classifier 1810 can be implemented on a personal computer or a mobile computing device such as a smartphone or tablet computer, for example, in a neural processing unit of such a device. Like the classification system 1700, the classification system 1800 will typically be a distributed system in which a remote neural network classifier 1810 communicates with a source neural network 1705, for example, via a data communication network.
[0162] In some implementations, the neural network classifier 1810 can be, for example, a deep neural network such as a convolutional neural network that includes convolutional layers, pooling layers, and fully connected layers. The convolutional layers can generate feature maps, for example, using linear convolutional filters and / or nonlinear activation functions. The pooling layers reduce the number of parameters and control overfitting. The computations performed by different layers in the image classifier 1820 can be defined in different ways in different implementations of the image classifier 1820.
[0163] In Figure 19 In the illustrated implementation, the classification system 1900 includes a source approximator 1905 and a linear classifier 1710. As discussed further below, the source approximator 1905 is a relatively simple neural network that is trained to receive input, at an input layer 1915 or elsewhere, and output a vector that approximates a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network. For example, the source approximator 1905 can be trained to approximate a recurrent source neural network such as, for example, a recurrent neural network that mimics a biological system and includes degrees of the morphology, chemistry, and other features of the biological system. In the illustrated implementation, the source approximator 1905 includes an input layer 1915 and an output layer 1920. The input layer 1915 can be coupled to receive input data. The output layer 1920 is coupled to output an approximation of a representation of activity within a neural network device for receipt by the input 1720 of the linear classifier. For example, the output layer 1920 can output a representation 1200Figure 12 , Figure 13 , Figure 14 ) of the approximation 1200'. Moreover, Figure 17 , Figure 18 the approximation 1200' of the representation 1200 schematically illustrated in Figure 19 , Figure 20 the approximation 1200' of the representation 1200 schematically illustrated in
[0164] Generally, the source approximator 1905 can perform inference on one or more computing devices with relatively limited computing performance. For example, the source approximator 1905 can be implemented on a personal computer or a mobile computing device such as a smartphone or tablet computer, e.g., in a neural processing unit of such a device. Generally and in contrast to the classification systems 1700, 1800, the classification system 1900 generally will be housed within a single housing, e.g., with the source approximator 1905 and the linear classifier 1710 implemented on the same data processing device or on data processing devices coupled through hardwired connections.
[0165] In Figure 20 , the classification system 2000 includes the source approximator 1905 and a neural network classifier 1810. The output layer 1920 of the source approximator 1905 is coupled to output the approximation 1200' of the representation of activity within the neural network device for receipt by the input 1820 of the neural network classifier 1810. Despite any differences between the approximation 1200' and the representation 1200, the neural network classifier 1810 still can classify the approximation 1200'. Generally and like the classification system 1900, the classification system 1900 generally will be housed within a single housing, e.g., with the source approximator 1905 and the neural network classifier 1810 implemented on the same data processing device or on data processing devices coupled through hardwired connections.
[0166] Figure 21is a schematic illustration of an edge device 2100 that includes a local artificial neural network that can be trained using a representation of the occurrence of a topology corresponding to activity in a source neural network. In this scenario, the local artificial neural network can be, for example, an artificial neural network that is executed entirely on one or more local processors that do not require a communication network to exchange data. Typically, the local processors will be connected by hardwired connections. In some instances, the local processors can be housed within a single enclosure, such as a single personal computer or a single handheld, mobile device. In some instances, the local processors can be controlled and accessed by a single individual or a limited number of individuals. In effect, by training a second, simpler and / or less highly trained but more idiosyncratic neural network using a representation of the occurrence of a topology in a more complex source neural network (e.g., using supervised learning or reinforcement learning techniques), an individual with limited computing resources and a limited number of training samples can train a neural network as needed. Storage requirements and computational complexity during training are reduced, and resources like battery life are conserved.
[0167] In the illustrated implementation, edge device 2100 is schematically illustrated as a security camera device that includes an optical imaging system 2110, image processing electronics 2115, a source approximator 2120, a representation classifier 2125, and a communication controller and interface 2130.
[0168] Optical imaging system 2110 can include, for example, one or more lenses (or even pinholes) and a CCD device. Image processing electronics 2115 can read the output of optical imaging system 2110 and can generally perform basic image processing functions. Communication controller and interface 2130 is a device configured to control the flow of information to and from device 2100. Among the operations that communication controller and interface 2130 can perform, as discussed further below, are transmitting images of interest to other devices and receiving training information from other devices. Accordingly, communication controller and interface 2130 can include both a data transmitter and receiver that can communicate over, for example, a data port 2135. Data port 2135 can be a wired port, a wireless port, an optical port, etc.
[0169] Source approximator 2120 is a relatively simple neural network that is trained to output a vector that approximates a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network. For example, source approximator 2120 can be trained to approximate a recurrent source neural network, such as, for example, a recurrent neural network that mimics a biological system and includes a degree of the morphology, chemistry, and other features of the biological system.
[0170] The representation classifier 2125 is a linear classifier or a neural network classifier that is coupled to receive approximations of representations of patterns of activity in the source neural network from the source approximator 2120 and output a classification result. The representation classifier 2125 can be, for example, a deep neural network, such as a convolutional neural network that includes convolutional layers, pooling layers, and fully connected layers. The convolutional layers can generate feature maps, e.g., using linear convolutional filters and / or non-linear activation functions. The pooling layers reduce the number of parameters and control overfitting. The computations performed by different layers in the representation classifier 2125 can be defined in different ways in different implementations of the representation classifier 2125.
[0171] In some implementations, in operation, the optical imaging system 2110 can generate raw digital images. The image processing electronics 2115 can read the raw images and will typically perform at least some basic image processing functions. The source approximator 2120 can receive images from the image processing electronics 2115 and perform an inference operation to output a vector that approximates a representation of a topology that occurs in patterns of activity in a relatively complex neural network. This approximation vector is input into the representation classifier 2125, which determines whether the approximation vector satisfies one or more sets of classification criteria. Examples include facial recognition and other machine vision operations. In the event that the representation classifier 2125 determines that the approximation vector satisfies one set of classification criteria, the representation classifier 2125 can instruct the communication controller and interface 2130 to transmit information about the image. For example, the communication controller and interface 2130 can transmit the image itself, the classification, and / or other information about the image.
[0172] Sometimes, it can be desirable to change the classification process. In these cases, the communication controller and interface 2130 can receive a training set. In some implementations, the training set can include raw or processed image data and representations of topologies that occur in patterns of activity in a relatively complex neural network. Such a training set can be used to retrain the source approximator 2120, e.g., using supervised learning or reinforcement learning techniques. In particular, the representations are used as target answer vectors and the desired outcome of the source approximator 2120 processing the raw or processed image data.
[0173] In other implementations, the training set can include representations of topologies that occur in patterns of activity in a relatively complex neural network and desired classifications of those representations of topologies. Such a training set can be used to retrain the neural network representation classifier 2125, e.g., using supervised learning or reinforcement learning techniques. In particular, the desired classifications are used as target answer vectors and the desired outcome of the representation classifier 2125 processing the representations of topologies.
[0174] Regardless of whether the source approximator 2120 or the representation classifier 2125 is retrained, the inferencing operations at the device 2100 can readily adapt to changing circumstances and targets without requiring large training data sets and time-intensive and computationally-intensive iterative training.
[0175] Figure 22 is a schematic illustration of a second edge device 2200 that includes a local artificial neural network that can be trained using representations of occurrences of topologies corresponding to activity in a source neural network. In the illustrated implementation, the second edge device 2200 is schematically illustrated as a mobile computing device such as a smartphone or tablet computer. The device 2200 includes an optical imaging system (e.g., on the back of the device 2200, not shown), image processing electronics 2215, a representation classifier 2225, a communication controller and interface 2230, and a data port 2235. These components can have features and perform actions corresponding to the actions of the optical imaging system 2110, the image processing electronics 2115, the representation classifier 2125, the communication controller and interface 2130, and the data port 2135 in the device 2100 Figure 21 ) with corresponding modifications.
[0176] The illustrated implementation of the device 2200 additionally includes one or more additional sensors 2240 and a multi-input source approximator 2245. The sensors 2240 can sense one of a number of features of the environment around the device 2200 or of the device 2200 itself. For example, in some implementations, the sensors 2240 can be an accelerometer that senses acceleration experienced by the device 2200. As another example, in some implementations, the sensors 2240 can be an acoustic sensor such as a microphone that senses noise in the environment of the device 2200. Yet another example of sensors 2240 includes chemical sensors (e.g., "artificial noses," etc.), humidity sensors, radiation sensors, etc. In some cases, the sensors 2240 are coupled to processing electronics that can read the output of the sensors 2240 (or other information such as, for example, a contact list or a map) and perform basic processing functions. Thus, different implementations of the sensors 2240 can have different "modalities" as the physical parameters of the physical sensing vary from sensor to sensor.
[0177] The multi-input source approximator 2245 is a relatively simple neural network that is trained to output a vector that approximates a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network. For example, the multi-input source approximator 2245 can be trained to approximate a recurrent source neural network such as, for example, a recurrent neural network that mimics a biological system and includes degrees of the morphology, chemistry, and other features of the biological system.
[0178] Unlike source approximator 2120, multi-input source approximator 2245 is coupled to receive raw or processed sensor data from multiple sensors and return an approximation of a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network based on that data. For example, multi-input source approximator 2245 can receive processed image data from image processing electronics 2215 as well as acoustic, acceleration, chemical, or other data from one or more sensors 2240. Multi-input source approximator 2245 can be, for example, a deep neural network such as a convolutional neural network that includes convolutional layers, pooling layers, and fully connected layers. The computations performed by different layers in multi-input source approximator 2245 can be specific to a single type of sensor data or multiple forms of sensor data.
[0179] Regardless of the particular organization of multi-input source approximator 2245, multi-input source approximator 2245 will return an approximation based on raw or processed sensor data from multiple sensors. For example, processed image data from image processing electronics 2215 and acoustic data from microphone sensors 2240 can be used by multi-input source approximator 2245 to approximate a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network that receives the same data.
[0180] At times, it can be desirable to change the classification process at device 2200. In these cases, communication controller and interface 2230 can receive a training set. In some implementations, the training set can include raw or processed images, sounds, chemical, or other data as well as representations of topologies that occur in a pattern of activity in a relatively more complex neural network. Such a training set can be used to retrain multi-input source approximator 2245, for example, using supervised or reinforcement learning techniques. In particular, the representations are used as target answer vectors and the desired outcome is for multi-input source approximator 2245 to process the raw or processed image or sensor data.
[0181] In other implementations, the training set can include representations of topologies that occur in a pattern of activity in a relatively more complex neural network as well as desired classifications of those representations of topologies. Such a training set can be used to retrain neural network representation classifier 2225, for example, using supervised or reinforcement learning techniques. In particular, the desired classifications are used as target answer vectors and the desired outcome is for neural network representation classifier 2225 to process the representations of topologies.
[0182] Regardless of whether the multi-input source approximator 2245 or the representation classifier 2225 is retrained, the inferencing operations at the device 2200 can readily adapt to changing circumstances and targets without requiring large training data sets and iterative training that is time intensive and computationally intensive.
[0183] Figure 23 is an illustrative example of a system 2300 in which a local neural network can be trained using a representation of the occurrence of a topology corresponding to activity in a source neural network. The target neural network is implemented on a relatively simple, less expensive data processing system, while the source neural network can be implemented on a relatively complex, more expensive data processing system.
[0184] The system 2300 includes a variety of local neural network devices 2305, telephone base stations 2310, wireless access points 2315, server systems 2320, and one or more data communication networks 2325.
[0185] The local neural network devices 2305 are devices that are configured to process data using a computationally less intensive target neural network. As illustrated, the local neural network devices 2305 can be implemented as any one of a mobile computing device, a video camera, an automobile, or a large variety of other appliances, fixtures, and mobile components, as well as different brands and models of devices within each category. Different local neural network devices 2305 can belong to different owners. In some implementations, access to the data processing functions of the local neural network devices 2305 will typically be restricted to the owners and / or designations of the owners.
[0186] The local neural network devices 2305 can each include one or more source approximators that are trained to output a vector that approximates a representation of a topology that occurs in a pattern of activity in a relatively more complex neural network. For example, the relatively more complex neural network can be a recurrent source neural network, such as, for example, a recurrent neural network that mimics a biological system and includes degrees of the morphology, chemistry, and other features of the biological system.
[0187] In some implementations, in addition to processing data using the source approximator, the local neural network device 2305 can be programmed to retrain the source approximator using representations of topologies that occur in patterns of activity in a relatively more complex neural network as target answer vectors. For example, the local neural network device 2305 can be programmed to perform one or more iterative training techniques (e.g., gradient descent or stochastic gradient descent). In other implementations, the source approximator in the local neural network device 2305 is trainable by, for example, a specialized training system or by a training system installed on a personal computer that can interact with the local neural network device 2305 to train the source approximator.
[0188] Each local neural network device 2305 includes one or more wireless or wired data communication components. In the illustrated implementation, each local neural network device 2305 includes at least one wireless data communication component, such as a mobile telephone transceiver, a wireless transceiver, or both. The mobile telephone transceiver is capable of exchanging data with a telephone base station 2310. The wireless transceiver is capable of exchanging data with a wireless access point 2315. Each local neural network device 2305 can also be capable of exchanging data with a peer mobile computing device.
[0189] The telephone base station 2310 and the wireless access point 2315 are connected for data communication with one or more data communication networks 2325, and can exchange information with a server system 2320 over the networks. Thus, the local neural network devices 2305 are generally also in data communication with the server system 2320. However, this need not be the case. For example, in implementations in which the local neural network devices 2305 are trained by other data processing devices, the local neural network devices 2305 need only be in data communication with these other data processing devices at least once.
[0190] The server system 2320 is a system of one or more data processing devices programmed to perform data processing activities in accordance with one or more machine-readable instruction sets. The activities can include providing a training set to a training system for the mobile computing devices 2305. As discussed above, the training system can be internal to the mobile local neural network devices 2305 themselves or on one or more other data processing devices. The training set can include representations corresponding to occurrences of topologies of activity in a source neural network and corresponding input data.
[0191] In some implementations, the server system 2320 also includes the source neural network. However, this need not be the case, and the server system 2320 can receive a training set from yet another system that implements the source neural network.
[0192] In operation, after a server system 2320 receives a training set (from a source neural network discovered at the server system 2320 itself or elsewhere), the server system 2320 can provide the training set to a trainer of the training mobile computing device 2305. The training set can be used to train a source approximator in the target local neural network device 2305 so that the target neural network approximates the operation of the source neural network.
[0193] Figure 24 、 Figure 25 、 Figure 26 、 Figure 27 are schematic illustrations of representations of occurrences of topologies in activity in neural networks used in four different systems 2400, 2500, 2600, 2700. The systems 2400, 2500, 2600, 2700 can be configured to perform any of a number of different operations. For example, the systems 2400, 2500, 2600, 2700 can perform an object localization operation, an object detection operation, an object segmentation operation, an object detection operation, a prediction operation, an action selection operation, etc.
[0194] An object localization operation localizes an object within an image. For example, a bounding box can be constructed around the object. In some cases, object localization can be combined with object recognition in which the localized object is labeled with an appropriate designation.
[0195] An object detection operation classifies image pixels as belonging to a particular class (e.g., belonging to an object of interest) or not belonging to a particular class. Typically, object detection is performed by grouping pixels and forming a bounding box around the group of pixels. The bounding box should fit closely around the object.
[0196] Object segmentation typically assigns a class label to each image pixel. Thus, object segmentation is performed on a pixel-by-pixel basis and typically requires only a single label to be assigned to each pixel, as opposed to a bounding box.
[0197] A prediction operation seeks to draw conclusions outside the range of observed data. Although a prediction operation can seek to foretell future occurrences (e.g., based on information about past and current states), a prediction operation can also seek to draw conclusions about past and current states based on incomplete information about those states.
[0198] An action selection operation seeks to select an action based on a set of conditions. Action selection operations have traditionally been broken down into different approaches, such as symbol-based systems (classical planning), distributed solutions, and reactive or dynamic planning.
[0199] The classification systems 2400, 2500 each perform a desired operation on a representation of a pattern of activity in a neural network. The systems 2600, 2700 each perform a desired operation on an approximation of a representation of a pattern of activity in a neural network. In the systems 2400, 2500, the represented patterns of activity occur in a source neural network device 1705 that is part of the system 2400, 2500 and are read from that source neural network device 1705. In contrast, in the systems 2400, 2500, the approximated patterns of activity occur in a source neural network device that is not part of the system 2400, 2500. However, the approximations of the representations of those patterns of activity are read from an approximator 1905 that is part of the system 2400, 2500.
[0200] In more detail, turning to Figure 24 , the system 2400 includes a source neural network 1705 and a linear processor 2410. The linear processor 2410 is a device that performs an operation based on a linear combination of features of a representation of a pattern of activity in a neural network (or an approximation of such a representation). The operation can be, for example, an object localization operation, an object detection operation, an object segmentation operation, a prediction operation, an action selection operation, etc.
[0201] The linear processor 2410 includes an input 2420 and an output 2425. The input 2420 is coupled to receive a representation of a pattern of activity in the source neural network 1705. The linear processor 2410 can receive the representation of the pattern of activity in the source neural network 1705 in a variety of ways. For example, the representation of the pattern of activity can be received as discrete events or as a continuous stream over a real-time or non-real-time communication channel. The output 2525 is coupled to output a result of processing from the linear processor 2410. In some implementations, the linear processor 2410 can be implemented on one or more computing devices with relatively limited computing performance. For example, the linear processor 2410 can be implemented on a personal computer or a mobile computing device such as a smartphone or tablet computer.
[0202] Turning to Figure 24 , the system 2400 includes a source neural network 1705 and a linear processor 2410. The linear processor 2410 is a device that performs an operation based on a linear combination of features of a representation of a pattern of activity in a neural network (or an approximation of such a representation). The operation can be, for example, an object localization operation, an object detection operation, an object segmentation operation, a prediction operation, an action selection operation, etc.
[0203] The linear processor 2410 includes an input 2420 and an output 2425. The input 2420 is coupled to receive a representation of a pattern of activity in the source neural network 1705. The linear processor 2410 can receive the representation of the pattern of activity in the source neural network 1705 in a wide variety of ways. For example, the representation of the pattern of activity can be received as discrete events or as a continuous stream over a real-time or non-real-time communication channel. The output 2525 is coupled to output the results of processing from the linear processor 2410. In some implementations, the linear processor 2410 can be implemented on one or more computing devices with relatively limited computational performance. For example, the linear processor 2410 can be implemented on a personal computer or a mobile computing device such as a smartphone or tablet computer.
[0204] In Figure 25 The classification system 2500 includes, in this example, the source neural network 1705 and a neural network 2510. The neural network 2510 is a neural network device configured to perform an operation based on a non-linear combination of features of a representation of a pattern of activity in the neural network (or an approximation of such a representation). The operation can be, for example, an object localization operation, an object detection operation, an object segmentation operation, a prediction operation, an action selection operation, and so on. In the illustrated implementation, the neural network 2510 is a feedforward network including an input layer 2520 and an output layer 2525. Like the linear processor 2410, the neural network 2510 can receive the representation of the pattern of activity in the source neural network 1705 in a wide variety of ways.
[0205] In some implementations, the neural network 2510 can perform inference on one or more computing devices with relatively limited computational performance. For example, the neural network 2510 can be implemented on a personal computer or a mobile computing device such as a smartphone or tablet computer, e.g., in a neural processing unit of such a device. Like the system 2400, the system 2500 will typically be a distributed system in which the remote neural network 2510 communicates with the source neural network 1705, e.g., via a data communication network. In some implementations, the neural network 2510 can be, for example, a deep neural network such as a convolutional neural network.
[0206] In Figure 26 The system 2600 includes, in this example, the source approximator 1905 and the linear processor 2410. Despite any differences between the approximation 1200' and the representation 1200, the processor 2410 can still perform operations on the approximation 1200'.
[0207] In Figure 27 The system 2700 includes, in this example, the source approximator 1905 and the neural network 2510. Despite any differences between the approximation 1200' and the representation 1200, the neural network 2510 can still perform operations on the approximation 1200'.
[0208] In some implementations, the systems 2600, 2700 can be implemented on an edge device, such as, for example, the edge devices 2100, 2200( Figure 21 、 Figure 22 )) In some implementations, the systems 2600, 2700 can be implemented as part of a system, such as the system 2300( Figure 23 )) in which a local neural network can be trained using a representation corresponding to an occurrence of a topology of activity in a source neural network.
[0209] Figure 28 is an illustrative example of a reinforcement learning system 2800 that includes an artificial neural network that can be trained using a representation corresponding to an occurrence of a topology of activity in a source neural network. Reinforcement learning is a type of machine learning in which an artificial neural network learns from feedback about the results of actions taken in response to decisions of the artificial neural network. A reinforcement learning system moves from one state to another in an environment by performing an action and receiving information that characterizes a new state and a reward and / or a regret that characterizes the success (or lack of success) of the action. Reinforcement learning seeks to maximize total reward (or minimize regret) through a learning process.
[0210] In the illustrated implementation, the artificial neural network in the reinforcement learning system 2800 is a deep neural network 2805 (or other deep learning architecture) that is trained using a reinforcement learning method. In some implementations, the deep neural network 2805 can be a local artificial neural network, such as the neural network 2510( Figure 25 、 Figure 27 )) and is implemented locally, for example, on a car, an airplane, a robot, or other device. However, this need not be the case, and in other implementations, the deep neural network 2805 can be implemented on a system of networked devices.
[0211] In addition to the source approximator 1905 and the deep neural network 2805, the reinforcement learning system 2800 includes an actuator 2810, one or more sensors 2815, and a teacher module 2820. In some implementations, the reinforcement learning system 2800 also includes one or more additional data sources 2825.
[0212] The actuator 2810 is a device that controls a mechanism or system that interacts with the environment 2830. In some implementations, the actuator 2810 controls a physical mechanism or system (e.g., steering of a car or positioning of a robot). In other implementations, the actuator 2810 can control a virtual mechanism or system (e.g., a virtual game board or an investment portfolio). Thus, the environment 2830 can also be physical or virtual.
[0213] The sensor 2815 is a device that measures a characteristic of the environment 2830. At least some of the measurements characterize an interaction between the controlled mechanism or system and other aspects of the environment 2830. For example, when the actuator 2810 steers a car, the sensor 2815 can measure one or more of the speed, direction, and acceleration of the car, the proximity of the car to other features, and the response of other features to the car. As another example, when the actuator 2810 controls an investment portfolio, the sensor 2815 can measure the value and risk associated with the portfolio.
[0214] In general, both the source approximator 1905 and the teacher module 2820 are coupled to receive at least some of the measurements obtained by the sensor 2815. For example, the source approximator 1905 can receive measurement data at the input layer 1915, and output an approximation 1200’ of a representation of the topology that occurs in the pattern of activity in the source neural network.
[0215] The teacher module 2820 is a device configured to interpret the measurements received from the sensor 2815 and to provide rewards and / or regrets to the deep neural network 2805. Rewards are positive and indicate successful control of the mechanism or system. Regrets are negative and indicate unsuccessful or less-than-optimal control. In general, the teacher module 2820 also provides a characterization of the measurements and the rewards / regrets of reinforcement learning. In general, the characterization of the measurements is an approximation of a representation of the topology that occurs in the pattern of activity in the source neural network, such as the approximation 1200’. For example, the teacher module 2820 can read the approximation 1200’ output from the source approximator 1905, and pair the read approximation 1200’ with a corresponding reward / regret value.
[0216] In many implementations, reinforcement learning does not occur in system 2800 in real-time or during active control of actuators 2810 by deep neural network 2805. Rather, training feedback can be collected by teacher module 2820 and used for reinforcement training when deep neural network 2805 is not actively instructing actuators 2810. For example, in some implementations, teacher module 2820 can be remote from deep neural network 2805 and only intermittently in data communication with deep neural network 2805. Regardless of whether reinforcement learning is intermittent or continuous, deep neural network 2805 can evolve, e.g., to optimize rewards and / or reduce regrets using information received from teacher module 2820.
[0217] In some implementations, system 2800 also includes one or more additional data sources 2825. Source approximator 1905 can also receive data from data sources 2825 at input layer 1915. In these instances, approximation 1200’ would be caused by processing both sensor data and data from data sources 2825.
[0218] In some implementations, data collected by one reinforcement learning system 2800 can be used for training or reinforcement learning by other systems, including other reinforcement learning systems. For example, characterization of measurements and reward / regret values can be provided by teacher module 2820 to a data exchange system that collects such data from a variety of reinforcement learning systems and redistributes data among them. Moreover, as discussed above, characterization of measurements can be an approximation of a representation of a topology that emerges in patterns of activity in a source neural network, such as approximation 1200’.
[0219] The particular operations performed by reinforcement learning system 2800 will of course depend on the particular operational scenario. For example, in scenarios where source approximator 1905, deep neural network 2805, actuators 2810, and sensors 2815 are part of an automobile, deep neural network 2805 can perform object localization and / or detection operations while maneuvering the automobile.
[0220] In implementations where data collected by reinforcement learning system 2800 is used for training or reinforcement learning by other systems, reward / regret values and approximation 1200’ that characterize a state of an environment when performing object localization and / or detection operations can be provided to a data exchange system. The data exchange system can then distribute reward / regret values and approximation 1200’ to other reinforcement learning systems 2800 associated with other vehicles for use in reinforcement learning at those other vehicles. For example, reinforcement learning can be used to improve object localization and / or detection operations at a second vehicle using reward / regret values and approximation 1200’.
[0221] However, the operations learned at other vehicles need not be identical to the operations performed by the deep neural network 2805. For example, the reward / regret values based on travel time and approximations 1200' caused by the input of sensor data that characterizes unexpectedly wet roadways at locations identified, e.g., by GPS data source 2825, can be used for route planning operations at another vehicle.
[0222] Embodiments of the operations and techniques described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0223] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0224] The term "data processing apparatus" encompasses all kinds of apparatus, devices, and machines for processing data, by way of example, includes a programmable processor, a computer, a system on a chip, or multiple ones of the same. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0225] A computer program, which can also be referred to or be a part of a program, software, an application, an app, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are
[0226] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit), and
[0227] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0228] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0229] Although this description contains many specific implementation details, these should not be construed as limiting the scope of any invention or of what can be claimed, but merely as describing implementations of particular embodiments of the invention. Certain features that are described in the context of separate embodiments can also be implemented in combination with each other. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0230] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order nor requiring all illustrated operations be performed to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products.
[0231] Accordingly, particular implementations of the subject matter have been described. Other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0232] A number of embodiments have been described. Nevertheless, it will be understood that various modification can be made. For example, while representation 1200 is a binary representation, in which each bit individually represents the presence or absence of a feature in the chart, other representations of information are possible. For example, a vector or matrix of multi-valued, non-binary digits can be used to represent, for example, the presence or absence of features and possibly other characteristics of those features. One example of such a feature is a weight of an edge of an activity that makes up a feature.
[0233]
[0234] Accordingly, other implementations are within the scope of the following claims.
Claims
1. A device comprising: A neural network trained to: generating, in response to a first input, an approximation of a first representation of a topological structure in a pattern of activity occurring in a source neural network in response to the first input, generating, in response to a second input, an approximation of a second representation of the topology in the pattern of activity occurring in the source neural network in response to the second input, and In response to a third input, an approximation of a third representation of the topology in the pattern of activity occurring in the source neural network in response to the third input is generated.
2. The apparatus of claim 1, wherein the topological structures all comprise two or more nodes in the source neural network and one or more edges between the nodes. The apparatus of claim 1 , wherein the topological structure comprises a simplex. The device of claim 1 , wherein the topological structure surrounds a cavity.
5. A device comprising: A neural network coupled to input a representation of a topological structure in a pattern of activity occurring in a source neural network in response to a plurality of different inputs, wherein the neural network is trained to process the representation and produce a responsive output.
6. The apparatus of claim 5, wherein the topological structures all include three or more nodes in the source neural network and three or more edges between the nodes. The apparatus of claim 5 , wherein the topological structure comprises a simplex.
8. A method implemented by a neural network device, the method comprising: a representation of a topological structure in a pattern of activity in an input source neural network, wherein the activity is responsive to input into the source neural network; processing the representation, wherein the processing is consistent with training the neural network to process a different such representation of the topology in the pattern of activity in the source neural network; as well as A result of the processing of the representation is output.
9. The method of claim 8, wherein the topological structures all include three or more nodes in the source neural network and three or more edges between the nodes.
10. The method of claim 8, wherein the topological structure comprises a simplex.