Recursive coupling of artificial learning units

By coupling the two artificial intelligence units and using modulation functions and asymmetric design, parallel processing of fast classification and in-depth analysis is achieved, solving the problem of slow real-time response speed in the existing technology, and improving the efficiency and adaptability of the system.

CN114556365BActive Publication Date: 2025-07-11FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080048132.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-21
Filing Date
2020-03-09
Publication Date
2025-07-11
Estimated Expiration
2040-03-09

AI Technical Summary

Technical Problem

Existing artificial intelligence systems need to be retrained or used different methods when facing multitasking or complex environments, resulting in slow real-time response speed and difficulty in quickly adapting to changes in high-dimensional space.

Method used

By coupling at least two AI units, using the modulation function to affect the parameters of the second AI unit, asymmetric complexity and classification memory design are formed, so that the first unit rapid classification and the second unit in-depth analysis can achieve parallel or time-dependent alternating evaluation of rapid classification and in-depth analysis.

Benefits of technology

It improves the system's real-time response speed and efficiency in the face of complex environments, and can make quick decisions and conduct in-depth analysis in a short time, similar to the combination of instinctive and conscious reactions in biological systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114556365B_ABST
    Figure CN114556365B_ABST
Patent Text Reader

Abstract

The present invention relates to a method in a system formed by at least two artificial learning units, the method comprising: inputting input values at at least a first artificial learning unit and a second artificial learning unit; obtaining a first output value from the first artificial learning unit; forming one or more modulation functions based on the output value of the first artificial learning unit; applying the formed one or more modulation functions to one or more parameters of the second artificial learning unit, wherein the one or more parameters affect the processing of the input value and the obtaining of the output value in the second artificial learning unit; and finally, obtaining a second output value from the second artificial learning unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for recursively coupling artificial intelligence units. Background Art

[0002] Artificial intelligence is now playing an increasingly important role in countless application fields. This was initially understood as any automation of intelligent behavior and machine learning. However, such systems are usually designed and trained for specific tasks. This form of artificial intelligence (AI) is commonly referred to as "weak AI", and it essentially simulates intelligent behavior in a fixed domain based on the application of computations and algorithms. Examples include systems that can identify certain patterns, such as safety systems in vehicles, or systems that can learn and implement certain rules, such as chess. At the same time, these systems are basically useless in other fields and must be completely retrained for other applications or even trained using completely different methods.

[0003] For the practical implementation of such artificial intelligence units, among other things, neural networks are also used. In principle, these networks replicate the functions of biological neurons at an abstract level. Several artificial neurons or nodes are interconnected, and they can receive, process, and transmit signals to other nodes. For example, for each node, a function, weights, and a threshold are defined to determine whether and to what extent a signal is passed to the node.

[0004] Generally, these nodes are considered in layers, such that each neural network has at least one output layer. Prior to this, other layers may exist as so-called hidden layers, thus forming a multi-layer network. These input values or features can also be regarded as layers. The connections between the nodes of different layers are called edges, and these edges are usually assigned a fixed processing direction. Depending on the network topology, it can be specified which node of a layer is linked to which node of the next layer. In this case, all nodes can be connected, however, for example, due to a learned weight value of 0, the signal cannot be further processed via a specific node.

[0005] The processing of signals in a neural network can be described by various functions. Below, this principle is described for a single neuron or a node of a neural network. From several different input values arriving at the node, the network input is formed by a propagation function (also called an input function). Usually, this propagation function includes a simple weighted sum, where each input value is given a related weight. However, in principle, other propagation functions are also possible. In this case, the weight w i can be specified as the weight matrix of the network.

[0006] The activation function f can depend on the threshold aktis applied to the network input of the nodes formed in this way. This function represents the relationship between the network input and the activity level of the neuron. Various activation functions are known, for example, a simple binary threshold function, the output of which is thus zero below the threshold and an identity above the threshold; the sigmoid function; or a piecewise linear function with a given slope. These functions are specified in the design of the neural network. The activation function f akt results in an activation state. Optionally, an additional output function f out or output function can be specified, which is applied to the output of the activation function and determines the final output value of the node. However, usually the result of the activation function is simply passed directly here as the output value, i.e., the identity is used as the output function. Depending on the nomenclature used, the activation function and the output function can also be combined into a transfer function f trans .

[0007] The output value of each node is then passed to the next layer of the neural network as the input value of the corresponding node in that layer, and the corresponding steps are repeated in that layer for processing using the corresponding functions and weights of the node. Depending on the topology of the network, there may also be backward edges to the previous layer or back to the output layer, thus forming a recurrent network.

[0008] In contrast, the network can change the weights w i used to weight each input value to adjust the output value and operation of the entire network, which is considered the "learning" of the neural network.

[0009] For this purpose, error backpropagation in the network is usually used, i.e., the output value is compared with the expected value, and this comparison is used to adapt the input value to minimize the error. Then the error feedback can be used to adjust the various parameters of the network accordingly, such as the step size (learning rate) or the weights of these node input values. Similarly, these input values can also be re-evaluated.

[0010] These networks can then be trained in training mode. The learning strategy used is also decisive for the possible applications of the neural network. In particular, the following variants are distinguished:

[0011] In supervised learning, an input pattern or training data set is given, and the output of the network is compared with the expected value.

[0012] Unsupervised learning leaves the discovery of correlations or rules to the system, thus only specifying these patterns to be learned. An intermediate variant is partially supervised learning, in which data sets without predefined classifications can also be used.

[0013] In reinforcement learning or Q-learning, an agent is created that can receive rewards and penalties for actions and tries to maximize the received rewards based on this, thereby adjusting its behavior.

[0014] An important application of neural networks is to classify input data or inputs into certain classes or categories, i.e., to identify correlations and assignments. These classes can be trained based on known data and are at least partially predefined, or they can be developed or learned independently by the network.

[0015] The basic operations and further specific details of such neural networks are known in the art, for example from R. Schwaiger, J. Steinwender: Neuronale Netze programmieren mit Python, Rheinwerk Computing, Bonn 2019.

[0016] A generally applicable AI system leads to a high-dimensional space if it is not trained only for a special task, and thus requires a training and test data set that increases exponentially. As a result, a real-time response quickly becomes impossible. Therefore, it is usually attempted to reduce the dimension and complexity of such systems. Different methods are being taken. For example, the complexity can be reduced by linking data sets, reducing degrees of freedom, and / or inputting known knowledge into the system. As another method, relevant data or interdependent data sets can be at least partially separated, for example by a method such as principal component analysis. By applying a filtering method to the features, data that is not prominent or negatively prominent when training the network can be eliminated, for example, by applying a test such as the chi-square test or other statistical tests. Finally, the selection of this training data itself can be done as an optimization problem in the AI network. This involves combining this training data in such a way that it can train a new network as quickly as possible.

[0017] More advanced methods include the so-called "convolutional neural network", which applies convolution in at least one layer of a multi-layer fully connected network instead of a simple matrix transformation. For this purpose, for example, the so-called "deep-dream" method is known, especially in the field of image recognition, where the weights are kept optimal in the training network, but instead, the input values (e.g., the input image) are modified according to the output values in a feedback loop. Thus, for example, what the system thinks it can recognize gradually fades or is inserted. This name refers to the fact that dreamlike images are created in this process. In this way, the internal processes and their directions of the neural network can be traced.

[0018] Obviously, these methods still show great differences in human intelligence. Although databases, text files, images, and audio files can, in principle, be compared with the way facts, languages, speech logic, sounds, images, and sequences of events are stored and processed in the brain, a significant difference in human intelligence, for example, is that it associates all this data with sensory and unconscious "soft" classification situations. Summary of the Invention

[0019] According to the present invention, a method for recursively coupling at least two artificial intelligence units with the features of the independent patent claims is proposed. Advantageous embodiments are the subject matter of the dependent claims and the following description.

[0020] In particular, according to one embodiment, a method is proposed in a system of at least two artificial intelligence units, the method comprising inputting input values into at least a first artificial intelligence unit and a second artificial intelligence unit, thereby obtaining a first output value of the first artificial intelligence unit. Based on the output value of the first artificial intelligence unit, one or more modulation functions are formed, and then the modulation function is applied to one or more parameters of the second artificial intelligence unit. In this regard, one or more parameters are parameters that in some way affect the processing of the input values and the obtaining of the output values in the second artificial intelligence unit. Furthermore, the output value of the second artificial intelligence unit is obtained. For example, these output values can represent the modulated output values of the second unit. In this way, two artificial intelligence units are coupled together without using direct feedback of the input or output values. Instead, one of these units is used to affect the function of the second unit by modulating parameters related to certain functions, resulting in a new coupling that leads to different results or output values compared to traditional learning units. In addition, by processing the input values of the two coupled units, results can be obtained in a shorter time or through a more in-depth analysis than in traditional systems, thereby improving the overall efficiency. In particular, fast classification of the problem at hand and consideration of rapid changes are achieved.

[0021] In an exemplary embodiment, at least one of the artificial intelligence units may include a neural network having a plurality of nodes, in particular one of the learning units to which the modulation function is applied. In this case, one or more parameters can be at least one of the following: the weighting of the nodes of the neural network, the activation function of the nodes, the output function of the nodes, the propagation function of the nodes. These are the basic components of a neural network for determining how data is processed in the network. Instead of defining new weights or functions for these nodes, modulation functions can be used to superimpose and modulate the existing self-learning and / or predefined functions of the network, depending on the result of the first artificial intelligence unit. In this case, this application of the modulation function can especially be carried out outside the training phase of the network, thus achieving the active coupling of two or more networks in the processing of input values.

[0022] According to an exemplary embodiment, a classification memory can be assigned to each artificial intelligence unit, where each artificial intelligence unit classifies input values into one or more categories stored in the classification memory, where each category is configured as one or more subordinate levels, and where the number of categories and / or levels in the first classification memory of the first artificial intelligence unit is less than the number of categories and / or levels in the second classification memory of the second artificial intelligence unit. By making the classification memories of two coupled artificial intelligence units asymmetric in this way, parallel or time-dependent alternating evaluations of input values with different objectives can be performed, such as a combination of a quick classification of the input values and a deep, slower analysis of the input values.

[0023] Alternatively or in addition to the asymmetric design of the classification memory, the complexity of the first artificial intelligence and the second artificial intelligence unit can also be designed such that, for example, the degree of complexity of the first artificial intelligence unit is significantly lower than that of the second artificial intelligence unit. In this regard, in the case of neural networks, for example, the first neural network can have significantly fewer nodes and / or layers and / or edges than the second neural network.

[0024] In a possible embodiment, the application of at least one modulation function can cause a time-dependent superposition of the parameters of the second artificial intelligence unit, where the at least one modulation function can include one of the following characteristics: a periodic function, a step function, a function with a brief increase in amplitude, a damped oscillation function, a beat frequency function that is a superposition of multiple periodic functions, a continuously increasing function, a continuously decreasing function. Combinations or time series of such functions are also conceivable. In this way, the relevant parameters of the learning unit can be superposed in a time-dependent manner such that, for example, these output values "jump" into the search space due to the modulation, which would not be achievable without the superposition.

[0025] Optionally, the second artificial intelligence unit can include a second neural network having a plurality of nodes, where the application of at least one modulation function deactivates at least a portion of the nodes. This type of deactivation can also be considered a "dropout" based on the output values of the first artificial intelligence unit and can also provide newly explored search regions in the classification as well as a reduced computational overhead, thereby accelerating the execution of the method.

[0026] In an exemplary embodiment, the method can further include determining the currently dominant artificial intelligence unit in the system and forming an overall output value of the system from the output values of the currently dominant unit. In this way, two or more networks in the system can be meaningfully coupled and synchronized.

[0027] In this regard, for example, the first artificial intelligence unit can be determined as the leading unit at least before one or more output values of the second artificial intelligence unit become available. In this way, it can be ensured that the system is always decision-safe, i.e., the reaction of the system is always possible (after the first run of the first artificial intelligence unit), even before all existing artificial intelligence units of the system have fully classified the input values.

[0028] In this case, the comparison of the current input value with the previous input value can also be further applied by at least one of the artificial intelligence units of the system, where if this comparison results in a deviation above a predetermined input threshold, the first artificial intelligence unit is set as the leading unit. In this way, it can be ensured that significantly changed input values (e.g., detecting a new situation through a sensor) immediately react to the new evaluation of these input values.

[0029] Additionally or alternatively, the current output value of the first artificial intelligence unit can also be compared with the previous output value of the first artificial intelligence unit, where if this comparison results in a deviation above a predetermined output threshold, the first artificial intelligence unit is determined as the leading unit. By evaluating the deviation of these output values, for example, in the case of a deviation category compared to the previous run, it is also possible to indirectly detect changes in these input values that are of a certain significance, thus making the new classification meaningful.

[0030] In some embodiments, the system can also include a timer that stores one or more predetermined time periods associated with one or more artificial intelligence units, and the timer is arranged to measure the elapse of a predetermined time period associated with one of these artificial intelligence units each time. Such elements form the possibility of synchronizing different units of the system, for example, as well as the possibility of controlling when the output value of a certain unit is expected or further processed. Therefore, the timer can be used to define an adjustable waiting time period for the entire system, within which a decision should be available as the overall output value of the system. This time can be, for example, a few ms, such as 30 ms or 50 ms, and this time may especially depend on the existing topology of the artificial intelligence units and the existing computing units (processors or other data processing devices).

[0031] Thus, for example, once an artificial intelligence unit is determined as the leading unit, the measurement of the predetermined time period assigned to this artificial intelligence unit can be started. In this way, it can be ensured that the unit develops a solution within a predetermined time, or alternatively, even abort the data processing.

[0032] In one possible embodiment, if a first time period in a timer predetermined for a first artificial intelligence unit has elapsed, a second artificial intelligence unit can be set as the master unit. This ensures that a reaction based on the first artificial intelligence unit is already possible before the input values ​​are analyzed by the other artificial intelligence units, which then analyze the data in more detail by the second unit.

[0033] In any embodiment, the input value may include, for example, one or more of the following: a measurement value detected by one or more sensors, data detected by a user interface, data retrieved from a memory, data received via a communication interface, data output by a computing unit. Thus, the input value may be, for example, image data captured by a camera, audio data, position data, physical measurements such as speed, distance measurements, resistance values, and generally any value captured by a suitable sensor. Similarly, a user may enter or select data via a keyboard or screen, and may optionally be associated with other data such as sensor data.

[0034] Other advantages and embodiments of the present invention will be apparent from the description and drawings.

[0035] It is to be understood that the features mentioned above and those still to be explained below can be used not only in the combination indicated in each case but also in other combinations or alone, without departing from the scope of the present invention.

[0036] The invention is schematically illustrated with reference to exemplary embodiments shown in the drawings and is described below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 shows a combination of two artificial intelligence units coupled together;

[0038] Figure 2 Various exemplary modulation functions are schematically shown;

[0039] Figure 3 shows the application of a dropout method in two coupled neural networks according to one embodiment;

[0040] Figure 4 shows an example with an additional timer Figure 1 The system shown;

[0041] Figure 5 Schematically shows a memory with associated classifications such as Figure 1 The system shown; and

[0042] Figure 6 An alternative system with three coupled artificial intelligence units is shown. Detailed implementation manner

[0043] Figure 1 An exemplary implementation of the artificial intelligence units 110 and 120 with two links is shown, which will be described in more detail below. In the following description, these artificial intelligence units are exemplarily designed as neural networks.

[0044] Therefore, a first artificial intelligence unit is provided here in the form of the first neural network 110, which can basically be used to classify the input signal X i and use the result of this classification to affect the second artificial intelligence unit 120 (here the second neural network). Preferably, the result of the first neural network is not used as the input value of the second neural network, but is used to affect the existing weights, step sizes, and functions of the network. In particular, these parameters of the second neural network may be affected such that these parameters are not completely redefined, but the original parameters of the second network 120 are modulated or superimposed based on the output signal of the first neural network 110. This means that the two neural networks preferably operate independently in other ways, such as training their basic values themselves, but may be coupled by superposition. In this regard, the two neural networks may be basically similar to each other in design, but for example, have significantly different complexity levels, such as the number of existing layers and classifications. In addition, each neural network has its own memory.

[0045] In a possible implementation, the first neural network 110 can be used as a classification network for roughly and quickly classifying input values, and then at the same time, based on the classification result, correspondingly affect the second network by modulating the parameters of the second network. For this purpose, the first neural network can be a network with a relatively small number of levels, which has a memory (with fewer categories K1, K2,...K n ), and these categories are preferably highly abstracted to achieve rough classification. For example, the first neural network can be limited to 10 categories, 50 categories, 100 categories, or 500 categories, and of course these numbers are only understood as rough examples. In this regard, the training of the first neural network can especially be performed independently of the further coupled neural networks. However, additionally or alternatively, a training phase in a coupled state with one or more coupled neural networks can also be used.

[0046] Thus, the first neural network is designed to provide an available output within a relatively short period of time, which can be used to meaningfully influence the second neural network. A weight sum function can be generated from the output value Output1 of the first neural network 110, and this weight sum function can be superimposed on the self-generated weight sum function of the second neural network 120. This means that the second neural network initially operates independently and does not fully adopt the output value of the first network or the parameters obtained therefrom. Additionally, the second neural network 120 can initially be independently trained in the usual way, thus having self-generated weights.

[0047] In this case, the second neural network can be designed to be much more complex than the first neural network, especially it can have more levels and / or memory categories. The degree of increased complexity of the second neural network compared to the first network can be determined differently according to the application. Here, the input value or input data of the second neural network is preferably the same as the input value of the first neural network, so that now a more complex analysis can be performed with the same data. However, alternatively, the output value of the first neural network can also be at least partially used as the input value of the second network. In particular, for example, in the case where the complexity of the second network is significantly different, a second network can be provided, where the original input value that is also used as the input value of the first network is provided to this second network as the input value, and in addition, the output value of the first network is used as the input value of the second network.

[0048] Figure 2 An exemplary different modulation function f is shown mod , and one or more parameters of the second neural network can be superimposed with this modulation function. In principle, this superimposition or modulation can be carried out in any way. For example, when the modulation function f mod_w is applied to the weight w i2 of a node, it can be provided that the weight matrix of the second network 120 is used as the independent variable of the modulation function, or a one-dimensional (even if different) function can be provided for each of the weights w i2 . When the modulation function f mod_f is applied to one of the descriptive functions of the second neural network, that is, applied to the transfer function f trans2 of the network 120, the activation function f akt2 , the propagation function or the output function f out2 , this can be done by combining the two functions, and the modulation function f mod_f can also be applied only to a part or all of the relevant descriptive functions (for example, applied to all activation functions f akt2) Modulation can be equally applied to all nodes of the network, or alternatively can be applied only to a part of the nodes, or can be differently modulated for each node. Similarly, for example, modulation can be applied separately or otherwise interleaved for each layer of the network.

[0049] In particular, the modulation function f mod can also be a time-dependent function such that the weights w of the second neural network are changed in a time-dependent manner i2 or function. However, static modulation functions for modulating the second neural network are also conceivable. In this case, the modulation is applied to the parameters of the second network 120 that have been initially defined for this second network (such as propagation functions or activation functions), or parameters independently obtained during the training phase, such as adaptive self-generated weights.

[0050] Figure 2 Eight different time-dependent modulation functions are shown as examples. Example a) shows a simple binary step function where a specified value of zero is specified within a specified time and then a value greater than zero is specified. Here, the second value can in principle be 1, but can also have different values such that the original parameter is additionally affected by a factor. In this way, for example, the weighting is turned on and off in a time-dependent manner, or the weighting is amplified in a time-dependent manner. Example b) shows the opposite case where a step function with a second value less than zero is predetermined. Similarly, as an alternative to the variants of Example a) and Example b), step functions can be envisaged that include two different values not equal to 0 such that the level is increased or decreased in a corresponding time-dependent manner.

[0051] Example c) shows a periodic modulation function that can also be applied to any parameter of the second network and in this way periodically amplifies or attenuates certain elements in a time-dependent manner. For example, for different nodes and / or different layers, different amplitudes and / or periods can also be selected respectively for such functions. Any periodic function can be used at this point, such as a sine function or even a discontinuous function. Depending on the type of cascading of these functions with the self-generated function of the second network, only positive or negative function values can be selected.

[0052] Example d) shows a slow, continuous and instantaneous increase and decrease of the level. On the other hand, example e) describes a short, approximately rectangular high level with otherwise low function values, which can optionally be zero. Similarly, example f) shows an irregular distribution and very short peaks or spikes, thus causing the level to increase within a very short time. Here, these peaks have different amplitudes and can have positive and negative values (relative to the base value). For the variants from example e) and example f), there can be a regular, periodic and completely irregular in time (e.g., randomly determined) distribution of the peaks or expansions. In this case, for example, a short level increase can be located within the decision cycle of the second neural network, while a longer and significant level change can extend over multiple decision cycles.

[0053] Figure 2 Example g) in further shows damped oscillations, which can also be arbitrarily designed with different damping and amplitudes. Finally, example h) shows a time series of different oscillations around the base value, where in particular the period lengths of the oscillations are different, while the amplitudes remain the same. Such a combination of different oscillations can also be designed as an additional superposition, i.e., a beat frequency.

[0054] Generally speaking, any modulation function can be conceived, Figure 2 and the functions shown in can only be understood as examples. In particular, any combination of the shown exemplary functions is possible. It should also be understood that, depending on the desired effect of the modulation function, the baselines shown in all examples can run at 0 or other base values. For a pure cascade of the modulation function with the corresponding modulation function, a base value of 0 and the corresponding increase in the function value can be used to ensure that the corresponding node only contributes to the processing in a time-dependent manner and is switched off at other times. On the other hand, in the case where the base value is 1, for example, for the Figure 2 example in a), the modulation function applied to the weights first reproduces the self-generated weights of the modulation network as the base value and then starts from a step higher value, accordingly increasing the weights. Thus, such functions are also used to modulate functions such as activation functions.

[0055] As described above, the modulation function can be formed based on the output values of the first artificial intelligence unit, i.e., in this example, based on the first neural network. The relationship between these output values and the modulation function thus formed can be arbitrary. For example, this correlation can be generated at least partially during the joint training phase of the coupled network. In other embodiments, it can be predetermined how to design the dependence between the modulation function and the output values of the first network. Optionally, it can also be decided that for certain output values, the modulation of the second network is not carried out first.

[0056] Alternatively or in addition to applying the modulation function to the weights and functions of the second neural network, a coupled dropout method can also be applied, which is shown in Figure 3 . This is generally a neural network training process in which only a portion of the neurons present in the hidden layer and the input layer are used in each training cycle, and the remaining portion is not used ("dropped out"). To this end, the prior art typically sets the dropout rate based on the feedback error of the network, which determines what portion of the total network consists of neurons that are turned off. Similarly, instead of neurons, some edges or connections between neurons can be turned off.

[0057] In an exemplary embodiment, such partial disconnection of neurons and / or edges can now also be used in the second neural network, where, based on the output value of the first neural network, the dropout parameters are now used not based on the error feedback of the network itself, but rather as in time-dependent modulation. In this regard, for example, the dropout rate of the second neural network can be determined based on the output value Output1 of the first neural network 310, which is then applied to the second network. The figure again shows two coupled networks 310, coupled network 320 as in Figure 1 , but now the neurons or nodes 326, node 328 of the second network 320 are schematically represented as circles. Here, the connecting edges are not shown, and the arrangement of the shown neurons is not intended to have any particular relationship to its actual topology. Through the dropout rate, a portion of the available neurons are now deactivated and thus not used. The active neurons 326 of the second network are shown shaded in the figure, while the unfilled neurons are intended to represent the dropped-out neurons 328.

[0058] Generally speaking, the coupled dropout described herein can also be understood as a modulation function f mod , either as a modulation function using 0 or 1 as weights or, for example, as the output function of each node. This can be based on the output value of the first network to determine which neurons 326, neurons 328 are turned off, or it can simply specify a rate and a random function can be used to determine which neurons are turned off. In this regard, the dropout rate can also be determined again based on the output value Output1 of the first network 310. In this regard, the dropout modulation function can also optionally result in time-dependent turning off, which would correspond to, for example, as shown in Figure 2 , the cascading of the dropout function and the modulation function. Similarly, a sequence of turning-off patterns that have proven effective in previous training can also be adopted, such that, for example, a cyclic pattern change is adopted for turning off in the second neural network 320.

[0059] Generally speaking, this discarding ensures that the neural network works faster. This discarding also prevents neighboring neurons from becoming too close in behavior. The coupling discarding described above can be used both in the joint training phase where two networks are coupled and in an already trained network.

[0060] In order to ensure that the coupled neural networks complement each other in a meaningful way, it is possible to determine which neural network dominates the overall system at any given time. The network whose output value determines the output of the overall system can be designated as the dominant network or dominant position. In the following, it is assumed that at any time, exactly one network of a set of two or more coupled networks is dominant, so that the output of the dominant network is equal to the output of the overall system. However, other embodiments are also conceivable in principle, such that, for example, rules are specified that describe the processing of the output values ​​of the dominant networks into the final overall output value in the case of more than one dominant network.

[0061] In an exemplary embodiment, a timer or a timing element can be implemented for this purpose, which defines a time specification for one or more coupled neural networks. In this case, the time specification is preferably understood as a maximum value or a time upper limit after which the output value of the corresponding network must be present, so that the output can also be present earlier. At the latest after the time specified for a specific network has expired, the output value of this network is then evaluated. The timer can thus control and / or change the dominance between coupled networks based on a fixed time specification.

[0062] Exemplary embodiments of this type are Figure 4 Here, the formation and coupling of the two neural networks 410 and 420 may correspond to Figure 1 The timer 440 now ensures that the output of the first neural network 410 is evaluated at the latest after a predetermined time, which is defined by a predetermined time parameter value. The required time can be determined, for example, from the input value X iThe measurement starts at the time when it is fed into the corresponding network. The selection of the predetermined time parameters of the network can thus be carried out specifically according to the complexity of the network, so that the actually expected available results can be expected within a predetermined time. In an example such as the previously described one, where the first neural network 410 is preferably formed by a network with a small number of hidden layers and a small number of classifications, a correspondingly shorter time can also be selected for this first network. Similarly, other considerations can be taken into account when selecting the time parameters for the network, such as the available hardware, which has a decisive influence on the computing time of the network and / or the application field considered by the coupled network. In addition, the predetermined timing parameters can be variable and can be modified or redefined, for example, according to the results from at least one of the coupled neural networks. It should be understood that such a time specification should at least include the time period required as the minimum time for traversing the corresponding network 410, network 420 once. For example, in Figure 4 it, a time span of 30 ms is specified for the first network, so that during the process run, this network dominates within the time from 0 ms to 30 ms from the start of the process. However, of course, other suitable values can also be selected for this time span.

[0063] During the time period (here 30 ms) specified by the time parameters of the first network 410, this first neural network will process the input value X in the normal way i . After the predetermined time has passed, the output Output1 of the first neural network 410 can be used to generate a function for superimposing or modulating the weights and functions of the second neural network itself. In addition, the output value of the first neural network can also be processed independently, as an alternative or in addition to being used to affect the second network 420, and for example, this output value is used as the fast output of the entire system.

[0064] Once the modulation functions f mod_f 、f mod_w have been applied to the second neural network 420, the timer 440 can start a new timing measurement, and now apply the second timing parameter predetermined for the second neural network 420.

[0065] In this regard, the second neural network 420 can also optionally independently utilize the input value X i , even before being modulated by the obtained modulation functions f mod_f 、f mod_w , so that for example, this input value can also be provided to the second neural network 420 even before the start of the second predetermined time period, and can be processed accordingly there. After the first time period has passed, then by applying the corresponding modulation functions f mod_f 、f mod_wTo superimpose the parameter values and functions of the second neural network. In this regard, one or more modulation functions can be formed for different parts of the second neural network 420, such as for the weights, output functions, propagation functions, and / or activation functions of the second neural network. In the case where the second neural network 420 is formed to be significantly more complex than the first neural network 410, for example, by having significantly more layers and nodes and / or by having a higher number of memory categories, the second neural network will require a relatively higher computational workload and thus also more time, such that in this case, a correspondingly longer second time period can be selected.

[0066] In this regard, optionally, even when another network is determined to be the dominant network in the entire system based on the current time span, each of the networks 410 and 420 can continue to continuously process and evaluate these input values. In particular, in the example of the two coupled networks shown, even when the second network is dominant, the first network can continuously evaluate the input values, so that after the second time period has passed and the second network has found a solution, the output value of the entire system can correspond to the output value of the second network. In this way, a fast classification network such as the first network 410 described herein, which always evaluates the available input values, can also perform short-term interventions until the output value found enters the total output to a certain extent. Such embodiments will be described in more detail below.

[0067] Therefore, through such time control by a predetermined time period in the timer, the entire system can make a decision earlier and, for example, can already act without having to complete the final evaluation and detailed analysis of the second neural network. For example, the situation in an autonomous driving system can be considered to be evaluated by a system having at least two coupled networks. Through the first unit or the first neural network, an early classification of "danger" can be achieved, which does not yet involve any further evaluation of the nature of the danger but can already lead to an immediate reaction, such as slowing down the vehicle's speed and braking and activating the sensor system. At the same time, based on this classification, that is, under the influence of the output value modulation of the first network, the second neural network conducts a more in-depth analysis of the situation, resulting in a further reaction or change of the entire system based on the output value of the second network.

[0068] It is also possible not to specify a time limit for each of the coupled networks, but only for one of the networks (or, if there are more than two coupled networks, also only for a subset of the coupled networks). For example, in the above example, the timer can be applied to the first fast classification neural network, while the second network has no fixed time limit, and vice versa. Such embodiments can also be combined with other methods for determining the current dominant network, which will be described in more detail below.

[0069] In all embodiments with an inserted timer, a precondition can be that the output value of the neural network that currently has an active timer is used as the output of the entire system. Due to a certain delay time in the time required for the network to reach the first solution for a given input value, within this delay time, the previous output value (of the first network or the second network) can still be used as the total output value.

[0070] If the timer is defined only for some of the coupled networks, for example, the timer is only active for the first network, then it can be defined that the output of the entire system generally always corresponds to the output of the second network, and if the timer is active for the first network, that is, the predefined time period is running actively and has not expired, then only the output of the first network is substituted.

[0071] In a system with more than two networks, reasonable synchronization between the networks can also be achieved by calibrating a predefined time span and changing the timer, especially if several networks with different tasks reach a result simultaneously, which in turn has an impact on one or more other networks. Similarly, by adjusting the predefined time periods and sequences, synchronization can also be achieved between several separate overall systems, each system including several coupled networks. In this case, the systems can be synchronized, for example, by time calibration and then run independently but synchronously according to their respective timer specifications.

[0072] Alternatively or in addition to changing the respective dominant neural network in the entire system based on the timer, each neural network itself can also make a decision to transfer the dominant position in a cooperative manner. This may mean, for example, that the first neural network of the entire system processes the input value and obtains a certain first solution or a certain output value.

[0073] Similar to changing the center of gravity with the help of the timer, here it can be specified that the output value of the total network corresponds to the output value of the currently dominant network in each case.

[0074] For this purpose, for example, the change in the input value can be evaluated. As long as the input value remains basically unchanged, the distribution of the dominant position between the coupled networks can also remain basically unchanged, and / or it can be determined only based on the timer. However, if the input value changes suddenly, a predefined dominant position can be established, which takes precedence over other dominant behaviors of the coupled networks. For example, for a suddenly changed input value, it can be determined that in any case, the dominant position will initially revert to the first neural network. This also restarts the optional timer for this first neural network, and the process is executed as described above. For example, these input values may change significantly if the sensor values detect a new environment, or if a previously evaluated process has been completed and a new process is now to be triggered.

[0075] The threshold can be specified in the form of a significance threshold, which can be used to determine whether a change in the input value should be regarded as significant and lead to a change in dominance. A single significance threshold can also be predefined for different input values or for each input value, or a general value, for example in the form of a percentage deviation, can be provided as a basis for evaluating changes in the input value. Similarly, instead of a fixed significance threshold, there can also be thresholds that change in a situation-dependent, timely or adaptive manner, or they can be functions, matrices or patterns on the basis of which the significance of a change can be evaluated.

[0076] Alternatively or in addition, a change in dominance between the coupled networks can depend on the output values found for each network. For example, according to an embodiment, the first neural network can evaluate the input value and / or its change. In this case, for the classes that the first neural network can be used for classification, a significance threshold can be predefined in each case, such that if the first neural network finds a significant change in the class found for the input data, the dominance immediately transfers to the first neural network, enabling a rapid re-evaluation of the situation and a reaction if necessary. In this way, it is also possible to prevent the situation where, although a significant change in the input situation detected by the first fast classification network has occurred, the analysis continues for an unnecessarily long time without the change being considered in depth by the second neural network.

[0077] In all the above examples, the output value of the entire system can be further used in any way, such as as a direct or indirect control signal for an actuator, as data stored for future use, or as a signal passed to an output unit. In all cases, the output value can initially also be further processed by additional functions and evaluations and / or combined with other data and values.

[0078] Figure 5 Again, a simple embodiment as in Figure 1 is shown, in which there are two networks 510, 520 with unidirectional coupling, and the classification memories 512, 522 of each network are now schematically shown. The type of classification K i used here is initially secondary and will be described in more detail below. In particular, the dimensions and structures of the two classification memories of the first network 512 and the second network 522 can be significantly different, such that two neural networks with different speeds and foci or centers of gravity are formed. Thus, for example, as has been briefly described, an interaction between a fast, rough classification network and a slower but more detailed analysis network can be achieved to form a coupled overall system.

[0079] In this example, the first neural network 510 consists of a relatively small number of classifications K1, K2,... K nFormed, this first neural network can also, for example, only follow a flat hierarchy, such that classification is performed in only one dimension. Preferably, such a first network 510 can also be formed by a relatively simple topology, i.e., having not too many n neurons and hidden layers. However, in principle, the network topology can be substantially independent of these classifications.

[0080] Then the second neural network 520 can have a significantly larger and / or more complex classification system. For example, as Figure 5 shown, the memory 522 or underlying classification can also be hierarchically structured in multiple levels 524. The total number m of classes K1, K2, ..., K m of the second network 520 can be very large, especially significantly larger than the number of classes n used by the first neural network 510. For example, the number of classes m, the number of classes n can differ by one or more orders of magnitude. Thus, an asymmetric distribution of the individual networks in the overall system is achieved.

[0081] Then the fast classification of the first neural network 510 can be used to quickly classify the input value. For this purpose, abstract summary classes can preferably be used. In one example, the classification of a sensed situation (e.g., based on sensor data such as image and audio data) can then initially be performed by the first neural network 510 as "large, potentially dangerous animal" without performing any further analysis for this purpose. This means that, for example, in the first network, further classification is not based on the animal species (wolf, dog) or as a dangerous predator, but only based on as broad general characteristics as possible, such as size, tooth detection, attack posture, and other characteristics. The data essentially corresponding to the output "dangerous" can then optionally have been passed to an appropriate external system for a preliminary and quick response, such as a warning system for the user or a specific actuator of an automated system. In addition, the output Output 1 of the first neural network 510 is used to generate the described modulation function for the second neural network 520.

[0082] The same input value X i, for example, the sensor values are also provided to the second neural network 520. In this case, according to the embodiment, these input values can be input immediately, that is, substantially simultaneously with the first network, or input with a delay. In this case, they are input before applying the modulation function or only when applying these modulation functions, that is, when the result of the first network is available. Preferably, to avoid delays, they should not be provided to the second neural network later, especially in the case of time-critical processes. Then, the second neural network also calculates a solution, and the original self-generated weights of this second network and its basis functions (such as the specified activation function and output function) can each be superimposed based on the modulation function formed by the output values of the first network. This allows the iterative operation of the second network to omit a large number of possible variants that there is no time to process in the case where the first network quickly detects a critical situation (for example, a dangerous situation). Although the analysis speed of the second neural network is slower, as described above, possible reactions can already be executed based on the first neural network. This corresponds to the first instinctive reaction in a biological system. Compared with the first network, the hierarchical and significantly larger memory of the second network allows for a precise analysis of the input values, in this example, a detailed classification of the category "dog", the corresponding breed, behavioral characteristics indicating dangerous or harmless situations, etc. If necessary, after the second neural network has obtained a result, then the previous reaction of the entire system can be overridden, for example, by reducing the level of the first classification "dangerous" again.

[0083] Overall, for such a coupled overall system with asymmetric classification, it can be envisaged, for example, that the category K of the first network 510 is quickly classified nPerforms a primary execution of abstract classification, such as new / known situations, dangerous / non-dangerous events, interesting / uninteresting features, required / unrequired decisions, etc., without going into depth. In this regard, this first classification does not necessarily correspond to the final result ultimately found by the second unit 520. However, the two-stage classification performed by at least one fast analysis unit and one deep analysis unit thus allows for a similar emotional or instinctive response of the overall artificial intelligence system. For example, if an object recognized through image recognition might be a snake, the "worst-case scenario" might preferably be the result of the first classification, regardless of whether the classification might be correct. The situation of evolutionary knowledge and instinctive responses existing in human intelligence can be replaced by a fast first classification with pre-programmed knowledge, so that appropriate default responses (keeping a distance, initiating movement, activating increased attention) can also be executed by the entire system and its actuators. Then, the additional modulation of the second learning unit based on this first classification can be understood as being similar to an emotion-related superposition, that is, for example, corresponding to a fear response that automatically initiates a conscious situation analysis different from a situation understood to be harmless. The superposition of the parameters of the second neural network performed by the modulation function can thereby lead to the necessary transformation into other classification spaces that would not otherwise be reached or would not be reached immediately by default.

[0084] Therefore, such systems can be used in various application fields, such as in all applications where critical decision-making situations occur. For example, driving systems, rescue or warning systems for different types of hazards, surgical systems, and generally complex and non-linear tasks.

[0085] In the embodiments described so far, only two artificial intelligence units are coupled together. However, this idea also applies in principle to more than two units, such that, for example, three or more artificial intelligence units can be coupled in an appropriate manner, whereby it can be determined which units can modulate the parameters of a particular one or more other units. Figure 6 An example is shown where three neural networks 610, neural network 620, neural network 630 (and / or other artificial intelligence units) can be provided, where the output value of the first network 610 generates a modulation function for the weights and / or functions of the second network 620, and where the output value of this second network in turn generates a modulation function for the weights and / or functions of the third network 630. In this way, arbitrarily long chains of artificial intelligence units can be formed, which influence each other in a coupled manner through superposition.

[0086] Similar to the earlier example with two neural networks, in one embodiment, all coupled networks can receive the same input value, and the processing can only be coupled through the modulation of the respective networks. However, similarly, it can be envisioned that in an embodiment, for example, in Figure 1After the two neural networks in [the example], a third neural network is provided, which receives the output values of the first network and / or the second network as input values. Optionally, the function and / or weights of the third neural network can also be modulated by modulation functions, such as those formed by the output values of the first network. These modulation functions can be the same as or different from the modulation functions formed for the second network. Optionally, for example, the output value of the third network can be used to form additional modulation functions, which are then recursively applied to the first network and / or the second network.

[0087] It should be understood that various further combinations of the corresponding coupled learning units are possible, where at least two connected units have a coupling through modulation functions that form descriptive parameters for these units, particularly in the case of neural networks for the weights and / or functions of the network. As the number of coupled units increases, more complex variations of modulation and coupling can be envisioned.

[0088] As already mentioned at the beginning, the embodiments described herein are described as examples with respect to neural networks, but in principle can also be transferred to other forms of machine learning. In this case, all variants are considered in which at least a second artificial intelligence unit can be influenced by a first artificial intelligence unit through superposition or modulation based on output values. By modifying the weights and functions of a neural network by superposition of the modulation functions from the foregoing examples, this can be replaced by a corresponding modulation of any suitable parameter that controls or describes the operation of such learning units. In each example, the term "learning unit" can be replaced by a special case of a neural network, and conversely, the described neural networks of the exemplary embodiments can also each be implemented in a general manner in the form of artificial intelligence units, even if not explicitly stated in the respective examples.

[0089] In addition to neural networks, the examples also include evolutionary algorithms, support vector machines (SVMs), decision trees, and special forms such as random forests or genetic algorithms.

[0090] Similarly, neural networks and other artificial intelligence units can be combined. In particular, for example, the first neural network in the previous example, which is shown as a fast classification unit, can be replaced by any other artificial intelligence unit. In this case, a method that is particularly suitable for quickly and roughly classifying features can also be selectively chosen. However, the output values of such a first learning unit can then be applied in the same way as described for the two neural networks to form a modulation function for a second artificial intelligence unit, which can also be a neural network in particular.

[0091] It should be understood that the above examples can be combined in any way. For example, in any of the described embodiments, there can also be as combined Figure 4The described timer. Similarly, in all examples, the learning unit may include a classification memory, as described in conjunction with Figure 5 as an example. All these variants are equally applicable to the coupling of several artificial intelligence units.

Claims

1. A method in a system of at least two artificial intelligence units, the method comprising: Input an input value (X i ) to at least a first artificial intelligence unit (110, 310, 410, 510, 610) and a second artificial intelligence unit (120, 320, 420, 520, 620), where the input value includes sensor values and / or audio data and / or image data; Obtaining a first output value (Output1) of the first artificial intelligence unit (110, 310, 410, 510, 610); Form one or more modulation functions (f mod_f , f mod_w ) based on the output value (Output1) of the first artificial intelligence unit; Applying one or more formed modulation functions to one or more parameters of the second artificial intelligence unit (120, 320, 420, 520, 620), wherein the one or more parameters affect the processing of input values and the obtaining of output values in the second artificial intelligence unit (120, 320, 420, 520, 620); Among them, the artificial intelligence units (110, 120; 310, 320; 410, 420; 510, 520; 610, 620, 630) include a neural network having a plurality of nodes, and among them, the one or more parameters are at least one of the following: the weights (w i ) of the nodes of the neural network, the activation function (f akt ) of the nodes, the output function (f out ) of the nodes, the propagation function of the nodes; Obtaining a second output value (Output2) of the second artificial intelligence unit; Determining a unit as the currently dominant unit in the system from at least two artificial intelligence units; and Forming a current total output value of the system from the output values (Output1, Output2) of the currently dominant unit, wherein the total output value represents the classification of the input values.

2. The method according to claim 1, wherein A classification memory (512, 522) is associated with each artificial intelligence unit, wherein each artificial intelligence unit classifies the input value into one or more categories (K1, K2,..., K n 、K m ), the one or more categories (K1, K2,..., K n 、K m ) are stored in the classification memory (512, 522), the categories are all configured as one or more related levels (524), and the number (n) of the categories and / or levels in the first classification memory (512) of the first artificial intelligence unit is less than the number (m) of the categories and / or levels in the second classification memory (522) of the second artificial intelligence unit.

3. The method according to claim 1 or claim 2, wherein Applying the at least one modulation function causes a temporal dependence superposition of the parameters of the second artificial intelligence unit, and wherein the at least one modulation function (f mod_f , f mod_w ) includes one of the following: a periodic function, a step function, a function with a transient increase in amplitude, a damped oscillation function, a beat frequency function that is a superposition of multiple periodic functions, a continuously increasing function, a continuously decreasing function.

4. The method according to claim 3, wherein, The second artificial intelligence unit includes a second neural network having a plurality of nodes, and wherein applying the at least one modulation function deactivates at least a portion of the nodes.

5. The method according to claim 3, wherein Setting the first artificial intelligence unit (110, 310, 410, 510, 610) as the dominant unit at least before one or more output values (Output2) of the second artificial intelligence unit (120, 320, 420, 520, 620) are available.

6. The method according to claim 3, Further comprising comparing a current input value with a previous input value by at least one of the artificial intelligence units in the system of the artificial intelligence units, Among them, If the deviation of the comparison result is higher than a predetermined input threshold, determining the first artificial intelligence unit (110, 310, 410, 510, 610) as the dominant unit.

7. The method according to claim 3, further comprising comparing a current output value of the first artificial intelligence unit with a previous output value of the first artificial intelligence unit, Among them, If the deviation of the comparison result is higher than a predetermined output threshold, determining the first artificial intelligence unit as the dominant unit.

8. The method according to claim 3, wherein, The system further includes a timer (440), the timer (440) storing one or more predetermined time periods associated with one or more of the artificial intelligence units (410, 420), and wherein the timer (440) is arranged to measure the elapse of a predetermined time period associated with a corresponding one of the artificial intelligence units (410, 420).

9. The method according to claim 8, wherein, When it is determined that an artificial intelligence unit is the dominant unit, start measuring the predetermined time period allocated to the artificial intelligence unit.

10. The method according to claim 8, wherein If a first time period predetermined for the first artificial intelligence unit (410) has elapsed in the timer, determining the second artificial intelligence unit (420) as the dominant unit.

11. The method according to claim 8, wherein, The input value (X i ) includes at least one of the following: a measurement value detected by one or more sensors, data detected by a user interface, data retrieved from a memory, data received via a communication interface, data output by a computing unit.

Citation Information

Patent Citations

  • Methods and apparatus for reinforcement learning

    CN105637540A

  • Fixed Point Neural Network Based On Floating Point Neural Network Quantization

    CN107636697A