Autonomous control of a device by means of a technical motivation value

The method for autonomous device control addresses the inefficiencies of existing AI methods by using randomized actuation and sensor readings to establish action instructions, allowing devices to learn and adapt autonomously without extensive training data.

WO2025124781A1PCT designated stage expired Publication Date: 2025-06-19SCHREIBER CARL ALBERT
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/079047
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-10-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing artificial intelligence methods require extensive training data and complex coordination for distributed learning, making them error-prone and inefficient for autonomous device control.

Method used

A method for autonomous device control that initializes through randomized actuation of effectors and reading of internal and external sensors, storing interrelationships between effector actuation and device/environment properties as action instructions, and controlling the device based on predefined target properties using these instructions.

Benefits of technology

Enables devices to learn independently and adaptively, minimizing the need for initial training data and improving the efficiency and accuracy of autonomous decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024079047_19062025_PF_FP_ABST
    Figure EP2024079047_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is directed to a process for the autonomous control of a device, wherein the device moves generally in a physical real world and no longer has to be restricted with respect to its physical design. The process provides the advantage that the device carries out autonomous learning and continuously improves the learned knowledge or behavior. In general, the disadvantage in the prior art that training data must first be created by selection and interpretation, as is the case in common artificial-intelligence methods, is overcome. In general, the process is universally usable and the device learns independently and constantly corrects its own knowledge base. Also proposed are a device designed to carry out the process and a system assembly comprising a plurality of the proposed devices. A computer program product and a computer-readable storage medium which carry out the process steps or cause a computer to carry out the process are also proposed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Autonomous control of a device using a technical motivation value

[0002] The present invention is directed to a method for the autonomous control of a device, wherein the device generally moves in a physical real world and does not need to be further specified with regard to its physical design. The method provides the advantage that the device carries out autonomous learning and continuously improves the learned knowledge or behavior. In general, the disadvantage of the prior art that training data must first be created, as is the case with common artificial intelligence methods, is overcome. In general, the method is universally applicable and the device learns automatically, constantly correcting its own knowledge base. Furthermore, a device is proposed which is configured to carry out the method, as well as a system arrangement comprising several of the proposed devices.Furthermore, a computer program product and a computer-readable storage medium are proposed which carry out the method steps or cause a computer to carry out the method.

[0003] Various methods from artificial intelligence (AI) are known from the state of the art. These involve providing training data and then having algorithms recognize regularities in a supervised learning phase. In this case, special algorithms can recognize regularities in the database and then extract them. This way, implicit knowledge is extracted from large amounts of data. Once the training phase is complete, the appropriately trained algorithms are applied to actual data and, from this, generate implicit dependencies for solving real-world problems. A disadvantage of this is that the training data must first be provided, which can already have an undesirable influence on the learning phase. This is not only error-prone but also complex.

[0004] Furthermore, swarm intelligence is known from the prior art, in which several devices are provided, which then solve a problem collaboratively in a distributed manner. Here, too, the control and coordination of the individual participants is complex and sometimes error-prone. This is especially true when distributed learning is to take place, which in turn requires coordination.

[0005] Artificial neural networks are also known from the state of the art. These networks, based on graph theory, provide neurons and connections, i.e., nodes and edges. It is known that these networks mimic functions of the human brain and can learn in the process. Edge weights can be varied, new edges can be added or old ones deleted, and existing nodes can be deactivated or new nodes can be added. This results in a dynamically learning overall system.

[0006] However, the use of an ANN in the manner presented above has at least four problems:

[0007] 1. For objects on which the ANN has not been trained, the ANN does not provide any information to call the routines associated with the object.

[0008] 2. The programmed properties or behaviors of the detected and identified objects may either be incorrect, have changed in the meantime, or are still unknown.

[0009] 3. One or more scientists, experts, etc., select the training data and specify the desired results. Even in unsupervised learning, the training data and hyperparameters are still specified by humans, and this cannot eliminate the possibility of errors.

[0010] 4. An ANN cannot correct errors. Every single piece of information is distributed across all of the ANN's parameters, similar to a hologram in which each data point contains information about the entire image. Therefore, an ANN must always be completely erased and retrained.

[0011] The device, referred to here as Ego, does not fail when encountering unknown objects, unlike CNs, such as in the video when the self-driving car suddenly sees two boxes on the road. ANN generally stands for artificial neural network.

[0012] In the current state of the art, there is a need to create autonomous control of devices in such a way that complex preparatory work such as the provision of selected and interpreted training data and the learning process can be prevented or minimized. There is therefore a need for a self-learning system that is also capable of collectively learning or estimating what other participants are likely to do or are even capable of doing. Thus, there is a need for a system that includes multiple participants, with each participant actually maintaining and developing knowledge about the other participants, which is referred to here as empathy.

[0013] Accordingly, it is an object of the present invention to propose a method for autonomously controlling a device, which acts independently solely on the basis of provided target information and in doing so builds up action knowledge. The proposed method should be able to recognize the scope of action of effectors of any device and create an improved method or at least an alternative method for autonomously controlling the device(s). Furthermore, it is an object of the present invention to propose a device for carrying out the method and a system arrangement comprising a plurality of the proposed devices. Furthermore, it is an object to propose a computer program product and a computer-readable storage medium which contain instructions which execute the method.

[0014] The problem is solved by the features of patent claim 1. Further advantageous embodiments are specified in the subclaims.

[0015] Accordingly, a method for autonomously controlling a device is proposed, comprising initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties; storing interrelationships existing between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation, device properties, and environmental properties into a target triplet;and controlling the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet;

[0016] The proposed method is used for the autonomous control, i.e. the generation of instructions, of a device that is generally physically embodied. This could be a car, a robot, or a production facility. There are no restrictions on mobility in this case, meaning the device can move on land, in the air, on water, or even underwater. Autonomous in this context means that the method ensures that the device learns independently and automatically recognizes which possible courses of action are available. These can be learned, and the sensors learn which action leads to which result from which starting point. In general, it is possible for the device to define its own targets, or the targets can be transmitted externally.Internal objectives might, for example, be maintaining operational capability. An internal objective might, for example, be to visit a charging station when the battery level is low. An external objective can be communicated by specifying what task or activity the device should perform.

[0017] In a preparatory process step, initialization occurs by randomly activating effectors. An effector is generally a physical entity that influences the real world. This can, for example, affect the device itself, such as steering, accelerating, braking, etc., and / or the targeted manipulation of objects by, for example, robot arms, directly or with tools, or by remote-acting instruments, be they projectiles or beams or, for example, fire-fighting agents such as water or extinguishing foam. In the example of an automobile, this could be the brake, the accelerator pedal, etc., but also headlights, indicators, etc. The effectors are activated, and the effects of these actions are then measured using internal and external sensors. This is logged and can then be used in the further course of the process.In this way, actions are learned and used effectively in later procedural steps.

[0018] Since the proposed device or method can be operated or controlled entirely without initial knowledge, the preparatory step involves randomized actuation, i.e., arbitrary actuation. This requires learning which actions are even possible and what effect they can have without having already performed or trained them with a goal in mind, so that what has been learned is not tied to specific external goals and possibly not applicable to other external goals later. In further iterations of the method, the learned actions are then carried out in a targeted manner. During randomized actuation, the system parameters of the effectors are checked, and a robot arm, for example, is moved into all possible positions. This is logged in each case, and it is recognized which action of the effector has what effect on the real world.In the example of a headlight, it is possible that it only provides the on and off parameters. The headlight is then switched on and off, and the external sensors measure the headlight's impact on the environment. The internal sensors measure the headlight's operating parameters, such as temperature development. With more advanced LED headlights, it is also possible to adjust both the color and the intensity. This is tested through randomized activation, and the internal and external results are then logged.

[0019] According to the invention, internal sensor units and external sensor units are proposed. Internal and external do not refer to the arrangement of the sensor units, but rather, according to the present invention, they mean that the sensor units measure external conditions with respect to the device or, analogously, internal conditions. Internal conditions are all system parameters of the device itself. External conditions are environmental variables, i.e., environmental variables surrounding the device. Thus, device properties are measured by means of the internal sensor units, and environmental properties are measured by means of the external sensor units. For example, if the device has a certain battery level, this is an internal parameter.If an object is transferred from A to B using a robot arm, this is an external state, while the state of the robot arm is, in turn, an internal state. Typically, internal and external parameters interact, and actions always result in, for example, a reduced battery level, while these actions, in turn, affect external properties. In the other direction, it is an interaction that, using effectors, the device can be brought closer to a charging station, which then initiates a charging process, which in turn influences the internal battery level.

[0020] The recorded interrelationships are then saved, thus logging how the activation of the effectors influences the internal and external parameters. The device properties are therefore saved along with the environmental properties, and this defines instructions. Activating the effectors is therefore an action that influences both the device itself and its environment. For example, if the device moves from a first geographical point to a second geographical point, this changes the device's environment, which is an external parameter, and this also changes the internal state of the device, for example, through a temperature development or a change in the battery level. In this way, a triplet can be saved that indicates what the internal state was, what the external state was, and which action is then carried out.This is converted into a new internal state and a new external state. Thus, an output triplet consists of the activation of the effectors, device properties, and environmental properties, and this is converted into a target triplet. The new internal state, the new external state, and possible actions can be specified in the target triplet.

[0021] The concept of triples is merely intended to illustrate the general transition of system states. Alternatively, internal states and external states can be transformed into a new internal state and a new external state using a function. Thus, the action consists of a mapping function from a first tuple to a second tuple. Those skilled in the art will recognize that the power of both descriptions is equivalent. Either an output tuple is transformed into a target tuple using a function, or an input triple is formed, which specifies internal states, external states, and an action. This action then leads to a target triple, which has a new internal state, a new external state, and another field.The remaining field can either remain empty, although it is preferable that at least one action be performed here that is now possible in this state. The last field can also be filled in such a way that, for example, a vector is introduced that contains identifiers for further actions.

[0022] To illustrate this with an example, the device can activate the effector called the motor drive and move forward. What is now logged is an output triplet of an internal state, namely a battery charge level, an external state, namely an image signal of a physical environment, and an action instruction, namely "drive." This output vector or the output triplet is then converted into a target triplet, which contains a new battery charge level, a new image signal of the environment, and actions that would now be possible, such as moving forward or reversing. During the forward movement, it is also possible to specify that, for example, braking or activating the headlights would be possible.In an alternative example, the output tuple of the battery state and the image signal of the environment is converted using the "Move" function into a target tuple that describes the new battery state and a new image signal of the environment. This creates a record of what happens internally and externally during a specific action. This knowledge can be reused in subsequent action steps toward a predefined goal.

[0023] In further method steps, target device properties or predefined target environment properties can be provided. Thus, the device can be controlled such that the at least one predefined target device property and / or the at least one predefined target environment property is established. The action instructions are known from the preceding method steps, and thus it is known how an internal state and an external state can be converted into a further internal state and a further external state. The target specifications, i.e., the target device property and / or the target environment property, serve this purpose. These target properties therefore determine what is to be achieved, and this is achieved by executing the action instructions.The device's current state is read out, and then stored actions are selected that, starting from the actual state, reach the target state. In a preferred case, executing an action instruction that immediately achieves the target specifications is sufficient. In a typical case, however, the initial state is successively transformed into the target state using several action properties. Thus, several action instructions are linked in such a way that the target specifications are ultimately achieved.

[0024] To illustrate this with an example, the method may have identified that the device can move from A to B to C to D. Thus, the instructions are stored which stipulate that the car can drive from A to B, from B to C and from C to D. It is also implicitly stored that the car can drive from A to C and from A to D. Furthermore, it is stored that the car can drive from C to D. These possibilities are now used and if a target specification is given which states that the car should drive from A to C, the device has two options to choose from, namely to drive from A to B to C or to drive from A to C. In general, it would also be possible to drive from A to D and then from D to C. Which selection is made can also be defined in the target specifications. For example, a maximum duration of the journey can be defined in time.In addition, a maximum energy consumption can be defined. In preparatory process steps, actions are carried out randomly and are then available at runtime. In general, the process can provide for randomized actuation of the effectors, but it can also additionally provide for the reuse of previously learned action instructions. This allows for a learning phase of randomized actuation, as well as an execution phase that uses previously stored actions. These can also be alternated in any order. Furthermore, in the execution phase, the stored action knowledge is refined using internal and external sensors, and new action instructions are generated that may be preferential to a different objective.Furthermore, the target specifications can change during the execution of an action in such a way that, for example, a battery condition becomes critical. The target for the system's own operability is increased to such an extent that the system recognizes that it is now necessary to drive to a charging station. This results in an overriding target specification, at least temporarily, namely to drive to a charging station, even if this does not achieve the goal of driving to the original target coordinate. As soon as a certain charge level is reached, the target for the system's own operability is downgraded again, and the target for driving to the destination is increased to such an extent that the journey can now be continued.

[0025] According to one aspect of the present invention, the method is iterated in a learning manner such that new instructions are constantly being recognized, each of which converts an output triplet into a target triplet, and these instructions are stored to control the device. This has the advantage that the method is established, entirely or at least partially, in such a way that new situations are constantly being recognized and instructions are created such that effectors are actuated. The method can thus be varied such that the actuation of the effectors is no longer randomized, but rather existing instructions can be combined, or it is also possible to randomize parts of existing instructions, so that new possible instructions are created, for which it is also clear which interactions they trigger.Furthermore, the method can also be established in such a way that, starting with the storage of interrelationships, it is established up to the control of the device. This means that even when known instructions are carried out, interrelationships are stored and new parameters can be identified. For example, it can be recognized that if the same instruction is carried out twice, different target parameters arise. This can be the case, for example, if a crosswind arises while driving a car. This is stored and then the external sensor detects that wind must have occurred here. This creates a new output triplet and, using the effectors, it can be randomly tested how to countersteer in such a situation.Thus, new interrelationships have been identified and if such an initial triple is identified again, it is now clear which action must be carried out in order to achieve the desired target triple.

[0026] According to a further aspect of the present invention, the reading of internal and / or external effectors is carried out by storing individual support values ​​and interpolating and / or extrapolating intermediate values. This has the advantage that the support values ​​can be selected according to their availability and that they can also be estimated in such a way that a measurement does not have to be available for every possible value. Rather, existing measurements can be used to estimate which values ​​would result under normal circumstances. This allows additional support values ​​to be calculated mathematically.

[0027] According to a further aspect of the present invention, the target device property and / or the target environment property are at least temporarily influenced by a device property and / or an environment property. This has the advantage that the target specifications can also be changed or are subject to prioritization. For example, if the vehicle is to travel from a first geographical point to a second geographical point, the approach to a charging station can be prioritized in the meantime, since otherwise the final destination would not be reached. This results in the advantage of always guaranteeing that the target specifications are met, and a fail-safe method is created.

[0028] According to a further aspect of the present invention, multiple action instructions are combined into a choreography that transforms an initial triplet into a target triplet. This has the advantage that even complex, i.e., compound actions can be executed. As the method progresses or becomes established, new action instructions are continually created, which are combined into increasingly efficient choreographies. Thus, the proposed method is iteratively improved.

[0029] According to one aspect of the present invention, the device is a manufacturing robot with two robot arms, each with an effector (gripping hand), which are designed in the sense of the device, ie they “know” how to grasp, move, manipulate, etc. objects.

[0030] The device is now demonstrated only once how the new workpieces, Workpiece 1 and Workpiece 2, should be grasped, turned, and then put together. The result is then marked as the target for the device. From now on, the device can repeat this process at any time: Show it once instead of extensive training.

[0031] If the two workpieces now appear on a conveyor belt, two alter egos are created for each. Due to the situation, the initial triplet, and the associated goal, the motivation value drops to negative. The device recognizes a call to action and combines the two, resulting in a motivation value that leads the device to a standstill or rest position. This is an example of how the device's ability to learn from observation can be used.

[0032] According to a further aspect of the present invention, the method is carried out on a first device and, based on external sensor units of the first device, a second device is detected which also carries out the method and whose generated interactions are stored on the first device or the interactions are stored remotely and made available by means of an interface for at least the first device and / or the second device. This has the advantage that the method is carried out on multiple instances and the capabilities of the second device are monitored by the first device. The behavior of the second device is therefore analyzed and stored on the first device or centrally, accessible to the first device.The results of performing the method on the second device are therefore read out and / or measured by the first device and stored centrally as potential own actions on the first device or described.

[0033] According to a further aspect of the present invention, the first device is controlled based on the detected interactions of the second device. This has the advantage that the first device can be controlled according to at least one predefined target device property and / or at least one predefined

[0034] Target environment property can be determined using at least the stored action instruction of the second device.

[0035] According to a further aspect of the present invention, interrelationships of the detected second device between its parameters and its actions are used to generate action instructions for the first device. This has the advantage of creating a network of devices that create or store action instructions or interactions among themselves by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties, and then share these instructions or interactions among themselves for individual or joint use. The conversion of the source triplet into target triplet can also be carried out collectively.Thus, corresponding choreographies are not limited to one device that carries out the process, but rather the choreographies can be carried out by several devices in a division of labor.

[0036] According to a further aspect of the present invention, an object is detected using external sensors and its detected parameters are compared with device properties and / or instructions of the devices by mapping them to a copy of the central template, whether trained by the device itself or adopted and implemented by a comparable device. This has the advantage that the device can essentially clone itself or can infer the properties of the object from its own properties. This can also be referred to as empathy. If, for example, a device is identified using the external sensors that has similar characteristics to the executing device, it is assumed that this newly detected device has similar capabilities to the executing device.To illustrate this with an example, we will consider a vehicle that implements the proposed invention. This first vehicle therefore implements the proposed method and detects a further, i.e. second, device using an external sensor, in this case an imaging unit. The first device or the first vehicle has learned that at a certain speed it can only brake to a limited extent and can only steer laterally to a limited extent. The first vehicle now detects that the second vehicle has similar dimensions and is traveling at a similar speed to the first vehicle. The instructions are then conceptually projected onto the second vehicle, and the first vehicle detects that the second vehicle can only brake to a limited extent and can only change direction to a very limited extent.If an overtaking maneuver is initiated on a highway, the first vehicle detects that the second vehicle could brake or also change lanes. Thus, based on its own behavior, or rather the possible exit triplets and the possible target triplets, it concludes that this object has similar characteristics. Furthermore, the external sensors can monitor the behavior of the detected second vehicle, and then update its own data memory, which stores the interrelationships. Thus, exit triplets and target triplets can also be detected, and in turn, conclusions can be drawn from the second vehicle to the first vehicle.

[0037] According to a further aspect of the present invention, the detected parameters are used to predict the behavior of the detected object. This has the advantage that it is recognized that the device itself executes a specific action instruction in a certain situation, so that the other detected object would very likely also execute the same action in the same situation. Thus, the device's own actions or executed action instructions are logged, and it is then assumed that the detected object could behave the same or at least similarly.

[0038] According to a further aspect of the present invention, interrelationships of the detected object between its parameters and its actions are used to generate action instructions for the device. This has the advantage that externally detected interrelationships with respect to the proposed device or the device controlled by the method can also be used. Thus, the device not only creates interrelationships that are evaluated, but rather, the external world can also be observed, and conclusions can then be drawn regarding the conversion of the initial triplet into the target triplet.

[0039] According to a further aspect of the present invention, an effector is a processor, a memory, a motor, a gripper arm, a locomotion unit, a drive, an extremity, a windshield wiper, a turn signal, a light, a steering system, and / or an executing unit. This has the advantage that, based on the proposed device or the device to be controlled according to the method, all possible hardware units can be present. The embodiment is merely exemplary and not exhaustive, so that any physical unit can be controlled autonomously according to the proposed method.

[0040] According to a further aspect of the present invention, internal sensor units detect an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a range, a position, a processor utilization, a memory utilization, and / or a system parameter of the device. This has the advantage that the internal sensor units can measure all parameters and states of the device that executes the method or that is executed or controlled by the method.

[0041] According to a further aspect of the present invention, external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal, and / or an external parameter. This has the advantage that the entire environment of the device can be analyzed and recorded. For this purpose, all possible sensors that detect the environment in some way are possible, either individually or in combination.

[0042] Since trained neural networks are proposed according to one aspect of the present invention, it is advantageous that the central template according to the invention takes over the functions of trained neural networks, but in its possibilities and flexibility far exceeds those of the NN.

[0043] The central template can either be trained independently or copied from a comparable device. In the case of artificial intelligence in a smart home, a central template can never learn to move independently, for example, to monitor the actions of an elderly person as part of health monitoring (keywords: empathy, fall control, etc.); it must be copied and implemented from another source, given the current state of the art.

[0044] The initial training or development of the central template is comparable to children's play and is not determined by the achievement of externally defined goals. It serves only to learn one's own actions: What can I do? Later, external goals can be achieved through combinations of the learned action options, whereby the path to this goal is or must be designed by the child. Otherwise, actions would be fixated on certain goals from the outset (opening the door), only to fail when other goals (vacuuming) could be achieved.

[0045] According to one aspect of the present invention, this occurs by testing one's own options such as accelerating, steering, etc. - similar to "babbling". However, the device simultaneously detects and stores (and this goes beyond "babbling") the changes (distances, speeds, acceleration, etc.) caused by this, so that it no longer sets its own actions (steering) as the target triplet, but rather the changes it can achieve in this way, and this goes beyond "babbling". Just as a child first learns (has to learn) to crawl and then begins to crawl deliberately here and there. Thus, for children as well as for the device, the target triplet becomes increasingly "abstract" compared to the child's own, "original" options for action, which thus become mere means for ever more complex target triplet. NN, on the other hand, are and remain limited to the trained initial and target triplet.

[0046] The device according to one aspect of the present invention sets itself goals or triplet goals in the juvenile, training phase, just like small children do, such as: I want to go there, I want to push there. This random, independent goal setting negatively influences the internal motivation value; in other words, it is "deflected." The device then attempts to change this value in the opposite direction, i.e., positively, through its own actions. A child would likely perceive this as (positive) excitement or tension (not as a negative feeling such as fear, for example). Technically speaking, each triplet goal can and is mapped to an internal value, which the device can use to control its own behavior. With NN, the human must specifically indicate the achievement during the training phase; otherwise, the backpropagation process would not function.

[0047] During the juvenile training phase, according to one aspect of the present invention, the device attempts to reach the goal by activating its options such as "accelerating (moving forward), steering, braking." The sensor system detects that, for example, "accelerating" but steering in the wrong direction does not affect the distance value in a way that positively changes the motivation value. Each action therefore immediately leads to a change in the motivation value, from which the device recognizes whether it is getting closer to the goal or not, so that it continues the action (in the case of a positive change) or changes it (in the case of no change or a negative change).Using the matrix of support values ​​and interpolation, the device can also detect whether a small change in the action (accelerating slightly more or slightly less) positively or negatively changes the motivation value, allowing it to independently optimize the actions for reaching the goal. In this sense, it can identify and select the faster or shorter route. For example, if the target it wants to reach is diagonally in front of the device to the right (detectable via the sensors and the internal polar coordinate system), it can detect that steering to the left (away from the target) negatively changes the change in the motivation value, while steering to the right positively changes it.

[0048] Using the stored support values, the device can calculate the respective effect on the motivation level within the framework of a simulation calculation before deciding on a specific action (steering left or right) and then immediately choose the optimal action (steering right in the above scenario) itself, without the need for prior training, as would be necessary with a NN. This is, of course, only possible if a sufficiently large database is available.

[0049] According to one aspect of the present invention, restrictive constraints can also be considered. For example, a road does not lead directly to the destination, which lies behind a freshly plowed field, but only at an angle of 45°. Now, the simulation of the straight path (largest motivation change) leads over the field with the risk of getting stuck (massive negative motivation value change) or continues along the road (small but still positive motivation change).

[0050] Thus, the device or method according to one aspect of the present invention learns or is capable of:

[0051] 1. Actions change the motivational value, whereby the genuine aim is to bring this value to a maximum positive value (whether this is 0, a largest or smallest possible value in a realization is irrelevant for the principle).

[0052] 2. Some actions influence the change in motivational value more strongly, others less strongly, so that the action(s) with the strongest influence are technically recognizable and selectable. 3. From the stored support values ​​and a new situation (initial triplet), the device can select, or "decide on," the optimal actions for achieving the goal in a simulation calculation, without having been specifically trained on the initial and target triplet.

[0053] 4. The ability to simulate not only allows a direct action to transform the initial triplet directly into the target triplet, but also allows the achievement of the target triplet to be divided into several individual steps, for example, in order not to violate certain constraints.

[0054] 5. If the device can calculate the consequences of its own actions through a technical simulation (when will I be where and what speed will I be at), the consequences of the actions of other, new, unknown objects can also be technically calculated via the simulation. The device copies its own object instance (the ego) with all data, methods, and functions and updates the respective copy (the alter ego) with the current sensor values.

[0055] An internal instruction is, for example, the independent actuation of the steering by the device in order to align the device, for example, with the current target - it results from the simulation and the resulting optimal trajectory for reaching the target triple.

[0056] According to one aspect of the present invention, an external action instruction is the external setting of a (new) target triple.

[0057] According to one aspect of the present invention, an initial triple is a current situation consisting of the device, the other dynamic objects, the constraints of the scenario, etc., which “causes” the device to find an optimal path (trajectory) to the target triple through simulation, i.e., the “theoretical” testing of various developments.

[0058] According to one aspect of the present invention, device properties are the own physical properties (mass, size, "horsepower",...), the ability to become active itself (brake, steer, accelerate,...), the ability to change the environment (push, shove, push,...).

[0059] According to one aspect of the present invention, environmental properties are (see also initial triple): a current situation consisting of the device, the other dynamic objects, the constraints of the scenario...

[0060] The target triplet can be described according to one aspect of the present invention as follows: in the juvenile training phase, these would initially be purely random direct actions, the effects of which (distances, collisions, acceleration, etc.) are stored in such a way that these effects then become direct positive (to be achieved) and negative (collisions: to be avoided at all costs) target triplet. As soon as a sufficiently large database exists, the device can set goals that are no longer directly recognizable, but can only be reached after, for example, by circumventing an obstacle: again something that children learn after a certain time: Overcoming the "out of sight, out of mind" mentality.

[0061] Thus, according to one aspect of the present invention, there is no failure in unknowns. This is a technical solution to a technical problem.

[0062] According to one aspect of the present invention, the device is controlled according to at least one predefined target device property and at least one predefined target environment property using at least one stored action instruction as follows: The at least one action instruction specifies how the at least one predefined target device property and the at least one predefined

[0063] Target environment property is set based on the source triple and the target triple

[0064] At least one action instruction converts the source triplet into the target triplet. There is an initial configuration, and this is converted into a target configuration using one or more action instructions—in the latter case, a chain of action instructions that convert individual source triplet into target triplet.

[0065] Now, the method or device has not yet learned how to determine the predefined target device property and the at least one predefined target environment property based on the learned transitions or

[0066] The device must therefore select the action instruction(s) from the machine-learned wealth of experience and the development forecasts of a dynamic situation that best achieves the target configuration (optimal trajectory). If the device approaches the target configuration, this action instruction can be used; the motivation value changes positively. If, instead, the device moves away from the target configuration, this action instruction cannot be used; the motivation value changes negatively. This can be evaluated or simulated based on the stored action instructions. This makes it possible, according to the invention, to independently select the action instruction(s) that work best, i.e. that are expected to lead to the best path (optimal trajectory) to the target configuration.

[0067] For example, when approaching the proposed device, a distance can be determined using the proposed method using sensors. This learning process transforms an initial triplet into a target triplet. The initial triplet in the present example is actuation (switching on the means of transport), device properties (the device can drive, steer, etc.), and environmental properties (the device is located at geographical point A). The target triplet is actuation (the device turns around), device properties (the device cannot drive due to an obstacle), and environmental properties (the device is located at geographical point B, but in front of an obstacle). The device now notices that the motivation value has changed negatively because an obstacle was encountered.The at least one action instruction is used which maximally improves a motivation value and the motivation value indicates how closely the conversion of the initial triple into the target triple achieves the predefined target device property and the at least one predefined target environment property.

[0068] Since these triplets are stored, this movement can be simulated, and the instruction that best achieves the target configuration is used. This allows for a planned and proactive selection / use of the instruction(s).

[0069] The motivation value can also take into account technical configurations of the device. For example, high energy consumption can also worsen the motivation value. Damage to the device can also negatively affect the motivation value. Further examples include excessive time requirements, meaning the shortest time requirement is used, or undesirable target configurations such as damage to external objects arise. A metric can be empirically provided that correlates the factors influencing the motivation value, and an empirical optimum is created, which is then implemented according to the action instruction. The effectors are controlled accordingly. The object is also achieved by a device configured to carry out a method according to one of the preceding claims,comprising an initialization unit configured for initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured for measuring internal device properties and reading of a plurality of external sensor units configured for measuring external environmental properties; a storage unit configured for storing interrelations existing between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction is an output triple of actuation,Device properties and environmental properties are converted into a target triplet; and a control unit configured to control the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one stored instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environmental property is set based on the initial triplet and the target triplet.

[0070] The problem is also solved by an arrangement comprising several devices which are linked by communication technology and exchange instructions.

[0071] The problem is also solved by a computer program product with control commands that implement the proposed method or operate the proposed device.

[0072] According to the invention, it is particularly advantageous that the method can be used to operate the proposed devices and units. Furthermore, the proposed devices and units are suitable for implementing the method according to the invention. Thus, each device implements structural features suitable for executing the corresponding method. However, the structural features can also be configured as method steps. The proposed method also provides steps for implementing the function of the structural features. Furthermore, physical components can also be provided virtually or in a virtualized form.

[0073] Further advantages, features and details of the invention will become apparent from the following description, in which aspects of the invention are described in detail with reference to the drawings. The features mentioned in the claims and in the description can each be essential to the invention individually or in any combination. Likewise, the features mentioned above and those further explained here can each be used individually or in groups in any combination. Parts or components with similar functions or that are identical are sometimes provided with the same reference numerals. The terms “left”, “right”, “top” and “bottom” used in the description of the exemplary embodiments refer to the drawings in an orientation with normally legible figure designations or normally legible reference numerals.The embodiments shown and described are not intended to be exhaustive, but rather are exemplary in nature to illustrate the invention. The detailed description is intended to inform those skilled in the art; therefore, known circuits, structures, and methods are not shown or explained in detail in order not to obscure the understanding of the present description. The figures show:

[0074] Figure 1: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention;

[0075] Figure 2: A diagram that can be used to estimate internal and / or external properties. This allows internal and / or external parameters to be measured and interpolated between the sampling points or to extrapolate to further sampling points.

[0076] Figure 3: a representation of a real-world situation in which a child wants to cross a street and the proposed device recognizes the situation and provides instruction;

[0077] Figure 4: a real-world situation in road traffic in which the proposed device or method according to one aspect of the present invention is applied; and

[0078] Figure 5: another real-world situation in road traffic, wherein the device practices cornering according to another aspect of the present invention.

[0079] Figure 6A: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention in the juvenile phase with the primary goal of specifying one's own actions; and

[0080] Figure 6B: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention in the adult, application phase for achieving one's own goals, taking into account and predicting the behaviors of other objects relevant in a situation.

[0081] Figure 1 shows a schematic flow diagram of a method for autonomously controlling a device, comprising initialization 100 by means of randomized actuation 101 of effectors and reading 102 of a plurality of internal sensor units configured to measure internal device properties and reading 103 of a plurality of external sensor units configured to measure external environmental properties; storing 104 of interrelationships existing between the actuation 101 of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation 101, device properties, and environmental properties into a target triplet;and controlling 105 the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored 104 action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet;

[0082] The model presented here, according to one aspect of the present invention, uses a central template that helps the AI ​​navigate the unknown and constantly changing world. Functionally, it replaces the NNs trained with pre-selected and interpreted data, meaning it is not pre-determined by specific external goals or pre-defined external objects, which leads to much greater freedom of perception and action. This approach allows for short-term behavioral predictions of the observed objects, thus assessing the development of situations and incorporating this into the design of one's own actions. It also enables learning from observation and empathy.Aspects of the invention are: central template as the center of the AI ​​for capturing all objects in a situation; data storage that enables the recognition of laws; behavioral predictions of all observed objects in a situation; empathy ability based on the central template.

[0083] Ability to learn from observations;

[0084] Identification of influencing factors that cannot be directly observed;

[0085] Recognition of natural laws and causal relationships; computer-friendly form of theories that can be self-created, stored, reused, and continuously improved; and / or self-assessment and evaluation of self-created theories as a basis for cognition, learning, and the selection of the best theory for achieving one's own goals.

[0086] As an example, a car that behaves according to these ideas is described. How does it assess a child or a box, such as overtaking on the highway, or how does it recognize the law of gravity and crosswinds (as an example of recognizing causality and non-manifestable objects). It further describes how learning from observation is possible through empathy. All of this shows that the possibilities of this class extend beyond driving cars. The essential basis, however, is a physically existing robot in a physically existing environment that directly or indirectly observes other objects.

[0087] Some aspects of the present invention are proposed below, which enable an exemplary implementation of the method or device and / or system arrangement. The following aspects are to be understood merely as examples and can be applied individually or in combination.

[0088] Possible hardware of the robot or the proposed device:

[0089] This describes the technical equipment that the robot or car may have according to one aspect of the present invention.

[0090] Effectors:

[0091] In simple cases, such as a car, the effectors would be, apart from windshield wipers, indicators, lights, etc. Braking, acceleration, and steering, with which the robot influences its environment, even if it's just its own change of location, is also a change in situation. Internal sensors:

[0092] The internal sensors measure and provide feedback on the body's own characteristics, such as height, weight, speed, acceleration, etc., based on the actions of the effectors. These are necessary for assessing the body's own actions and are prerequisites for learning. This allows the body to recognize its own actions, such as the strength of braking and the effect (the extent of negative acceleration), make connections, and store memory for independent learning, meaning it can continuously improve its actions.

[0093] As a physical entity in a physical environment, the robot is subject to the laws of physics. A measured acceleration, i.e., a change in speed, depends not only on the mass but also on how far the accelerator pedal is pressed or how hard the brake is applied. If these laws are not to be hard-coded, but rather the robot is to discover, store, and use them independently, it requires internal sensors to learn and store the consequences of the use of effectors. The internal sensors of a car, for example, therefore record data such as size, speed, acceleration, etc.

[0094] External sensors:

[0095] According to one aspect of the present invention, the external sensor system detects the environment (distances, objects, etc.), for example, using lidar, image processing, etc. Of course, the robot must be appropriately informed about its environment. What exactly is required will be described in more detail below. However, it is not intended to predetermine which systems are best suited, as the information requirements are lower due to the use of the central template. Not all theoretically available information needs to be detected and processed.

[0096] Possible robot software:

[0097] Here, the parts are described with which the presented Kl achieves its intelligence according to one aspect of the present invention.

[0098] Data storage: Data storage is a highly available data store for data tuples composed of data from internal sensors, effectors, actions, and their motivational values ​​(see below). The store appropriately links executed actions and the associated actual, experienced, and observed results. It therefore only stores internal data in tables and not, for example, pixels of one or more images. Thus, there is only data whose internal meaning is known.

[0099] Another part of this module, according to one aspect of the present invention, is the repeated checking of the internal consistency of the data. This means that it must not contain any contradictions, structures that lead to loops or circles, etc. The goal is that only one decision can ever be derived from the data. The routine that checks this should also remove stored redundant data, not only to keep the database as lean as possible, but also to prevent overadaptation. What is considered redundant depends on the type of theory formation.

[0100] Figure 2 shows the theory formation according to one aspect of the present invention:

[0101] The central template can simulate, save, and use any functional interrelationship for prognosis using the simple means described here.

[0102] Braking or acceleration functionally depend on the accelerated mass and the applied force. However, this function is initially unknown, because the mathematical function of the velocity change dV depends on the conditions Vn (current speed, weight, etc.) and the use of the effectors Em (force on the accelerator pedal, steering, etc.), thus: dV = f(Vn, Em) (1)

[0103] This function should be learned independently rather than programmed, as this allows the learned function to be expanded, corrected, and improved later in the same way. The effective acceleration when braking (or accelerating) is linked via the internal sensors of the measured acceleration a to the (mathematically) independent variable of force F (how hard the respective pedal is pressed) and the mass m according to the law F = m * a. The principle by which the robot itself finds this equation is simple. To illustrate, let's take a (different, arbitrary) quadratic function y = -((x - 4)) 2)+9, which is to be reproduced. Suppose that initially only the two measured values ​​at points 11 and 12 exist. Now suppose that at the point x=3 (point 2) a forecast p is required, which leads to a value of y=3.8 via the straight line a between 11 and 12. However, since the error compared to the ex post measured value y=8.0 is much too large, this point 2R at x=3.0 and y=8.0 is stored in memory. Subsequently, for a new forecast either the degrees b1 between [11, 2R] or b2 between [2R, 12] would be used. Over time, many more points arise in the above example (31, 32, 33, ..), via which the functional relationship between x and y can be reproduced with any desired accuracy using the respective straight lines along the data tuples (c1, c2, c3, ..). The accuracy is only limited by the lack of experience (yet) and the size of the memory.Of course, this can be extended to any number of variables, if the computing and storage capacities allow it; instead of straight lines, hyperplanes a1x1 + ... + anxn = const are used for the forecast by means of piecewise linear interpolation or extrapolation.

[0104] However, to keep the number of stored corner or data tuples as small as possible, they can also be removed: If an ex post analysis reveals that a data tuple is so close to or even on the line / hyperplane between the two neighboring tuples, then this tuple can be deleted as redundant. This not only reduces the memory requirements but also the computational effort, and over time the behavior adapts better and better to the actual, albeit still unknown, mathematical function, which then leads to increasingly more efficient behavior over time.

[0105] This approach also prevents overadaptation, which must be considered when selecting and sizing training data for an ANN. Redundant experiences are thus prevented from being repeatedly stored, which would completely overshadow the rare events, the so-called "black swans," with potentially severe consequences.

[0106] Storing functional relationships between variables has another advantage: it also stores every possible inverse function, as the stored data tuples do not distinguish which values ​​were the 'dependent' and which were the 'independent' when they occurred. In the above function, x is the independent variable, and the corresponding y-value can be determined using the function y = -((x-4) 2)+9. The x-value would have to be searched for in the database, and if the x-value isn't explicitly stored, the y-value can be approximately determined using the next two x-values. However, one can also simply search for the y-value in the memory and determine an x-value using piecewise linear interpolation or extrapolation—this is much easier than trying to determine the inverse function of, for example, the equation above.

[0107] In this very simple way, any functional relationship can be simulated with arbitrary precision. The only prerequisite is sufficient action training, as in a complex sport. Of course, the accepted error E can also be changed and adjusted over time, whether it needs to be reduced to increase the required precision, or it can be increased to reduce storage and computational effort, since this in turn influences the number of stored key points and data points. It is a continuous optimization in an ongoing process.

[0108] Within the scope of this invention, according to one aspect of the present invention, a theory about the rules and laws of a current situation is therefore defined as a self-contained, consistent data set from which the robot can calculate the predictions of this theory using piecewise linear interpolation or extrapolation. The data set consists of its own measurement results and possibly also of the parameters of created alter egos. In this way, the robot can independently determine, save, reuse, and improve each of its theories. As part of the self-assessment (see below), it is also able to select and apply the specifically best theory (data set) with the highest competence (see below) in a given situation.

[0109] Actions:

[0110] According to one aspect of the present invention, an action represents an ordered sequence of effector deployments to achieve a goal. Actions can be designed using the robot's known functional relationships between effectors, internal parameters, and known effects. Initially, in the juvenile learning phase (see below), simple actions are performed. The robot learns to accelerate and brake, then accelerate, wait, and hit a barrier (externally braked with a 'risk of injury'), accelerate, steer, brake, and so on. These simple actions are then combined into increasingly complex 'choreographies,' which are then further optimized, for example, by making braking smoother through decreasing pedal pressure with decreasing speed, and by making cornering smoother. Optimization is achieved by a motivation value (see below).), which represents a kind of internal evaluation of the performed action. Thus, from the stored (partial) choreographies, the one with the highest motivational value is selected. In this way, the student can independently and continuously improve the use of their effectors. Furthermore, based on the current situation and the experiences gained during the action, the student can predict their own situation in the future, which forms the basis for their own action planning.

[0111] Motivational values:

[0112] Motivational values ​​reflect a kind of reward. On the one hand, they are determined internally after a self-imposed goal has been achieved (successful); on the other hand, a motivational value can also be transmitted externally (well done). For example, if a route is covered with strong steering movements, low speed, and yet high lateral acceleration, this sequence of actions receives a lower motivational value than the result of 'training,' after which the same route is covered faster but with lower lateral acceleration. This, of course, also has to be stored in memory.

[0113] The motivational values ​​thus represent a form of non-determined, intrinsic goals. They can be linked to long-term goals, such as the end point of a journey, or short-term intermediate goals, such as visiting a gas station when the tank is empty. As the tank fills, the motivational value for 'refueling' slowly increases until it exceeds that of the long-term goal. Pursuit of the long-term goal is interrupted in favor of refueling, and then resumed.

[0114] Target system:

[0115] According to one aspect of the present invention, the robot is always "switched on," although it can of course have an off switch. It is therefore always performing an activity. However, this also includes doing nothing (charging) or, in a resting state, optimizing its own long-term memory of action options (sleeping). The action chosen is always the one with the highest motivation value. If an action not currently being performed achieves a higher motivation value than the one currently being performed, it is canceled, terminated, or interrupted so that the one with the highest motivation can be performed. If the motivation of such a car on the way from Munich to Hamburg to refuel or recharge its batteries becomes greater than reaching Hamburg, it will interrupt the journey at the next opportunity and head for a gas station.

[0116] In the juvenile phase (see below), i.e., the learning phase, of such a robot or car, the motivations for driving, braking, accelerating, and cornering over short distances will receive high motivational values, simply because the internal long-term memory signals that learning is needed, or rather, reducing errors in order to refine one's own actions. Later, in the adult phase (see below), when such behavior causes no or only insignificant changes in the data memory, and error reduction approaches zero, such behavior becomes "boring," receiving only low motivational values, and then, for example, energy conservation is preferred.

[0117] Thus, according to one aspect of the present invention, this target system avoids a commitment to reality, which would mean an interpretation of reality that could possibly prove to be wrong in the future.

[0118] The central template, the ego:

[0119] The central template represents the core and basis of this AI according to one aspect of the present invention. It is called Ego because it places itself at the center of its data processing and initially observes and evaluates everything from its own perspective. Therefore, no data or meaning (e.g., identifying images) needs to be provided or taught to the Ego; it creates everything itself. According to the invention, the central template or Ego in the device can have been created by the device itself through the independent learning of its own abilities or can have been copied and implemented from a comparable device.

[0120] Subjective polar coordinate system:

[0121] All data from this class's observations and experiences are subjective, or rather, according to one aspect of the present invention, are recorded from the perspective of the ego. The spatial concept also follows this idea. A Cartesian coordinate system with an assumed origin of the observed environment somewhere is not designed, in which each detected object is assigned its respective coordinates. Instead, a polar coordinate system can be used. This not only automatically determines the mathematical origin in the ego, but also the ego's subjective orientation, including front-back, top-bottom, right-left, and the motion vector. The surrounding (empty) space is assumed to be given. The observed objects are only recorded via the distance from the observer's own viewpoint, i.e., the origin of the polar coordinate system, and the angle. Empty space in the physical sense is therefore not assumed to be a separate entity or dimension.It is there and simply an option to move there, unless another object is there, standing, or moving (toward it), in which case the 'place' would be occupied, or the space would not be empty. A further advantage of subjectivist polar coordinates arises below for the concept of alter egos (see below).

[0122] Juvenile phase:

[0123] Within strict guidelines, such as the integrity of the observed objects and the robot itself, maintaining operational readiness (battery charge level), objectives such as smooth acceleration and braking, optimized cornering, and others, the robot should / can learn the use of its effectors and the resulting consequences itself, primarily to coordinate its effectors with its internal and external sensors and, through continuous training (100-105), occasionally interrupted by rest periods, for example, to recharge the batteries, and to maintain the internal database (redundancy, consistency, etc.), to refine them until the detected error has been sufficiently reduced from planning and results. For a car, this would primarily be the use of accelerating, steering, braking, reaching a target point (success) or missing it (error), etc.First, the ego sets itself simple goals for simple actions, which are carried out and stored with the results, which are later combined into increasingly complex "choreographies." This training phase comes to an end when the actions of the self- or externally set goals ("drive there") can no longer be optimized, i.e., when the mathematical module can no longer reduce the error or can only marginally reduce it and / or the database can hardly be optimized any further (number and position of the points in memory). As the error reduction decreases ever more slowly, the juvenile phase draws closer to its end. An ANN, on the other hand, requires a large amount of specifically selected and interpreted training data, whereas the ego only needs a real environment—one could say a playroom—where it cannot cause any damage, in order to try out and develop its own abilities.It follows intrinsic goals and not externally imposed ones like an artificial intelligence (ANN), and thus does not interpret the environment or the world. Adult phase:

[0124] In the juvenile phase, the device has learned the orderly sequence of its effector deployments for achieving its set goals. What is the value of the knowledge and experience gained in the juvenile phase? All data about the world that the ego has learned in the world in which it moves are its own data, its own observations, and its own experiences. In a higher scientific or philosophical sense, one must state that these data are true from the ego's perspective, in the sense of "This is so" or "This was so." This thus represents a natural and solid starting point for knowledge. Naturally, the ego must maintain its internal data. In addition to the optimization already mentioned above, it must also ensure that they are free of contradictions and errors. Otherwise, the manufacturer or operator of the robot could not assume that the robot would achieve its intended goals with its designed actions.

[0125] Changes in the environment that one experiences and causes oneself, such as changing location by accelerating, steering, and braking, are experienced and stored by linking internal and external data. There is no causal knowledge that goes beyond one's own effectors, and this does not require any causal knowledge in the sense of questions such as why does pressing the accelerator lead to acceleration, steering to cornering, or braking to negative acceleration. The ego of a car only needs "there are three options" to deliberately change location. On this level, the important thing is "This is how it is," not the "Why." A pigeon on the road certainly has no idea how a car or a cyclist works. But if such a "road user" were to pass it far enough, it would stay put; if the direction of movement were to change in its direction, it would fly away.She doesn't know why, but she is aware of the possible behaviors and therefore pays close attention.

[0126] Alter ego:

[0127] When the (adult) ego detects an object in the observable environment through external sensory processing, the ego simply creates a copy of itself and adjusts the copy's parameters according to the observed properties. The ego is thus the central template for all observable and unobservable objects and phenomena. Hence the names "ego" and, consequently, "alter ego" for the copies.

[0128] By using the ego as the central template, there are no longer any unknown objects. There are only unknown parameters, which can be measured via external sensors, estimated from one's own experience (keyword: prejudice), and whose possible range can be narrowed down by further observations.

[0129] Just as the ego calculates its behavior using its methods, stored data, and parameters and can predict its own situation, the behavior of alter egos can also be calculated using the same methods, data, and adapted parameters. The assumptions thus made about the alter egos are limited to the same physical laws and to the fact that short-term goals can be derived from the observed orientation and direction of movement. Be it a box on the street, a kangaroo in Australia, or an elderly lady with a walker.

[0130] Programmatically, the ego should exist as an object. Then, for a newly emerging object, a copy of the ego can be easily created, and the parameters of the copy of the new object can then be adjusted using external sensors and one's own experience. There could be a list of the alter egos of the current situation into which the new copy is inserted or created:

[0131] Class CEgo { ...};

[0132] CEgo Ego( ... );

[0133] / / Beginning of education & learning, juvenile phase

[0134] / / Beginning of the adult phase

[0135] CEgo AlterEgos[];

[0136] / / new object detected

[0137] AlterEgos[i] = Ego.clone(); / / Clone the ego for the object

[0138] AlterEgos[i].adjust(...); / / Parameter adjustment of the new object The ego has thus internally created an alter ego as a calculable image of an observed object in the environment and, as shown in Figure 6B, it can form the inner loop for the continuous updating of the observed situation consisting of object recognition 311, updating of the observable object parameters 312, prediction of the behavior of the objects 313, adaptation of the sequence of effector deployments 314 and a second, outer loop after reaching the goal and entering a rest and (also) loading phase 315, until it continues with the initialization 310 for a new goal.

[0139] It doesn't matter whether the observed object is familiar or completely new. If it is new, the range of possible feature values ​​is wider, which requires increased attention (sampling frequency). This has the following advantages:

[0140] Every observed object is fundamentally known!

[0141] Behavior prediction of all external objects using your own methods and data!

[0142] Humans no longer have to intervene to program in what is missing.

[0143] The ego continually learns and improves its behavior.

[0144] The ego is capable of empathy and is aware of the current situation (see below).

[0145] The ego can learn from observing alter egos (see below).

[0146] Alter egos can embody unknown causalities and rules (see below).

[0147] Figure 3 shows scenario 1 “one child” according to one aspect of the present invention.

[0148] When the ego, for example the kl of a simple car, encounters another road user, in this example a child, it creates a copy of itself and adjusts the parameters (distance d to the ego, orientation, speed vK, ...):

[0149] Thus, the alter ego can use the ego's methods to predict where it will be and when. And thus, by comparing its own calculations of when and where it will be, the ego can determine whether a potentially dangerous situation might arise, in order to react accordingly—i.e., to brake. Thus, using its own methods and experiences, the child's short-term behavior can be predicted. NOTHING needs to be specially programmed or trained for the child that would not be present when this child first encounters it. Figure 4 shows Scenario 2, "Highway," according to one aspect of the present invention.

[0150] Another example that more clearly demonstrates how powerful this simple idea of ​​alter egos is is a situation on a highway with a truck and two cars behind it, and a car in the overtaking lane, of which the second, PKWEgo, is our ego. Based on its own knowledge of the force law (F=m*a) of braking force and mass, with the parameters of the observed cars and the truck—i.e., with the assumed mass, measured speed, and assumed braking effect—the ego can, using its own methods alone, predict the braking distances of all road users in a potential accident situation, thus determining its own safety distance. But let's first address the ego's "thoughts" or assumptions about other road users for predicting their actions:

[0151] The truck is traveling at a certain speed, which the ego (the passenger car ego) has learned (without needing to be specifically trained) that such tall road users rarely exceed. Therefore, the truck's alter ego won't predict a change in speed.

[0152] The PKWEgo, the focus of the analysis, would be able to drive faster and overtake the LKWO, taking other road users into account. Overtaking would bring the PKWEgo to its destination faster, thus providing the PKWEgo with a higher motivational value at the moment, as it prepares for the overtaking maneuver.

[0153] Car1, another alter ego, inherently thinks the same as the PWKEgo, because the ego would act the same way in its place. The ego therefore assumes that Car1 also wants to overtake LKWO. Of course, Car1's alter ego "sees" the PKWEgo and will carry out its (the ego's) assumed overtaking maneuver, taking the PKWEgo's existence into account. For example, it will activate its indicators and steer with particular attention to the PKWEgo and possibly other road users in the overtaking lane. That is what the PKWEgo's ego "thinks," and how it would predict Car1's behavior.

[0154] Another possible road user, Car 2, is approaching at high speed in the overtaking lane. The ego also creates an alter ego for this, predicting its actions based on the clear road. As shown above, the program, the ego, can predict the entire situation with all relevant road users and plan its own sequence of effector deployments accordingly, preventing a dangerous development.

[0155] Alter ego, consciousness:

[0156] Thus, the ego internally maintains a complete picture of the observable environment, including all recognized objects, and the ability to predict their behavior in the short term. This would be a possible and programmatically feasible definition of consciousness—a genuine one, not a feigned or imitated one.

[0157] The short-term behavioral forecast based on the ability to put oneself in the place of the observed objects means here to consider the situation with all the data converted for the alter egos and to calculate the development of the situation using the calculated probable behavior of the observed objects.

[0158] The ego in the motorway situation above can, without the need for extra routines to be programmed, predict the behavior of all observed objects, assuming that all road users want to move forward as quickly as possible and without an accident, continuously adapt the predictions to the observed actual behavior (braking, accelerating, steering, using indicators, flashing headlights, etc.), and overtake the truck itself at an appropriate moment.

[0159] The ego can easily assess the possible behavior of all participants in this situation by applying its own methods and experiences (data) with the appropriate parameters. Isn't that what humans do?

[0160] The examples make it clear that the basis, the foundation, and the starting point of this class is always the ego, the central template. The more differentiated its capabilities in internal and external sensory perception and the better its training in the juvenile phase, the more intelligent its behavior and the greater its potential to predict the behavior of observed objects. If one also gave the ego a parameter for its own vulnerability—in the case of a car, one would probably speak more of malleability—the ego could also assess the risks of strong or weak contact. BUT, the ego only ever needs to be built, programmed, and trained (once), although the latter part can of course also be transferred from a trained ego. Observed objects do not need to be identified (CN Ns and their training data), nor does the behavior of the identified objects need to be programmed.One could, of course, argue that ego isn't enough to predict the behavior of other road users. But then, one only has to consider how humans try to assess the behavior of other road users. Humans, too, don't know the long-term goals of other road users on a highway; they can only guess at their short-term goals (to move forward as quickly as possible without causing an accident) and prepare accordingly.

[0161] The robot’s capabilities:

[0162] Empathy:

[0163] Above, we demonstrated how the ego, with the help of alter egos, can place itself in the situations of the observed objects in its environment and calculate their behavior from their perspective. One could call this empathy, if the term is not reserved exclusively for humans.

[0164] Learning from observation:

[0165] Learning from observation means observing behavior that is potentially better than one's own and, if possible, adopting it. The basis is the comparison between another's behavior and one's own, and this comparison is made possible by the empathy defined above.

[0166] Here, ego empathy is realized through alter egos, through which the ego places itself in the shoes of other observed objects in order to predict their behavior from their perspective on the environment. Suppose there is a discrepancy between the predicted and observed behavior. And if the alter ego now changes the sequence of effectors' deployment in such a way that the observed behavior is replicated, the ego has the opportunity to copy the observed behavior if it offers an advantage.

[0167] Figure 5 shows Scenario 3 with different "cornering behavior." Let's imagine the ego has learned to always corner at the same distance from the right edge of the road, the line LE. Now it observes a car ahead that cuts the corner by turning earlier but less sharply, thereby also reducing lateral acceleration and thus cornering faster, on the dotted line LB.

[0168] Seen in this light, the empathy described here is a prerequisite for learning from observing and imitating the behavior of others. (What's wrong with assuming something similar also applies to humans—learning through imitation based on empathy?)

[0169] If the data from the external sensors, converted to the situation and position of the alter ego, is linked with the ego's stored, simple and more complex action sequences, which were copied into the alter ego, the ego is able to learn from observation: It can recognize the difference between the self-planned action sequence of its own effectors (always maintaining the same distance from the edge of the road) and the observed one (earlier braking and turning, smaller steering movements, higher cornering speed, and earlier acceleration) and determine that it could thus drive through the corners faster. This form of learning is certainly faster than the usual trial-and-error process.

[0170] Recognizing causal laws:

[0171] In the theory building section, we explained how any observed functional relationship can be replicated. Let's use this ability to recognize natural laws, for which we can also use alter egos if necessary and appropriate. Of course, they don't lie or drive on a road, but they cause changes that can be detected by external and / or internal sensors.

[0172] Let's imagine an experiment in which our small ego observes a falling apple. It creates an alter ego and notices the accelerated movement toward Earth. The ego's first assumption will be that the apple itself caused the acceleration; like the car ego, it has its own "gas pedal" to accelerate. Then the ego will observe that no apple moves on the ground itself, and that other things also fall to Earth and then stop moving. Furthermore, the ego itself might have learned the law of gravity. On a downhill road, without pressing the gas pedal, it would have noticed and learned an acceleration, but always only "downhill," whereas "uphill" requires more gas. The steeper the angle α, the greater the force, according to this familiar formula (with g = acceleration due to gravity, 9.81 m / sec2):

[0173] F = m * g * sin(a) (2)

[0174] We remember that the ego does not explicitly know this equation or the equation for free fall, i.e. the law of gravity (sin(a=90°) = 1), but through the data points and piecewise linear interpolation or extrapolation it has implicit knowledge about this relationship between the relevant quantities.

[0175] The ego thus behaves as if it knew that, as a physicist would say, as a body with a heavy mass, it is subject to the above law of gravity. When it then sees another object and creates its alter ego, the alter ego is also subject to this law of gravity. In this way, this knowledge essentially acquires the status of a law of nature that affects all bodies without needing to be specially programmed. If the behavior of such an ego were observed from the outside, one could not tell whether it truly knows the law of gravity, as a physicist does, or whether it is merely feigning this insight. Even if the ego sees a pigeon sitting on the street that flies away as soon as the ego approaches, it can recognize that the flapping of its wings and the upward acceleration correspond.

[0176] Another example would be a strong crosswind. Imagine the ego is a large empty van. During normal straight-line driving, the lateral acceleration is zero. But then, while driving straight, the van suddenly shifts sideways, and the lateral acceleration sensor triggers. In this case, the ego can simply create a new alter ego for the unexpected force and the unknown cause and link it to this lateral acceleration. If the ego can then also record data from the environment, such as in the forest or on the plains, or the movement of branches and add it to the data logger, the ego has not only detected a previously unknown quantity (crosswind), it has also identified a possible causal relationship. If the ego could also see and assess the extent of the movement of the branches, it could even roughly estimate the strength of the force acting on the side and thus the likely influence on the driving behavior.In this way, the ego has created its own theory with a new variable. Similar to the physicists who introduced dark matter and dark energy, although they are neither visible nor observable, or like Nobel laureate Peter Higgs, who conceived a particle in the 1960s that was then first detected at CERN in 2012. The alter ego, the copy of the self, can of course never discover the true nature of a phenomenon with this approach in the sense of a search for truth, but it is sufficient for a competent grasp of the situation.

[0177] Self-assessment:

[0178] The first form of gaining knowledge has already been described. In the section on "Recognizing Causal Laws," the example of a crosswind was developed. A force (initially) unknown to the ego suddenly creates a lateral acceleration on a straight stretch of road, a measurable effect on the robot or the ego. For this purpose, it creates an alter ego. However, further information is initially missing that would signal to the ego when and with what intensity the phenomenon occurs. Only when the ego observes a temporal relationship between the movements of the branches, sometimes stronger, sometimes weaker, and its own lateral acceleration, sometimes none, and sometimes none, can the ego connect the two and assume the same cause for both. What is happening here? The ego cannot recognize the crosswind itself, but it can recognize its own reaction to it (lateral acceleration) and the reaction of the branches.In this way, it creates an element that sometimes has an effect on the ego and sometimes not, but which not only has an effect on the ego, but also on other observable objects.

[0179] The second form of gaining knowledge is based on comparing one's own behavior with that of other objects. It is premised on the assumption that the behavior of the observed objects is 'rational' and not random, although this could also be recognized as such, which would then make the comparison no longer a criterion for any potential gain in knowledge. Two different behavioral comparisons are conducted, allowing the ego to determine whether its own competence (see below) in a situation is inferior, equal, or superior to that of the observed objects. This means that the observed objects can perceive more, the same, or less of the situation. An observed van suddenly slows down on a straight road. It appears to see something unknown to the observing ego; it perceives less than the object. If inferior competence is determined, this can be understood as a reason toto specifically examine these situations in order to at least raise one's own competence to the level of others. The comparison is not whether the behavior is right or wrong, which would only lead to the well-known problems of truth theories, but only whether and when a behavior changes and whether behaviors repeat themselves. The first behavioral comparison concerns different situations: Do the observed objects exhibit different behaviors in situations that are different for the ego? The second comparison concerns the same situations: Does the behavior of the observed objects repeat itself in situations that are the same for the ego, or does it vary? From these observations of several situations over a longer period of time, the ego can conclude that its own competence is inferior.If objects repeatedly observed in situations that are the same for the ego show different behavior and their behavior also varies in situations that are different for the ego, then one's own competence is superior. If objects repeatedly observed in situations that are the same for the ego repeat their behavior but do not vary in situations that are different for the ego, then one's own competence is equivalent. If objects repeatedly observed in situations that are the same for the ego repeat their behavior but vary their behavior in situations that are different for the ego.

[0180] With this simple comparison, the ego can recognize how competent one's own competence is in certain situations compared to other objects. Mind you, this is not an attempt to compare oneself with the "truth of the real world."

[0181] From these comparisons, the ego can evaluate itself and, for example, determine whether it should behave more cautiously in certain situations and try to identify what is (still) unknowable. It can be understood as an internal mandate to conduct targeted research as part of "constant learning and the continued expansion of insight and knowledge" (see above).

[0182] Competence:

[0183] The results of these comparisons can also be used externally to assess the competence of the robot's abilities, which helps determine its potential applications. Unlike the truth of knowledge about the real world that the ego has acquired, there are no linguistic problems with competence when it comes to comparing competencies or determining greater or lesser competence.

[0184] However, the concept of truth in relation to theories has another important aspect that competence must also fulfill. Theories enable predictions. Therefore, it is important to know the quality of a theory and, consequently, its predictions. In general, the attribution of truth to a theory serves as a criterion for being able to apply it, or for being permitted to apply it. This has the advantage that the criterion of truth automatically excludes all competing theories, thus eliminating the problem of deciding which theory to use. The concept of competence must achieve something similar.

[0185] Competence is defined and used here as the grasp of the relevant influencing factors of a situation or similar situations. Thus, it is initially a limited criterion, the level of which can be determined through a comparative process, in contrast to the truth of a theory, which assumes temporally and spatially unlimited validity. However, this can neither be proven nor does the criterion allow for the comparison of different theories.

[0186] Thus, competence is better suited to describing the capabilities of a robot than the concept of truth and, unlike truth, it can be determined by the robot itself for self-assessment in a formal procedure.

[0187] Figure 6A shows a schematic flow diagram of a method for autonomously controlling a device in the juvenile or training phase. The goal and end of this phase is the sufficiently precise use and application of the effectors, resulting in the smallest possible error. After initialization 200, there is a randomized actuation 201 of effectors and the reading 202 of a plurality of internal sensor units configured to measure internal device properties and a reading 203 of a plurality of external sensor units configured to measure external environmental properties.storing 204 interrelationships that exist between the actuation 201 of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation 201, device properties and environmental properties into a target triplet; controlling 205 the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one stored 204 action instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environmental property is set based on the output triplet and the target triplet;

[0188] Figure 6A shows the juvenile phase:

[0189] 200 Initialization as preparation for learning

[0190] 201 randomized activation of the effectors

[0191] 202 Reading 202 a variety of internal sensor units

[0192] 203 internal device properties and a readout 203 from a variety of external sensor units

[0193] 204 Saving 204 of Interrelationships

[0194] 205 Determine the error of the movement precision, immediately, if too large continue at 201

[0195] Optional: Rest period for charging and ex post for data preparation:

[0196] Errors, redundancy, consistency,...

[0197] Figure 6B shows the adult phase:

[0198] 310 Initialization as preparation to achieve a longer-term goal

[0199] 311 Recognizing the objects 311 in the current situation

[0200] 312 Creating or updating alter egos or their parameters 312

[0201] 313 Behavioral prediction 313 of the observed objects using the current alter egos

[0202] 314 Calculation of the optimal use of effectors 314 for further target tracking

[0203] 315 Rest period 315 for charging and ex post data processing:

[0204] Errors, redundancy, consistency,... Here, robots, devices, systems and ego are used synonymously.

Claims

Patent claims 1. A method for autonomously controlling a device, comprising: - initializing (100) by means of randomized actuation (101) of effectors and reading (102) of a plurality of internal sensor units configured to measure internal device properties and reading (103) of a plurality of external sensor units configured to measure external environmental properties; - storing (104) interrelationships existing between the actuation (101) of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation (101), device properties and environmental properties into a target triplet; and - controlling (105) the device according to at least one predefined target device property and at least one predefined target environment property using at least one stored (104) action instruction which specifies how the at least one predefined target device property and the at least one predefined target environment property are set based on the initial triplet and the target triplet, wherein the at least one action instruction is used which maximally increases a motivation value and the motivation value specifies how closely the conversion of the initial triplet to the target triple reaches the at least one predefined target device property and the at least one predefined target environment property.

2. Method according to claim 1, characterized in that the method is iterated in a learning randomized manner such that new action instructions are always recognized, which each convert an output triple into a target triple and these action instructions are stored for controlling (105) the device.

3. Method according to claim 1 or 2, characterized in that the reading (102, 103) of internal and / or external effectors is carried out in such a way that individual support values are stored and intermediate values are interpolated and / or extrapolated.

4. Method according to one of the preceding claims, characterized in that the target device property and / or the target environment property are at least temporarily dependent on a device property and / or an environment property is influenced.

5. Method according to one of the preceding claims, characterized in that several instructions are combined to form a choreography which converts an initial triple into a target triple.

6. Method according to one of the preceding claims, characterized in that the method is carried out on a first device and, based on external sensor units of the first device, a second device is recognized which also carries out the method and whose generated interactions are stored on the first device.

7. The method according to claim 6, characterized in that the first device is controlled (105) based on the detected interactions of the second device.

8. Method according to one of claims 6 or 7, characterized in that interrelationships of the detected second device between its parameters and its actions are used to create action instructions of the first device.

9. Method according to one of the preceding claims, characterized in that instructions for action are simulated before their implementation and the maximum motivation value is determined.

10. Method according to one of the preceding claims, characterized in that internal sensor units detect a motivation value, an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a range, a position, a processor utilization, a memory utilization and / or a system parameter of the device.

11. Method according to one of the preceding claims, characterized in that empirically technical factors influencing the motivation value are determined and these are correlated with a stored metric for determining the maximum motivation value.

12. Apparatus adapted to carry out a method according to one of the preceding claims, comprising: - an initialization unit configured to initialize (100) by means of randomized actuation (101) of effectors and reading (102) of a plurality of internal sensor units configured to measure internal device properties and a reading (103) of a plurality of external sensor units configured to measure external environmental properties; - a storage unit configured to store (104) interrelationships existing between the actuation (101) of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation (101), device properties and environmental properties into a target triplet; and - a control unit configured to control (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored (104) action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the initial triplet and the target triplet, wherein the at least one action instruction is used which maximally increases a motivation value and the motivation value specifies how closely the conversion of the initial triplet to the target triple reaches the at least one predefined target device property and the at least one predefined target environment property.

13. System arrangement comprising a plurality of devices according to claim 12 which are communicatively coupled and exchange instructions.

14. A computer program product comprising instructions which, when the program is executed by at least one computer, cause the computer to carry out the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium comprising instructions which, when executed by at least one computer, cause the computer to carry out the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Autonomous driving of a device

    EP4403315A1

  • Autonomous vehicle system

    US20220126864A1

  • Annotation-Free Conscious Learning Robots Using Sensorimotor Training and Autonomous Imitation

    US20220339781A1