Autonomous control of a device

EP4633874A1Pending Publication Date: 2025-10-22SCHREIBER CARL ALBERT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023734524
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-06-19
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing methods for artificial intelligence and autonomous systems require extensive preparatory work, including the creation and specification of training data, which can be error-prone and time-consuming, and are not effective in handling unknown objects or environments, as they lack the ability to autonomously learn and adapt.

Method used

A method for autonomously controlling devices through randomized actuation of effectors, reading internal and external sensor data, and storing interrelationships to generate action instructions, allowing the device to learn and improve without initial knowledge, and to recognize and respond to new situations.

Benefits of technology

Enables devices to autonomously learn and adapt, reducing the need for preparatory work, improving action recognition and decision-making in dynamic environments, and allowing for continuous improvement of knowledge and behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method for autonomously controlling a device. The device in general moves in a physical real world and does not need to be limited in terms of the physical design of the device. The method is advantageous in that the device carries out an autonomous learning process and the learned knowledge or behavior is continuously improved. In general, the disadvantage in the prior art that training data must first be produced, as is the case in current artificial intelligence methods, is overcome. In general, the method is universally applicable, and the device learns automatically and constantly corrects its own knowledge base in the process. The invention additionally relates to a device which is designed to carry out the method, to a system assembly having multiple proposed devices, to a computer program product, and to a computer-readable storage medium which carries out the steps or prompts a computer to carry out the method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Autonomous control of a device

[0002] The present invention is directed to a method for the autonomous control of a device, wherein the device generally moves in a physical real world and does not need to be further specified with regard to its physical design. The method provides the advantage that the device carries out autonomous learning and continuously improves the learned knowledge or behavior. In general, the disadvantage of the prior art that training data must first be created, as is the case with common artificial intelligence methods, is overcome. In general, the method is universally applicable and the device learns automatically, constantly correcting its own knowledge base. Furthermore, a device is proposed which is configured to carry out the method, as well as a system arrangement comprising several of the proposed devices.Furthermore, a computer program product and a computer-readable storage medium are proposed which carry out the method steps or cause a computer to carry out the method.

[0003] Various methods from artificial intelligence (AI) are known from the state of the art. These involve providing training data and then having algorithms recognize regularities in a supervised learning phase. In this case, special algorithms can recognize regularities in the database and then extract them. This way, implicit knowledge is extracted from large amounts of data. Once the training phase is complete, the appropriately trained algorithms are applied to actual data and, from this, generate implicit dependencies for solving real-world problems. A disadvantage of this is that the training data must first be provided, which can already have an undesirable influence on the learning phase. This is not only error-prone but also complex.

[0004] Furthermore, swarm intelligence is known from the prior art, in which several devices are provided, which then solve a problem collaboratively in a distributed manner. Here, too, the control and coordination of the individual participants is complex and sometimes error-prone. This is especially true when distributed learning is to take place, which in turn requires coordination.

[0005] Artificial neural networks are also known from the state of the art. These networks, based on graph theory, provide neurons and connections, i.e., nodes and edges. It is known that these networks mimic functions of the human brain and can learn in the process. Edge weights can be varied, new edges can be added or old ones deleted, and existing nodes can be deactivated or new nodes can be added. This results in a dynamically learning overall system.

[0006] However, the use of an ANN in the manner presented above has at least four problems:

[0007] 1 . For objects on which the ANN has not been trained, the ANN does not provide any information to call the routines associated with the object.

[0008] 2. The programmed properties or behaviors of the detected and identified objects may either be incorrect, have changed in the meantime, or are still unknown.

[0009] 3. One or more scientists, experts, etc., select the training data and specify the desired results. Even in unsupervised learning, the training data and hyperparameters are still specified by humans, and this cannot eliminate the possibility of errors.

[0010] 4. An ANN cannot correct errors. Every single piece of information is distributed across all of the ANN's parameters, similar to a hologram in which each data point contains information about the entire image. Therefore, an ANN must always be completely erased and retrained.

[0011] This means there's no recognition failure (unlike with ANNs) when the ego encounters a previously unknown object, such as in the video when the self-driving car suddenly sees two boxes on the road. ANN generally stands for artificial neural network.

[0012] There is a need in the prior art to create autonomous control of devices in such a way that complex preparatory work such as the provision of training data and teaching can be prevented or minimized. There is therefore a need for a self-learning system that is also capable of collectively learning or estimating what other participants are likely to do or are even capable of doing. There is therefore a need for a system that comprises multiple participants, with each participant actually maintaining knowledge about the other participants and developing this knowledge further, which is referred to herein as empathy. Accordingly, it is an object of the present invention to propose a method for autonomously controlling a device that acts independently solely on the basis of provided target information and in doing so builds up action knowledge.The proposed method is intended to be able to detect the scope of action of effectors of any device and to create an improved method or at least an alternative method for autonomously controlling the device(s). Furthermore, it is an object of the present invention to propose a device for carrying out the method and a system arrangement comprising a plurality of the proposed devices. Furthermore, it is an object to propose a computer program product and a computer-readable storage medium that contain instructions that execute the method.

[0013] The problem is solved by the features of patent claim 1. Further advantageous embodiments are specified in the subclaims.

[0014] Accordingly, a method for autonomously controlling a device is proposed, comprising initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties; storing interrelationships existing between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation, device properties, and environmental properties into a target triplet;and controlling the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet;

[0015] The proposed method is used for the autonomous control, i.e. the generation of instructions, of a device that is generally physically embodied. This could be a car, a robot, or a production facility. There are no restrictions on mobility in this case, meaning the device can move on land, in the air, on water, or even underwater. Autonomous in this context means that the method ensures that the device learns independently and automatically recognizes which possible courses of action are available. These can be learned, and the sensors learn which action leads to which result from which starting point. In general, it is possible for the device to define its own targets, or the targets can be transmitted externally.Internal objectives might, for example, be maintaining operational capability. An internal objective might, for example, be to visit a charging station when the battery level is low. An external objective can be communicated by specifying what task or activity the device should perform.

[0016] In a preparatory process step, initialization occurs by randomly activating effectors. An effector is generally a physical entity that influences the real world. This can, for example, affect the device itself, such as steering, accelerating, braking, etc., and / or the targeted manipulation of objects by, for example, robot arms, directly or with tools, or by remote-acting instruments, be they projectiles or beams or, for example, fire-fighting agents such as water or extinguishing foam. In the example of an automobile, this could be the brake, the accelerator pedal, etc., but also headlights, indicators, etc. The effectors are activated, and the effects of these actions are then measured using internal and external sensors. This is logged and can then be used in the further course of the process.In this way, actions are learned and used effectively in later procedural steps.

[0017] Since the proposed device or method can be implemented or controlled entirely without initial knowledge, the preparatory step involves randomized actuation, i.e., arbitrary actuation, since it must first be learned which action leads to which result. In further iterations of the method, the learned actions are then carried out in a targeted manner. During randomized actuation, the system parameters of the effectors are checked and, for example, a robot arm is moved into all possible positions. This is logged in each case, and it is recognized which effect of which effector action has what effect on the real world. In the example of a headlight, it is possible that this only provides the on and off parameters.The headlights are then switched on and off, and external sensors measure the headlight's impact on the environment, while internal sensors measure operating parameters of the headlight, such as temperature development. Advanced LED headlights also allow adjustment of both the color and intensity. This is tested by randomly activating the headlights, and internal and external results are then logged.

[0018] According to the invention, internal sensor units and external sensor units are proposed. Internal and external do not refer to the arrangement of the sensor units, but rather, according to the present invention, they mean that the sensor units measure external conditions with respect to the device or, analogously, internal conditions. Internal conditions are all system parameters of the device itself. External conditions are environmental variables, i.e., environmental variables surrounding the device. Thus, device properties are measured by means of the internal sensor units, and environmental properties are measured by means of the external sensor units. For example, if the device has a certain battery level, this is an internal parameter.If an object is transferred from A to B using a robot arm, this is an external state, while the state of the robot arm is, in turn, an internal state. Typically, internal and external parameters interact, and actions always result in, for example, a reduced battery level, while these actions, in turn, affect external properties. In the other direction, it is an interaction that, using effectors, the device can be brought closer to a charging station, which then initiates a charging process, which in turn influences the internal battery level.

[0019] The recorded interrelationships are then saved, thus logging how the activation of the effectors influences the internal and external parameters. The device properties are therefore saved along with the environmental properties, and thus instructions for action are defined. Activating the effectors is therefore an action that influences both the device itself and its environment. For example, if the device moves from a first geographical point to a second geographical point, this changes the device's environment, which is an external parameter, and this also changes the internal state of the device, for example, through a temperature development or a change in the battery level. In this way, a triplet can be saved that indicates what the internal state was, what the external state was, and which action is then carried out.This is converted into a new internal state and a new external state. Thus, an output triplet consists of the activation of the effectors, device properties, and environmental properties, and this is converted into a target triplet. The new internal state, the new external state, and possible actions can be specified in the target triplet.

[0020] The concept of triples is merely intended to illustrate the general transition of system states. Alternatively, internal states and external states can be transformed into a new internal state and a new external state using a function. Thus, the action consists of a mapping function from a first tuple to a second tuple. Those skilled in the art will recognize that the power of both descriptions is equivalent. Either an output tuple is transformed into a target tuple using a function, or an input triple is formed, which specifies internal states, external states, and an action. This action then leads to a target triple, which has a new internal state, a new external state, and another field.The remaining field can either remain empty, although it is preferable that at least one action be performed here that is now possible in this state. The last field can also be filled in such a way that, for example, a vector is introduced that contains identifiers for further actions.

[0021] To illustrate this with an example, the device can activate the effector called the motor drive and move forward. What is now logged is an output triplet of an internal state, namely a battery charge level, an external state, namely an image signal of a physical environment, and an action instruction, namely "drive." This output vector or the output triplet is then converted into a target triplet, which contains a new battery charge level, a new image signal of the environment, and actions that would now be possible, such as moving forward or reversing. During the forward movement, it is also possible to specify that, for example, braking or activating the headlights would be possible.

[0022] In an alternative example, the output tuple of the battery state and the image signal of the environment is converted using the "Move" function into a target tuple that describes the new battery state and a new image signal of the environment. This creates a record of what happens internally and externally during a specific action. This knowledge can be reused in subsequent action steps toward a predefined goal.

[0023] In further method steps, target device properties or predefined target environment properties can be provided. Thus, the device can be controlled such that the at least one predefined target device property and / or the at least one predefined target environment property is established. The action instructions are known from the preceding method steps, and thus it is known how an internal state and an external state can be converted into a further internal state and a further external state. The target specifications, i.e., the target device property and / or the target environment property, serve this purpose. These target properties therefore determine what is to be achieved, and this is achieved by executing the action instructions.The device's current state is read out, and then stored actions are selected that, starting from the actual state, reach the target state. In a preferred case, executing an action instruction that immediately achieves the target specifications is sufficient. In a typical case, however, the initial state is successively transformed into the target state using several action properties. Thus, several action instructions are linked in such a way that the target specifications are ultimately achieved.

[0024] To illustrate this with an example, the method may have identified that the device can move from A to B to C to D. Thus, the instructions are stored which stipulate that the car can drive from A to B, from B to C and from C to D. It is also implicitly stored that the car can drive from A to C and from A to D. Furthermore, it is stored that the car can drive from C to D. These possibilities are now used and if a target specification is given which states that the car should drive from A to C, the device has two options to choose from, namely to drive from A to B to C or to drive from A to C. In general, it would also be possible to drive from A to D and then from D to C. Which selection is made can also be defined in the target specifications. For example, a maximum duration of the journey can be defined in time.In addition, a maximum energy consumption can be defined. In preparatory process steps, actions are carried out randomly and are then available at runtime. In general, the process can provide for randomized actuation of the effectors, but it can also additionally provide for the reuse of previously learned action instructions. This allows for a learning phase of randomized actuation, as well as an execution phase that uses previously stored actions. These can also be alternated in any order. Furthermore, in the execution phase, the stored action knowledge is refined using internal and external sensors, and new action instructions are generated that may be preferential to a different objective.Furthermore, the target specifications can change during the execution of an action in such a way that, for example, a battery condition becomes critical. The target for the system's own operability is increased to such an extent that the system recognizes that it is now necessary to drive to a charging station. This results in an overriding target specification, at least temporarily, namely to drive to a charging station, even if this does not achieve the goal of driving to the original target coordinate. As soon as a certain charge level is reached, the target for the system's own operability is downgraded again, and the target for driving to the destination is increased to such an extent that the journey can now be continued.

[0025] According to one aspect of the present invention, the method is iterated in a learning manner such that new instructions are constantly being recognized, each of which converts an output triplet into a target triplet, and these instructions are stored to control the device. This has the advantage that the method is established, in whole or in part, in such a way that new situations are constantly being recognized and instructions are created such that effectors are actuated. The method can thus be varied such that the actuation of the effectors is no longer randomized, but rather existing instructions can be combined, or it is also possible to randomize parts of existing instructions, so that new possible instructions are created, for which it is also clear which interactions they trigger.Furthermore, the method can also be established in such a way that, starting with the storage of interrelationships, it is established up to the control of the device. This means that even when known instructions are carried out, interrelationships are stored and new parameters can be identified. For example, it can be recognized that if the same instruction is carried out twice, different target parameters arise. This can be the case, for example, if a crosswind arises while driving a car. This is stored and then the external sensor detects that wind must have occurred here. This creates a new output triplet and, using the effectors, it can be randomly tested how to countersteer in such a situation.Thus, new interrelationships have been identified and if such an initial triple is identified again, it is now clear which action must be carried out in order to achieve the desired target triple.

[0026] According to a further aspect of the present invention, the reading of internal and / or external effectors is carried out by storing individual support values ​​and interpolating and / or extrapolating intermediate values. This has the advantage that the support values ​​can be selected according to their availability and that they can also be estimated in such a way that a measurement does not have to be available for every possible value. Rather, existing measurements can be used to estimate which values ​​would result under normal circumstances. This allows additional support values ​​to be calculated mathematically.

[0027] According to a further aspect of the present invention, the target device property and / or the target environment property are at least temporarily influenced by a device property and / or an environment property. This has the advantage that the target specifications can also be changed or are subject to prioritization. For example, if the vehicle is to travel from a first geographical point to a second geographical point, the approach to a charging station can be prioritized in the meantime, since otherwise the final destination would not be reached. This results in the advantage of always guaranteeing that the target specifications are met, and a fail-safe method is created.

[0028] According to a further aspect of the present invention, multiple action instructions are combined into a choreography that transforms an initial triplet into a target triplet. This has the advantage that even complex, i.e., compound actions can be executed. As the method progresses or becomes established, new action instructions are continually created, which are combined into increasingly efficient choreographies. Thus, the proposed method is iteratively improved.

[0029] According to a further aspect of the present invention, an object is detected using external sensors and its detected parameters are compared with device properties and / or action instructions of the devices. This has the advantage that the device can essentially clone itself or can infer the properties of the object from its own properties. This can also be referred to as empathy. If, for example, a device is identified using the external sensors that has similar characteristics to the executing device, it is assumed that this newly detected device has similar capabilities to the executing device. To illustrate this in an example, a vehicle is used that implements the proposed invention.This first vehicle therefore carries out the proposed method and, using an external sensor—in this case an imaging unit—detects a further, second device. The first device, or rather the first vehicle, has learned that at a certain speed it can only brake to a limited extent and can only steer laterally to a limited extent. The first vehicle now detects that the second vehicle has similar dimensions and is traveling at a similar speed to the first vehicle. The instructions are then conceptually projected onto the second vehicle, and the first vehicle detects that the second vehicle can only brake to a limited extent and can only change direction to a very limited extent. If an overtaking maneuver is now initiated on a motorway, the first vehicle detects that the second vehicle could brake or also change lanes.Thus, based on its own behavior, or rather the possible starting triplets and the possible target triplets, it is concluded that this object has similar properties. Furthermore, the behavior of the detected second vehicle can be monitored using external sensors, and the system's own data storage, which stores the interrelationships, can be updated. Thus, starting triplets and target triplets can also be recognized, and in turn, conclusions can be drawn from the second vehicle to the first vehicle.

[0030] According to a further aspect of the present invention, the detected parameters are used to predict the behavior of the detected object. This has the advantage that it is recognized that the device itself carries out a certain action instruction in a certain situation, so that the other detected object would very likely also carry out the same action in the same situation. The device's own actions or executed action instructions are therefore logged and it is then assumed that the detected object could behave in the same or at least similarly. According to a further aspect of the present invention, interrelationships of the detected object between its parameters and its actions are used to create action instructions for the device. This has the advantage that externally recognized interrelationships with regard to the proposed device orof the device controlled by the method. Thus, the device not only creates interactions that are evaluated, but rather the external world can also be observed, and conclusions can then be drawn about the conversion of the source triplet into the target triplet.

[0031] According to a further aspect of the present invention, an effector is a processor, a memory, a motor, a gripper arm, a locomotion unit, a drive, an extremity, a windshield wiper, a turn signal, a light, a steering system, and / or an executing unit. This has the advantage that, based on the proposed device or the device to be controlled according to the method, all possible hardware units can be present. The embodiment is merely exemplary and not exhaustive, so that any physical unit can be controlled autonomously according to the proposed method.

[0032] According to a further aspect of the present invention, internal sensor units detect an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a range, a position, a processor utilization, a memory utilization, and / or a system parameter of the device. This has the advantage that the internal sensor units can measure all parameters and states of the device that executes the method or that is executed or controlled by the method.

[0033] According to a further aspect of the present invention, external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal, and / or an external parameter. This has the advantage that the entire environment of the device can be analyzed and recorded. For this purpose, all possible sensors that detect the environment in some way are possible, either individually or in combination.

[0034] The object is also achieved by a device arranged to carry out a

[0035] Method according to one of the preceding claims, comprising a

[0036] Initialization unit configured for initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties; a storage unit configured to store interrelationships existing between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation, device properties, and environmental properties into a target triplet;and a control unit configured to control the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet;

[0037] The problem is also solved by an arrangement comprising several devices which are linked by communication technology and exchange instructions.

[0038] The problem is also solved by a computer program product with control commands that implement the proposed method or operate the proposed device.

[0039] According to the invention, it is particularly advantageous that the method can be used to operate the proposed devices and units. Furthermore, the proposed devices and units are suitable for implementing the method according to the invention. Thus, each device implements structural features suitable for executing the corresponding method. However, the structural features can also be configured as method steps. The proposed method also provides steps for implementing the function of the structural features. Furthermore, physical components can also be provided virtually or in a virtualized form.

[0040] Further advantages, features and details of the invention will become apparent from the following description, in which aspects of the invention are described in detail with reference to the drawings. The features mentioned in the claims and in the description can each be essential to the invention individually or in any combination. Likewise, the features mentioned above and those further explained here can each be used individually or in groups in any combination. Parts or components with similar functions or that are identical are sometimes provided with the same reference numerals. The terms “left”, “right”, “top” and “bottom” used in the description of the exemplary embodiments refer to the drawings in an orientation with normally legible figure designations or normally legible reference numerals.The embodiments shown and described are not intended to be exhaustive, but rather are exemplary in nature to illustrate the invention. The detailed description is intended to inform those skilled in the art; therefore, known circuits, structures, and methods are not shown or explained in detail in order not to obscure the understanding of the present description. The figures show:

[0041] Figure 1: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention;

[0042] Figure 2: A diagram that can be used to estimate internal and / or external properties. This allows internal and / or external parameters to be measured and interpolated between the sampling points or to extrapolate to further sampling points.

[0043] Figure 3: a representation of a real-world situation in which a child wants to cross a street and the proposed device recognizes the situation and provides instruction;

[0044] Figure 4: a real-world situation in road traffic in which the proposed device or method according to one aspect of the present invention is applied; and

[0045] Figure 5: another real-world situation in road traffic, wherein the device practices cornering according to another aspect of the present invention.

[0046] Figure 6A: a schematic flow diagram of a method for autonomously controlling a device according to an aspect of the present invention in the juvenile phase with the primary goal of specifying one's own actions; and Figure 6B: a schematic flow diagram of a method for autonomously controlling a device according to an aspect of the present invention in the adult phase, the application phase for achieving one's own goals, taking into account and predicting the behaviors of other objects relevant in a situation.

[0047] Figure 1 shows a schematic flow diagram of a method for autonomously controlling a device, comprising initialization 100 by means of randomized actuation 101 of effectors and reading 102 of a plurality of internal sensor units configured to measure internal device properties and reading 103 of a plurality of external sensor units configured to measure external environmental properties; storing 104 of interrelationships that exist between the actuation 101 of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation 101, device properties and environmental properties into a target triplet; and actuation 105 of the device according to at least one predefined target device property and / or at least one predefined

[0048] Target environment property using at least one stored 104 action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet.

[0049] The model presented here, according to one aspect of the present invention, uses a central template that helps the child navigate the unknown and constantly changing world. This approach allows for short-term behavioral predictions of observed objects, thus assessing the development of situations and incorporating them into the design of one's own actions. It also enables learning from observation and empathy. Aspects of the invention include:

[0050] - Central template as the center of the AI ​​for recording all objects of a situation; data storage that enables the recognition of laws;

[0051] - Behavioral predictions of all observed objects in a situation; empathy ability based on the central template; ability to learn from observations;

[0052] Identification of influencing factors that cannot be directly observed;

[0053] Recognition of natural laws and causal relationships; computer-friendly form of theories that can be self-created, stored, reused, and continuously improved; and / or self-assessment and evaluation of self-created theories as a basis for cognition, learning, and the selection of the best theory for achieving one's own goals.

[0054] As an example, a car that behaves according to these ideas is described. How does it assess a child or a box, such as overtaking on the highway, or how does it recognize the law of gravity and crosswinds (as an example of recognizing causality and non-manifestable objects). It further describes how learning from observation is possible through empathy. All of this shows that the possibilities of this class extend beyond driving cars. The essential basis, however, is a physically existing robot in a physically existing environment that directly or indirectly observes other objects.

[0055] Some aspects of the present invention are proposed below, which enable an exemplary implementation of the method or device and / or system arrangement. The following aspects are to be understood merely as examples and can be applied individually or in combination.

[0056] Possible hardware of the robot or the proposed device:

[0057] This describes the technical equipment that the robot or car may have according to one aspect of the present invention.

[0058] Effectors:

[0059] In simple cases, such as a car, the effectors would be, apart from windshield wipers, turn signals, and lights. Braking, acceleration, and steering, with which the robot influences its environment, even if it's just changing its own location, are also a change of situation.

[0060] Internal sensors:

[0061] The internal sensors measure and provide feedback on the body's own characteristics, such as height, weight, speed, acceleration, etc., based on the actions of the effectors. These are necessary for assessing the body's own actions and are prerequisites for learning. This allows the body to recognize its own actions, such as the strength of braking and the effect (the extent of negative acceleration), make connections, and store memory for independent learning, meaning it can continuously improve its actions.

[0062] As a physical entity in a physical environment, the robot is subject to the laws of physics. A measured acceleration, i.e., a change in speed, depends not only on the mass but also on how far the accelerator pedal is pressed or how hard the brake is applied. If these laws are not to be hard-coded, but rather the robot is to discover, store, and use them independently, it requires internal sensors to learn and store the consequences of the use of effectors. The internal sensors of a car, for example, therefore record data such as size, speed, acceleration, etc.

[0063] External sensors:

[0064] According to one aspect of the present invention, the external sensor system detects the environment (distances, objects, etc.), for example, using lidar, image processing, etc. Of course, the robot must be appropriately informed about its environment. What exactly is required will be described in more detail below. However, it is not intended to predetermine which systems are best suited, as the information requirements are lower due to the use of the central template. Not all theoretically available information needs to be detected and processed.

[0065] Possible robot software:

[0066] Here, the parts are described with which the presented Kl achieves its intelligence according to one aspect of the present invention.

[0067] Data storage:

[0068] Data storage is a highly available data store for data tuples composed of data from internal sensors, effectors, actions, and their motivational values ​​(see below). The store appropriately links executed actions and the associated actual, experienced, and observed results. It therefore only stores internal data in tables and not, for example, pixels of one or more images. Thus, there is only data whose internal meaning is known.

[0069] Another part of this module, according to one aspect of the present invention, is the repeated checking of the internal consistency of the data. This means that it must not contain any contradictions, structures that lead to loops or circles, etc. The goal is that only one decision can ever be derived from the data. The routine that checks this should also remove stored redundant data, not only to keep the database as lean as possible, but also to prevent overadaptation. What is considered redundant depends on the type of theory formation.

[0070] Figure 2 shows the theory formation according to one aspect of the present invention:

[0071] The central template can simulate, save, and use any functional interrelationship for prognosis using the simple means described here.

[0072] Braking or acceleration functionally depend on the accelerated mass and the applied force. However, this function is initially unknown, because the mathematical function of the velocity change dV depends on the conditions Vn (current speed, weight, etc.) and the use of the effectors Em (force on the accelerator pedal, steering, etc.), thus: dV = f(Vn, Em) (1)

[0073] This function should be self-learned rather than programmed, as this allows the learned function to be expanded, corrected, and improved later in the same way. The effective acceleration during braking (or accelerating) is linked via the internal sensors of the measured acceleration a to the (mathematically) independent variable of force F (how hard the respective pedal is pressed) and mass m according to the law F = m * a. The principle by which the robot itself finds this equation is simple.

[0074] To illustrate, let us take a (different, arbitrary) quadratic function y = -((x- 4) 2 )+9, which is to be reproduced.

[0075] Suppose that initially only the two measured values ​​at points 11 and 12 exist. Now suppose that at the point x=3 (point 2) a forecast p is required, which leads to a value of y=3.8 via the straight line a between 11 and 12. However, since the error compared to the ex post measured value y=8.0 is much too large, this point 2R at x=3.0 and y=8.0 is stored in memory. Subsequently, a new forecast would use either the degrees b1 between [11, 2R] or b2 between [2R, 12]. Over time, many more points arise in the above example (31, 32, 33,..), via which the functional relationship between x and y can be reproduced with any desired accuracy using the respective straight lines along the data tuples (c1, c2, c3, ..). The accuracy is only limited by the lack of experience (yet) and the size of the memory.Of course, this can be extended to any number of variables, if the computing and storage capacities allow it; instead of straight lines, hyperplanes a1x1 + ... + anxn = const are used for the forecast by means of piecewise linear interpolation or extrapolation.

[0076] However, to keep the number of stored corner or data tuples as small as possible, they can also be removed: If an ex post analysis reveals that a data tuple is so close to or even on the line / hyperplane between the two neighboring tuples, then this tuple can be deleted as redundant. This not only reduces the memory requirements but also the computational effort, and over time the behavior adapts better and better to the actual, albeit still unknown, mathematical function, which then leads to increasingly more efficient behavior over time.

[0077] This approach also prevents overadaptation, which must be considered when selecting and sizing training data for an ANN. Redundant experiences are thus prevented from being repeatedly stored, which would completely overshadow the rare events, the so-called "black swans," with potentially severe consequences.

[0078] Storing functional relationships between variables has another advantage: it also stores every possible inverse function, as the stored data tuples do not distinguish which values ​​were the 'dependent' and which were the 'independent' when they occurred. In the above function, x is the independent variable, and the corresponding y-value can be determined using the function y = -((x-4) 2)+9. The x-value would have to be searched for in the database, and if the x-value is not explicitly stored, the y-value can be approximately determined using the next two x-values. However, one can also simply search for the y-value in the memory and determine an x-value using piecewise linear interpolation or extrapolation – this is much easier than trying to determine the inverse function of, for example, the equation above. In this very simple way, any functional relationship can be simulated with any degree of accuracy; the only prerequisite is sufficient action training, as in a complex sport.Of course, the accepted error E can also be changed and adjusted over time, whether it needs to be reduced to increase the required precision, or whether it can be increased to reduce the memory and computational effort, since this in turn influences the number of stored corner and data points; it is a continuous optimization in the ongoing process.

[0079] Within the scope of this invention, according to one aspect of the present invention, a theory about the rules and laws of a current situation is therefore defined as a self-contained, consistent data set from which the robot can calculate the predictions of this theory using piecewise linear interpolation or extrapolation. The data set consists of its own measurement results and possibly also of the parameters of created alter egos. In this way, the robot can independently determine, save, reuse, and improve each of its theories. As part of the self-assessment (see below), it is also able to select and apply the specifically best theory (data set) with the highest competence (see below) in a given situation.

[0080] Actions:

[0081] According to one aspect of the present invention, an action represents an ordered sequence of effector deployments to achieve a goal. Actions can be designed using the robot's known functional relationships between effectors, internal parameters, and known effects. Initially, in the juvenile learning phase (see below), simple actions are performed. The robot learns to accelerate and brake, then accelerate, wait, and hit a barrier (externally braked with a 'risk of injury'), accelerate, steer, brake, and so on. These simple actions are then combined into increasingly complex 'choreographies,' which are then further optimized, for example, by making braking smoother through decreasing pedal pressure with decreasing speed, and by making cornering smoother. Optimization is achieved by a motivation value (see below).), which represents a kind of internal evaluation of the performed action. Thus, from the stored (partial) choreographies, the one with the highest motivational value is selected. In this way, the student can independently and continuously improve the use of their effectors. Furthermore, based on the current situation and the experiences gained during the action, the student can predict their own situation in the future, which forms the basis for their own action planning.

[0082] Motivational values:

[0083] Motivational values ​​reflect a kind of reward. On the one hand, they are determined internally after a self-imposed goal has been achieved (successful); on the other hand, a motivational value can also be transmitted externally (well done). For example, if a route is covered with strong steering movements, low speed, and yet high lateral acceleration, this sequence of actions receives a lower motivational value than the result of "training," after which the same route is covered faster but with lower lateral acceleration. This, of course, also has to be stored in memory.

[0084] The motivational values ​​thus represent a form of non-fixed, intrinsic goals. They can be linked to long-term goals such as the end point of a journey or short-term intermediate goals such as visiting a gas station when the tank is empty. As the tank fills, the motivational value for "refueling" slowly increases until it exceeds that of the long-term goal. Pursuit of the long-term goal is interrupted in favor of refueling and then resumed.

[0085] Target system:

[0086] According to one aspect of the present invention, the robot is always "switched on," although it can of course have an off switch. It is therefore always performing an activity. However, this also includes doing nothing (charging) or optimizing its own long-term memory of action options in a rest state (sleeping). The action chosen is always the one with the highest motivation value. If an action that is not currently being performed achieves a higher motivation value than the one currently being performed, it is canceled, ended, or interrupted so that the one with the highest motivation can be performed. If the motivation of such a car on the way from Munich to Hamburg to refuel or charge the batteries becomes greater than reaching Hamburg, it will interrupt the journey at the next opportunity and head for a gas station.

[0087] In the juvenile phase (see below), i.e., the learning phase, of such a robot or car, the motivations for driving, braking, accelerating, and cornering over short distances will receive high motivational values, simply because the internal long-term memory signals that learning is needed, or rather, the error reduction should be reduced in order to refine one's own actions. Later, in the adult phase (see below), when such behavior causes no or only insignificant changes in the data memory, and the error reduction approaches zero, such behavior becomes "boring." 1 , it receives only low motivation values ​​and then, for example, energy saving is preferred.

[0088] Thus, according to one aspect of the present invention, this target system avoids a commitment to reality, which would mean an interpretation of reality that could possibly prove to be wrong in the future.

[0089] The central template, the ego:

[0090] The central template represents the core and basis of this AI according to one aspect of the present invention. It is called ego because it places itself at the center of its data processing and initially observes and evaluates everything from its own perspective. Therefore, the ego does not need to be given any data or its meaning (e.g., identifying images) or trained; it creates everything itself.

[0091] Subjective polar coordinate system:

[0092] All data from this class's observations and experiences are subjective, or rather, according to one aspect of the present invention, are recorded from the perspective of the ego. The spatial concept also follows this idea. A Cartesian coordinate system with an assumed origin of the observed environment somewhere is not designed, in which each detected object is assigned its respective coordinates. Instead, a polar coordinate system can be used. This not only automatically determines the mathematical origin in the ego, but also the ego's subjective orientation, including front-back, top-bottom, right-left, and the motion vector. The surrounding (empty) space is assumed to be given. The observed objects are only recorded via the distance from the observer's own viewpoint, i.e., the origin of the polar coordinate system, and the angle. Empty space in the physical sense is therefore not assumed to be a separate entity or dimension.It is there and simply an option to move there, unless another object is there, standing, or moving (toward it), in which case the 'place' would be occupied, or the space would not be empty. A further advantage of subjectivist polar coordinates arises below for the concept of alter egos (see below).

[0093] Juvenile phase:

[0094] Within strict guidelines, such as the integrity of the observed objects and the robot itself, maintaining operational readiness (battery charge level), objectives such as smooth acceleration and braking, optimized cornering, and others, the robot should / can learn the use of its effectors and the resulting consequences itself, primarily to coordinate its effectors with its internal and external sensors and, through continuous training (100-105), occasionally interrupted by rest periods, for example, to recharge the batteries, and to maintain the internal database (redundancy, consistency, etc.), to refine them until the detected error has been sufficiently reduced from planning and results. For a car, this would primarily be the use of accelerating, steering, braking, reaching a target point (success) or missing it (error), etc.First, the ego sets itself simple goals for simple actions, which are carried out and stored with the results, which are later combined into increasingly complex 'choreographies'. This training phase comes to an end when the actions of the self- or externally set goals ("drive there") can no longer be optimized, i.e., when the mathematical module can no longer reduce the error or can only marginally reduce it and / or the database can hardly be optimized any further (number and position of the points in memory). As the error reduction decreases ever more slowly, the juvenile phase draws closer to its end. An ANN, on the other hand, requires a large amount of specifically selected training data, whereas the ego only needs a real environment—one could say a playroom—where it cannot cause any damage, in order to try out and develop its own abilities.It follows intrinsic goals and not externally predetermined ones like an ANN and therefore does not interpret the environment or the world.

[0095] Adult phase:

[0096] In the juvenile phase, the device has learned the orderly sequence of its effector deployments for achieving its set goals. What is the value of the knowledge and experience gained in the juvenile phase? All data about the world that the ego has learned in the world in which it moves are its own data, its own observations, and its own experiences. In a higher scientific or philosophical sense, one must state that these data are true from the ego's perspective, in the sense of "This is so" or "This was so." This thus represents a natural and solid starting point for knowledge. Naturally, the ego must maintain its internal data. In addition to the optimization already mentioned above, it must also ensure that they are free of contradictions and errors. Otherwise, the manufacturer or operator of the robot could not assume that the robot would achieve its intended goals with its designed actions.

[0097] Changes in the environment that one experiences and causes oneself, such as changing location by accelerating, steering, and braking, are experienced and stored by linking internal and external data. There is no causal knowledge that goes beyond one's own effectors, and this does not require any knowledge of the causes in the sense of questions such as why does pressing the accelerator lead to acceleration, steering to cornering, or braking to negative acceleration. The ego of a car only needs "there are three options" to deliberately change location. On this level, the important thing is "This is how it is," not the "Why." A pigeon on the road certainly has no idea how a car or a cyclist works. But if such a "road user" were to drive past it far enough, it would stay put; if the direction of movement were to change in its direction, it would fly away.She doesn't know why, but she is aware of the possible behaviors and therefore pays close attention.

[0098] Alter ego:

[0099] When the (adult) ego detects an object in the observable environment through external sensory processing, the ego simply creates a copy of itself and adjusts the copy's parameters according to the observed properties. The ego is thus the central template for all observable and unobservable objects and phenomena. Hence the names "ego" and, consequently, "alter ego" for the copies.

[0100] By using the ego as the central template, there are no longer any unknown objects. There are only unknown parameters, which can be measured via external sensors, estimated from one's own experience (keyword: prejudice), and whose possible range can be narrowed down by further observations.

[0101] Just as the ego calculates its behavior using its methods, stored data, and parameters and can predict its own situation, the behavior of alter egos can also be calculated using the same methods, data, and adapted parameters. The assumptions thus made about the alter egos are limited to the same physical laws and to the fact that short-term goals can be derived from the observed orientation and direction of movement. Be it a box on the street, a kangaroo in Australia, or an elderly lady with a walker.

[0102] Programmatically, the ego should exist as an object. Then, for a newly emerging object, a copy of the ego can be easily created, and the parameters of the copy of the new object can then be adjusted using external sensors and one's own experience. There could be a list of the alter egos of the current situation into which the new copy is inserted or created:

[0103] Class CEgo { ...};

[0104] CEgo Ego( ... );

[0105] / / Beginning of education & learning, juvenile phase

[0106] / / Beginning of the adult phase

[0107] CEgo AlterEgos[];

[0108] / / new object detected

[0109] AlterEgos[i] = Ego.clone(); / / Clone the ego for the object

[0110] AlterEgos[i].adjust(...); / / Parameter adjustment of the new object

[0111] The ego has thus internally created an alter ego as a computable image of an observed object in the environment. As shown in Figure 6B, it can continuously update the observed situation by forming the inner loop consisting of object recognition 311, updating the observable object parameters 312, predicting the behavior of the objects 313, adjusting the sequence of effector deployments 314, and a second, outer loop after reaching the target and entering a rest and (also) loading phase 315, until it continues with initialization 310 for a new target. It does not matter whether the observed object is familiar or completely new. If it is new, the range of possible feature expressions is larger, which requires increased attention (sampling frequency). This has the following advantages:

[0112] Every observed object is fundamentally known!

[0113] Behavior prediction of all external objects using your own methods and data!

[0114] Humans no longer have to intervene to program in what is missing.

[0115] The ego continually learns and improves its behavior.

[0116] The ego is capable of empathy and is aware of the current situation (see below).

[0117] The ego can learn from observing alter egos (see below).

[0118] Alter egos can embody unknown causalities and rules (see below).

[0119] Figure 3 shows scenario 1 “one child” according to one aspect of the present invention.

[0120] When the ego, for example the kl of a simple car, encounters another road user, in this example a child, it creates a copy of itself and adjusts the parameters (distance d to the ego, orientation, speed vK, ...):

[0121] In this way, the alter ego can use the ego's methods to predict where it will be and when. And in this way, the ego can compare its own calculations of when and where it will be, and determine whether a potentially dangerous situation might arise, so that it can react accordingly—that is, slow down. In this way, the child's short-term behavior can be predicted using its own methods and experiences. Nothing needs to be specially programmed or trained for the child that wouldn't be present when this child first encounters it.

[0122] Figure 4 shows Scenario 2 “Highway” according to one aspect of the present invention.

[0123] Another example that more clearly demonstrates how powerful this simple idea of ​​alter egos is is a situation on a highway with a truck and two cars behind it, and a car in the overtaking lane, of which the second, PKWEgo, is our ego. Based on its own knowledge of the force law (F=m*a) of braking force and mass, with the parameters of the observed cars and the truck—i.e., with the assumed mass, measured speed, and assumed braking effect—the ego can, using its own methods alone, predict the braking distances of all road users in a potential accident situation, thus determining its own safety distance. But let's first address the ego's "thoughts," or rather, its assumptions about other road users, for predicting their actions:

[0124] The truck is traveling at a certain speed, which the ego (the passenger car ego) has learned (without needing to be specifically trained) that such tall road users rarely exceed. Therefore, the truck's alter ego won't predict a change in speed.

[0125] The PKWEgo, the focus of the analysis, would be able to drive faster and overtake the LKWO, taking other road users into account. Overtaking would bring the PKWEgo to its destination faster, thus providing the PKWEgo with a higher motivational value at the moment, as it prepares for the overtaking maneuver.

[0126] Car1, another alter ego, inherently thinks the same as cargo, because the ego would act the same way in its place. The ego therefore assumes that car1 also wants to overtake truck. Of course, car1's alter ego "sees" cargo and will carry out its (ego's) assumed overtaking maneuver, taking the cargo's existence into account. For example, it will use its turn signals and steer with particular attention to cargo and possible other road users in the overtaking lane. That is what cargo's ego "thinks," and how it would predict car1's behavior.

[0127] Another possible road user, Car 2, is approaching at high speed in the overtaking lane. The ego also creates an alter ego for this, predicting its actions based on the clear road.

[0128] - As shown above, the program, the ego, can predict the entire situation with all relevant road users and plan its own sequence of effector deployments accordingly, so that no dangerous development occurs.

[0129] Alter ego, consciousness: Thus, the ego internally has a complete picture of the observable environment, including all recognized objects, and the ability to predict their behavior in the short term. This would be a possible and programmatically feasible definition of consciousness: a genuine one, not a feigned or imitated one.

[0130] The short-term behavioral forecast based on the ability to put oneself in the place of the observed objects means here to consider the situation with all the data converted for the alter egos and to calculate the development of the situation using the calculated probable behavior of the observed objects.

[0131] The ego in the motorway situation above can, without the need for extra routines to be programmed, predict the behavior of all observed objects, assuming that all road users want to move forward as quickly as possible and without an accident, continuously adapt the predictions to the observed actual behavior (braking, accelerating, steering, using indicators, flashing headlights, etc.), and overtake the truck itself at an appropriate moment.

[0132] The ego can easily assess the possible behavior of all participants in this situation by applying its own methods and experiences (data) with the appropriate parameters. Isn't that what humans do?

[0133] The examples make it clear that the basis, the foundation, and the starting point of this class is always the ego, the central template. The more differentiated its capabilities in internal and external sensory perception and the better its training in the juvenile phase, the more intelligent its behavior and the greater its potential to predict the behavior of observed objects. If one also gave the ego a parameter for its own vulnerability—in the case of a car, one would probably speak more of malleability—the ego could also assess the risks of strong or weak contact. BUT, the ego only ever needs to be built, programmed, and trained (once), although the latter part can of course also be transferred from a trained ego. Observed objects do not need to be identified (ANNs and their training data), nor does the behavior of the identified objects need to be programmed.One could, of course, argue that ego isn't enough to predict the behavior of other road users. But just consider how humans try to assess the behavior of other road users. Humans, too, don't know the long-term goals of other road users on a highway; they can only guess at their short-term goals (to move forward as quickly as possible without causing an accident) and prepare accordingly.

[0134] The robot’s capabilities:

[0135] Empathy:

[0136] Above, we demonstrated how the ego, with the help of alter egos, can place itself in the situations of the observed objects in its environment and calculate their behavior from their perspective. One could call this empathy, if the term is not reserved exclusively for humans.

[0137] Learning from observation:

[0138] Learning from observation means observing behavior that is potentially better than one's own and, if possible, adopting it. The basis is the comparison between another's behavior and one's own, and this comparison is made possible by the empathy defined above.

[0139] Here, ego empathy is realized through alter egos, through which the ego places itself in the shoes of other observed objects in order to predict their behavior from their perspective on the environment. Suppose there is a discrepancy between the predicted and observed behavior. And if the alter ego now changes the sequence of effectors' deployment in such a way that the observed behavior is replicated, the ego has the opportunity to copy the observed behavior if it offers an advantage.

[0140] Figure 5 shows scenario 3 different “curve behavior”.

[0141] Let's imagine the ego has learned to always corner at the same distance from the right edge of the road, the line LE. Now it observes a car ahead that cuts the corner by turning earlier but less sharply, thereby also reducing lateral acceleration and thus cornering faster, on the dotted line LB.

[0142] Seen in this light, the empathy described here is a prerequisite for learning from observing and imitating the behavior of others. (What's wrong with assuming something similar also applies to humans—learning through imitation based on empathy?)

[0143] If the data from the external sensors, converted to the situation and position of the alter ego, is linked with the ego's stored, simple and more complex action sequences, which were copied into the alter ego, the ego is able to learn from observation: It can recognize the difference between the self-planned action sequence of its own effectors (always maintaining the same distance from the edge of the road) and the observed one (earlier braking and turning, smaller steering movements, higher cornering speed, and earlier acceleration) and determine that it could thus drive through the corners faster. This form of learning is certainly faster than the usual trial-and-error process.

[0144] Recognizing causal laws:

[0145] In the theory building section, we explained how any observed functional relationship can be replicated. Let's use this ability to recognize natural laws, for which we can also use alter egos if necessary and appropriate. Of course, they don't lie or drive on a road, but they cause changes that can be detected by external and / or internal sensors.

[0146] Let's imagine an experiment in which our small ego observes a falling apple. It creates an alter ego and notices the accelerated movement toward Earth. The ego's first assumption will be that the apple itself caused the acceleration; like the car ego, it has its own "gas pedal" to accelerate. Then the ego will observe that no apple moves on the ground itself, and that other things also fall to Earth and then stop moving. Furthermore, the ego itself might have learned the law of gravity. On a downhill road, without pressing the gas pedal, it would have noticed and learned an acceleration, but always only "downhill," whereas it needs to accelerate more "uphill." The steeper the angle α, the greater the force, according to this familiar formula (with g = acceleration due to gravity, 9.81 m / sec2):

[0147] F = m * g * sin(a) (2)

[0148] We remember that the ego does not explicitly know this equation or the equation for free fall, i.e. the law of gravity (sin(a=90°) = 1 ), but through the data points and piecewise linear interpolation or extrapolation it has implicit knowledge about this relationship between the relevant quantities.

[0149] The ego thus behaves as if it knew that, as a physicist would say, as a body with a heavy mass, it is subject to the above law of gravity. When it then sees another object and creates its alter ego, the alter ego is also subject to this law of gravity. In this way, this knowledge essentially acquires the status of a law of nature that affects all bodies without needing to be specially programmed. If the behavior of such an ego were observed from the outside, one could not tell whether it truly knows the law of gravity, as a physicist does, or whether it is merely feigning this insight. Even if the ego sees a pigeon sitting on the street that flies away as soon as the ego approaches, it can recognize that the flapping of its wings and the upward acceleration correspond.

[0150] Another example would be a strong crosswind. Imagine the ego is a large empty van. During normal straight-line driving, the lateral acceleration is zero. But then, while driving straight, the van suddenly shifts sideways, and the lateral acceleration sensor triggers. In this case, the ego can simply create a new alter ego for the unexpected force and the unknown cause and link it to this lateral acceleration. If the ego can then also record data from the environment, such as in the forest or on the plains, or the movement of branches and add it to the data logger, the ego has not only detected a previously unknown quantity (crosswind), it has also identified a possible causal relationship. If the ego could also see and assess the extent of the movement of the branches, it could even roughly estimate the strength of the force acting on the side and thus the likely influence on the driving behavior.In this way, the ego has created its own theory with a new variable. Similar to the physicists who introduced dark matter and dark energy, although they are neither visible nor observable, or like Nobel laureate Peter Higgs, who conceived a particle in the 1960s that was then first detected at CERN in 2012. The alter ego, the copy of the self, can of course never discover the true nature of a phenomenon with this approach in the sense of a search for truth, but it is sufficient for a competent grasp of the situation.

[0151] Self-assessment:

[0152] The first form of gaining knowledge has already been described. In the section on "Recognizing Causal Laws," the example of a crosswind was developed. A force (initially) unknown to the ego suddenly creates a lateral acceleration on a straight stretch of road, a measurable effect on the robot or the ego. For this purpose, it creates an alter ego. However, further information is initially missing that would signal to the ego when and with what intensity the phenomenon occurs. Only when the ego observes a temporal relationship between the movements of the branches, sometimes stronger, sometimes weaker, and its own lateral acceleration, sometimes none, and sometimes none, can the ego connect the two and assume the same cause for both. What is happening here? The ego cannot recognize the crosswind itself, but it can recognize its own reaction to it (lateral acceleration) and the reaction of the branches.In this way, it creates an element that sometimes has an effect on the ego and sometimes not, but which not only has an effect on the ego, but also on other observable objects.

[0153] The second form of gaining knowledge is based on comparing one's own behavior with that of other objects. It is premised on the assumption that the behavior of the observed objects is 'rational' and not random, although this could also be recognized as such, in which case the comparison would no longer be a criterion for any potential gain in knowledge. Two different behavioral comparisons are conducted, allowing the ego to determine whether its own competence (see below) in a situation is inferior, equal, or superior to that of the observed objects. This means that the observed objects can perceive more, the same, or less of the situation. An observed van suddenly slows down on a straight road. It appears to see something unknown to the observing ego; it perceives less than the object. If inferior competence is determined, this can be understood as a reason toto specifically examine these situations in order to at least raise one's own competence to the level of others. The comparison is not whether the behavior is right or wrong, which would only lead to the well-known problems of truth theories, but only whether and when a behavior changes and whether behaviors repeat themselves. The first behavioral comparison concerns different situations: Do the observed objects exhibit different behaviors in situations that are different for the ego? The second comparison concerns the same situations: Does the behavior of the observed objects repeat itself in situations that are the same for the ego, or does it vary? From these observations of several situations over a longer period of time, the ego can conclude that its own competence is inferior.If objects repeatedly observed in situations that are the same for the ego show different behavior and their behavior also varies in situations that are different for the ego, then one's own competence is superior. If objects repeatedly observed in situations that are the same for the ego repeat their behavior but do not vary in situations that are different for the ego, then one's own competence is equivalent. If objects repeatedly observed in situations that are the same for the ego repeat their behavior but vary their behavior in situations that are different for the ego.

[0154] With this simple comparison, the ego can recognize how competent one's own competence is in certain situations compared to other objects. Mind you, this is not an attempt to compare oneself with the "truth of the real world."

[0155] From these comparisons, the ego can evaluate itself and, for example, determine whether it should behave more cautiously in certain situations and try to identify what is (still) unknowable. It can be understood as an internal mandate to conduct targeted research as part of "constant learning and the continued expansion of insight and knowledge" (see above).

[0156] Competence:

[0157] The results of these comparisons can also be used externally to assess the competence of the robot's abilities, which helps determine its potential applications. Unlike the truth of knowledge about the real world that the ego has acquired, there are no linguistic problems with competence when it comes to comparing competencies or determining greater or lesser competence.

[0158] However, the concept of truth in relation to theories has another important aspect that competence must also fulfill. Theories enable predictions. Therefore, it is important to know the quality of a theory and, consequently, its predictions. In general, the attribution of truth to a theory serves as a criterion for being able to apply it, or for being allowed to apply it, with the advantage that the criterion of truth automatically excludes all competing theories, thus eliminating the problem of deciding which theory to use. The concept of competence must achieve something comparable. Competence is defined and used here as the understanding of the relevant influencing factors of a situation or similar situations. It is therefore initially a limited criterion, the level of which can, however, be determined through a comparative process, in contrast to the truth of a theory, which assumes validity unlimited in time and space.However, this cannot be proven, nor does the criterion allow for the comparison of different theories.

[0159] Thus, competence is better suited to describing the capabilities of a robot than the concept of truth and, unlike truth, it can be determined by the robot itself for self-assessment in a formal procedure.

[0160] Figure 6A shows a schematic flow diagram of a method for autonomously controlling a device in the juvenile or developmental phase. The goal and end of this phase is the sufficiently precise use and deployment of the effectors, resulting in the smallest possible error.After initialization 200, there is a randomized actuation 201 of effectors and the reading 202 of a plurality of internal sensor units configured to measure internal device properties and a reading 203 of a plurality of external sensor units configured to measure external environmental properties; a storage 204 of interrelations which exist between the actuation 201 of the effectors and the read-out device properties and environmental properties as action instructions in such a way that an action instruction converts an output triplet of actuation 201, device properties and environmental properties into a target triplet; a control 205 of the device according to at least one predefined target device property and / or at least one predefined.

[0161] Target environment property using at least one stored 204 action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the source triplet and the target triplet.

[0162] Figure 6A shows the juvenile phase:

[0163] 200 Initialization as preparation for learning

[0164] 201 randomized activation of the effectors

[0165] 202 Reading 202 a plurality of internal sensor units 203 internal device properties and a reading 203 of a plurality of external sensor units

[0166] 204 Saving 204 of Interrelationships

[0167] 205 Determine the error of the movement precision, immediately, if too large continue at 201

[0168] Optional: Rest period for charging and ex post for data preparation:

[0169] Errors, redundancy, consistency,...

[0170] Figure 6B shows the adult phase:

[0171] 310 Initialization as preparation to achieve a longer-term goal 311 Recognizing the objects 311 in the current situation

[0172] 312 Creating or updating alter egos or their parameters 312

[0173] 313 Behavioral prediction 313 of the observed objects using the current alter egos

[0174] 314 Calculation of the optimal use of effectors 314 for further target tracking

[0175] 315 Rest period 315 for charging and ex post data processing:

[0176] Errors, redundancy, consistency,...

[0177] Here, robots, devices, systems and ego are used synonymously.

Claims

Patent claims 1. A method for autonomously controlling a device, comprising: - initializing (100) by means of randomized actuation (101) of effectors and reading (102) of a plurality of internal sensor units configured to measure internal device properties and reading (103) of a plurality of external sensor units configured to measure external environmental properties; - storing (104) interrelationships existing between the actuation (101) of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation (101), device properties and environmental properties into a target triplet; and - controlling (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored (104) action instruction which specifies how the at least one predefined target device property and / or the at least one predefined Target environment property is set based on the source triple and the target triple.

2. Method according to claim 1, characterized in that the method is iterated in a learning manner in such a way that new action instructions are always recognized, which each convert an output triple into a target triple and these action instructions are stored for controlling (105) the device.

3. Method according to claim 1 or 2, characterized in that the reading (102, 103) of internal and / or external effectors is carried out in such a way that individual support values ​​are stored and intermediate values ​​are interpolated and / or extrapolated.

4. Method according to one of the preceding claims, characterized in that the target device property and / or the target environment property is at least temporarily influenced by a device property and / or an environment property.

5. Method according to one of the preceding claims, characterized in that several instructions are combined to form a choreography which Source triplet converted into a target triplet.

6. Method according to one of the preceding claims, characterized in that an object is detected using external sensor units and its detected parameters are compared with device properties and / or instructions of the devices.

7. Method according to claim 6, characterized in that the detected parameters are used to predict a behavior of the detected object.

8. Method according to one of claims 6 or 7, characterized in that interrelationships of the detected object between its parameters and its actions are used to create action instructions of the device.

9. Method according to one of the preceding claims, characterized in that an effector is present as a processor, a memory, a motor, a gripper arm, a locomotion unit, a drive, an extremity, a windscreen wiper, an indicator, a light, a steering and / or an executing unit.

10. Method according to one of the preceding claims, characterized in that internal sensor units detect an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a range, a position, a processor utilization, a memory utilization and / or a system parameter of the device.

11. Method according to one of the preceding claims, characterized in that external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal and / or an external parameter.

12. Apparatus adapted to carry out a method according to one of the preceding claims, comprising: - an initialization unit configured to initialize (100) by means of randomized actuation (101) of effectors and reading (102) of a plurality of internal sensor units configured to measure internal device properties and a reading (103) of a plurality of external sensor units configured to Measurement of external environmental properties; - a storage unit configured to store (104) interrelationships existing between the actuation (101) of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation (101), device properties and environmental properties into a target triplet; and - a control unit configured to control (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored (104) action instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet.

13. System arrangement comprising a plurality of devices according to claim 12 which are communicatively coupled and exchange instructions.

14. A computer program product comprising instructions which, when the program is executed by at least one computer, cause the computer to carry out the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium comprising instructions which, when executed by at least one computer, cause the computer to carry out the steps of the method according to any one of claims 1 to 11.