AUTONOMOUS CONTROL OF A DEVICE

DE502022005200D1Active Publication Date: 2025-09-11SCHREIBER CARL ALBERT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502022005200
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-09-11
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

Existing artificial intelligence methods require extensive training data selection and preparation, which limits their adaptability and can introduce errors, and distributed learning systems face complexity and coordination challenges.

Method used

A method for autonomously controlling devices through randomized actuation of effectors, using internal and external sensors to log interrelationships between effector actions and environmental/internal properties, converting these into action instructions, and iteratively learning and refining these instructions to achieve predefined targets.

Benefits of technology

Enables devices to learn and adapt without initial training data, recognizing new situations and optimizing actions autonomously, ensuring effective and flexible control.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention is directed to a method for the autonomous control of a device, wherein the device generally moves in a physical real world and does not need to be further specified with regard to its physical design. The method provides the advantage that the device carries out autonomous learning and continuously improves the learned knowledge or behavior. In general, the disadvantage of the prior art that training data must first be created, as is the case with common artificial intelligence methods, is overcome. In general, the method is universally applicable and the device learns automatically, constantly correcting its own knowledge base. Furthermore, a device is proposed which is configured to carry out the method, as well as a system arrangement comprising several of the proposed devices.Furthermore, a computer program product and a computer-readable storage medium are proposed which carry out the method steps or cause a computer to carry out the method.

[0002] TAKAHASHI KUNIYUKl ET AL: "Dynamic motion learning for multi-DOF flexible-joint robots using active-passive motor babbling through deep learning", ADVANCED ROBOTICS, Vol. 31, No. 18, September 17, 2017, pages 1002–1015, XP055840673, presents a learning strategy for robots with flexible joints with multiple degrees of freedom to perform dynamic motion tasks. Although robots with flexible joints offer several potential advantages, such as exploiting intrinsic dynamics and passively adapting to environmental changes with mechanical compliance, controlling such robots is challenging due to the increasing complexity of their dynamics.

[0003] Various methods from artificial intelligence (K1), also known in English as artificial intelligence (AI), are known from the state of the art. These involve selecting and providing training data, and then having algorithms recognize regularities in a supervised learning phase. By specifying initial data and target values ​​and using special algorithms, regularities can be recognized and then utilized. This makes it possible to harness implicit knowledge within large data sets. Once the training phase is complete, the appropriately trained algorithms are applied to actual data, which in turn generate implicit dependencies for solving real-world problems.A disadvantage here is that the training data must first be selected and specified, which limits the resulting artificial intelligence to the training data, and undesirable influences can already be exerted on the learning phase, which is not only error-prone but also time-consuming.

[0004] Furthermore, swarm intelligence is known from the prior art, in which several devices are provided, which then solve a problem collaboratively in a distributed manner. Here, too, the control and coordination of the individual participants is complex and sometimes error-prone. This is especially true when distributed learning is to take place, which in turn requires coordination.

[0005] Artificial neural networks are also known from the state of the art. These networks, based on graph theory, provide neurons and connections, i.e., nodes and edges. It is known that these networks mimic functions of the human brain and can learn in the process. Edge weights can be varied, new edges can be added or old ones deleted, and existing nodes can be deactivated or new nodes can be added. This results in a dynamically learning overall system.

[0006] However, the use of an ANN in the manner presented above has at least four problems: 1. For objects that the ANN has not been trained on, the ANN does not provide any information to call the routines assigned to the object. 2. The programmed properties or behaviors of the detected and identified objects may be incorrect, have changed in the meantime, or are still unknown. 3. One or more scientists, experts, etc. select the training data and specify the results to be achieved. Even with unsupervised learning, the training data and hyperparameters are still specified by humans, and this cannot rule out errors. 4. An ANN cannot correct errors. Every single piece of information is distributed across all of the ANN's parameters, similar to a hologram in which each data point contains information about the entire image. An ANN must therefore always be completely deleted and completely retrained.

[0007] This means there's no recognition failure (unlike with ANNs) when the ego encounters a previously unknown object, such as in the video when the self-driving car suddenly sees two boxes on the road. ANN generally stands for artificial neural network.

[0008] In the current state of the art, there is a need to create autonomous control of devices in such a way that complex preparatory work such as providing training data and learning can be prevented or minimized. There is therefore a need for a self-learning system that is also capable of collectively learning or estimating what other participants are likely to do or are even capable of doing. Thus, there is a need for a system that includes multiple participants, with each participant actually maintaining and developing knowledge about the other participants, which is referred to here as empathy.

[0009] Accordingly, it is an object of the present invention to propose a method for autonomously controlling a device, which acts independently solely on the basis of provided target information and in doing so builds up action knowledge. The proposed method should be able to recognize the scope of action of effectors of any device and create an improved method or at least an alternative method for autonomously controlling the device(s). Furthermore, it is an object of the present invention to propose a device for carrying out the method and a system arrangement comprising a plurality of the proposed devices. Furthermore, it is an object to propose a computer program product and a computer-readable storage medium which contain instructions which execute the method.

[0010] The problem is solved by the features of patent claim 1. Further advantageous embodiments are specified in the subclaims.

[0011] Accordingly, a method for autonomously controlling a device is proposed, comprising initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties; storing interrelationships existing between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation, device properties, and environmental properties into a target triplet;and controlling the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet, wherein the method is iterated in a learning manner in which new instructions are constantly recognized, each of which converts an output triplet into a target triplet, and these instructions are stored for controlling the device.

[0012] The proposed method is used for the autonomous control, i.e. the generation of instructions, of a device that is generally physically embodied. This could be a car, a robot, or a production facility. There are no restrictions on mobility in this case, meaning the device can move on land, in the air, on water, or even underwater. Autonomous in this context means that the method ensures that the device learns independently and automatically recognizes which possible courses of action are available. These can be learned, and the sensors learn which action leads to which result from which starting point. In general, it is possible for the device to define its own targets, or the targets can be transmitted externally.Internal objectives might, for example, be maintaining operational capability. An internal objective might, for example, be to visit a charging station when the battery level is low. An external objective can be communicated by specifying what task or activity the device should perform.

[0013] In a preparatory process step, initialization occurs by randomly activating effectors. An effector is generally a physical entity that influences the real world. This can, for example, affect the device itself, such as steering, accelerating, braking, etc., and / or the targeted manipulation of objects by, for example, robot arms, directly or with tools, or by remote-acting instruments, be they projectiles or beams or, for example, fire-fighting agents such as water or extinguishing foam. In the example of an automobile, this could be the brake, the accelerator pedal, etc., but also headlights, indicators, etc. The effectors are activated, and the effects of these actions are then measured using internal and external sensors. This is logged and can then be used in the further course of the process.In this way, actions are learned and used effectively in later procedural steps.

[0014] Since the proposed device or method can be implemented or controlled entirely without any initial knowledge, the preparatory step involves randomized actuation, i.e., arbitrary actuation, since it must first be learned which action leads to which result. In further iterations of the method, the learned actions are then carried out in a targeted manner. During randomized actuation, the system parameters of the effectors are checked and, for example, a robot arm is moved into all possible positions. This is logged in each case, and it is recognized which effect of each effector action on the real world what effect it has. In the example of a headlight, it is possible that this only provides the on and off parameters.The headlights are then switched on and off, and external sensors measure the headlight's impact on the environment, while internal sensors measure operating parameters of the headlight, such as temperature development. Advanced LED headlights also allow adjustment of both the color and intensity. This is tested through random activation, and internal and external results are then logged.

[0015] According to the invention, internal sensor units and external sensor units are proposed. Internal and external do not refer to the arrangement of the sensor units, but rather, according to the present invention, they mean that the sensor units measure external conditions with respect to the device or, analogously, internal conditions. Internal conditions are all system parameters of the device itself. External conditions are environmental variables, i.e., environmental variables surrounding the device. Thus, device properties are measured by means of the internal sensor units, and environmental properties are measured by means of the external sensor units. For example, if the device has a certain battery level, this is an internal parameter.If an object is transferred from A to B using a robot arm, this is an external state, while the state of the robot arm is, in turn, an internal state. Typically, internal and external parameters interact, and actions always result in, for example, a reduced battery level, while these actions, in turn, affect external properties. In the other direction, it is an interaction that, using effectors, the device can be brought closer to a charging station, which then initiates a charging process, which in turn influences the internal battery level.

[0016] The recorded interrelationships are then saved, thus logging how the activation of the effectors influences the internal and external parameters. The device properties are therefore saved along with the environmental properties, and thus instructions for action are defined. Activating the effectors is therefore an action that influences both the device itself and its environment. For example, if the device moves from a first geographical point to a second geographical point, this changes the device's environment, which is an external parameter, and this also changes the internal state of the device, for example, through a temperature development or a change in the battery level. In this way, a triplet can be saved that indicates what the internal state was, what the external state was, and which action is then carried out.This is converted into a new internal state and a new external state. Thus, an output triplet consists of the activation of the effectors, device properties, and environmental properties, and this is converted into a target triplet. The new internal state, the new external state, and possible actions can be specified in the target triplet.

[0017] The concept of triples is merely intended to illustrate the general transition of system states. Alternatively, internal states and external states can be transformed into a new internal state and a new external state using a function. Thus, the action consists of a mapping function from a first tuple to a second tuple. Those skilled in the art will recognize that the power of both descriptions is equivalent. Either an output tuple is transformed into a target tuple using a function, or an input triple is formed, which specifies internal states, external states, and an action. This action then leads to a target triple, which has a new internal state, a new external state, and another field.The remaining field can either remain empty, although it is preferable that at least one action be performed here that is now possible in this state. The last field can also be filled in such a way that, for example, a vector is introduced that contains identifiers for further actions.

[0018] To illustrate this with an example, the device can activate the effector called the motor drive and move forward. What is now logged is an output triplet of an internal state, namely a battery charge level, an external state, namely an image signal of a physical environment, and an action instruction, namely "drive." This output vector or output triplet is then converted into a target triplet, which contains a new battery charge level, a new image signal of the environment, and actions that would now be possible, such as moving forward or reversing. During the forward movement, it is also possible to specify that, for example, braking or activating the headlights would be possible.

[0019] In an alternative example, the output tuple of the battery state and the image signal of the environment is converted using the "Move" function into a target tuple that describes the new battery state and a new image signal of the environment. This creates a record of what happens internally and externally during a specific action. This knowledge can be reused in subsequent action steps toward a predefined goal.

[0020] In further method steps, target device properties or predefined target environment properties can be provided. Thus, the device can be controlled such that the at least one predefined target device property and / or the at least one predefined target environment property is established. The action instructions are known from the preceding method steps, and thus it is known how an internal state and an external state can be converted into a further internal state and a further external state. The target specifications, i.e., the target device property and / or the target environment property, serve this purpose. These target properties therefore determine what is to be achieved, and this is achieved by executing the action instructions.The device's current state is read out, and then stored actions are selected that, starting from the actual state, reach the target state. In a preferred case, executing an action instruction that immediately achieves the target specifications is sufficient. In a typical case, however, the initial state is successively transformed into the target state using several action properties. Thus, several action instructions are linked in such a way that the target specifications are ultimately achieved.

[0021] To illustrate this with an example, the method may have identified that the device can move from A to B to C to D. Thus, the instructions are stored which stipulate that the car can drive from A to B, from B to C and from C to D. It is also implicitly stored that the car can drive from A to C and from A to D. Furthermore, it is stored that the car can drive from C to D. These possibilities are now used and if a target specification is given which states that the car should drive from A to C, the device has two options to choose from, namely to drive from A to B to C or to drive from A to C. In general, it would also be possible to drive from A to D and then from D to C. Which selection is made can also be defined in the target specifications. For example, a maximum duration of the journey can be defined in time.In addition, a maximum energy consumption can be defined. In preparatory process steps, actions are carried out randomly and are then available at runtime. In general, the process can provide for randomized actuation of the effectors, but it can also additionally provide for the reuse of previously learned action instructions. This allows for a learning phase of randomized actuation, as well as an execution phase that uses previously stored actions. These can also be alternated in any order. Furthermore, in the execution phase, the stored action knowledge is refined using internal and external sensors, and new action instructions are generated that may be preferential to a different objective.Furthermore, the target specifications can change during the execution of an action in such a way that, for example, a battery condition becomes critical. The target for the system's own operability is increased to such an extent that the system recognizes that it is now necessary to drive to a charging station. This results in an overriding target specification, at least temporarily, namely to drive to a charging station, even if this does not achieve the goal of driving to the original target coordinate. As soon as a certain charge level is reached, the target for the system's own operability is downgraded again, and the target for driving to the destination is increased to such an extent that the journey can now be continued.

[0022] According to the present invention, the method is iterated in a learning manner such that new instructions are constantly being recognized, each of which converts an output triplet into a target triplet, and these instructions are stored to control the device. This has the advantage that the method is established, entirely or at least partially, in such a way that new situations are constantly being recognized and instructions are created such that effectors are activated. The method can thus be varied such that the activation of the effectors is no longer randomized, but rather existing instructions can be combined, or it is also possible to randomize parts of existing instructions so that new possible instructions are created, for which it is also clear which interactions they trigger.Furthermore, the method can also be established in such a way that, starting with the storage of interrelationships, it is established up to the control of the device. This means that even when known instructions are carried out, interrelationships are stored and new parameters can be identified. For example, it can be recognized that if the same instruction is carried out twice, different target parameters arise. This can be the case, for example, if a crosswind arises while driving a car. This is stored and then the external sensor detects that wind must have occurred here. This creates a new output triplet and, using the effectors, it can be randomly tested how to countersteer in such a situation.Thus, new interrelationships have been identified and if such an initial triple is identified again, it is now clear which action must be carried out in order to achieve the desired target triple.

[0023] According to a further aspect of the present invention, the reading of internal and / or external effectors is carried out by storing individual support values ​​and interpolating and / or extrapolating intermediate values. This has the advantage that the support values ​​can be selected according to their availability and that they can also be estimated in such a way that a measurement does not have to be available for every possible value. Rather, existing measurements can be used to estimate which values ​​would result under normal circumstances. This allows additional support values ​​to be calculated mathematically.

[0024] According to a further aspect of the present invention, the target device property and / or the target environment property are at least temporarily influenced by a device property and / or an environment property. This has the advantage that the target specifications can also be changed or are subject to prioritization. For example, if the vehicle is to travel from a first geographical point to a second geographical point, the approach to a charging station can be prioritized in the meantime, since otherwise the final destination would not be reached. This results in the advantage of always guaranteeing that the target specifications are met, and a fail-safe method is created.

[0025] According to a further aspect of the present invention, multiple action instructions are combined into a choreography that transforms an initial triplet into a target triplet. This has the advantage that even complex, i.e., compound actions can be executed. As the method progresses or becomes established, new action instructions are continually created, which are combined into increasingly efficient choreographies. Thus, the proposed method is iteratively improved.

[0026] According to a further aspect of the present invention, an object is detected using external sensors and its detected parameters are compared with device properties and / or action instructions of the devices. This has the advantage that the device can essentially clone itself or can infer the properties of the object from its own properties. This can also be referred to as empathy. If, for example, a device is identified using the external sensors that has similar characteristics to the executing device, it is assumed that this newly detected device has similar capabilities to the executing device. To illustrate this in an example, a vehicle is used that implements the proposed invention.This first vehicle therefore carries out the proposed method and, using an external sensor—in this case an imaging unit—detects a further, second device. The first device, or rather the first vehicle, has learned that at a certain speed it can only brake to a limited extent and can only steer laterally to a limited extent. The first vehicle now detects that the second vehicle has similar dimensions and is traveling at a similar speed to the first vehicle. The instructions are then conceptually projected onto the second vehicle, and the first vehicle detects that the second vehicle can only brake to a limited extent and can only change direction to a very limited extent. If an overtaking maneuver is now initiated on a motorway, the first vehicle detects that the second vehicle could brake or also change lanes.Thus, from one's own behavior or the possible starting triplets and the possible target triplets, it is concluded that this object has similar properties.

[0027] Furthermore, the external sensors can be used to monitor the behavior of the detected second vehicle, and then update the system's own data memory, which stores the interrelationships. This allows source triplets and target triplets to be detected, and in turn, conclusions can be drawn from the second vehicle to the first vehicle.

[0028] According to a further aspect of the present invention, the detected parameters are used to predict the behavior of the detected object. This has the advantage that it is recognized that the device itself executes a specific action instruction in a certain situation, so that the other detected object would very likely also execute the same action in the same situation. Thus, the device's own actions or executed action instructions are logged, and it is then assumed that the detected object could behave the same or at least similarly.

[0029] According to a further aspect of the present invention, interrelationships of the detected object between its parameters and its actions are used to generate action instructions for the device. This has the advantage that externally detected interrelationships with respect to the proposed device or the device controlled by the method can also be used. Thus, the device not only creates interrelationships that are evaluated, but rather, the external world can also be observed, and conclusions can then be drawn regarding the conversion of the initial triplet into the target triplet.

[0030] According to a further aspect of the present invention, an effector is a motor, a gripper arm, a locomotion unit, a drive, a windshield wiper, a turn signal, a light, a steering system, or an executing unit. This has the advantage that, based on the proposed device or the device to be controlled according to the method, all possible hardware units can be present. This embodiment is merely exemplary and not exhaustive, so that any physical unit can be controlled autonomously according to the proposed method.

[0031] According to a further aspect of the present invention, internal sensor units detect an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a position, and / or a system parameter of the device. This has the advantage that the internal sensor units can measure all parameters and states of the device that executes the method or that is executed or controlled by the method.

[0032] According to a further aspect of the present invention, external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal, and / or an external parameter. This has the advantage that the entire environment of the device can be analyzed and recorded. For this purpose, all possible sensors that detect the environment in some way are possible, either individually or in combination.

[0033] The object is also achieved by a device configured to carry out a method according to one of the preceding claims, comprising an initialization unit configured to initialize by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and reading of a plurality of external sensor units configured to measure external environmental properties; a storage unit configured to store interrelationships that exist between the actuation of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation, device properties and environmental properties into a target triplet;and a control unit configured to control the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored instruction specifying how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet, wherein the method to be executed by the device is iterated in a learning manner in which new instructions are constantly recognized, each of which converts an output triplet into a target triplet, and these instructions are stored for controlling the device.

[0034] The task can also be solved by an arrangement comprising several devices which are linked by communication technology and exchange instructions.

[0035] According to the invention, it is particularly advantageous that the method can be used to operate the proposed devices and units. Furthermore, the proposed devices and units are suitable for implementing the method according to the invention. Thus, each device implements structural features suitable for executing the corresponding method. However, the structural features can also be configured as method steps. The proposed method also provides steps for implementing the function of the structural features. Furthermore, physical components can also be provided virtually or in a virtualized form.

[0036] Further advantages, features, and details of the invention will become apparent from the following description, in which aspects of the invention are described in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. Likewise, the features mentioned above and those further explained here may be used individually or in combinations. Parts or components with similar functions or identical components are sometimes provided with the same reference numerals. The terms "left," "right," "top," and "bottom" used in the description of the exemplary embodiments refer to the drawings in an orientation with a normally legible figure designation or normally legible reference numerals.The embodiments shown and described are not intended to be exhaustive, but rather are exemplary in nature to illustrate the invention. The detailed description is intended to inform those skilled in the art; therefore, known circuits, structures, and methods are not shown or explained in detail in order not to obscure the understanding of the present description. The figures show: . Figure 1: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention; Figure 2: a diagram by means of which internal and / or external properties can be estimated. Thus, internal and / or external parameters can be measured, and interpolation can be performed between the support points, or further support points can be extrapolated; Figure 3: a representation of a real-world situation in which a child wants to cross a street and the proposed device recognizes the situation and provides instruction; Figure 4: a real-world situation in road traffic in which the proposed device or method according to one aspect of the present invention is used; and Figure 5: another real-world situation in road traffic, wherein the device practices cornering according to another aspect of the present invention.Figure 6A: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention in the juvenile phase with the primary goal of specifying one's own actions; and Figure 6B: a schematic flow diagram of a method for autonomously controlling a device according to one aspect of the present invention in the adult phase, the application phase for achieving one's own goals, taking into account and predicting the behaviors of other objects relevant in a situation.

[0037] Figure 1shows, in a schematic flow diagram, a method for autonomously controlling a device, comprising initialization 100 by means of randomized actuation 101 of effectors and reading 102 of a plurality of internal sensor units configured to measure internal device properties and reading 103 of a plurality of external sensor units configured to measure external environmental properties; storing 104 of interrelationships existing between the actuation 101 of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation 101, device properties, and environmental properties into a target triplet;and controlling 105 the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored 104 instruction which specifies how the at least one predefined target device property and / or the at least one predefined target environment property is set based on the output triplet and the target triplet, wherein the method is iterated in a learning manner in which new instructions are constantly recognized, each of which converts an output triplet into a target triplet, and these instructions are stored for controlling the device.;

[0038] The model presented here, according to one aspect of the present invention, uses a central template that helps the AI ​​navigate the unknown and constantly changing world. This approach allows for short-term behavioral predictions of observed objects, thus estimating the development of situations and incorporating them into the design of its own actions. It also enables learning from observation and empathy. Aspects of the invention include: Central template as the center of the AI ​​for recording all objects in a situation; data storage that enables the recognition of laws; behavioral predictions of all observed objects in a situation; empathy based on the central template; ability to learn from observations; recognition of influencing factors that are not directly observable; recognition of natural laws and causal dependencies; computer-friendly form of theories that can be self-created, stored, reused, and continuously improved; and / or self-assessment and the evaluation of self-created theories as the basis for cognition, learning, and the selection of the best theory for achieving one's own goals.

[0039] As an example, a car that behaves according to these ideas is described. How does it assess a child or a box, such as an overtaking maneuver on the highway, or how does it recognize the law of gravity and crosswinds (as an example of recognizing causality and non-manifestable objects). It further describes how learning from observation is possible through empathy. All of this shows that the capabilities of this AI extend beyond driving cars. The essential basis, however, is a physically existing robot in a physically existing environment that directly or indirectly observes other objects.

[0040] Some aspects of the present invention are proposed below, which enable an exemplary implementation of the method or device and / or system arrangement. The following aspects are to be understood merely as examples and can be applied individually or in combination.

[0041] Possible hardware of the robot or the proposed device: This describes the technical equipment that the robot or the car may have according to one aspect of the present invention. Effectors:

[0042] In simple cases, such as a car, the effectors would be, apart from windshield wipers, indicators, lights, etc. Braking, accelerating, and steering, with which the robot influences its environment, even if it's just changing its own location, is also a change of situation. Internal sensors:

[0043] The internal sensors measure and provide feedback on the body's own characteristics, such as height, weight, speed, acceleration, etc., based on the actions of the effectors. These are necessary for assessing the body's own actions and are prerequisites for learning. This allows the body to recognize its own actions, such as the strength of braking and the effect (the extent of negative acceleration), make connections, and store memory for independent learning, meaning it can continuously improve its actions.

[0044] As a physical entity in a physical environment, the robot is subject to the laws of physics. A measured acceleration, i.e., a change in speed, depends not only on the mass but also on how far the accelerator pedal is pressed or how hard the brake is applied. If these laws are not to be hard-coded, but rather the robot is to discover, store, and use them independently, it requires internal sensors to learn and store the consequences of the use of effectors. The internal sensors of a car, for example, therefore record data such as size, speed, acceleration, etc. External sensors:

[0045] According to one aspect of the present invention, the external sensor system detects the environment (distances, objects, etc.), for example, using lidar, image processing, etc. Of course, the robot must be appropriately informed about its environment. What exactly is required will be described in more detail below. However, it is not intended to predetermine which systems are best suited, as the information requirements are lower due to the use of the central template. Not all theoretically available information needs to be detected and processed. Possible robot software:

[0046] Here, the parts with which the presented AI achieves its intelligence according to one aspect of the present invention are described. Data storage:

[0047] Data storage is a highly available data store for data tuples composed of data from internal sensors, effectors, actions, and their motivational values ​​(see below). The store appropriately links executed actions and the associated actual, experienced, and observed results. It therefore only stores internal data in tables and not, for example, pixels of one or more images. Thus, there is only data whose internal meaning is known.

[0048] Another part of this module, according to one aspect of the present invention, is the repeated checking of the internal consistency of the data. This means that it must not contain any contradictions, structures that lead to loops or circles, etc. The goal is that only one decision can ever be derived from the data. The routine that checks this should also remove stored redundant data, not only to keep the database as lean as possible, but also to prevent overadaptation. What is considered redundant depends on the type of theory formation.

[0049] Figure 2 shows the theory formation according to one aspect of the present invention: The central template can simulate, store and use for prognosis any functional interaction using the simple means described here.

[0050] Braking or acceleration functionally depend on the accelerated mass and the applied force. However, this function is initially unknown, because the mathematical function of the velocity change dV depends on the conditions Vn (current speed, weight, etc.) and the use of the effectors Em (force on the accelerator pedal, steering, etc.), thus: dV = f Vn , Em

[0051] This function should be learned automatically rather than programmed in, so that what has been learned can be expanded, corrected, and improved later in the same way. The effective acceleration when braking (or accelerating) is linked via the internal sensors of the measured acceleration a to the (mathematically) independent variable of force F (how hard the respective pedal is pressed) and the mass m according to the law F = m * a. The principle by which the robot itself finds this equation is simple. To illustrate, let's take a (different, arbitrary) quadratic function y = -((x-4) 2< )+9, which we want to simulate.

[0052] Suppose that initially only the two measured values ​​at points 11 and 12 exist. Now suppose that at point x=3 (point 2) a forecast p is required, which leads to a value of y=3.8 via the straight line a between 11 and 12. However, since the error compared to the ex post measured value y=8.0 is much too large, this point 2R at x=3.0 and y=8.0 is stored in memory. Subsequently, a new forecast would use either the degrees b1 between [11, 2R] or b2 between [2R, 12]. Over time, many more points emerge in the above example (31, 32, 33, ..), via which the functional relationship between x and y can be reproduced with arbitrary accuracy by the respective straight lines along the data tuples (c1, c2, c3, ..). The accuracy is only limited by the (as yet) lack of experience and the size of the memory.Of course, this can be extended to any number of variables, if the computing and storage capacities allow it; instead of straight lines, hyperplanes a1x1 + ... + anxn = const are used for the forecast by means of piecewise linear interpolation or extrapolation.

[0053] However, to keep the number of stored corner or data tuples as small as possible, they can also be removed: If an ex post analysis reveals that a data tuple is so close to or even on the line / hyperplane between the two neighboring tuples, then this tuple can be deleted as redundant. This not only reduces the memory requirements but also the computational effort, and over time the behavior adapts better and better to the actual, albeit still unknown, mathematical function, which then leads to increasingly more efficient behavior over time.

[0054] This approach also prevents overadaptation, which must be considered when selecting and sizing training data for an ANN. Redundant experiences are thus not repeatedly stored, which would completely overshadow the rare events, the so-called "black swans," with potentially severe consequences.

[0055] Storing functional relationships between variables has another advantage: it also stores every possible inverse function, as the stored data tuples do not distinguish between which values ​​were the 'dependent' and which were the 'independent' when they occurred. In the above function, x is the independent variable, and the corresponding y-value can be determined using the function y = -((x-4) 2< )+9. The x-value would have to be searched for in the database, and if the x-value is not explicitly stored, the y-value can be approximately determined using the next two x-values. However, one can also simply search for the y-value in memory and determine an x-value using piecewise linear interpolation or extrapolation - this is much easier than trying to determine the inverse function of, for example, the above equation.

[0056] In this very simple way, any functional relationship can be simulated with arbitrary precision. The only prerequisite is sufficient action training, as in a complex sport. Of course, the accepted error ε can also be changed and adjusted over time, whether it needs to be reduced to increase the required precision, or it can be increased to reduce storage and computational effort, since this in turn influences the number of stored key points and data points. It is a continuous optimization in an ongoing process.

[0057] Within the scope of this invention, according to one aspect of the present invention, a theory about the rules and laws of a current situation is therefore defined as a self-contained, consistent data set from which the robot can calculate the predictions of this theory using piecewise linear interpolation or extrapolation. The data set consists of its own measurement results and possibly also of the parameters of created alter egos. In this way, the robot can independently determine, save, reuse, and improve each of its theories. As part of the self-assessment (see below), it is also able to select and apply the specifically best theory (data set) with the highest competence (see below) in a given situation. Actions:

[0058] According to one aspect of the present invention, an action represents an ordered sequence of effector deployments to achieve a goal. Actions can be designed using the robot's known functional relationships between effectors, internal parameters, and known effects. Initially, in the juvenile learning phase (see below), simple actions are performed. The AI ​​learns to accelerate and brake, then accelerate, wait, and hit a barrier (externally braked with a 'risk of injury'), accelerate, steer, brake, and so on. These simple actions are then combined into increasingly complex 'choreographies,' which are then further optimized, for example, by making braking smoother through decreasing pedal pressure as the speed decreases, and by making cornering smoother. The optimization is achieved by a motivation value (see below).), which represents a kind of internal evaluation of the performed action. Thus, the one with the highest motivational value is selected from the stored (partial) 'choreographies'. This allows the AI ​​to independently and continuously improve the use of its effectors. Furthermore, based on the current situation and the action experiences gained, the AI ​​can predict its own situation at future points in time, which forms the basis for its own action planning. Motivational values:

[0059] Motivational values ​​reflect a kind of reward. On the one hand, they are determined internally after a self-imposed goal has been achieved (successful); on the other hand, a motivational value can also be transmitted externally (well done). For example, if a route is covered with strong steering movements, low speed, and yet high lateral acceleration, this sequence of actions receives a lower motivational value than the result of 'training,' after which the same route is covered faster but with lower lateral acceleration. This, of course, also has to be stored in memory.

[0060] Motivational values ​​thus represent a form of non-fixed, intrinsic goals. They can be linked to long-term goals, such as the end point of a journey, or short-term intermediate goals, such as visiting a gas station when the tank is empty. As the tank fills, the motivation value for 'fueling' slowly increases until it exceeds that of the long-term goal. Pursuit of the long-term goal is interrupted in favor of refueling, and then resumed. Target system:

[0061] According to one aspect of the present invention, the robot is always 'switched on,' although it can of course have an off switch. It is therefore always performing an activity. However, this also includes doing nothing (charging) or optimizing its own long-term memory of action options in a rest state (sleeping). The action chosen is always the one with the currently highest motivation value. If an action not currently being performed achieves a higher motivation value than the one currently being performed, it is aborted, ended, or interrupted so that the one with the highest motivation can be performed. If the motivation of such a car on the way from Munich to Hamburg to refuel or recharge its batteries becomes greater than reaching Hamburg, it will interrupt the journey at the next opportunity and head for a gas station.

[0062] In the juvenile phase (see below), i.e., the learning phase, of such a robot or car, the motivations for driving, braking, accelerating, and cornering over short distances will receive high motivational values, simply because the internal long-term memory signals that learning is needed, or rather, reducing errors in order to refine one's own actions. Later, in the adult phase (see below), when such behavior causes no or only insignificant changes in the data memory, and error reduction approaches zero, such behavior becomes 'boring,' it receives only low motivational values, and then, for example, energy conservation is preferred.

[0063] Thus, according to one aspect of the present invention, this target system avoids a commitment to reality, which would mean an interpretation of reality that could possibly prove to be wrong in the future. The central template, the ego:

[0064] The central template represents the core and basis of this AI according to one aspect of the present invention. It is called Ego because it places itself at the center of its data processing and initially observes and evaluates everything from its own perspective. Therefore, the Ego does not need to be given any data or its meaning (e.g., identifying images) or trained; it creates everything itself. Subjective polar coordinate system:

[0065] All data from this AI's observations and experiences are subjective, or rather, according to one aspect of the present invention, are recorded from the perspective of the ego. The spatial concept also follows this idea. A Cartesian coordinate system with an assumed origin of the observed environment, in which each detected object is assigned its respective coordinates, is not designed. Instead, a polar coordinate system can be used. This not only automatically determines the mathematical origin in the ego, but also the ego's subjective orientation, including front-back, up-down, right-left, and the motion vector. The surrounding (empty) space is assumed to be given. The observed objects are recorded only via their distance from the ego's own viewpoint, i.e., the origin of the polar coordinate system, and the angle. Empty space in the physical sense is therefore not assumed to be a separate entity or dimension.It is there and simply an option to move there, unless another object is there, standing, or moving (toward it), in which case the 'place' would be occupied, or the space would not be empty. A further advantage of subjectivist polar coordinates arises below for the concept of alter egos (see below). Juvenile phase:

[0066] Within strict guidelines, such as the integrity of the observed objects and the robot itself, maintaining operational readiness (battery charge level), objectives such as smooth acceleration and braking, optimized cornering, and others, the robot should / can learn the use of its effectors and the resulting consequences itself, primarily to coordinate its effectors with its internal and external sensors and, through continuous training (100-105), occasionally interrupted by rest periods, for example, to recharge the batteries, and to maintain the internal database (redundancy, consistency, etc.), to refine them until the detected error has been sufficiently reduced from planning and results. For a car, this would primarily be the use of accelerating, steering, braking, reaching a target point (success) or missing it (error), etc.First, the ego sets itself simple goals for simple actions, which are then executed and stored with the results, which are later combined into increasingly complex choreographies. This training phase ends when the actions of the self- or externally set goals ("drive there") can no longer be optimized, i.e., when the mathematical module cannot reduce the error or can only marginally reduce it, and / or the database can hardly be optimized any further (number and location of the points in memory). As the error reduction decreases more and more slowly, the juvenile phase is increasingly drawing to a close. An artificial intelligence (ANN), on the other hand, requires a large amount of specifically selected training data, whereas the ego only needs a real environment—one could say a playroom—where it cannot cause damage, to test and develop its own abilities.It follows intrinsic goals and not externally predetermined ones like an ANN and therefore does not interpret the environment or the world. Adult phase:

[0067] In the juvenile phase, the device has learned the orderly sequence of its effector deployments for achieving its set goals. What is the value of the knowledge and experience gained in the juvenile phase? All data about the world that the ego has learned in the world in which it moves are its own data, its own observations, and its own experiences. In a higher scientific-theoretical or philosophical sense, one must state that these data are true from the ego's perspective, in the sense of "This is so" or "This was so." This thus represents a natural and solid starting point for knowledge. Naturally, the ego must maintain its internal data. In addition to the optimization already mentioned above, it must also ensure that they are consistent and error-free. Otherwise, the manufacturer or operator of the robot could not assume that the robot would achieve its intended goals with its designed actions.

[0068] Changes in the environment that one experiences and causes oneself, such as changing location by accelerating, steering, and braking, are experienced and stored by linking internal and external data. There is no causal knowledge that goes beyond one's own effectors, and this does not require any causal knowledge in the sense of questions such as why does pressing the accelerator lead to acceleration, steering to cornering, or braking to negative acceleration. The ego of a car only needs "there are three options" to change location in a targeted manner. On this level, the important thing is the "how it is," not the "why." A pigeon on the road certainly has no idea how a car or a cyclist works. But if such a "road user" were to pass it far enough, it would stay put; if the direction of movement were to change in its direction, it would fly away.She doesn't know why, but she is aware of the possible behaviors and therefore pays close attention. Alter ego:

[0069] When the (adult) ego detects an object in the observable environment through external sensory processing, the ego simply creates a copy of itself and adjusts the copy's parameters according to the observed properties. The ego is thus the central template for all observable and unobservable objects and phenomena. Hence the names "ego" and, consequently, "alter ego" for the copies.

[0070] By using the ego as the central template, there are no longer any unknown objects. There are only unknown parameters, which can be measured via external sensors, estimated from one's own experience (keyword: prejudice), and whose possible range can be narrowed down by further observations.

[0071] Just as the ego calculates its behavior using its methods, stored data, and parameters and can predict its own situation, the behavior of alter egos can also be calculated using the same methods, data, and adapted parameters. The assumptions thus made about the alter egos are limited to the same physical laws and to the fact that short-term goals can be derived from the observed orientation and direction of movement. Be it a box on the street, a kangaroo in Australia, or an elderly lady with a walker.

[0072] Programmatically, the ego should exist as an object. Then, for a newly emerging object, a copy of the ego can be easily created, and the parameters of the copy of the new object can then be adjusted using external sensors and one's own experience. There could be a list of the alter egos of the current situation into which the new copy is inserted or created: Class CEgo { ...}; CEgo Ego( ... ); / / Beginning of education & learning, juvenile phase / / Beginning of the adult phase CEgo AlterEgos[]; / / New object recognized AlterEgos[i] = Ego.clone(); / / Clone the ego for the object AlterEgos[i].adjust(...); / / Parameter adjustment of the new object

[0073] Thus, the ego has internally created an alter ego as a calculable image of an observed object in the environment and it can, as in Figure 6Bshown, for the continuous updating of the observed situation, the inner loop consists of object recognition 311, updating of the observable object parameters 312, prediction of the behavior of the objects 313, adaptation of the sequence of effector deployments 314 and a second, outer loop after reaching the target and entering a rest and (also) loading phase 315, until it continues with the initialization 310 for a new target.

[0074] It doesn't matter whether the observed object is familiar or completely new. If it is new, the range of possible feature values ​​is wider, which requires increased attention (sampling frequency). This has the following advantages: Every observed object is known intrinsically! Behavior prediction of all external objects using proprietary methods and data! Humans no longer need to intervene to program in missing information. The ego continuously learns and improves its behavior. The ego is capable of empathy and is aware of the current situation (see below). The ego can learn from observing alter egos (see below). Alter egos can embody unknown causalities and rules (see below).

[0075] Figure 3 shows scenario 1 "one child" according to one aspect of the present invention.

[0076] When the ego, for example the AI ​​of a simple car, encounters another road user, in this example a child, it creates a copy of itself and adjusts the parameters (distance d from the ego, orientation, speed vK, ...): In this way the alter ego can use the ego's methods to predict where it will be and when. And in this way the ego can, by comparing its own calculations of when and where it will be, determine whether a potentially dangerous situation could arise in order to react accordingly, i.e. to brake. In this way the child's short-term behavior can be predicted using its own methods and experience. NOTHING needs to be specially programmed or trained for the child that would not be present when this AI encounters a child for the first time.

[0077] Figure 4 shows Scenario 2 "Highway" according to one aspect of the present invention.

[0078] Another example that more clearly demonstrates the power of this simple idea of ​​alter egos is a situation on a highway with a truck and two cars behind it, and a car in the overtaking lane, of which the second, PKWEgo, is our ego. Based on its own knowledge of the force law (F=m*a) of braking force and mass, with the parameters of the observed cars and the truck—i.e., with the assumed mass, measured speed, and assumed braking effect—the ego can, using its own methods alone, predict the braking distances of all road users in a potential accident situation, thus determining its own safety distance. But let's first address the ego's 'thoughts' or assumptions about other road users for predicting their actions: Truck0 is traveling at a certain speed, which the ego (cargo) has learned (without needing to be specifically trained) that such tall road users rarely exceed. The truck's alter ego will therefore not predict a change in speed. Cargo, the center of attention, would be able to drive faster and would overtake truck0, taking other road users into account. Overtaking would get cargo to its destination faster and thus receive a higher motivational value at the time; it prepares the overtaking maneuver. Car1, another alter ego, systematically thinks the same as cargo, because the ego would act the same way in its place. The ego therefore assumes that car1 also wants to overtake truck0.Of course, the alter ego of PKWEgo 'sees' PKWEgo and will carry out its (the ego's) assumed overtaking maneuver, taking the PKWEgo's existence into account. For example, it will activate its indicator and steer with particular attention to PKWEgo and any other road users in the overtaking lane. This is what the ego of PKWEgo 'thinks' and how it would predict PKWEgo's behavior. A possible additional road user, PKWEgo, is approaching at high speed in the overtaking lane. The ego also creates an alter ego for this, predicting its actions based on the clear road. As shown above, the program, the ego, can predict the entire situation with all relevant road users and plan its own sequence of effector deployments accordingly, ensuring that no dangerous developments occur. Alter ego, consciousness:

[0079] Thus, the ego internally maintains a complete picture of the observable environment, including all recognized objects, and the ability to predict their behavior in the short term. This would be a possible and programmatically feasible definition of consciousness—a genuine one, not a feigned or imitated one.

[0080] The short-term behavioral forecast based on the ability to put oneself in the place of the observed objects means here to consider the situation with all the data converted for the alter egos and to calculate the development of the situation using the calculated probable behavior of the observed objects.

[0081] The ego in the highway situation above can be used for all observed objects without having to program any extra routines: predict their behavior with the assumed condition that all road users want to move forward as quickly as possible and without an accident, continuously adapt the predictions to the observed actual behavior (braking, accelerating, steering, using indicators, flashing headlights, etc.) and overtake the truck themselves at an appropriate moment.

[0082] The ego can easily assess the possible behavior of all participants in this situation by applying its own methods and experiences (data) with the appropriate parameters. Isn't that exactly what humans do?

[0083] The examples make it clear that the basis, foundation, and starting point of this AI is always the ego, the central template. The more differentiated its capabilities in internal and external sensory processing and the better its training in the juvenile phase, the more intelligent its behavior and the greater its potential to predict the behavior of observed objects. If one also gives the ego a parameter for its own vulnerability—in the case of a car, one would probably speak more of malleability—the ego could also assess the risks of strong or weak contact. BUT, the ego only ever needs to be built, programmed, and trained (once), although the latter part can of course also be transferred from a trained ego. Observed objects do not need to be identified (ANNs and their training data), nor does the behavior of the identified objects need to be programmed.One could, of course, argue that ego isn't enough to predict the behavior of other road users. But then, one only has to consider how humans try to assess the behavior of other road users. Humans, too, don't know the long-term goals of other road users on a highway; they can only guess at their short-term goals (to move forward as quickly as possible without causing an accident) and prepare accordingly. The robot’s capabilities: Empathy:

[0084] Above, we demonstrated how the ego, with the help of alter egos, can place itself in the situations of the observed objects in its environment in order to calculate their behavior from their perspective. One could call this empathy, if one doesn't reserve the term exclusively for humans. Learning from observation:

[0085] Learning from observation means observing behavior that is potentially better than one's own and, if possible, adopting it. The basis is the comparison between another's behavior and one's own, and this comparison is made possible by the empathy defined above.

[0086] Here, ego empathy is realized through alter egos, through which the ego places itself in the shoes of other observed objects in order to predict their behavior from their perspective on the environment. Suppose there is a discrepancy between the predicted and observed behavior. And if the alter ego now changes the sequence of effectors' deployment in such a way that the observed behavior is replicated, the ego has the opportunity to copy the observed behavior if it offers an advantage.

[0087] Figure 5 Scenario 3 shows different "curve behavior".

[0088] Let's imagine the ego has learned to always corner at the same distance from the right edge of the road, the line LE. Now it observes a car ahead that cuts the corner by turning earlier but less sharply, thereby also reducing lateral acceleration and thus cornering faster, on the dotted line LB.

[0089] Seen in this light, the empathy described here is a prerequisite for learning from observing and imitating the behavior of others. (What's wrong with assuming something similar also applies to humans—learning through imitation based on empathy?)

[0090] If the data from the external sensors, converted to the situation and position of the alter ego, is linked with the ego's stored, simple and more complex action sequences, which were copied into the alter ego, the ego is able to learn from observation: It can recognize the difference between the self-planned action sequence of its own effectors (always maintaining the same distance from the edge of the road) and the observed one (earlier braking and turning, smaller steering movement and higher cornering speed, and earlier acceleration) and determine that it could thus drive through the corners faster. This form of learning is certainly faster than the usual trial and error. Recognizing causal laws:

[0091] In the theory building section, we explained how any observed functional relationship can be replicated. Let's use this ability to recognize natural laws, for which we can also use alter egos if necessary and appropriate. Of course, they don't lie or drive on a road, but they cause changes that can be detected by external and / or internal sensors.

[0092] Let's imagine an experiment in which our Kl-ego observes a falling apple. It creates an alter ego and notices the accelerated movement toward Earth. The ego's first assumption will be that the apple itself caused the acceleration; like the car-ego, it has its own 'gas pedal' to accelerate. Then the ego will observe that no apple moves on the ground itself, and that other things also fall to Earth and then stop moving. Furthermore, the ego itself might have learned the law of gravity. On a downhill road, without pressing the gas pedal, it would have noticed and learned an acceleration, but always only 'downhill,' whereas 'uphill,' it needs to accelerate more. The steeper the angle α, the greater the force, according to this familiar formula (with g = acceleration due to gravity, 9.81 m / sec2): F = m * g * sin α

[0093] We remember that the ego does not explicitly know this equation or the equation for free fall, i.e. the law of gravity (sin(α=90°) = 1), but through the data points and piecewise linear interpolation or extrapolation it has an implicit knowledge about this relationship between the relevant quantities.

[0094] The ego thus behaves as if it knew that, as a physicist would say, as a body with a heavy mass, it is subject to the above law of gravity. When it then sees another object and creates its alter ego, the alter ego is also subject to this law of gravity. In this way, this knowledge essentially acquires the status of a law of nature that affects all bodies without needing to be specially programmed. If the behavior of such an ego were observed from the outside, one could not tell whether it truly knows the law of gravity, as a physicist does, or whether it is merely feigning this insight. Even if the ego sees a pigeon sitting on the street that flies away as soon as the ego approaches, it can recognize that the flapping of its wings and the upward acceleration correspond.

[0095] Another example would be a strong crosswind. Imagine the ego is a large empty van. During normal straight-line driving, the lateral acceleration is zero. But then, while driving straight, the van suddenly shifts sideways, and the lateral acceleration sensor triggers. In this case, the ego can simply create a new alter ego for the unexpected force and the unknown cause and link it to this lateral acceleration. If the ego can then also record data from the environment, such as in the forest or on the plains, or the movement of branches and add it to the data logger, the ego has not only detected a previously unknown quantity (crosswind), it has also identified a possible causal relationship. If the ego could also see and assess the extent of the movement of the branches, it could even roughly estimate the strength of the force acting on the side and thus the likely influence on the driving behavior.In this way, the ego has created its own theory with a new variable. Similar to the physicists who introduced dark matter and dark energy, although they are neither visible nor observable, or like Nobel laureate Peter Higgs, who conceived a particle in the 1960s that was then first detected at CERN in 2012. The alter ego, the copy of the self, can of course never discover the true nature of a phenomenon with this approach in the sense of a search for truth, but it is sufficient for a competent grasp of the situation. Self-assessment:

[0096] The first form of gaining knowledge has already been described. In the section on "Recognizing Causal Laws," the example of a crosswind was developed. A force (initially) unknown to the ego suddenly generates a lateral acceleration on a straight stretch of road, a measurable effect on the robot or the ego. For this purpose, it creates an alter ego. However, further information is initially missing that would signal to the ego when and with what intensity the phenomenon occurs. Only when the ego observes a temporal relationship between the movements of the branches, sometimes stronger, sometimes weaker, and its own lateral acceleration, sometimes none, and sometimes none, can the ego connect the two and assume the same cause for both. What is happening here? The ego cannot recognize the crosswind itself, but it can recognize its own reaction to it (lateral acceleration) and the reaction of the branches.In this way, it creates an element that sometimes has an effect on the ego and sometimes not, but which not only has an effect on the ego, but also on other observable objects.

[0097] The second form of gaining knowledge is based on comparing one's own behavior with that of other objects. It assumes that the behavior of the observed objects is 'rational' and not random, although this could also be recognized as such, in which case the comparison would no longer be a criterion for any potential gain in knowledge. Two different behavioral comparisons are carried out, allowing the ego to determine whether its own competence (see below) in a situation is inferior, equal, or superior to that of the observed objects. This means that the observed objects can perceive more, the same, or less of the situation. A van being observed suddenly slows down on a straight road. It appears to see something unknown to the observing ego; it perceives less than the object.If inferior competence has been identified, this can be seen as an opportunity to specifically examine these situations in order to at least raise one's own competence to the level of the others. The comparison is not about whether the behavior is right or wrong - that would only lead to the well-known problems of truth theories - but only about whether and when behavior changes and whether behaviors are repeated. The first comparison of behavior concerns different situations: Do the observed objects display different behaviors in situations that are different for the ego? The second comparison concerns identical situations: Does the behavior of the observed objects repeat itself in situations that are the same for the ego, or does it vary? From these observations of several situations over a longer period of time, the ego can conclude, . that one's own competence is inferior if objects repeatedly observed in situations that are the same for the ego show different behavior and their behavior also varies in situations that are different for the ego, that one's own competence is superior if objects repeatedly observed in situations that are the same for the ego repeat their behavior but their behavior does not vary in situations that are different for the ego, that one's own competence is equivalent if objects repeatedly observed in situations that are the same for the ego repeat their behavior but their behavior varies in situations that are different for the ego.

[0098] With this simple comparison, the ego can recognize how competent one's own competence is in certain situations compared to other objects. Mind you, this is not an attempt to compare oneself with the "truth of the real world."

[0099] From these comparisons, the ego can evaluate itself and, for example, determine whether it should behave more cautiously in certain situations and try to identify what is (still) unknowable. It can be understood as an internal mandate to conduct targeted research as part of "constant learning and the continued expansion of insight and knowledge" (see above). Competence:

[0100] The results of these comparisons can also be used externally to assess the competence of the robot's abilities, which helps determine its potential applications. Unlike the truth of knowledge about the real world that the ego has acquired, there are no linguistic problems with competence when it comes to comparing competencies or determining greater or lesser competence.

[0101] However, the concept of truth in relation to theories has another important aspect that competence must also fulfill. Theories enable predictions. Therefore, it is important to know the quality of a theory and, consequently, its predictions. In general, the attribution of truth to a theory serves as a criterion for being able to apply it, or for being permitted to apply it. This has the advantage that the criterion of truth automatically excludes all competing theories, thus eliminating the problem of deciding which theory to use. The concept of competence must achieve something similar.

[0102] Competence is defined and used here as the grasp of the relevant influencing factors of a situation or similar situations. Thus, it is initially a limited criterion, the level of which can be determined through a comparative process, in contrast to the truth of a theory, which assumes temporally and spatially unlimited validity. However, this can neither be proven nor does the criterion allow for the comparison of different theories.

[0103] Thus, competence is better suited to describing the capabilities of a robot than the concept of truth and, unlike truth, it can be determined by the robot itself for self-assessment in a formal procedure.

[0104] Figure 6Ashows a schematic flow diagram of a method for autonomously controlling a device in the juvenile or training phase. The goal and end of the phase is the sufficiently precise use and application of the effectors, with the smallest possible error. After initialization 200, there is a randomized actuation 201 of effectors and the reading 202 of a plurality of internal sensor units configured to measure internal device properties and a reading 203 of a plurality of external sensor units configured to measure external environmental properties.storing 204 interrelationships that exist between the actuation 201 of the effectors and the read-out device properties and environmental properties as action instructions, such that an action instruction converts an output triplet of actuation 201, device properties, and environmental properties into a target triplet; controlling 205 the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one stored 204 action instruction that specifies how the at least one predefined target device property and / or the at least one predefined target environmental property is set based on the output triplet and the target triplet;

[0105] Figure 6A shows the juvenile phase: 200Initialization as preparation for learning 201Randomized actuation of the effectors 202Reading 202 a plurality of internal sensor units 203Internal device properties and a reading 203 from a plurality of external sensor units 204Saving 204 of interrelationships 205Determining the error in movement precision, immediately if too large, continue at 201

[0106] Optional: Rest period for recharging and ex post data preparation: errors, redundancy, consistency, etc.

[0107] Figure 6B shows the adult phase: 310 Initialization as preparation for achieving a longer-term goal 311 Recognition of objects 311 in the current situation 312 Creation or updating of alter egos or their parameters 312 313 Behavioral prediction 313 of the observed objects using the current alter egos 314 Calculation of the optimal use of effectors 314 for further goal pursuit 315 Rest phase 315 for recharging and ex post data processing: errors, redundancy, consistency, etc.

Claims

1. A method for autonomously controlling a device, the method comprising the following steps: - initialising (100) by means of randomised actuation (101) of effectors and reading (102) a plurality of internal sensor units set up to measure internal device properties as internal states and a reading (103) of external environmental properties as external states by a plurality of external sensor units set up for measurement, - whereby an object is detected by means of the external sensor units and its detected parameters are compared with device properties and instructions for action of the devices and the detected parameters are used to predict the behaviour of the detected object, - whereby an action instruction as a mapping function transfers an output triplet of actuation of the effectors, device properties and environmental properties into a target triplet that specifies a new internal state, a new external state and possible actions that are possible in the new states, - storage (104) of interrelationships existing between the actuation (101) of the effectors and the read-out device properties and environmental properties as instructions for action; and - controlling (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one memorised (104) action instruction, which indicates how the at least one predefined target device property and / or the at least one predefined target environment property is set using the output triplet and the target triplet, the method being iterated in a learning manner, in which new action instructions are constantly recognised, which in each case convert an output triplet into a target triplet and these action instructions are stored for controlling the device.

2. The method according to claim 1, characterised in that the method is iterated in a learning manner in such a way that new action instructions are always recognised, which in each case convert an output triplet into a target triplet and these action instructions are stored for actuating (105) the device.

3. The method according to claim 1 or 2, characterised in that the readout (102, 103) of internal and / or external effectors is carried out in such a way that individual support values are stored and intermediate values are interpolated and / or extrapolated.

4. The method according to one of the preceding claims, characterised in that the target device property and / or the target environment property is at least temporarily influenced by a device property and / or an environment property.

5. The method according to one of the preceding claims, characterised in that a plurality of action instructions are combined to form a choreography which converts an output triplet into a target triplet.

6. The method according to one of the preceding claims, characterised in that an object is detected by means of external sensor units and its detected parameters are compared with device properties and / or instructions for action of the devices.

7. The method according to claim 6, characterised in that the detected parameters are used to predict the behaviour of the detected object.

8. The method according to one of claims 6 or 7, characterised in that interrelationships of the detected object between its parameters and its actions are used to generate instructions for action of the device.

9. The method according to one of the preceding claims, characterised in that an effector is present as a motor, a gripper arm, a locomotion unit, a drive, a windscreen wiper, an indicator, a light, a steering or an executive unit.

10. The method according to one of the preceding claims, characterised in that internal sensor units detect an effector state, an effector setting, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a position, and / or a further system parameter of the device.

11. The method according to one of the preceding claims, characterised in that external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal and / or a further external parameter.

12. An apparatus adapted to carry out a method according to any one of the preceding claims, comprising: - an initialisation unit set up for initialising (100) by means of randomised actuation (101) of effectors and reading out (102) a plurality of internal sensor units set up for measuring internal device properties as internal states and reading out (103) external environmental properties as external states by a plurality of external sensor units set up for measurement, - wherein the device is designed to detect an object by means of the external sensor units and to compare its detected parameters with device properties and instructions for action of the devices and to use the detected parameters to predict the behaviour of the detected object, - whereby an action instruction as a mapping function transfers an output triplet of actuation of the effectors, device properties and environmental properties into a target triplet that specifies a new internal state, a new external state and possible actions that are possible in the new states, - a memory unit set up for storing (104) correlations which exist between the actuation (101) of the effectors and the read-out device properties and environmental properties as instructions for action, - wherein the device is further adapted to be controlled according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored action instruction indicating how to set the at least one predefined target device property and / or the at least one predefined target environment property using the output triplet and the target triplet, and - wherein the device is further designed so that the method is iterated in a learning manner, in which new action instructions are constantly recognised, each of which converts an output triplet into a target triplet and these action instructions are stored for controlling the device.