Autonomous control of a device
Patent Information
- Application Number
- BR112025011902
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-06-19
- Publication Date
- 2026-08-04
- Estimated Expiration
- 2043-06-19
Smart Images

Figure 00000047_0000 
Figure 00000048_0000 
Figure 00000049_0000
Description
1 / 41 Autonomous control of a device
[01] The present invention is directed to a method for autonomously controlling a device, such that the device generally moves in a real physical world and does not need to be further specified with respect to its physical configuration. The method creates the advantage that the device performs autonomous learning and continuously improves the learned knowledge or behavior. In general, this overcomes the disadvantage in the prior art that training data must first be created, as is the case with conventional artificial intelligence methods. In general, the method can be used universally and the device learns automatically and constantly corrects its own knowledge base. Furthermore, a device configured to perform the method is proposed, as well as a system arrangement comprising several of the proposed devices.Furthermore, a computer program product and a computer-readable storage medium are proposed, which either execute the steps of the method or cause a computer to execute the method.
[02] TAKAHASHI KUNIYUKI ET AL: Dynamic motion learning for multi-DOF flexible-joint robots using active-passive motor babbling through deep learning, ADVANCED ROBOTICS, vol. 31, no. 18, September 17, 2017, pages 1002-1015, XP055840673 shows a learning strategy for multi-degree-of-freedom flexible-joint robots to perform dynamic motion tasks. Although flexible-joint robots offer several potential advantages, such as exploiting intrinsic dynamics and passively adapting to environmental changes with mechanical compliance, controlling these robots is challenging due to the increasing complexity of their dynamics.
[03] Several AI methods of artificial intelligence are known from the state of the art, which provide training data to be selected and supplied and algorithms to identify regularities in a training phase. Here, it is possible that special algorithms recognize Petition 870250048900, dated 11 / 06 / 2025, page 7 / 180 2 / 41 regularities are identified through the specification of initial data and target values, and these are then used. This means that implicit knowledge can be used on large amounts of data. After the training phase is complete, the properly trained algorithms are applied to real data and, in turn, generate implicit dependencies to solve real-world problems. A disadvantage here is that the training data must first be selected and specified, which limits the resulting artificial intelligence to the training data, and undesirable influence on the learning phase can already be exerted here, which is not only error-prone but also expensive.
[04] Collective intelligence is also known from the state of the art, where multiple devices are provided and then collaboratively solve a problem in a distributed manner. Here too, the control and coordination of individual participants is complex and sometimes error-prone. This is particularly true when distributed learning occurs, which, in turn, must be coordinated.
[05] Artificial neural networks, which provide neurons and connections, i.e., vertices and edges, based on graph theory, are also known from the state of the art. These networks are known to mimic the functions of the human brain and are also capable of learning. Edge weights can be varied, new edges can be added or old ones deleted, and existing vertices can be deactivated or new vertices can be added. This results in a dynamic overall learning system.
[06] However, using an ANN in the manner described above presents at least four problems: 1. For objects on which the ANN has not been trained, the ANN provides no information whatsoever to trigger the routines assigned to the object. 2. The programmed properties or behaviors of the detected and identified objects may be incorrect or have changed in that time. Petition 870250048900, dated 11 / 06 / 2025, page 8 / 180 3 / 41 halftime or still unknown. 3. One or more scientists, experts, etc. select the training data and specify the results to be achieved. Even with unsupervised learning, the training data and hyperparameters are still specified by humans, and this cannot rule out errors. 4. An ANN cannot correct errors. Each individual piece of information is distributed across all parameters of the ANN, similar to a hologram where each data point contains information about the entire image. An ANN must therefore always be completely wiped and retrained.
[07] This means that there is no recognition failure (unlike ANNs) when the Ego encounters a previously unknown object, as in the video when the autonomous car suddenly sees two boxes in the roadway. ANN usually stands for artificial neural network.
[08] In the state of the art, there is a need to create autonomous control of devices so that time-consuming preparatory work, such as providing training and teaching data, can be avoided or minimized. There is therefore a need for a self-learning system that is also capable of learning collectively or estimating what other participants are likely to do or can do. Therefore, there is also a need for a system that understands multiple participants, so that each participant in fact also has knowledge about the other participants and develops this further, which is referred to here as empathy.
[09] Consequently, it is a task of the present invention to propose a method for the autonomous control of a device, which acts independently only on the basis of the target information provided and accumulates action knowledge in the process. The proposed method must be able to recognize the scope of action of the effectors of any device and create a Petition 870250048900, dated 11 / 06 / 2025, page 9 / 180 4 / 41 improved method or at least an alternative method for autonomously controlling the device or devices. Furthermore, it is a task of the present invention to propose a device for carrying out the method, as well as a system arrangement comprising several of the proposed devices. Additionally, it is a task to propose a computer program product and a computer-readable storage medium containing instructions that execute the method.
[10] The problem is solved with the features of claim 1. Additional advantageous features are presented in the sub-claims.
[11] Thus, a method is proposed for autonomously controlling a device, comprising initialization by means of randomized actuation of effectors and reading a plurality of internal sensor units configured to measure the internal properties of the device and reading a plurality of external sensor units configured to measure external environmental properties; storage of existing correlations between the actuation of the effectors and the reading device properties and environmental properties as action instructions, such that an action instruction converts an initial triplet of actuation, device properties and environmental properties into a target triplet;and the activation of the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored action instruction indicating how to set at least one predefined target device property and / or at least one predefined target environment property based on the initial triplet and the target triplet, wherein the method is iterated in a learning manner, in which new action instructions are constantly recognized, each of which converts an initial triplet into a target triplet and such action instructions are stored to control the device.
[12] The proposed method is used to autonomously control, that is, to create instructions for action, a device that is generally physical. This could be an automobile, a robot, or a production plant. In terms of Petition 870250048900, dated 11 / 06 / 2025, page 10 / 180 5 / 41 Mobility: No restrictions are foreseen here, meaning the device can move on land, in the air, on water, or even underwater. Autonomous in this context means that the method ensures the device learns automatically and also automatically recognizes which possible actions are available. These can be learned, and sensors are used to learn which action starts at a certain starting point and leads to a specific result. In general, the device can provide its own targets or the targets can be transmitted externally. The device's own targets could, for example, be maintaining operational capability. An internal target could be, for example, visiting a charging station when the battery level is low. An external target could be transmitted specifying which task or activity the device should perform.
[13] In one stage of the preparatory process, initialization occurs through the randomized activation of effectors. An effector is generally a physical unit that has an influence on the real world. This can, for example, affect the device itself, such as steering, acceleration, braking, etc., and / or the targeted manipulation of objects by, for example, robot arms, directly or with tools, or by instruments with a remote effect, whether they are projectiles or jets or, for example, fire-fighting agents, such as water or extinguishing foams. In the car example, this could be the brake, the accelerator pedal, etc., but also headlights, indicators, etc. The effectors are activated and then the effects of these actions are measured using internal and external sensors. This is recorded and can then be used throughout the process. In this way, the actions are learned and used in a targeted manner in later stages of the process.
[14] Since the proposed device or method can be performed or controlled entirely without prior knowledge, the preparatory step is a randomized actuation, that is, an arbitrary actuation, since first the learning must be which action leads to which result. In other iterations of the process, the learned actions are Petition 870250048900, dated 11 / 06 / 2025, page 11 / 180 6 / 41 then triggered in a targeted manner. During randomized actuation, the system parameters of the end effectors are checked and, for example, a robot arm is moved to all possible positions. This is recorded in each case, and it is recognized which action of the end effector has which effect in the real world. In the example of a spotlight, it is possible that it only provides the on and off parameters. Then, it is turned on and off, and the external sensor system is used to measure the influence of the spotlight on the environment, and the internal sensor system is used to measure the operational parameters of the spotlight, such as temperature development. With advanced LED spotlights, it is also possible to adjust the color and intensity. This is tested through randomized actuation, and the internal and external results are then recorded.
[15] According to the invention, internal sensor units and external sensor units are proposed. Internal and external do not refer to the arrangement of the sensor units, but rather, according to the present invention, to the fact that the sensor units measure external conditions in relation to the device or, analogously, internal conditions. Internal conditions are all the system parameters of the device itself. External conditions are environmental variables, i.e., of objects around the device. This means that internal sensor units are used to measure the properties of the device and external sensor units are used to measure environmental properties. For example, if the device has a certain battery state, this is an internal parameter. If an object is transferred from A to B by means of a robot arm, this is an external state, so that the state of the robot arm is, in turn, an internal state.Therefore, internal and external parameters typically interact, and actions always result in a reduced battery state, for example, while these actions, in turn, affect external properties. Conversely, it is an interaction where the device can be brought to a charging station via effectors, and then a charging process is initiated, which, in turn, influences the internal state of the battery. Petition 870250048900, dated 11 / 06 / 2025, page 12 / 180 7 / 41
[16] Thus, the recorded interactions are saved and a record is made of how the triggering of the effectors influences the internal and external parameters. The device properties are therefore stored along with the environmental properties and the instructions for action are defined. Triggering the effectors is therefore an action that influences both the device itself and the environment. For example, if the device moves from a first geographic point to a second geographic point, this changes the device's environment, which is an external parameter, and also changes the device's internal state, for example, through a change in temperature or a change in battery state. A triplet can therefore be stored, which indicates what the internal state was, what the external state was, and what action is then performed. This is converted into a new internal state and a new external state.This means that there is an initial triplet from the triggering of the effectors, the device properties, and the environmental properties, and this is transferred to a target triplet. The new internal state, the new external state, and the possible actions can be specified in the target triplet.
[17] The concept of triplets is intended only to illustrate the general transition of the system's states. In general, it is also possible that internal states and external states are transformed into a new internal state and a new external state by means of a function. Therefore, the action consists of a function mapping a first tuple to a second tuple. The expert will recognize that the power of both descriptions is equivalent. An initial tuple is converted into a target tuple by means of a function or an input triplet is formed, which specifies internal states, external states and an action, such that such action leads to a target triplet, which has a new internal state, a new external state and an additional field. The additional field may remain empty, although it is preferable that at least one action be performed here, now being possible in this state.The last field can also be filled in so that, for example, a vector is entered here containing the identifiers of additional actions.
[18] To illustrate this with an example, the device can trigger the end effector Petition 870250048900, dated 11 / 06 / 2025, page 13 / 180 8 / 41 called motor activation and drive forward. What is now recorded is an initial triplet of an internal state, i.e., a battery charge level, an external state, i.e., an image signal from a physical environment, and an action instruction, i.e., activation. This initial vector or initial triplet is now converted into a target triplet, which has a new battery charge level, a new image signal from the environment, and actions that would now be possible, such as driving forward or backward. During the forward movement action, it is also possible to specify that braking or headlight activation would be possible, for example.
[19] In an alternative example, the initial battery state tuple and the environment image signal are converted into a target tuple using the Move function, which describes the new battery state and a new environment image signal. This records what happens internally and externally during a given action. This knowledge can be reused in other action steps towards a predefined target.
[20] In other stages of the process, predefined target device properties or target environment properties can be provided. Thus, the device can be controlled so that at least one predefined target device property and / or at least one predefined target environment property is defined. The instructions for action are known from the previous stages of the method and, therefore, it is known how an internal state and an external state can be converted into an additional internal state and an additional external state. The target specifications, i.e., the target device property and / or the target environment property, are used for this purpose. Therefore, these target properties determine what should be achieved and this is done by executing the action instructions. Then, the device state is read and the stored actions are selected, based on the actual state, in order to achieve the target state.In a preferred case, it is sufficient to execute an action instruction that immediately meets the target specifications. However, in a typical case, the initial state will be successively transferred to the target state through various properties. Petition 870250048900, dated 11 / 06 / 2025, p. 14 / 180 9 / 41 of action. In other words, several action instructions are linked in such a way that the goals are ultimately achieved.
[21] To illustrate this with an example, the procedure may have identified that the device can move from A to B to C to D. This means that the instructions are stored so that the car can move from A to B, from B to C, and from C to D. It is also implicitly memorized that the car can drive from A to C and from A to D. In addition, it is also stored that the car can drive from C to D. These options are now used, and if there is a destination specification stating that the car should drive from A to C, the device has two options to choose from: driving from A to B to C or driving from A to C. In general, it would also be possible to drive from A to D and then from D to C. The selection made can also be defined in the destination specifications. For example, a maximum trip duration can be defined in terms of time. In addition, a maximum energy consumption can be defined.Therefore, actions are randomized in stages of the preparatory process and then become available at runtime. In general, the procedure can provide randomized triggering of effectors, but it can also add the reuse of previously learned instructions. This means that a randomized triggering learning phase can be performed, as well as an execution phase that uses already saved actions. This can also be alternated in any order. Furthermore, the stored action knowledge is also refined in the execution phase through internal and external sensors, and new action instructions are generated, which may be favored over a different target.Furthermore, targets can also change during the execution of an action so that, for example, a battery condition becomes critical and, in this respect, the target of the vehicle's own operability is increased so that the process recognizes that it is now necessary to drive to a charging station. This results, at least temporarily, in an overlapping target, namely, moving to a charging station, even if this does not achieve the objective. Petition 870250048900, dated 11 / 06 / 2025, p. 15 / 180 10 / 41 objective of moving to the original target coordinate. Once a certain load level is reached, the vehicle's own operability target is lowered again and the target for moving to the destination is increased so that the journey can now be continued.
[22] According to one aspect of the present invention, the method is iterated in a learning manner so that new instructions for action are always recognized, which in each case convert an initial triplet into a target triplet and such instructions for action are stored to trigger the device. This has the advantage that the method is established wholly or at least partially so that new situations are always recognized and instructions for action are created in such a way that the effectors are triggered. The method can be varied so that the triggering of the effectors is no longer randomized, but the existing instructions for action can be combined or it is also possible that parts of the existing instructions for action are randomized so that new possible instructions for action are created in relation to which it is also clear which interactions trigger.Furthermore, the procedure can also be established so that, starting from the storage of interactions, control of the device is established. This means that interactions are also saved when known instructions are executed and new parameters can be identified. For example, it can be recognized that different target parameters are created when the same action instruction is executed twice. This might be the case, for example, when a crosswind occurs while driving a car. This is saved, and the external sensor recognizes that a wind must have occurred here. This creates a new initial triplet, and the effectors can be used to randomly test how to perform countersteering in such a situation. This means that new interrelationships have been recognized, and if such an initial triplet is identified again, it is now clear what action should be performed to achieve the desired target triplet.
[23] According to a further aspect of the present invention, the internal and / or external effectors are read in such a way that the individual support values Petition 870250048900, dated 11 / 06 / 2025, page 16 / 180 11 / 41 are stored and intermediate values are interpolated and / or extrapolated. This provides the advantage that support values can be selected according to their availability and that these can also be estimated so that a measurement does not need to be available at all possible values. Instead, existing measurements can be used to estimate what values would result under normal circumstances. This means that other support values can be calculated mathematically.
[24] According to a further aspect of the present invention, the property of the target device and / or the property of the target environment is at least temporarily influenced by a property of the device and / or a property of the environment. This provides the advantage that the target specifications can also be altered or prioritized. If, for example, a second geographic point is reached from a first geographic point, the approach to a charging station can be prioritized in the meantime, otherwise the final destination would not be reached. This, therefore, provides the advantage that it is always guaranteed that the targets will be reached and a failsafe procedure is created.
[25] According to a further aspect of the present invention, several action instructions are combined to form a choreography that converts an initial triplet into a target triplet. This provides the advantage that even complex, i.e., compound actions can be performed and, as the method progresses or is established, new action instructions are constantly created, which are combined to form increasingly efficient choreographies. Therefore, the proposed procedure is iteratively improved.
[26] According to a further aspect of the present invention, an object is detected using external sensors and its detected parameters are compared with the device properties and / or instructions for the devices. This provides the advantage that the device can Petition 870250048900, dated 11 / 06 / 2025, page 17 / 180 12 / 41 virtually clone itself or infer the properties of the object from its own properties. This can also be called empathy. If, for example, external sensors identify a device that has similar characteristics to the executing device, it is assumed that such a newly recognized device has capabilities similar to the executing device. To illustrate this with an example, a vehicle carrying out the proposed invention is used. This first vehicle thus performs the proposed method and recognizes an additional device, namely the second, by means of an external sensor, in this case an imaging unit. The first device or the first vehicle has learned that it can only brake to a limited extent at a certain speed and can only drive transversely to a limited extent. The first vehicle now recognizes that the second vehicle has similar dimensions and is moving at a speed similar to the first vehicle.The instructions are then conceptually projected onto the second vehicle, and the first vehicle recognizes that the second vehicle can only brake to a limited extent and can only change direction to a very limited extent. If an overtaking maneuver is now initiated on a road, the first vehicle recognizes that the second vehicle can brake or also change lanes. It is therefore concluded, from its own behavior or from the possible initial triplet and the possible target triplet, that this object has similar characteristics. Furthermore, external sensors can be used to monitor the behavior of the detected second vehicle and then update the vehicle's own data memory, which stores the interactions. This means that the initial triplets and the target triplets can also be recognized, and conclusions can be drawn from the second vehicle to the first vehicle.
[27] According to a further aspect of the present invention, the detected parameters are used to predict the behavior of the detected object. This has the advantage that it is recognized that the device itself executes a certain action instruction in a given situation, so that the other detected object would most likely also execute the same. Petition 870250048900, dated 11 / 06 / 2025, page 18 / 180 13 / 41 same action in the same situation. Therefore, the user's own actions or executed instructions are recorded, and then it is assumed that the detected object could behave in the same way or at least in a similar way.
[28] According to a further aspect of the present invention, the interrelationships of the detected object between its parameters and its actions are used to create instructions for the device. This provides the advantage that externally recognized correlations related to the proposed device or to the device that is controlled by the method can also be used. Thus, the device not only creates interrelationships that are evaluated, but also the external world can be observed and conclusions about the transfer of initial triplets into target triplets can be recognized.
[29] According to a further aspect of the present invention, an effector is present as a processor, a memory, a motor, a gripper arm, a locomotion unit, a drive, an end, a windshield wiper, an indicator, a light, a direction and / or an execution unit. This provides the advantage that all possible hardware units can be present using the proposed device or the device to be controlled according to the method. The design is only exemplary and not exhaustive, so that any physical unit can be autonomously controlled according to the proposed method.
[30] According to a further aspect of the present invention, the internal sensor units detect an effector state, an effector configuration, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a range, a position, a processor utilization, a memory utilization and / or a system parameter of the device. This provides the advantage that the internal sensor units can measure all the parameters and states of the device that the method executes or that is executed or controlled by the method.
[31] According to a further aspect of the present invention, the units Petition 870250048900, dated 11 / 06 / 2025, page 19 / 180 14 / 41 external sensors detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal, and / or an external parameter. This provides the advantage that the entire device environment can be analyzed and recorded. All possible sensors that recognize the environment in any way can be used individually or in combination.
[32] The problem is also solved by a device configured to perform a method according to one of the preceding claims, comprising a startup unit configured for initialization by means of randomized actuation of effectors and reading of a plurality of internal sensor units configured to measure internal device properties and a reading of a plurality of external sensor units configured to measure external environmental properties; a memory unit configured to store interrelationships that exist between the actuation of the effectors and the reading of device properties and environmental properties as action instructions, such that an action instruction converts an initial triplet of actuation, device properties and environmental properties into a target triplet;and a control unit arranged to control the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored action instruction that indicates how at least one predefined target device property and / or at least one predefined target environment property is defined based on the initial triplet and the target triplet, wherein the method to be executed by the device is iterated in a learning manner, in which new action instructions are constantly recognized, each of which converts an initial triplet into a target triplet and such action instructions are stored to control the device.
[33] The problem is also solved by an arrangement comprising several devices that are coupled by communication technology and exchange action instructions. Petition 870250048900, dated 11 / 06 / 2025, page 20 / 180 15 / 41
[34] The task is also solved by a computer program product with control commands that implement the proposed method or operate the proposed device.
[35] According to the invention, it is particularly advantageous that the method can be used to operate the proposed devices and units. Furthermore, the proposed devices and units are suitable for carrying out the method according to the invention. Thus, in each case, the device implements suitable structural features to carry out the corresponding method. However, the structural features can also be designed as process steps. The proposed method also provides steps to implement the function of the structural features. In addition, the physical components can also be provided virtually or virtualized.
[36] Further advantages, features and details of the invention are evident from the following description, in which aspects of the invention are described in detail with reference to the drawings. The features mentioned in the claims and in the description may each be essential to the invention individually or in any combination. Similarly, the aforementioned features and the features described further in this document may be used individually or in any combination. Functionally similar or identical parts or components are sometimes provided with the same reference signs. The terms left, right, top and bottom used in the description of the embodiments refer to the drawings in an orientation with a normally legible figure designation or normally legible reference signs.The embodiments shown and described should not be understood as conclusive, but are illustrative in nature to explain the invention. The detailed description serves to inform the specialist; therefore, known configurations, structures, and methods are not shown or explained in detail in the description in order not to impede the understanding of the present description. The figures show: Petition 870250048900, dated 11 / 06 / 2025, page 21 / 180 16 / 41 Figure 1: a schematic flowchart of a method for autonomously controlling a device according to an aspect of the present invention; Figure 2: A diagram that can be used to estimate internal and / or external properties. Internal and / or external parameters can be measured and interpolated between support points or extrapolated from other support points; Figure 3: a representation of a real-world situation in which a child wants to cross a road and the proposed device recognizes the situation and instructs accordingly; Figure 4: a real-world situation in a road traffic situation in which the proposed device or method according to an aspect of the present invention is used; and Figure 5: Another real-world situation in a road traffic situation where the device is making a turn in accordance with another aspect of the present invention; Figure 6A: a schematic flowchart of a method for autonomously controlling a device according to an aspect of the present invention in the juvenile phase with the main objective of making one's own actions more precise; and Figure 6B: a schematic flowchart of a method for autonomously controlling a device according to an aspect of the present invention in the adult, the application phase to achieve its own objectives, taking into account and predicting the behavior of other relevant objects in a situation.
[37] Figure 1 presents in a schematic flowchart a method for autonomous control of a device, comprising initialization 100 by means of randomized actuation 101 of effectors and reading 102 of a plurality of internal sensor units configured to measure internal device properties and a reading 103 of a plurality of Petition 870250048900, dated 11 / 06 / 2025, page 22 / 180 17 / 41 external sensor units configured to measure external environmental properties; store 104 existing correlations between the actuation 101 of the effectors and the readout device properties and environmental properties as action instructions, so that an action instruction converts an initial triplet of actuation 101, device properties and environmental properties into a target triplet;and the activation 105 of the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored action instruction 104, which indicates how at least one predefined target device property and / or at least one predefined target environment property is defined using the initial triplet and the target triplet, wherein the process is iterated in a learning manner, in which new action instructions are constantly recognized, which in each case convert an initial triplet into a target triplet and these action instructions are stored to control the device.
[38] According to one aspect of the present invention, the model presented here uses a central model with which AI finds its way in the unknown and in the constantly changing world. This approach makes it possible to create short-term behavioral predictions of observed objects in order to estimate the development of situations and take this into account when designing its own actions. It also makes it possible to learn through observation and empathy. Aspects of the invention are: - Central model as the AI hub for registering all objects in a given situation; - Data storage that enables laws to be recognized; - Behavioral predictions of all observed objects in a given situation; - Capacity for empathy based on the central model; - Ability to learn through observation; - Recognizing influencing factors that are not directly observable; Petition 870250048900, dated 11 / 06 / 2025, page 23 / 180 18 / 41 - To recognize natural laws and causal dependencies; - A computerized form of theories that can be created, stored, reused, and continuously improved; and / or Self-assessment and evaluation of self-generated theories as a basis for cognition, learning, and selection of the best theory to achieve one's own goals.
[39] A car that behaves according to these ideas is described as an example. The way it judges a child or a box, judges an overtaking maneuver on the road, or recognizes the law of gravity and crosswinds (as an example of recognizing causalities and non-manifest objects). It also describes how empathy can be used to learn from observation. All this demonstrates that the possibilities of this AI go beyond driving cars. However, the essential basis is a physically existing robot in a physically existing environment that observes other objects directly or indirectly.
[40] Next, some aspects of the present invention are proposed which allow for an exemplary implementation of the method or device and / or system arrangement. The following aspects should be understood as merely illustrative and may be applied individually or in combination. Possible hardware for the proposed robot or device:
[41] The technical equipment that the robot or car may have according to an aspect of the present invention is described herein.
[42] Effectors:
[43] In simple cases, such as a car, for example, the effectors would be, in addition to windshield wipers, indicators, lights, braking, acceleration and steering, with which the robot influences its environment, even if it is only its own change of position, they are also changes of situation. Petition 870250048900, dated 11 / 06 / 2025, page 24 / 180 19 / 41
[44] Internal sensors:
[45] The internal sensor system measures and provides feedback on the user's own characteristics, such as height, weight, speed, acceleration, etc., based on the actions of the effectors. These are necessary for evaluating one's own actions and as a prerequisite for learning. This enables the recognition of one's own actions, for example, the braking force and the effect (the extent of negative acceleration), the realization and storage in memory of connections to learn independently, i.e., to continuously improve one's own actions.
[46] As a physical entity in a physical environment, the robot is subject to the laws of physics. A measured acceleration, that is, a change in velocity, depends not only on mass, but also on the extent to which the accelerator pedal is pressed or with what force the brake is applied. If these laws are not permanently programmed and await identification, storage and use by the robot, it needs internal sensors to learn and store the consequences of the use of the effectors. The internal sensors of a car, for example, record data such as size, speed, acceleration, etc.
[47] External sensors:
[48] According to one aspect of the present invention, the external sensor system detects the environment (distances, objects, etc.), for example, by means of lidar, image processing, etc. Obviously, the robot must have adequate knowledge of the environment. What exactly is needed is described in more detail below. However, it is not necessary to determine in advance which systems are most suitable, as the information requirements are lower due to the use of the central model. Not everything that is theoretically available in terms of information needs to be recorded and processed.
[49] Possible robot software:
[50] The parts by which the AI presented achieves its intelligence according Petition 870250048900, dated 11 / 06 / 2025, page 25 / 180 20 / 41 with an aspect of the present invention are described in this document.
[51] Data storage:
[52] The data storage is a highly available data memory for data tuples consisting of data from internal sensors, effectors, actions and their motivation values (see below). The memory appropriately combines the actions performed and the associated actual, presented and observed results. Therefore, it stores only internal data in tables and not, for example, pixels from one or more images. There is, therefore, only data whose internal meaning is known.
[53] According to one aspect of the present invention, an additional part of this module is the repeated verification of the internal consistency of the data. This means that they should not contain contradictions, structures that lead to loops or circles, etc. The objective is that only one decision can result from the data. The routine that verifies this should also remove redundant stored data, not only to keep the database as lean as possible, but also to avoid over-adaptation. What is considered redundant depends on the type of theory formation.
[54] Figure 2 presents the formation of the theory according to an aspect of the present invention:
[55] Using the simple means described here, the central model can reproduce any functional correlation, save it and use it for prediction.
[56] Braking or acceleration depends functionally on the accelerated mass and the applied force. However, this function is initially unknown because the mathematical function of the change in velocity dV depends on the preconditions Vn (current velocity, weight, etc.) and the use of the effectors Em (force on the accelerator pedal, steering, etc.), that is: dV = f(Vn, Em) (1)
[57] This function must be self-learned and not programmed, as this allows what has been learned to be expanded, corrected and improved. Petition 870250048900, dated 11 / 06 / 2025, page 26 / 180 21 / 41 subsequently in the same way. The effective acceleration during braking (or acceleration) is linked to the (mathematically) independent variable of force F (with what force the pedal is pressed) and mass m according to the law F = m * a through the internal acceleration sensor system measured a. The principle by which the robot itself finds this equation is simple. To illustrate this, let's adopt a quadratic (different, arbitrary) function y = -((x⁴)²) + 9, which must be modeled.
[58] Let us initially assume that there are only the two measured values at points 11 and 12. Now, let us assume that a prediction p is needed at the point x=3 (point 2), which leads to a value of y=3.8 by means of the straight line a between 11 and 12. However, since the error for the measured value ex post y=8.0 is very large, this point 2R at x=3.0 and y=8.0 is carried to memory. As a result, the degrees b1 between [11,2R] or b2 between [2R,12] would be used for a new prediction. Over time, many more points (31, 32, 33, ...) are created in the example above, by means of which the functional relationship of x and y can be modeled with any precision by the respective straight lines along the data tuples (c1, c2, c3, ...). The precision is limited only by the (yet) unacquired experience and the size of the memory. Logically, this can be extended to any number of variables if computing and memory capabilities allow; instead of straight lines, hyperplanes a1x1 + ...+ anxn = const are used for prediction by means of extrapolation or piecewise linear interpolation.
[59] However, in order to minimize the amount of corner or data tuples stored, they can also be removed again: If an ex post analysis reveals that a data tuple is so close to or even on the straight line / hyperplane between the two neighboring tuples, then this tuple can be deleted as redundant. This not only reduces the memory requirement but also the computational effort, and the behavior adapts better and better over time to the actual, albeit still unknown, mathematical function, leading to increasingly efficient behavior over time. Petition 870250048900, dated 11 / 06 / 2025, page 27 / 180 22 / 41
[60] This procedure also avoids over-adaptation, which must be taken into account when selecting and defining the scope of training data for an ANN. Therefore, redundant experiments are not stored repeatedly, which would completely obscure events that rarely occur, the so-called “black swans”, with potentially serious consequences.
[61] Storing functional correlations of variables has another advantage, as it also stores all possible inverse functions, since the stored data tuples do not distinguish which values were 'dependent' and which were 'independent' when they occurred. In the function above, x is the independent variable and the corresponding y value can be determined using the function y = -((x-4)2)+9. The x value would have to be looked up in the database and, if the x value is not explicitly stored, the y value can be approximately determined using the next two x values. However, you can also simply look up the y value in memory and determine an x value using extrapolation or piecewise linear interpolation – this is much easier than trying to determine the inverse function of the equation above, for example.
[62] In this very simple way, any functional interrelation can be modeled with any degree of precision, provided that the education of the action is sufficient, as in a complex sport. Naturally, the accepted error ε can also be altered and adapted over time, either when it needs to be reduced to increase the required precision, or when it can be increased to reduce memory and computational effort, as this again influences the number of support points and stored data; it is a continuous optimization in the ongoing process.
[63] According to one aspect of the present invention, a theory of the rules and laws of a current situation is therefore defined as an independent and consistent set of data from which the predictions of that theory can be calculated by the robot by means of extrapolation or piecewise linear interpolation. The data set consists of its own measurement results. Petition 870250048900, dated 11 / 06 / 2025, page 28 / 180 23 / 41 and possibly also in the parameters of created Alter Egos. In this way, the robot can determine, save, reuse, and improve each of its theories. As part of the self-assessment (see below), it is also able to select and apply the best specific theory (data set) in a situation, the one with the highest competence (see below).
[64] Actions:
[65] According to one aspect of the present invention, an action is an ordered sequence of effector operations to achieve a goal. With the aid of functional relationships between effectors, internal parameters, and effects known to the robot, actions can be designed. Initially, in the juvenile learning phase (see below), simple actions are performed. The AI learns to accelerate and brake, then accelerate, wait, and hit a barrier (externally braked with a risk of injury), accelerate, steer, brake, and so on. These simple actions are then assembled to form increasingly complex “choreographies” and are then further optimized, for example, by making braking smoother, reducing pedal pressure as speed decreases, and, for example, making turns more fluid. Optimization is achieved by a motivation value (see below), which represents a kind of internal evaluation of the action performed.From the stored (partial) “choreographies,” the one with the highest motivational value is selected. Thus, AI can continuously improve the use of its effectors. Furthermore, references to the current situation and past actions can be used to predict the user's own situation at future points in time, forming the basis for the user's own action planning.
[66] Motivational values:
[67] Motivational values reflect a kind of reward. On the one hand, they are determined internally after a self-imposed goal has been achieved (successful); on the other hand, a motivational value can also be transmitted externally (well done). If, for example, a route is traveled with strong steering movements, low speed and yet high Petition 870250048900, dated 11 / 06 / 2025, page 29 / 180 24 / 41 lateral accelerations, this sequence of actions receives a lower motivation value than the result of 'education', after which the same route is covered more quickly and still with lower lateral accelerations. Of course, this also needs to be memorized.
[68] Therefore, motivational values represent a form of unspecified intrinsic goals. They can be linked to long-term goals, such as the end point of a journey, or short-term intermediate goals, such as visiting a gas station when the tank is empty: as the tank decreases, the motivational value to refuel slowly increases until it is greater than that of the long-term goal. The pursuit of the long-distance destination is interrupted in favor of refueling and then resumed.
[69] Target system:
[70] According to one aspect of the present invention, the robot is always switched on, although it may, of course, have an off switch. Therefore, it is always performing an activity. However, this also includes doing nothing (recharging) or optimizing its own long-term action memory options (rest) in an inactive state. The action chosen is always the one that currently has the highest motivation value. If an action that is not currently being performed reaches a higher motivation value than the action that is currently being performed, it is canceled, terminated, or interrupted so that the action with the higher motivation can be performed. If the motivation of such a car on the way from Munich to Hamburg to refuel or recharge its batteries becomes greater than reaching Hamburg, it will interrupt the journey at the next opportunity and proceed to a gas station.
[71] In the juvenile phase (see below), that is, in the learning phase, of such a robot or car, the motivations for driving, braking, accelerating and turning over short distances receive high motivational values, simply because long-term internal memory signals that it must learn or reduce the error of Petition 870250048900, dated 11 / 06 / 2025, page 30 / 180 25 / 41 action to make one's own actions more precise. Later, in adulthood (see below), when such behavior causes no changes or only insignificant changes in data memory, the reduction of errors approaches zero, such behavior becomes "boring," receives only low motivational values, and then, for example, energy saving is preferred.
[72] Thus, according to one aspect of the present invention, this target system avoids a fixation on reality, which would mean an interpretation that could prove to be wrong in the future.
[73] The central model, the Ego:
[74] The central model represents the core and basis of this AI according to one aspect of the present invention. It is called Ego because it places itself at the center of its data processing and, first and foremost, sees and evaluates everything from its own perspective. Therefore, Ego does not need to receive or train any data with its meaning (e.g., image identification); it creates everything on its own.
[75] Subjective polar coordinate system:
[76] All data from the observations and experiments of this AI are subjective or made from the perspective of the Ego according to an aspect of the present invention. The spatial concept also follows this idea. No Cartesian coordinate system with a presumed origin somewhere in the observed environment is designed, in which the respective coordinates are assigned to each recognized object. Instead, a polar coordinate system can be used. This not only automatically determines the mathematical origin of the Ego, but also the subjective orientation of the Ego with front-back, top-bottom, right-left, including the motion vector. The surrounding (empty) space is taken for granted. The observed objects are recognized only by means of the distance to the viewpoint itself, i.e., the origin of the polar coordinate system and the angle. Empty space in the physical sense is therefore not considered an entity or Petition 870250048900, dated 11 / 06 / 2025, page 31 / 180 26 / 41 separate dimension. It is there and is simply an option to move there, unless there is, remains or moves another object (there), then the “place” would be occupied or the space would not be empty. An additional advantage of subjectivist polar coordinates arises next for the concept of Alter Egos (see below).
[77] Juvenile phase:
[78] Within unconditional requirements, such as the integrity of the observed objects and the robot itself, the maintenance of operational readiness (battery charge status), objectives such as uniform acceleration and braking and optimized turns and others, the robot must / can learn the use of its effectors and the resulting consequences on its own, above all to coordinate its effectors with its internal and external sensors and adjust them through increasing education 100-105, occasionally interrupted by rest phases, for example, to recharge the batteries and maintain the internal database (redundancy, consistency, ...) until the recognized error of the planning and the results becomes sufficiently small. In the case of a car, this would mainly involve accelerating, driving, braking, reaching a destination (success) or missing it (error), etc.Initially, the Ego establishes simple goals for simple actions that are performed and memorized with results, which are later assembled to form increasingly complicated "choreographies." This educational phase ends when the actions of the self-defined or externally defined goals (going there) cannot be optimized, that is, when the mathematical module cannot or can only marginally reduce the error and / or the database can hardly be optimized (number and position of points in memory). If the error reduction slows down more and more, the juvenile phase gets closer and closer to its end. An ANN, on the other hand, requires a large amount of specifically selected training data, while the Ego only needs a real environment, such as a playroom, where it cannot cause any harm to prove and exercise its own abilities. It follows intrinsic and not externally predetermined goals like an ANN and, therefore, Petition 870250048900, dated 11 / 06 / 2025, page 32 / 180 27 / 41 no interpretation of the environment or the world.
[79] Adult stage:
[80] In the juvenile phase, the device learned the ordered sequence of its effector applications to achieve the established objectives. What is the value of the knowledge and experience acquired in the juvenile phase? All the data about the world that the Ego learned in the world in which it moves are its own data, its own observations, and its own experiences. In a superordinate scientific-theoretical or philosophical sense, it must be asserted that this data is true from the Ego's point of view, in the sense of that-it-is-so or that-it-was-so. This represents a natural and fixed starting point for cognition. The Ego must, of course, maintain its internal data. In addition to the optimization already mentioned above, it must also ensure that it is free from contradictions and errors. Otherwise, the robot's manufacturer or operator could not presume that the robot will achieve the desired objectives with its conceived actions.
[81] Self-experienced and self-induced changes in the environment, such as a change of location through acceleration, steering, and braking, are experienced and stored by connecting internal and external data. There is and is not a need to go beyond one's own effectors to understand the causes in the sense of questions such as why pressing the accelerator pedal leads to acceleration, changing direction to curves, braking to negative acceleration. A car's Ego only needs three options to change location in a directed manner. At this level, the as-is is important here, not the why. A pigeon on the street certainly has no idea how a car or a cyclist works. However, if this "road user" passes far enough away, it would remain seated, but if the direction of movement were changed towards it, it would fly away. It doesn't know why, but it is aware of the behavioral possibilities and therefore pays close attention.
[82] Alter Ego: Petition 870250048900, dated 11 / 06 / 2025, page 33 / 180 28 / 41
[83] When the (adult) Ego recognizes an object in the observable environment through external sensors, the Ego simply creates a copy of itself and adjusts the parameters of the copy according to the observed properties. The Ego is therefore the central model for all observable and unobservable objects and phenomena. Hence the names Ego and, consequently, Alter Ego for the copies.
[84] By using the Ego as the central model, there are no more unknown objects. There are only unknown parameters, but they can be measured by means of external sensors, estimated from personal experience (keyword: bias) and their possible range can be limited by additional observations.
[85] Just as the Ego calculates its behavior using its stored methods, data, and parameters and can predict its own situation, the behavior of Alter Egos can also be calculated using the same adapted methods, data, and parameters. The assumptions made about Alter Egos are limited to the same physical laws and the fact that short-term goals can be derived from the observed orientation and direction of movement. Whether the observed object is a box in the street, a kangaroo in Australia, or a lady with a walker.
[86] In program terms, the Ego must exist as an object. Then, a copy of the Ego can simply be created for a new object that appears and the parameters of the new object copy can then be adapted with the help of external sensors and the user's own experience. There can be a list with the Alter Egos of the current situation in which the new copy is inserted or created: Blind class { ...}; Blind Ego( ... ); Beginning of education and learning, youth phase Petition 870250048900, dated 11 / 06 / 2025, page 34 / 180 29 / 41 / / Beginning of adulthood Blind AlterEgos[]; ··· / / newly recognized object AlterEgos[i] = Ego.clone(); / / Clone the Ego to the object AlterEgos[i].adjust(···); / / Adjustment of the new object parameter
[87] Thus, the Ego internally created an Alter Ego as a calculable image of an observed object in the environment and, as shown in Figure 6B, it can form the internal object recognition circuit 311, updating of observable object parameters 312, prediction of object behavior 313, adaptation of the sequence of effector operations 314 and a second external circuit after reaching the target and entering a rest state and (also) carrying the phase 315, until it continues with the initialization 310 for a new target.
[88] It does not matter whether the observed object is familiar or completely new. If it is new, the range of possible characteristic values is greater, which requires greater attention (scan frequency). This provides the following advantages: Every observed object is basically known! Predicting the behavior of all external objects using your own methods and data! Humans no longer need to intervene to program what is missing. The ego is constantly learning and improving its behavior. The Ego is capable of empathy and awareness of the current situation (see below). Petition 870250048900, dated 11 / 06 / 2025, page 35 / 180 30 / 41 The Ego can learn by observing Alter Egos (see below). Alter Egos may incorporate unknown causalities and rules (see below).
[89] Figure 3 presents scenario 1 a child according to an aspect of the present invention.
[90] When the Ego, for example the AI of a simple car, encounters another road user, in this example a child, it creates a copy of itself and adjusts the parameters (distance d to the Ego, orientation, speed vK, ...):
[91] Thus, the Alter Ego can use the Ego’s methods to predict where and when it will be. And by comparing this with its own calculation of when it will be in such a place, the Ego can determine if a potentially dangerous situation might arise in order to respond accordingly, i.e., slow down. In this way, the child’s short-term behavior can be predicted using its own methods and experience. NOTHING needs to be specially programmed or trained for the child that is not present when this AI first encounters a child.
[92] Figure 4 presents scenario 2 “Road” according to an aspect of the present invention.
[93] Another example that demonstrates more clearly how powerful this simple idea of Alter Egos is, is a situation on a highway with a truck and two cars behind it and a car in the fast lane, in which the second car is our Ego. Based on its own knowledge of the law of force (F=m*a) of braking force, mass with the parameters of the observed cars and the truck, i.e., with presumed mass, measured speed and presumed braking effect, the Ego can initially predict the braking distances of all road users in a possible accident situation using only its own methods to determine its own safe distance. However, let's first deal with the Ego's "thoughts" or assumptions about the other road users in order to predict their actions: Petition 870250048900, dated 11 / 06 / 2025, page 36 / 180 31 / 41 - The LKW0 drives at a certain speed, which the Ego (PKWEgo) has learned (without needing special training) that these heavy road users rarely exceed. Therefore, the truck's Alter Ego does not foresee any change in speed. - The PKWEgo, the center of observation, would be able to drive faster and overtake the LKW0, taking into account other road users. Overtaking would bring the PKWEgo to its destination more quickly and therefore would have a greater motivational value at the moment it prepares to overtake. - PKW1, another Alter Ego, thinks the same as PKWEgo in terms of the system, because the Ego would act in its place. Therefore, the Ego assumes that PKW1 also wants to overtake LKW0. Clearly, the Alter Ego of car 1 “sees” PKWEgo and will perform its presumed overtaking maneuver (from the Ego), taking into account the existence of PKWEgo, for example, by setting its indicators and paying special attention to PKWEgo and other possible road users in the overtaking lane. This would be what PKWEgo's Ego “thinks” and how it would predict PKW1's behavior. Another potential lane user, PKW2, is approaching at high speed in the fast lane. Ego also creates an Alter Ego for this and elaborates on its action prediction based on the clear lane. As shown above, the program, Ego, can predict the entire situation with all relevant users of the roadway and plan its own sequence of effector operations accordingly so that no dangerous developments occur. Petition 870250048900, dated 11 / 06 / 2025, page 37 / 180 32 / 41
[94] Alter Ego, consciousness:
[95] The Ego, therefore, has a complete internal understanding of the observable environment with all recognized objects and the ability to predict their short-term behavior. This would be a definition of consciousness that is possible and implementable in a programmatic, real, non-false or imitated way.
[96] Short-term behavioral prediction based on the ability to put oneself in the place of the observed objects means looking at the situation with all the data converted to Alter Egos and calculating the development of the situation using the calculated probable behavior of the observed objects.
[97] The Ego in the road situation above can be used for all observed objects without having to program extra routines: - predict their behavior based on the premise that all road users want to move as quickly as possible and without accidents; - Continuously adjust predictions based on observed real-world behavior (braking, acceleration, steering, adjustment indicators, headlight flasher activation, etc.); and - overtake the truck at the appropriate time.
[98] The possible behavior of all involved in this situation can be easily evaluated by the Ego applying its own methods and experiences (data) with the appropriate parameters. Don't humans do the same?
[99] The examples make it clear that the basis, the foundation, and the starting point of this AI is always the Ego, the central model. The more differentiated its internal and external sensory capacities are, and the better its education in the juvenile phase, the more intelligent its behavior will be and the greater its potential to predict the behavior of observed objects. If the Ego were also given a parameter for its own vulnerability - in Petition 870250048900, dated 11 / 06 / 2025, page 38 / 180 33 / 41 In the case of a car, one would probably talk about deformability – the Ego could also assess the risks of strong or weak contact. HOWEVER, ONLY the Ego has to be developed, programmed, and educated (once), so the last part can also be transferred by an instructed Ego. The observed objects do not need to be identified (ANNs and their training data), nor does the behavior of the identified objects need to be programmed. Of course, one could argue that the Ego is not sufficient to predict the behavior of other road users. But you only need to think about how humans try to predict the behavior of other road users. Humans also do not know the long-term goals of road users on a highway; they can only guess the short-term goals of other road users (to move as fast as possible without an accident) and prepare accordingly. Robot features:
[100] Empathy:
[101] It has been shown above how the Ego can use Alter Egos to put itself in the place of the observed objects in its environment in order to calculate its behavior from their point of view. You can call this empathy if you don't reserve the term exclusively for humans.
[102] Learning through observation:
[103] Learning through observation means observing behaviors that may be better than our own and, if possible, adopting them. The basis is the comparison between the behavior of other people and one's own behavior, and this comparison is possible thanks to the empathy defined above.
[104] The Ego's empathy is realized here by the Alter Egos, through which the Ego places itself in the situation of other observed objects in order to predict their behavior from their view of the environment. Assuming now that there is a difference between the predicted and the observed behavior. And if the Alter Ego now alters the sequence of actions of the effectors so that the Petition 870250048900, dated 11 / 06 / 2025, page 39 / 180 34 / 41 If observed behavior is reproduced, the Ego has the opportunity to copy the observed behavior if doing so constitutes an advantage.
[105] Figure 5 shows the curve behavior different from scenario 3.
[106] Let's imagine that Ego has learned to always take a curve at the same distance from the right edge of the road, the LE line. Now, he observes a car ahead that cuts the curve turning earlier, but less sharply, which also reduces lateral acceleration and therefore takes the curve faster, on the dotted curve LB.
[107] In this light, the empathy described here is a prerequisite for learning by observing and imitating the behavior of others. (There is no reason not to assume that the same applies to human beings, that is, learning through imitation based on empathy).
[108] If the data from the external sensors, converted to the situation and position of the Alter Ego, are now linked to the stored sequences of simple and more complex actions of the Ego, which have been copied to the Alter Ego, the Ego is able to learn through observation: It can recognize the difference between the self-planned sequence of actions of its own effectors (always the same distance to the edge of the road) and the observed sequence (anticipated braking and turning, less steering movement and higher speed on turns and earlier acceleration) and realize that it could drive faster on turns in this way. This form of learning is certainly faster than the usual trial-and-error approach.
[109] Recognition of causal laws:
[110] The section on theory formation explained how any observed functional relationship can be replicated. We will use this ability to recognize natural laws, for which Alter Egos are also used, when necessary and appropriate. Obviously, they are not lying down or driving on a road, but they cause changes that can be detected by Petition 870250048900, dated 11 / 06 / 2025, page 40 / 180 35 / 41 external and / or internal sensors.
[111] Let's imagine an experiment in which our AI Ego observes an apple falling. It creates an Alter Ego and perceives the accelerated movement towards the ground. Now, the Ego's first premise will be that the apple itself caused the acceleration; like the car's Ego, it has its own accelerator pedal to accelerate itself. But then the Ego will observe that no apple moves on the ground itself and that other things also fall to the ground and then stop moving. Furthermore, the Ego itself must have learned the law of gravity. On an inclined road, without pressing the accelerator pedal, it would have noticed and learned an acceleration, but only going downhill, as going uphill it has to accelerate more. The steeper the angle α, the greater the force according to the formula we know (with g = acceleration due to gravity, 9.81 m / s2): F = m * g * sin(α) (2)
[112] Remember, the Ego, of course, does not explicitly know this equation or the equation for free fall, that is, the law of gravity (sin(α=90°) = 1), but through the data points and piecewise linear interpolation or extrapolation it has an implicit knowledge of this relationship between the relevant variables.
[113] The Ego, therefore, behaves as if it knew that, as a physicist would say, like a body with a heavy mass, it is subject to the law of attraction of masses above. If it now sees another object and creates its Alter Ego, the Alter Ego is also subject to this attraction of masses. In this way, this perception assumes the state of a law of nature that affects all bodies without needing to be programmed into them. If the behavior of such an Ego were seen from the outside, it would be impossible to tell whether it really knows the law of gravity, like a physicist, or whether it is merely feigning this perception. Even if the Ego sees a pigeon sitting on the street, and that it flies away as soon as the Ego approaches, it can recognize that the beating of the wings and the upward acceleration correspond. Petition 870250048900, dated 11 / 06 / 2025, page 41 / 180 36 / 41
[114] Another example would be a strong crosswind. Suppose the Ego is a large empty van. During normal straight-line driving, lateral acceleration is zero. But then the van suddenly shifts to the side while traveling straight, and the lateral acceleration sensor is triggered. In this case, the Ego can simply create a new Alter Ego for the unexpected force and unknown cause and link it to this lateral acceleration. If the Ego can also record environmental data, such as in the forest or on the plain, or the movement of branches and add them to the data memory, the Ego has not only recognized a previously unknown variable (crosswind) but has also recognized a possible causal connection. If it can also see and judge the extent of branch movement, the Ego could even estimate the magnitude of the force acting on the side and thus the likely influence on driving behavior. In this way, the Ego itself has created a theory with a new variable.This is similar to how physicists introduced dark matter and dark energy, although not visible or observable, or how Nobel Prize winner Peter Higgs conceived of a particle in the 1960s that was then first detected at CERN in 2012. Of course, the Alter Ego, that is, the copy of oneself, can never recognize the true nature of a phenomenon with this approach in the sense of a search for truth, but it is sufficient for a competent understanding of the situation.
[115] Self-assessment:
[116] The first way of acquiring knowledge has already been described. The crosswind example was developed in the section Recognition of Causal Laws. A force (initially) unknown to the Ego suddenly generates a straight-line lateral acceleration, a measurable effect on the robot or the Ego. It creates an Alter Ego for this. However, initially there is no further information signaling to the Ego when the phenomenon occurs and how strong it is. Only when the Ego observes a temporal correlation between the movements of the branches, sometimes stronger and sometimes weaker, and its own lateral acceleration, sometimes nonexistent and sometimes existent, can the Ego link the two and Petition 870250048900, dated 11 / 06 / 2025, page 42 / 180 37 / 41 assume the same cause for both. What happens here? The Ego cannot recognize the crosswind itself, but it can recognize its own reaction to it (lateral acceleration) and the reaction of the branches. This creates an element that sometimes has an effect on the Ego and sometimes does not, but which not only has an effect on the Ego, but also on other observable objects.
[117] The second way of acquiring knowledge is based on comparing one's own behavior with that of other objects. The prerequisite is the premise that the behavior of the observed objects is “rational” and not random, although it may also be recognized as such, so that comparison would no longer be a requirement for a possible gain of knowledge. Two different behavioral comparisons are made so that the Ego can determine whether its own competence (see below) is inferior, equal, or superior to that of the observed objects in a given situation. This means that the observed objects may have greater, equal, or inferior perception of the situation. An observed van suddenly decelerates on a straight road. It appears to see something that is unknown to the observing Ego, and recognizes less than the object.If inferior competence has been identified, this can be seen as an opportunity to specifically analyze those situations in order to at least raise one's own competence to the level of others. The comparison is not about whether the behavior is right or wrong, which would only lead to the well-known problems of theories of truth, but only in relation to whether and when a behavior changes and whether behaviors are repeated. The first comparison of behavior concerns different situations: Do the objects observed in different situations also exhibit different behaviors for the Ego? The second comparison concerns the same situations: Does the behavior of the observed objects repeat itself in situations that are the same for the Ego, or does it vary? The Ego can draw conclusions from these observations of various situations over a long period of time. - that one's own competence is inferior if the objects observed are repeatedly in situations that are the same for the Ego. Petition 870250048900, dated 11 / 06 / 2025, page 43 / 180 38 / 41 exhibit different behaviors and also vary their behavior in situations that are different for the Ego. - that competence itself is superior if the observed objects repeatedly repeat their behavior in situations that are the same for the Ego, but do not vary their behavior in situations that are different for the Ego, - that competence itself is equivalent if the observed objects repeatedly repeat their behavior in situations that are the same for the Ego, but vary their behavior in situations that are different for the Ego.
[118] With this simple comparison, the Ego can recognize how great its own competence is in certain situations compared with other objects. Remember, this is not an attempt to compare with the truth of the real world.
[119] From these comparisons, the Ego can evaluate itself and deduce, for example, whether it should behave more cautiously in certain situations and try to identify what is (still) not recognizable. This can be understood as an internal mission to conduct targeted research as part of continuous learning and the continuous expansion of cognition and knowledge (see above).
[120] Competence:
[121] However, the result of these comparisons can also be used to externally assess the competence of the robot’s capabilities in order to decide on its possible applications. In contrast to the truth of the knowledge about the real world that the Ego has acquired, there are no linguistic problems in comparing competences and determining superior or inferior competence.
[122] However, the concept of truth in relation to theories has another important aspect that competence must also fulfill. Theories make predictions possible. Therefore, it is important to know the quality of Petition 870250048900, dated 11 / 06 / 2025, page 44 / 180 39 / 41 a theory and, subsequently, its predictions. In general, attributing truth to a theory serves as a criterion for the ability to apply it, to be able to apply it, with the advantage that the criterion of truth automatically excludes all competing theories and thus eliminates the problem of deciding which theory to use. The concept of competence must reach something comparable.
[123] Competence is defined and used here as the ability to understand the relevant influencing factors of a situation or similar situations. Therefore, it is initially a limited criterion, but its level can be determined in a comparative procedure, in contrast to the truth of a theory, which presumes unlimited validity in terms of time and space. However, this cannot be proven nor does the criterion admit that different theories can be compared with each other.
[124] This means that competence is more suitable for describing a robot's capabilities than the concept of truth and, unlike truth, can be determined by the robot itself in a formal self-assessment process.
[125] Figure 6A presents a schematic flowchart of a procedure for autonomous control of a device in the juvenile or education phase. The objective and end of the phase are to use and implement the effectors with sufficient precision and minimize errors.Initialization 200 is followed by randomized triggering 201 of effectors and reading 202 of a plurality of internal sensor units configured to measure internal device properties and reading 203 of a plurality of external sensor units configured to measure external environmental properties; a storage 204 of correlations that exist between triggering 201 of the effectors and reading device properties and environmental properties as action instructions, such that an action instruction is an initial trigger triplet 201 of device properties and environmental properties in a target triplet and a trigger 205 of the device according to at least one predefined target device property and / or by. Petition 870250048900, dated 11 / 06 / 2025, page 45 / 180 40 / 41 minus one predefined target environment property using at least one stored action 204 statement that specifies how at least one predefined target device property and / or at least one predefined target environment property is set based on the initial triplet and the target triplet.
[126] Figure 6A shows the juvenile phase: 200 Initiation as preparation for learning 201 Randomized activation of effectors 202 Reading 202 from a large number of internal sensor units 203 internal device characteristics and a reading 203 from a plurality of external sensor units 204 Rescue of 204 interrelationships 205 Determination of the error in the precision of the movement, immediately, and when very large, continuing in 201 Optional: Rest phase for recharging and ex-post for data preparation: Errors, redundancy, consistency, ...
[127] Figure 6B shows the adult stage: 310 Initialization as preparation for achieving a long-term goal 311 Recognition of objects 311 in the current situation 312 Creation or update of Alter Egos or their parameters 312 313 Behavioral prediction of observed objects using current Alter Egos 314 Calculation of the ideal implementation of effectors 314 for the further pursuit of their own objectives 315 Resting phase 315 for reloading and preparing ex-post data: Errors, redundancy, consistency. Petition 870250048900, dated 11 / 06 / 2025, page 46 / 180 41 / 41
[128] In this document, robots, devices, systems and Ego are used synonymously. Petition 870250048900, dated 11 / 06 / 2025, page 47 / 180
Claims
1 / 5 CLAIMS 1. A method for autonomously controlling a device, characterized in that it comprises the following steps: - Initialization (100) by means of randomized actuation (101) of effectors and reading (102) of a plurality of internal sensor units configured to measure internal properties of the device as internal states and a reading (103) of external environmental properties as external states by a plurality of external sensor units configured for measurement, wherein an object is detected by means of the external sensor units and its detected parameters are compared with the device properties and instructions for device action and the detected parameters are used to predict the behavior of the detected object, wherein an action instruction as a mapping function transfers an initial triplet of effector actuation,device properties and environmental properties for a target triplet that specifies a new internal state, a new external state, and possible actions that are possible in the new states, - storage (104) of existing correlations between the actuation (101) of the effectors and the read device properties and environmental properties as instructions for action; and - control (105) of the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one memorized action instruction (104), which indicates how at least one predefined target device property and / or at least one predefined target environment property is defined using the initial triplet and the target triplet, the method being iterated in a learning manner, in which new action instructions are constantly recognized,which in each case convert an initial triplet into a target triplet, and these action instructions are stored to control the device.
2. Method, according to claim 1, characterized in that the method is iterated in a learning manner so that new action instructions are always recognized, which in each case convert an initial triplet into a target triplet and such action instructions are stored for triggering (105) the device.
3. Method, according to any one of claims 1 or 2, characterized in that the reading (102, 103) of internal and / or external effectors is performed in such a way that the individual support values are stored and the intermediate values are interpolated and / or extrapolated.
4. A method, according to any of the preceding claims, characterized in that the property of the target device and / or the property of the target environment is altered by a device property and / or an environment property according to motivational values of each short-term or long-term goal.
5. A method, according to any of the preceding claims, characterized in that several action instructions are combined to form a choreography that converts an initial triplet into a target triplet.
6. A method, according to any of the preceding claims, characterized in that an object is detected by means of external sensor units and its detected parameters are compared with the device properties and / or instructions for the device's action.
7. Method, according to claim 6, characterized by the fact that the detected parameters are used to predict the behavior of the detected object.
8. A method, according to any one of claims 6 or 7, characterized in that the correlations of the detected object between its parameters and its actions are used to generate instructions for the device's action by identifying differences between a self-planned sequence of actions of the device and the sequence of actions observed in the detected object.
9. A method, according to any of the preceding claims, characterized in that an effector is present, such as a motor, a gripper arm, a motion unit, a drive, a windshield wiper, a flashing light, a light, a steering system, or an execution unit.
10. A method, according to any of the preceding claims, characterized in that the internal sensor units detect an effector state, an effector configuration, an effector position, a size, a weight, a dimension, a speed, an acceleration, a deceleration, a battery level, a position and / or an additional system parameter of the device.
11. A method, according to any of the preceding claims, characterized in that the external sensor units detect a distance, an object, a lidar signal, an image signal, a position, a size, a dimension, an acoustic signal and / or an additional external parameter.
12. Apparatus adapted to perform a method as defined in any of the preceding claims, characterized in that it comprises: - a startup unit configured to initialize (100) by means of random actuation (101) of effectors and to read (102) a plurality of internal sensor units configured to measure internal properties of the device as internal states and a reading (103) of external environmental properties as external states by a plurality of external sensor units configured for measurement, wherein an object is detected by means of the external sensor units and its detected parameters are compared with the device properties and instructions for device action and the detected parameters are used to predict the behavior of the detected object,wherein an action instruction as a mapping function transfers an initial triplet of effector triggering, device properties, and environmental properties to a target triplet that specifies a new internal state, a new external state, and possible actions that are possible in the new states, - a memory unit disposed for storage (104) of existing correlations between the triggering (101) of the effectors and the reading device properties and environmental properties as instructions for action; and - a control unit disposed for control (105) of the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one memorized action instruction (104), which indicates how at least one predefined target device property and / or at least one predefined target environment property is defined using the initial triplet and the target triplet,so that the method to be executed by the device is iterated in a learning manner, in which new action instructions are always recognized, which in each case convert an initial triplet into a target triplet and these action instructions are stored to control the device.