Autonomous control of devices

The method allows devices to autonomously learn and adapt by recording correlations between effector actions and environmental characteristics, overcoming the limitations of traditional AI methods by iteratively improving their operational instructions, ensuring effective operation in complex environments.

JP2026502836AActive Publication Date: 2026-01-27シュレイバーカール アルバート
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025535008
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-06-19
Publication Date
2026-01-27
Estimated Expiration
2043-06-19

Smart Images

  • Figure 2026502836000001_ABST
    Figure 2026502836000001_ABST
Patent Text Reader

Abstract

The present invention is directed to a method for autonomously controlling a device, whereby the device moves generally in the physical real world and no longer needs to be further restricted with respect to its physical configuration. The method yields the advantage that the device autonomously learns and continuously improves its learned knowledge or behavior. Generally, the method overcomes the drawback of the prior art that training data must first be created, as in traditional artificial intelligence methods. Generally, the method can be used universally, and the device automatically learns and constantly modifies its own knowledge base. Furthermore, devices set up to perform the method are proposed, as well as system configurations comprising some of the proposed devices. Furthermore, computer program products and computer-readable storage media are proposed for performing the method steps or for causing a computer to perform the method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is directed to a method for autonomously controlling a device, where the device generally moves in the physical real world and does not need to be further specified regarding its physical configuration. The method yields the advantage that the device autonomously learns and continuously improves its learned knowledge or behavior. Generally, the drawback of the prior art, in which training data must first be created, as in traditional artificial intelligence methods, is overcome. Generally, the method can be used universally, and the device automatically learns and continually modifies its own knowledge base. Furthermore, devices set up to perform the method are proposed, as well as system configurations comprising some of the proposed devices. Furthermore, computer program products and computer-readable storage media are proposed for performing the method steps or for causing a computer to perform the method. [Background technology]

[0002] Takahashi Kuniyuki et al., "Dynamic motion learning for multi-DOF flexible-joint robots using active-passive motor babbling through deep learning," ADVANCED ROBOTICS, Vol. 31, No. 18, September 17, 2017, pp. 1002-1015, XP055840673, presents a learning strategy for multi-DOF flexible-joint robots to perform dynamic motion tasks. Robots with flexible joints offer several potential advantages, such as leveraging their inherent dynamics and passively adapting to environmental changes through mechanical compliance. However, controlling such robots is challenging due to the increased complexity of their dynamics.

[0003] Various methods from Artificial Intelligence (AI) are known from the state of the art, and these methods rely on training data Selection and The algorithm is provided training Recognizing patterns in phases I expect . Here, special algorithms can recognize regularities through the specification of initial data and target values, and then use these regularities, which means that tacit knowledge can be utilized within large amounts of data. Once the training phase is complete, the properly trained algorithm is applied to real data, which in turn generates implicit dependencies to solve real-world problems. Training data must first be selected and specified, which limits the resulting artificial intelligence to the training data; Already at this stage there is the possibility of undesirable influence on the learning phase, which is not only error-prone but also costly.

[0004] Swarm intelligence is also known from the state of the art, whereby several devices are provided which then cooperate to solve problems in a distributed manner. Here too, the control and coordination of the individual participants can be complex and error-prone. This is especially true when distributed learning takes place, which must be coordinated.

[0005] Artificial neural networks are also known from the state of the art, which provide neurons and connections, i.e. nodes and edges, based on graph theory. These networks are known to mimic the functioning of the human brain and are also capable of learning. Edge weights can be changed, new edges can be added or old edges can be deleted, existing nodes can be deactivated, and new nodes can be added. This results in a dynamic learning holistic system.

[0006] However, there are at least four problems with using KNN in the above method: 1. For objects that the KNN has not been trained on, the KNN does not provide any information to call the routine assigned to the object. 2. The programmed characteristics or behavior of the detected and identified objects may be incorrect, may have changed in the meantime, or may still be unknown. 3. One or more scientists, experts, etc. select the training data and specify the results to be achieved. Even in unsupervised learning, the training data and hyperparameters are still specified by humans, which cannot eliminate errors. 4. ANNs cannot correct their errors. Each individual piece of information is distributed across all parameters of the ANN, similar to a hologram, where each data point contains information about the entire image. Therefore, the ANN must always be completely erased and completely retrained.

[0007] This means that (unlike KNN) there will be no recognition failures when the self encounters a previously unknown object, such as when a self-driving car suddenly sees two boxes on the roadway. KNN, commonly abbreviated as artificial neural network, stands for artificial neural network.

[0008] The state of the art requires creating autonomous control of devices so that time-consuming preparatory work such as providing training data and teaching can be prevented or minimized. Therefore, a self-learning system is needed that can also collectively learn or estimate what other participants are likely to do or are capable of. Therefore, a system involving several participants is needed, which we refer to as empathy, where each participant actually has knowledge about the others and can further develop this knowledge. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] TAKAHASHI KUNIYUKI et al., "Dynamic motion learning for multi-DOF flexible-joint robots using active-passive motor babbling through deep learning," ADVANCED ROBOTICS, Vol. 31, No. 18, September 17, 2017, pp. 1002-1015, XP055840673 Summary of the Invention [Problem to be solved by the invention]

[0010] Therefore, the present invention proposes a method for autonomous control of a device that operates independently based solely on provided target information, accumulating operational knowledge in the process. The proposed method should be able to recognize the operating range of an effector of any device and create an improved or at least alternative method for autonomously controlling one or more devices. Furthermore, the present invention proposes devices for performing the method, as well as system configurations including several of the proposed devices. Furthermore, the present invention proposes a computer program product and a computer-readable storage medium containing instructions for performing the method. [Means for solving the problem]

[0011] This problem is solved by the features of claim 1. Further advantageous embodiments are given in the dependent claims.

[0012] Thus, there is provided a method for autonomously controlling a device, comprising: initializing by random actuation of an effector and readouts of a plurality of internal sensor units set up to measure internal device characteristics and a plurality of external sensor units set up to measure external environmental characteristics; storing as operating instructions correlations existing between the actuation of the effector and the readout device characteristics and environmental characteristics, such that the operating instructions transform output triplets of the actuations, device characteristics and environmental characteristics into target triplets; and operating the device according to at least one predefined target device characteristic and / or at least one predefined target environmental characteristic using the at least one stored operating instruction instructing how to set at least one predefined target device characteristic and / or at least one predefined target environmental characteristic based on the output triplet and the target triplet, The method is repeated in a learning manner in which new operating instructions are continually recognized, each of which transforms an initial triplet into a target triplet, and these operating instructions are stored to control the device. A method is proposed that includes:

[0013] The proposed method is generally used to autonomously control physical devices, i.e., to create instructions for their actions. This may be a car, a robot, or a production plant. Regarding mobility, no limitations are envisioned here, meaning that the device can move on land, in the air, or on water, or even underwater. Autonomous in this context means that the method ensures that the device automatically learns and recognizes which possible actions are available. It can learn these and use sensors to learn which actions lead to which results from which starting points. In general, the device can provide its own goals, or goals can be transmitted externally. The device's own goal may be, for example, maintaining its operational capabilities. An internal goal could be, for example, visiting a charging station when the battery level is low. An external goal can be communicated by specifying which task or activity the device should perform.

[0014] In a preparation process step, initialization is performed by randomly actuating effectors. Effectors are generally physical units that affect the real world. This can affect the device itself, for example, steering, accelerating, braking, and / or the targeted manipulation of an object, for example, by a robotic arm, directly or with a tool, or by a remote-effect implement, such as a projectile or jet, or a fire extinguishing agent, for example, water or foam. In the car example, this could be the brake, accelerator pedal, etc., but also headlights, indicators, etc. Effectors are actuated, and the effects of these actions are then measured using internal and external sensors. This is recorded and can then be used further in the process. In this way, actions can be learned and used in a targeted manner in later process steps.

[0015] Because the proposed device or method can be fully executed or controlled without prior knowledge, the preparatory step is random actuation, i.e., arbitrary actuation, since it must first learn which actions lead to which results. In further iterations of the process, the learned actions are then actuated in a targeted manner. During random actuation, the system parameters of the effector are checked, for example, a robotic arm is moved to all possible positions. This is recorded each time, and the effect of each effector action on the real world is recognized. In the example of a spotlight, it is possible to provide only on and off parameters. It is then switched on and off, and external sensors are used to measure the headlight's environmental impact, and internal sensors are used to measure the headlight's operating parameters, such as temperature development. In advanced LED headlights, it is also possible to adjust both color and intensity. This is tested by random actuation, and the internal and external results are then recorded.

[0016] According to the present invention, an internal sensor unit and an external sensor unit are proposed. Internal and external do not refer to the placement of the sensor units; rather, according to the present invention, the sensor units measure external conditions related to the device, or equivalently, internal conditions. The internal conditions are all system parameters of the device itself. The external conditions are environmental variables, i.e., objects surrounding the device. This means that the internal sensor units are used to measure device characteristics, and the external sensor units are used to measure environmental characteristics. For example, if a device has a certain battery state, this is an internal parameter. If an object is transported from A to B by a robotic arm, this is an external condition, and the state of the robotic arm is therefore an internal condition. Thus, typically, internal and external parameters interact, and while actions always result in a decrease in battery condition, these actions also affect external characteristics, for example. In another direction, the device can be brought to a charging station by an effector, and then a charging process is initiated, which further affects the internal battery state.

[0017] The recorded interactions are then saved, recording how the activation of the effector affects internal and external parameters. Thus, device characteristics are stored along with environmental characteristics, and operational instructions are defined. Activating an effector is thus an action that affects both the device itself and the environment. For example, if a device moves from a first geographical location to a second geographical location, this changes the device's environment, which is an external parameter, and this also changes the device's internal state, for example, through a change in temperature or a change in battery status. This means that triples can be stored that indicate what the internal state was, what the external state was, and then what action to perform. This is translated into a new internal state and a new external state. This means that there is an output triplet from the activation of the effector, the device characteristics, and the environmental characteristics, which is transferred to a target triplet. The new internal state, the new external state, and the possible action can be specified in the target triplet.

[0018] The notion of a triple is intended merely to indicate a general transition of the system state. In general, it is also possible for an internal state and an external state to be transformed into a new internal state and a new external state by a function. Thus, an action consists of a mapping function from a first tuple to a second tuple. Those skilled in the art will recognize that both descriptions are equally powerful. Either an output tuple is transformed into a target tuple by a function, or an input triple is formed that specifies an internal state, an external state, and an action, which then leads to a target triple with a new internal state, a new external state, and further fields. The further fields can be left empty, but it is preferred that at least one action currently possible in this state is executed here. The last field can also be filled in, for example, so that a vector containing identifiers of further actions is introduced here.

[0019] To illustrate this with an example, a device can activate an effector called a motor driver to drive forward. What is recorded here is the internal state, i.e., battery charge level, the external state, i.e., image signals of the physical environment, and the action command, i.e., the output triplet of the "driver." This output vector or output triplet is then converted into a goal triplet with the new battery charge level, the new image signals from the environment, and the currently possible action, such as driving forward or reversing. During the forward movement action, it is also possible to specify, for example, braking or activating the headlights.

[0020] Alternatively, the output tuples of battery state and environmental image signals are transformed into goal tuples using a "movement" function that describes the new battery state and new environmental image signals. This records what happens internally and externally during a particular operation. This knowledge can be reused in further operational steps toward a predefined goal.

[0021] In a further process step, target device characteristics or predefined target environmental characteristics can be provided. Thus, the device can be controlled so that at least one predefined target device characteristic and / or at least one predefined target environmental characteristic is set. Operational instructions are known from the previous method steps, and thus a method for transforming internal and external states into further internal and external states is known. Target specifications, i.e., target device characteristics and / or target environmental characteristics, are used for this purpose. These target characteristics thus determine what should be achieved, which is achieved by executing the operational instructions. Thus, the device state is read, and then a stored operation that achieves the target state based on the actual state is selected. In preferred cases, it is sufficient to execute operational instructions that immediately achieve the target specification. However, in typical cases, an initial state will be sequentially transformed into a target state by several operational characteristics. In other words, several operational instructions are linked so that the goal is ultimately achieved.

[0022] To illustrate this with an example, a procedure can identify that a device can travel from A to B to C to D. This means that instructions are stored that the car can drive from A to B, B to C, and C to D. It is also implicitly stored that the car can drive from A to C and A to D. It is also stored that the car can drive from C to D. These options are used here, and if there is a destination specification that states that the car should drive from A to C, the device has two options to choose from: drive from A to B to C or drive from A to C. In general, it would also be possible to drive from A to D and then from D to C. The option to choose can also be defined in the destination specification. For example, the maximum duration of the journey can be defined in terms of time. Furthermore, a maximum energy consumption can be defined. Thus, actions are randomized in a preparation process step, and then these actions are available at runtime. In general, a procedure can provide random actuation of effectors, but can also provide reuse of previously learned instructions. This means that a randomly operated learning phase can be performed, as well as an execution phase using already saved operations. These can be alternated in any order. Furthermore, the stored operational knowledge can also be refined in the execution phase by internal and external sensors, generating new operational instructions that may be advantageous for different goals. Furthermore, the goal can also change during the execution of an operation, for example, if the battery state becomes critical. In this regard, the goal of the vehicle's own operability is promoted so that the process recognizes the need to drive to a charging station. This results, at least temporarily, in an additional goal, i.e., driving to a charging station, even if the original goal of moving to the target coordinates is not achieved. As soon as a certain charge level is reached, the goal of the vehicle's own operability is again downgraded, and the goal of driving to the destination is promoted so that the journey continues.

[0023] According to one aspect of the present invention, the method is iterated in a learning manner so that new operational instructions for transforming the output triplet into the target triplet are constantly recognized, and these operational instructions are stored for operating the device. This has the advantage that the method is established in whole or at least in part so that new situations are constantly recognized and operational instructions for the effectors are constantly created. The method can also be modified so that the activation of the effectors is no longer randomized but rather existing operational instructions can be combined, or some of the existing operational instructions can be randomized so that new possible operational instructions are created and it is clear which interactions they trigger. Furthermore, the procedure can be established so that it starts from storing interactions and continues to control the device. This means that when known instructions are executed, the interactions are also saved and new parameters can be identified. For example, it can be recognized that if the same operational instruction is executed twice, different target parameters will be created. This may be the case, for example, if a crosswind occurs during a vehicle's journey. This is saved, and an external sensor then recognizes that wind should have occurred here. This creates a new output triple, and the effector can be used to randomly try how to react in such a situation. This means that a new interaction has been recognized, and if such an output triple is identified again, it will now be clear what action must be taken to achieve the desired goal triple.

[0024] According to a further aspect of the invention, the internal and / or external effectors are read out so that individual support values ​​are stored and intermediate values ​​are interpolated and / or extrapolated. This has the advantage that the support values ​​can be selected according to their availability and can also be estimated so that measurements are not required to be available at all possible values. Instead, existing measurements can be used to estimate which values ​​will occur under normal circumstances. This means that further support values ​​can be mathematically calculated.

[0025] According to a further aspect of the present invention, the target device characteristics and / or the target environmental characteristics are at least temporarily influenced by the device characteristics and / or the environmental characteristics. This has the advantage that the target specifications can be changed or prioritized. For example, when a vehicle travels from a first geographical location to a second geographical location, access to a charging station can be prioritized during the journey if the final destination would not be reached otherwise. This therefore has the advantage that reaching the target is always guaranteed, creating a fail-safe procedure.

[0026] According to a further aspect of the invention, several operation instructions are combined to form a setup for transforming a source triplet into a target triplet. This has the advantage that even complex or compound operations can be performed, and as the method progresses or is established, new operation instructions are constantly created and combined to form increasingly efficient setups. Thus, the proposed procedure is iteratively improved.

[0027] According to a further aspect of the present invention, an object is detected using an external sensor, and its detected parameters are compared with device characteristics and / or instructions for the device. This has the advantage that the device can virtually clone itself or infer the object's characteristics from its own characteristics. This can also be called empathy. For example, if an external sensor identifies a device with similar characteristics to a running device, the newly recognized device is assumed to have similar capabilities. To illustrate this, we use a vehicle implementing the proposed invention. Thus, this first vehicle executes the proposed method and recognizes a further, i.e., second, device through an external sensor, in this case an imaging unit. The first device or first vehicle learns that it can only brake to a limited extent at a certain speed and can only steer laterally to a limited extent. Here, the first vehicle recognizes that the second vehicle has similar dimensions and is traveling at a similar speed to the first vehicle. In this case, commands are conceptually projected onto the second vehicle, and the first vehicle recognizes that the second vehicle can only brake to a limited extent and can only change direction to a very limited extent. If an overtaking maneuver is initiated on a highway, the first vehicle recognizes that the second vehicle may brake or change lanes. Therefore, based on its own behavior or the possible exit triplets and possible target triplets, it can conclude that this object has similar characteristics. Furthermore, external sensors can be used to monitor the detected behavior of the second vehicle and then update the vehicle's own data memory, which stores the interactions. This means that the output triplets and target triplets can also be recognized, allowing conclusions to be drawn from the second vehicle to the first vehicle.

[0028] According to a further aspect of the invention, the detected parameters are used to predict the behavior of the detected object. This has the advantage that it is recognized that the own device will perform a certain operating command in a particular situation, so that the other detected object will almost certainly perform the same operation in the same situation. Thus, the user's own actions or executed commands are recorded, and it is assumed that the detected object is likely to behave in the same way, or at least in a similar way.

[0029] According to a further aspect of the invention, the detected object interrelationships between its parameters and its behavior are used to generate instructions for the device. This has the advantage that externally recognized correlations related to the proposed device or devices controlled by the method can also be used. Thus, the device not only generates correlations to be evaluated, but can also observe the external world and then recognize conclusions regarding the transformation of output triples into target triples.

[0030] According to a further aspect of the invention, the effectors may be processors, memories, motors, gripper arms, movement units, drives, edges, windshield wipers, indicators, lights, steering and / or execution units. This has the advantage that all possible hardware units can be realized using the proposed device or devices to be controlled according to the method. The design is merely exemplary and not exhaustive, so any physical unit can be autonomously controlled according to the proposed method.

[0031] According to a further aspect of the invention, the internal sensor unit detects the state of the device's effectors, effector settings, effector position, size, weight, dimensions, speed, acceleration, deceleration, battery level, range, location, processor utilization, memory utilization, and / or system parameters, which has the advantage that the internal sensor unit can measure all parameters and states of the device on which the method is performed or which is performed or controlled by the method.

[0032] According to a further aspect of the invention, the external sensor unit detects distance, objects, lidar signals, image signals, position, size, dimensions, acoustic signals, and / or external parameters. This has the advantage that the complete environment of the device can be analyzed and recorded. All possible sensors that perceive the environment in any way can be used, individually or in combination.

[0033] The problem is a device adapted to perform the method of any one of the preceding claims, comprising an initialization unit adapted to initialize by random actuation of an effector and by readouts of a plurality of internal sensor units adapted to measure internal device properties and by readouts of a plurality of external sensor units adapted to measure external environmental properties; a memory unit set up to store interrelationships existing between the actuation of the effector and the readout device properties and environmental properties as operating instructions, such that the operating instructions transform output triplets of the actuations, device properties and environmental properties into target triplets; and a control unit arranged to control the device using at least one stored operating instruction instructing how at least one predefined target device property and / or at least one predefined target environmental property is set using the output triplet and the target triplet in accordance with the at least one predefined target device property and / or at least one predefined target environmental property, The method to be performed by the device is repeated in a learning manner in which new operating instructions are continually recognized, each of which transforms an initial triplet into a target triplet, and these operating instructions are stored to control the device. The problem is also solved by a device comprising: a control unit;

[0034] The problem is also solved by an arrangement comprising several devices coupled by communication techniques and exchanging operating instructions.

[0035] The problem is also solved by a computer program product having control commands for implementing the proposed method or for operating the proposed device.

[0036] According to the invention, it is particularly advantageous that the methods can be used to operate the proposed devices and units. Furthermore, the proposed devices and units are suitable for carrying out the methods according to the invention. Thus, in each case, the devices implement structural features that are suitable for carrying out the corresponding methods. However, the structural features can also be designed as process steps. The proposed methods also provide steps for implementing the functionality of the structural features. Furthermore, physical components can also be provided virtually or virtualized.

[0037] Further advantages, features, and details of the present invention will become apparent from the following description, in which aspects of the invention are described in detail with reference to the drawings. Each of the features described in the claims and the specification may be essential to the present invention individually or in any combination. Similarly, the features described above and further described herein may be used individually or in any combination. Functionally similar or identical parts or components may be designated by the same reference numerals. The terms "left," "right," "top," and "bottom" used in describing the embodiments refer to the drawings in their normally legible drawing designations or in their orientation with normally legible reference numerals. The illustrated and described embodiments should not be construed as definitive, but are merely exemplary in nature for illustrating the present invention. The detailed description serves to inform those skilled in the art; therefore, known circuits, structures, and methods are not shown or described in detail herein so as not to obscure an understanding of the present specification. [Brief explanation of the drawings]

[0038] [Figure 1] 1 is a schematic flow chart of a method for autonomously controlling a device according to an aspect of the present invention; [Figure 2] A diagram that can be used to estimate internal and / or external characteristics. In this way, internal and / or external parameters can be measured and interpolated between sampling points or extrapolated to further sampling points. [Figure 3] It is a representation of a real-world situation where a child is trying to cross a road, and the proposed device recognizes the situation and teaches from there. [Figure 4] 1 is a real-world scenario in road traffic conditions in which the proposed device or method according to an aspect of the present invention is used; [Figure 5] 10 is a further real-world scenario in road traffic conditions, where the device is performing cornering, according to a further aspect of the present invention; [Figure 6A] 1 is a schematic flow chart of a method for autonomously controlling a device according to an aspect of the present invention in a premature phase whose main purpose is to clarify its own behavior; [Figure 6B] 1 is a schematic flowchart of a method for autonomously controlling a device according to an aspect of the present invention in its maturation, application phase, to achieve its own goals while taking into account and predicting the behavior of other objects involved in the situation; DETAILED DESCRIPTION OF THE INVENTION

[0039] FIG. 1 illustrates a method for autonomous control of a device, comprising: initialization 100 by random actuation 101 of an effector and readout 102 of a plurality of internal sensor units set up to measure internal device characteristics and readout 103 of a plurality of external sensor units set up to measure external environmental characteristics; storing 104 as operation instructions correlations existing between the actuation 101 of the effector and the readout device characteristics and environmental characteristics, such that the operation instructions transform output triplets of the actuation 101, device characteristics and environmental characteristics into target triplets; and operating 105 the device using at least one stored 104 operation instruction instructing how at least one predefined target device characteristic and / or at least one predefined target environmental characteristic is set using the output triplet and the target triplet in accordance with at least one predefined target device characteristic and / or at least one predefined target environmental characteristic; The process is repeated in a learning fashion, where new operating instructions are continually recognized, these operating instructions transform initial triples into target triples each time, and these operating instructions are stored to control the device. , and a method including:

[0040] According to one aspect of the present invention, the model presented here uses a central template for AI to find its way in an unknown and constantly changing world. This approach allows it to create short-term behavior predictions of observed objects in order to estimate how a situation will unfold and take this into account when designing the device's own behavior. This approach also allows it to learn from observation and empathy. Aspects of the present invention include: A central template as the AI ​​center for recording all objects in a situation; Data storage that allows laws to be recognized; Predicting the behavior of all observed objects in a situation; Empathy based on a central template, The ability to learn from observation, Recognition of influencing factors that cannot be directly observed, Recognition of natural laws and causal dependencies; a computerized form of theory that can be created, stored, reused, and continually improved; and / or Self-evaluation and self-generated theory evaluation as a basis for cognition, learning, and choosing the best theory to achieve one's goals is.

[0041] As an example, we will describe a car that behaves according to these ideas. We will explain how the car discerns between a child or a box, how it discerns an overtaking maneuver on the highway, or how it recognizes the laws of gravity and crosswinds (as examples of causality and non-obvious object recognition). We will also explain how it can learn from observation using empathy. All of this shows that the potential of this AI goes beyond driving a car. However, the essential foundation is a robot that physically exists in a physically present environment observing other objects directly or indirectly.

[0042] In the following, several aspects of the present invention are proposed that allow for exemplary implementation of methods or devices and / or system configurations. The following aspects should be understood as merely examples and can be applied individually or in combination.

[0043] Possible hardware for the robot or proposed device: The technical equipment that a robot or vehicle according to one aspect of the present invention may have will now be described.

[0044] Effector: In the simple case of a car, for example, effectors might be windshield wipers, indicators, lights, etc. Braking, accelerating, and steering, which affect the robot's environment, are also situation changes, even if they only change the position of the robot itself.

[0045] Internal Sensor: An internal sensor system for measuring and providing feedback on its own characteristics, such as height, weight, speed, and acceleration, based on the effector's movements. Internal sensors are necessary for the evaluation of the user's movements and as a prerequisite for learning. This allows the user's movements, such as the strength and effect of braking (degree of negative acceleration), to be recognized, connected, and stored in memory for independent learning, i.e., for continuous improvement of the user's movements.

[0046] As a physical entity in a physical environment, a robot is subject to the laws of physics. Its measured acceleration, or change in velocity, depends not only on its mass but also on how hard the accelerator pedal is pressed or how hard the brakes are applied. If these laws were not to be permanently programmed, but rather the robot were to discover, memorize, and use them on its own, the robot would need internal sensors to learn and memorize the results of using its effectors. Thus, the internal sensors of a car record data such as size, speed, and acceleration.

[0047] External Sensors: According to one aspect of the invention, an external sensor system detects the environment (distance, objects, etc.), for example by means of LIDAR, image processing, etc. Naturally, the robot must be provided with knowledge of the environment in an appropriate way. What exactly is required is described in more detail below. However, since the use of a central template lowers the information requirements, it should not be predetermined which system is most suitable. It is not necessary that all theoretically available information be recorded and processed.

[0048] Possible software for the robot: The portions of the proposed AI that implement its intelligence in accordance with one aspect of the present invention are described herein.

[0049] Data storage: The data storage is a highly available data memory for data tuples consisting of data from internal sensors, effectors, actions, and their motivational values ​​(see below). The memory appropriately pairs actions taken with associated actual, experienced, and observed results. Thus, the memory stores only internal data in tables, not, for example, the pixels of one or more images. Thus, only data whose internal meaning is known exists.

[0050] According to one aspect of the present invention, a further part of this module is the repeated checking of the internal consistency of the data. This means that the data must not contain any contradictions, structures that lead to loops, circles, etc. The goal is that only one decision can result from the data. The routine that checks this must also remove stored redundant data, not only to keep the database as lean as possible, but also to prevent overfitting. What is considered redundant depends on the type of theorizing.

[0051] FIG. 2 illustrates a theorization according to one embodiment of the present invention.

[0052] The central template can, using the simple means described here, reproduce any functional correlation, store it, and use it for prediction.

[0053] Braking or acceleration is functionally dependent on the accelerating mass and the applied force. However, the mathematical function of the velocity change dV is initially unknown, since it depends on the preconditions Vn (current speed, weight, etc.) and the effector used Em (force on the accelerator pedal, steering, etc.). dV=f(Vn,Em) (1)

[0054] This function should be learned by the user, not programmed, as this allows what has been learned to be extended, modified and improved later in the same way. The effective acceleration when braking (or accelerating) is linked to the (mathematical) independent variables of force F (how hard the pedal is pressed) and mass m by the law F=m*a via an internal sensor system measuring the acceleration a. The principle by which the robot itself finds this equation is simple. To illustrate this, a (different, arbitrary) quadratic function y=-((x-4)) should be modelled. 2 )+9.

[0055] Suppose initially there are only two measurements, points 11 and 12. Now, suppose a prediction p is needed at point x=3 (point 2), which corresponds to a value of y=3.8 via line a between 11 and 12. However, the error for the subsequent measurement y=8.0 is too large, so point 2R at x=3.0 and y=8.0 is brought into memory. As a result, either frequency b1 between [11, 2R] or frequency b2 between [2R, 12] will be used for the new prediction. Over time, in the above example, many additional points (31, 32, 33, ...) will be created, through which the functional relationship between x and y can be modeled with arbitrary accuracy using respective lines along the data tuples (c1, c2, c3, ...). Accuracy is limited only by the experience (yet) ungained and the size of memory. Naturally, this can be extended to any number of variables, computational power and memory capacity permitting, and instead of a straight line a hyperplane a1x1+...+a1nxn=const is used for prediction by piecewise linear interpolation or extrapolation.

[0056] However, to minimize the number of corners or data tuples stored, they can also be removed again. That is, if a post-mortem analysis reveals that a data tuple is very close or lies on a line / hyperplane between two adjacent tuples, this tuple can be deleted as redundant. This not only reduces memory requirements, but also allows the computational effort and behavior to adapt better and better over time to the actual, though still unknown, mathematical function, which then leads to increasingly efficient behavior over time.

[0057] This procedure also prevents overfitting, which must be taken into account when selecting and defining the training data for the CNN, so that redundant experiences are not repeatedly memorized, completely obscuring rare events, so-called "black swans," with potentially serious consequences.

[0058] Storing the functional correlation of variables has another advantage, since in this way all possible inverse functions are also stored, and the stored data tuples do not distinguish which values ​​were "dependent" and which were "independent" when they occurred. In the above function, x is the independent variable, and the corresponding y value is the function y=-((x-4) 2 ) + 9. The x value must be looked up in a database, and if the x value is not explicitly stored, the y value can be approximately determined using the next two x values. However, it is also possible to simply look up the y value in memory and use piecewise linear interpolation or extrapolation to determine the x value, which is much easier than, for example, trying to determine the inverse of the above equation.

[0059] In this way, very simply, any functional context can be modeled with any accuracy, as long as the movements are well trained, as in the case of complex sports. Of course, the error tolerance ε can also be changed and adapted over time, whether it needs to be reduced to increase the required accuracy, or whether it can be increased to reduce memory and computational effort, which affects the number of corners and data points stored and is therefore a continuous optimization in an ongoing process.

[0060] Thus, according to one aspect of the present invention, a theory of rules and laws for a current situation is defined as a self-contained, consistent data set from which the predictions of this theory can be calculated by the robot by piecewise linear interpolation or extrapolation. The data set consists of the robot's own measurements and possibly other self-created parameters. In this way, the robot can determine, store, reuse, and improve each of its theories. As part of its self-evaluation (see below), it can also select and apply the specific theory (data set) that is best for the situation, i.e., the one with the highest capabilities (see below).

[0061] Operation: According to one aspect of the present invention, a behavior is an ordered sequence of effector actions to achieve a goal. The behavior can be designed using functional relationships between effectors, internal parameters, and known effects known to the robot. Initially, simple behaviors are performed in an immature learning phase (see below). The AI ​​learns to accelerate and brake, then accelerate, wait, and crash into a protective barrier (externally braked due to "risk of injury"), accelerate, steer, and brake, etc. These simple behaviors are then combined to form increasingly complex "setups," which are then further optimized, for example, by making braking more gradual by reducing pedal pressure as speed decreases, and by making cornering smoother. Optimization is achieved through a motivation value (see below), which represents a kind of internal evaluation of the performed behavior. Among the memorized (partial) "setups," the one with the highest motivation value is selected. In this way, the AI ​​can continuously improve its use of its effectors. Furthermore, based on current criteria and performed behavior experience, the robot can predict its own situation at future times, which forms the basis for the robot's own behavior planning.

[0062] Motivational Value: Motivational values ​​reflect a kind of reward. On the one hand, motivational values ​​are determined internally after a self-imposed goal has been achieved (success), and on the other hand, motivational values ​​can also be transmitted externally (good job). For example, if a route is driven with strong steering movements, low speed, and high lateral acceleration, this sequence of movements will receive a lower motivational value than the result of "training," after which the same route is driven faster and with lower lateral acceleration. Naturally, this too must be memorized.

[0063] Motivational values ​​therefore represent a form of unspecified internal goal. They can be linked to long-term goals, such as the end of a journey, or to short-term intermediate goals, such as visiting a gas station when the tank is empty; as the tank gets low, the motivational value of "refueling" slowly increases until it becomes greater than the motivational value of the long-term goal. The pursuit of a long-distance destination is interrupted to refuel and then resumed.

[0064] Objective System: According to one aspect of the present invention, the robot is always "switched on," although it may of course have an off switch. Thus, the robot is always active. However, this also includes doing nothing (recharging) or being at rest to optimize the robot's long-term memory of its own action options (sleep). The action selected is always the action with the currently highest motivation value. If an action not currently being performed reaches a higher motivation value than the currently being performed action, the currently being performed action is stopped, terminated, or interrupted so that the action with the highest motivation can be performed. If the motivation to refuel or recharge the batteries of such a car en route from Munich to Hamburg becomes greater than the motivation to reach Hamburg, the car will interrupt its journey at the next opportunity and head to a gas station.

[0065] In the immature phase (see below) or learning phase of such a robot or car, motivations for driving, braking, accelerating, and cornering over short distances are given high motivation values ​​simply because the internal long-term memory signals that the robot should learn or reduce its motion errors to make its own motion more accurate. Then, in the mature phase (see below), if such behavior causes no or only a small change in the data memory, the error reduction approaches zero and such behavior becomes "boring" and receives a low motivation value, in which case, for example, energy conservation is preferred.

[0066] Thus, in accordance with one aspect of the present invention, the goal system avoids fixation on reality, which means interpretations of reality that may prove to be incorrect in the future.

[0067] Central template, self: The central template represents the core and foundation of this AI according to one aspect of the present invention. It is called the Self because it places itself at the center of its data processing, seeing and evaluating everything first and foremost from its own perspective. Thus, the Self does not need to be fed or trained with data that has its meaning (e.g., identifying an image), but creates everything by itself.

[0068] Subjective polar coordinate system: All data of this AI's observations and experiences are subjective or created from the perspective of the self, according to one aspect of the present invention. The concept of space also follows this idea. A Cartesian coordinate system with a hypothetical origin somewhere in the observed environment, with respective coordinates assigned to each perceived object, is not designed. Instead, a polar coordinate system can be used. This automatically determines not only the self's mathematical origin, but also the self's subjective orientation—front / back, up / down, left / right, including motion vectors. The surrounding (empty) space is taken for granted. Observed objects are perceived only by the self's perspective, i.e., their distance and angle to the origin of the polar coordinate system. Thus, empty space in the physical sense is not assumed to be a separate entity or dimension. The self exists there, there is merely a choice to move there, and the "place" is occupied, or the space is not empty, unless another object exists, stands, or moves there. A further advantage of subjectivist polar coordinates arises below for another concept of self (see below).

[0069] Immature phase: Within unconditional requirements such as the integrity of the observed object and the robot itself, maintaining a state of readiness (battery charge), smooth acceleration and braking, optimized cornering, etc., the robot must / can learn to use its effectors and the results obtained, adjusting them with its internal and external sensors and fine-tuning them through incremental training 100-105, sometimes interrupted by rest phases, for example, to recharge batteries or maintain internal databases (redundancy, consistency, etc.), until the perceived errors from the plan and the results are sufficiently small. In the case of a car, this mainly includes accelerating, steering, braking, reaching a destination (success) or not (failure), etc. Initially, it sets itself simple goals of simple actions that are executed and memorized along with their results, which are later combined to form increasingly complex "setups." This training phase ends when the behavior of the self-set or externally set goal ("running there") cannot be further optimized, i.e., when the mathematical module cannot reduce the error or can only reduce the error slightly and / or cannot optimize the database (number and location of points in memory). If the error reduction decreases more slowly, the immature phase approaches its end more closely. KNN, on the other hand, requires a large amount of specially selected training data, while the self only needs a real environment, a so-called playroom, where it cannot cause any damage to test and train its own capabilities. Since the self follows an internal goal rather than an externally predetermined goal like KNN, there is no interpretation of the environment or the world.

[0070] Mature phase: During the immature phase, the device learns an ordered sequence of effector applications to achieve a set goal. What is the value of the knowledge and experience gained during this phase? All data about the world the self learns about the world it navigates through is its own data, its own observations and experiences. In a higher-level scientific or philosophical sense, this data must be true from the self's perspective—what is or what was. This therefore represents a natural, fixed starting point for cognition. Naturally, the self must maintain its internal data. In addition to the optimizations already mentioned above, it must also ensure the absence of contradictions and errors. Otherwise, the robot's manufacturer or operator cannot assume that the robot will achieve the desired goal with its designed movements.

[0071] Self-experienced and self-induced changes in the environment, such as changes in position due to acceleration, steering, and braking, are experienced and memorized by connecting internal and external data. There is no need to go further behind the self's effectors to understand causes in the sense of questions such as why pressing the gas pedal leads to acceleration, steering to cornering, or braking to negative acceleration. The car's self only needs one of three options to change position in a desired way. At this level, the "if" matters, not the "why." A pigeon on the street has no idea how cars or bicycles work. However, if such a "road user" passes far enough away from the pigeon, the pigeon will stay alight, but if the direction of movement changes toward the pigeon, the pigeon will fly away. The pigeon does not know why, but it is aware of the possibilities for behavior and therefore pays close attention.

[0072] Another self: When the (mature) self recognizes an object in the observable environment through external sensors, it simply creates a copy of itself and adjusts the copy's parameters according to the observed characteristics. The self is thus the central template of all observable and unobservable objects and phenomena. Hence the name self, and consequently, another self of the copy.

[0073] By using the self as a central template, there are no longer any unknown objects: there are only unknown parameters, which can be measured via external sensors, estimated from personal experience (keyword: preconceptions), and whose possible range can be limited by further observation.

[0074] Just as a self can use its methods, stored data, and parameters to calculate its behavior and predict its own situation, the behavior of another self can also be calculated using the same methods, data, and adapted parameters. The assumptions made about the other self are limited to the same laws of physics and the fact that short-term goals can be derived from observed orientation and direction of movement, whether the observed object is a box on the street, a kangaroo in Australia, or an elderly woman using a walker.

[0075] For a program, the self should exist as an object. In that case, a copy of the self can simply be created for each new object that appears, and the parameters of the new object copy can then be adapted with the help of external sensors and the user's own experience. There can be a list with another self of the current situation into which a new copy is inserted or created. Class CEgo {...}; CEgo Ego(...); / / The beginning of teaching and learning, the immature phase … / / Beginning of the maturation phase CEgo AlterEgos[]; … / / A new object was recognized AlterEgos[i]=Ego.clone(); / / Clone the object's self AlterEgos[i].adjust(...); / / Adjust parameters of new object …

[0076] Thus, the self internally creates another self as a computable image of the observed objects in the environment, and as shown in Figure 6B, the self can form an inner loop of object recognition 311, updating observable object parameters 312, predicting object behavior 313, adapting the sequence of effector actions 314, and a second outer loop after reaching the goal and entering a rest and (also) filling phase 315, followed by initialization 310 for the new goal.

[0077] It doesn't matter whether the observed object is familiar or completely new: if it is new, the range of possible property values ​​is larger and more attention (scanning frequency) is required. This has the following advantages: All observed objects are essentially known! Predict the behavior of all external objects using unique methods and data! Humans no longer need to program what they are missing. The self is constantly learning and improving its behavior. The self is capable of empathy and is aware of the current situation (see below). A self can learn from observing another self (see below). The other self can embody unknown causal relationships and rules (see below).

[0078] FIG. 3 illustrates Scenario 1 "Child" according to one embodiment of the present invention.

[0079] When a self, say a simple car AI, encounters another road user, in this case a child, it creates a copy of itself and adjusts its parameters (distance to self d, orientation, speed vK, ...).

[0080] In this way, the other self can use its own methods to predict where it will be and when. Then, by comparing this with its own calculations of where it will be and when, the self can determine whether a potentially dangerous situation may arise in order to react accordingly, i.e., slow down. In this way, the child's short-term behavior can be predicted using its own methods and experience. Nothing that would not be present when this AI first encounters a child needs to be specially programmed or trained for children.

[0081] FIG. 4 illustrates Scenario 2 "Highway" according to one embodiment of the present invention.

[0082] Another example that more clearly illustrates how powerful this simple idea of ​​another self is is a situation on a highway with a truck, two cars behind it, and one car in the fast lane, the second of which is the self of the present invention. Based on the force law of braking force (F=m*a), the self's own knowledge of the observed car and large truck parameters of mass, i.e., assumed mass, measured speed, and assumed braking effect, the self can predict the braking distance of all road users in a possible accident situation, first using its own method to determine its own safety distance. However, first, we will discuss the "ideas" or assumptions the self makes about other road users in order to predict their behavior. HGV0 travels at a particular speed that its self (PKWEgo) has learned (without having to be specially trained to do so) that such heavy road users rarely exceed. Therefore, the heavy truck's other self does not anticipate any speed changes. PKWEgo, the focus of observation, is able to drive faster, taking other road users into account, and overtakes HGV 0. Overtaking allows PKWEgo to reach its destination more quickly and therefore has a higher motivation value at the time PKWEgo prepares to overtake. Another self, PKW1, thinks the same as PWKEgo because of the system, since it is acting on its behalf. It therefore assumes that PKW1 is also trying to overtake LKW0. Naturally, Car1's other self "sees" PKWEgo and takes PKWEgo's presence into account and performs its (its) assumed overtaking maneuver by, for example, setting its indicators and paying particular attention to PKWEgo and any other road users in the overtaking lane. This is what PKWEgo's self "thinks" and how it predicts Car1's behavior. Another possible road user, Car 2, is approaching at high speed in the fast lane. Self also creates another self for it and creates its movement predictions based on the vacant lane. As described above, the program itself is able to anticipate the entire situation with all relevant road users and plan its own sequence of effector actions accordingly so that no dangerous developments occur.

[0083] Another Self, Consciousness: Thus, the self has a complete internal picture of the observable environment with all perceived objects and the ability to predict their behavior in the short term. This is a possible and programmatically realizable definition of consciousness, real, not faked or imitated.

[0084] Short-term behavior prediction based on the ability to put oneself in the place of the observed object means looking at the situation with all the transformed data about the other self and calculating the unfolding of the situation using the calculated possible behavior of the observed object.

[0085] The self in the above highway situation can be used as follows for all observed objects without the need to program extra routines. predicting the behavior of all road users based on the assumption that they are trying to move as quickly and without accident as possible; Continually adjust predictions based on observed actual behavior (braking, acceleration, steering, indicator settings, headlight flasher activation, etc.) Overtake large trucks at the appropriate time.

[0086] The possible behavior of all road users involved in this situation can be easily evaluated by the self by applying its own methods and experience (data) to them with the appropriate parameters. Wouldn't humans do the same?

[0087] These examples reveal that the foundation, basis, and starting point of this AI is always the self, the central template. The more differentiated its internal and external sensory capabilities are and the better its training in the immature phase, the more intelligent its behavior and the greater its potential to predict the behavior of observed objects. Given its own vulnerability parameters—in the case of a car, perhaps its deformability—the self can also assess the risk of a strong or weak contact. However, only the self needs to be constructed, programmed, and trained (once), and this last part can naturally be transferred by the trained self. Observed objects do not need to be identified (KNNs and their training data), nor do the behaviors of identified objects need to be programmed. Of course, one could argue that the self is not sufficient to predict the behavior of other road users. However, one only needs to consider how humans would evaluate their behavior. Humans, too, do not know the long-term goals of road users on the highway; they can only guess their short-term goals (moving forward as quickly as possible without accidents) and prepare accordingly.

[0088] Robot Abilities: empathy: We have shown above how one self can use another self to put itself in the place of observed objects in its environment in order to calculate their behavior from its point of view. This could be called empathy, if we were not to use the term only for humans.

[0089] Learning from observation: Learning from observation means observing behavior that is possibly better than one's own and, if possible, adopting it. It is based on comparing the behavior of others with one's own, a comparison made possible by empathy as defined above.

[0090] The self's empathy is realized here through another self, who places itself in the situation of other observed objects in order to predict their behavior from their perspective of the environment. Here, we assume that there is a difference between the predicted behavior and the observed behavior. If the other self then modifies the effector's action sequence so that the observed behavior is reproduced, the self has the opportunity to copy the observed behavior if it provides an advantage.

[0091] Figure 5 shows the different "curve behavior" for scenario 3.

[0092] Assume that the self has learned to always turn the same distance from the right edge of the road, line LE. Now the self observes the car ahead of it shortcut the curve by turning earlier but more gently on the dotted bend LB, which also reduces lateral acceleration and therefore travels around the curve faster.

[0093] From this perspective, empathy as described herein is a prerequisite for learning from observing and imitating the behavior of others (there is no reason not to assume that the same applies to humans: learning through empathy-based imitation).

[0094] If data from external sensors translated into the other self's situation and position is now linked to the self's memorized simple and more complex motion sequences, which are copied to the other self, the self can learn from observation as follows: The self can recognize the difference between the self-planned motion sequence of its own effectors (always the same distance to the road edge) and the observed sequence (earlier braking and turning, smaller steering movements and higher cornering speeds, and faster acceleration), and thus determine that it can travel faster through curves. This form of learning is reliably faster than the usual trial-and-error approach.

[0095] Recognition of causal laws: In the theorizing section, we explained how observed functional relationships can be replicated. We also use this ability to recognize natural laws and, if necessary and appropriate, to modify our other selves accordingly. Natural laws, of course, do not lie or run on roads, but they do cause changes that can be detected by external and / or internal sensors.

[0096] Consider an experiment in which the AI ​​self of the present invention observes a falling apple. It creates another self and notices its accelerating motion toward Earth. Here, the self's initial assumption is that the apple itself caused the acceleration, and that, like the car self, the apple has its own "gas pedal" to accelerate itself. However, in that case, the self observes that the apple does not move on its own across the ground, and that other things also fall to the ground and then stop moving there. Furthermore, the self must have learned the law of gravity. On a sloping road, the self would notice and learn acceleration without pressing the gas pedal, but only "downhill," while "uphill" requires more acceleration. The steeper the angle α, the greater the force according to the known formula (g = acceleration due to gravity, 9.81 m / sec²). F=m*g*sin(α) (2)

[0097] Note that the self, of course, does not explicitly know this equation or the equation of free fall, i.e., the law of gravity (sin(α=90°)=1), but has implicit knowledge of this relationship between the relevant variables through data points and piecewise linear interpolation or extrapolation.

[0098] Thus, the self behaves as if it were, as physicists would say, a massive object, knowing that it is subject to the above-mentioned law of mass gravitation. If the self now sees another object and creates another self, that other self is also subject to this mass gravitation. In this way, this realization assumes the status of a natural law that affects all objects without having to be programmed into them. If such behavior were to be observed from the outside, it would be impossible to know whether the self actually knows the law of gravity as physicists do, or whether it is merely faking this insight. Even if the self sees a pigeon alighting in the street that flies away as soon as the self approaches, the self can recognize the correspondence between the flapping of its wings and its upward acceleration.

[0099] Another example is a strong crosswind. The self is a large, empty van. During normal straight-ahead driving, lateral acceleration is zero. However, the van then suddenly shifts sideways while traveling straight, triggering the lateral acceleration sensor. In this case, the self can simply create a new, separate self for the unexpected force and unknown cause and link it to this lateral acceleration. If the self can then record data about the environment, such as in a forest or on a plain, or about the movement of branches and add that data to its data memory, the self not only recognizes the previously unknown variable (the crosswind), but also recognizes possible causal relationships. If the self can see and determine the extent of the branch movement, the self can even estimate the strength of the force acting to the side and therefore its possible effect on driving behavior. In this way, the self has created a theory with a new variable—invisible and unobservable, similar to the physicists who introduced dark matter and dark energy, or Nobel Prize-winning Peter Higgs, who conceived the particle in the 1960s and later detected at CERN in 2012. The other self, that is, the copy of the self, can of course never use this method to perceive the essence of the phenomenon in the sense of searching for truth, but it is sufficient for an adequate grasp of the situation.

[0100] self-evaluation: The first form of knowledge has already been explained. The example of the crosswind was developed in the section "Recognizing Causal Laws." A force (initially) unknown to the self suddenly produces a measurable effect on the robot or the self: a lateral acceleration in a straight line. The self creates another self for this. However, no further information initially exists to tell the self when this phenomenon occurs or how strong it is. Only when the self observes the temporal correlation between the branch's movement, sometimes stronger and sometimes weaker, and its own lateral acceleration, sometimes absent, can it link the two and attribute the same cause to both. What happens here? The self cannot perceive the crosswind itself, but it can perceive its own reaction to the crosswind (lateral acceleration) and the branch's reaction. This creates an element that sometimes affects the self, sometimes not, and that affects not only the self but also other observable objects.

[0101] The second form of knowledge acquisition is based on a comparison of one's own behavior with that of other objects. While the prerequisite is the assumption that the behavior of the observed object is "rational" and non-randomized, it can also be recognized as such, making comparison no longer a requirement for the acquisition of possible knowledge. Two different behavioral comparisons are performed, allowing the self to determine whether its own abilities (see below) are inferior, equal, or superior to those of the observed object in a given situation. This means that the observed object is able to perceive the situation more, similarly, or less. The observed van suddenly slows down on a straight road. The self appears to be seeing something unknown to the observing self, its awareness less than that of the object. If inferior abilities are identified, this can be seen as an opportunity to specifically analyze these situations in order to at least raise its own abilities to the level of the other. The comparison is not about whether behavior is right or wrong, which leads to well-known problems in truth theory, but only about whether and when behavior changes and whether behavior is repeated. The first comparison of behavior concerns different situations. That is, does the observed object in different situations also behave differently for the self? The second comparison concerns the same situation: does the behavior of the observed object replicate it in situations that are the same for the self, or does it change? The self can then draw conclusions from these observations of several situations over a longer period of time. If an object repeatedly observed in a situation that is the same for the self behaves differently, and if it also changes its behavior in a situation that is different for the self, then the self itself is inferior. If the individual repeats these behaviors in situations where the repeatedly observed object is the same for the individual, but does not change these behaviors in situations where the object is different for the individual, then the individual's own abilities are superior. If a person repeats the behavior in situations where the repeatedly observed object is the same for the self, but changes the behavior in situations where the object is different for the self, then the self's own abilities are equivalent. The conclusion can be drawn.

[0102] This simple comparison allows the self to recognize how powerful it is in a particular situation compared to other objects. Note that this is not an attempt to compare with "real-world truth."

[0103] From these comparisons, the self can evaluate itself, for example, infer whether it should behave more carefully in certain situations, and try to identify what it cannot (yet) recognize. This can be understood as an internal obligation to conduct targeted research as part of "continuous learning and the ongoing expansion of cognition and knowledge" (see above).

[0104] Expertise: However, the results of these comparisons can also be used to externally assess the capabilities of a robot in order to determine possible uses for the robot.In contrast to the truth of self-acquired knowledge about the real world, there is no linguistic problem with comparing capabilities and judging which are greater or less.

[0105] However, the concept of truth in relation to a theory has another important aspect that competence must also satisfy: theories enable predictions. It is therefore important to know the quality of a theory, and subsequently its predictions. In general, the attribution of truth to a theory serves as a criterion for being able to apply it, and for being allowed to apply it; the truth criterion has the advantage of automatically ruling out all competing theories, thus removing the problem of deciding which theory to use. The concept of competence must achieve something equivalent to this.

[0106] Expertise is defined and used here as the ability to grasp the relevant influencing factors of a situation or similar situations. Therefore, while this is initially a limited criterion, its level can be determined in comparative procedures, as opposed to the truth of a theory, which assumes unlimited validity in time and space. However, this cannot be proven, nor does this criterion allow different theories to be compared with each other.

[0107] This means that capabilities are better suited to describing a robot's capabilities than the concept of truth, and, unlike truth, can be determined by the robot itself in a formal process for self-assessment.

[0108] FIG. 6A shows a schematic flowchart of a procedure for autonomous control of a device during the pre-maturity or training phase. The goal and end of the phase is to use and deploy the effector with sufficient accuracy and minimize errors. Initialization 200 is followed by random actuation 201 of the effector and reading 202 of multiple internal sensor units set up to measure internal device characteristics and reading 203 of multiple external sensor units set up to measure external environmental characteristics; storing 204 as operation instructions a correlation that exists between the actuation 201 of the effector and the readout device and environmental characteristics, such that the operation instructions are to a target triplet of output triplet of the actuation 201, device characteristics, and environmental characteristics; and operating 205 the device using at least one stored 204 operation instruction in accordance with at least one pre-defined target device characteristic and / or at least one pre-defined target environmental characteristic, specifying how at least one pre-defined target device characteristic and / or at least one pre-defined target environmental characteristic is set based on the output triplet and the target triplet.

[0109] Figure 6A shows the immature phase. 200 Initialization in preparation for learning 201 Random Effector Activation 202 Reading out multiple internal sensor units 202 203 Internal device characteristics and readout from multiple external sensor units 203 204 Preserving Interrelationships204 205 Determine movement accuracy error, immediately, if too large continue with 201 Optional: Rest phase for recharging and post-processing for data preparation: Error, redundancy, consistency,...

[0110] Figure 6B shows the maturation phase. 310 Initialization in preparation for achieving long-term goals 311 Recognizing objects in the current situation 311 312 Creating or updating another self or its parameters 312 313 Predicting the behavior of observed objects using their current alternate selves313 314 Computing optimal effector deployment to further one's own goals 315 Rest Phase for Recharge and Post-Event Data Preparation: Error, redundancy, consistency,...

[0111] In this specification, robot, device, system, and self are used interchangeably.

Claims

1. 1. A method for autonomous control of a device, comprising: initialization (100) by random actuation of an effector (101) and reading out a plurality of internal sensor units (102) arranged to measure internal device characteristics and a plurality of external sensor units (103) arranged to measure external environment characteristics; storing (104) as an operating instruction the interrelationship that exists between the actuation (101) of the effector and the read-out device and environmental characteristics, such that the operating instruction transforms an output triplet of actuation (101), device and environmental characteristics into a target triplet; Controlling (105) the device using at least one stored (104) operating instruction in accordance with at least one predefined target device characteristic and / or at least one predefined target environmental characteristic, the operating instruction instructing how the at least one predefined target device characteristic and / or the at least one predefined target environmental characteristic is set using the output triplet and the target triplet; A method comprising:

2. 2. The method of claim 1, wherein the method is iterated in a learning manner so that new operating instructions for transforming an output triplet into a target triplet are constantly recognized, and these operating instructions are stored for operating (105) the device.

3. 3. The method according to claim 1 or 2, characterized in that the reading (102, 103) of the internal and / or external effectors is performed in such a way that individual support values ​​are stored and intermediate values ​​are interpolated and / or extrapolated.

4. 4. The method according to claim 1, wherein the target device characteristic and / or the target environmental characteristic are at least temporarily influenced by a device characteristic and / or an environmental characteristic.

5. 5. A method according to any one of claims 1 to 4, characterized in that several operation instructions are combined to form a procedure for transforming an output triplet into a target triplet.

6. 6. The method according to any one of claims 1 to 5, characterized in that an object is detected by an external sensor unit and its detected parameters are compared with device properties and / or operation instructions of the device.

7. 7. The method of claim 6, wherein the detected parameters are used to predict the behavior of the detected object.

8. 8. A method according to claim 6 or 7, characterized in that the correlation of the detected object between its parameters and its operation is used to generate instructions for the operation of the device.

9. 9. The method according to any one of claims 1 to 8, characterized in that the effectors are present as processors, memories, motors, gripper arms, movement units, drives, edges, windscreen wipers, flashers, lights, steering and / or execution units.

10. 10. The method of claim 1, wherein an internal sensor unit detects the state of an effector, effector settings, effector position, size, weight, dimensions, speed, acceleration, deceleration, battery level, range, position, processor utilization, memory utilization, and / or system parameters of the device.

11. 11. The method according to claim 1, wherein the external sensor unit detects distances, objects, lidar signals, image signals, positions, sizes, dimensions, acoustic signals, and / or external parameters.

12. 12. An apparatus adapted to carry out the method according to any one of claims 1 to 11, comprising: an initialization unit arranged to initialize (100) by random actuation (101) of an effector and reading (102) a plurality of internal sensor units arranged to measure internal device characteristics and reading (103) a plurality of external sensor units arranged to measure external environment characteristics; a memory unit arranged to store (104) as an operation command the correlation existing between said operation (101) of said effector and said read-out device characteristics and environmental characteristics, such that said operation command transfers an output triplet of operation (101), device characteristics and environmental characteristics to a target triplet; a control unit arranged to control (105) said device according to at least one predefined target device characteristic and / or at least one predefined target environmental characteristic using at least one stored (104) operating instruction instructing how said at least one predefined target device characteristic and / or said at least one predefined target environmental characteristic is set based on said output triplet and said target triplet; An apparatus comprising:

13. 13. A system arrangement comprising several devices according to claim 12, coupled by communication technology, for exchanging operating instructions.

14. A computer program product comprising instructions which, when executed by at least one computer, cause said computer to perform the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium comprising instructions that, when executed by at least one computer, cause the computer to perform the steps of the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Device and method for controlling a robot device

    DE102020212658A1

  • Model-free control of dynamical systems with deep reservoir computing

    WO2020159947A1