Autonomous control of devices
By randomly driving actuators and sensor measurements, recording and iteratively learning the correlation between actuator actions and environmental properties, the problem of large amounts of training data required for autonomous control of devices in existing technologies is solved, and autonomous learning and error correction in unknown environments are achieved, reducing costs and complexity.
Patent Information
- Application Number
- CN202380092260.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-06-19
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-06-19
AI Technical Summary
In the existing technology, autonomous control methods for devices require a large amount of training data and are difficult to autonomously learn and correct errors in unknown environments, resulting in high costs and increased complexity.
By randomly driving the actuator, combining internal and external sensor measurements, recording the correlation between the actuator action and the environment and device properties, forming action instructions, and iterative learning to autonomously control the device.
It achieves autonomous learning and error correction without the need for initial knowledge, improves the device's autonomous control capabilities in unknown environments, and reduces the complexity and cost of preparation work.
Smart Images

Figure CN120603683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for autonomously controlling a device, wherein the device generally moves in the real physical world and does not need to be further specified in terms of its physical configuration. The advantage achieved by the method is that the device learns autonomously and continuously improves the learned knowledge or behavior. Overall, the disadvantage of the prior art that training data must first be created, as is usually the case with traditional artificial intelligence methods, is overcome. Overall, the method has universal applicability, and the device automatically learns and continuously corrects its own knowledge base. In addition, a device for implementing the method is proposed, as well as a system arrangement comprising a plurality of the proposed devices. In addition, a computer program product and a computer-readable storage medium are proposed, which are used to perform the method steps or to cause a computer to perform the method. Background Art
[0002] Various artificial intelligence (AI) methods are known in the art. These methods require training data and then let the algorithm identify patterns during a learning phase. Specific algorithms can identify and extract patterns from databases. This approach allows implicit knowledge to be extracted from large amounts of data. After the training phase, the appropriately trained algorithm is applied to real-world data to generate implicit dependencies to solve real-world problems. A drawback of this approach is that training data must be available first, and this can already negatively impact the learning phase, making it error-prone and costly.
[0003] Swarm intelligence is also known in the art, where multiple devices collaborate to solve problems in a distributed manner. In this case, the control and coordination of the individual participants is also complex and sometimes error-prone. This is particularly true when distributed learning, which requires further coordination, is required.
[0004] Artificial neural networks are also known from the prior art. They are based on graph theory, providing neurons and connections, i.e., nodes and edges. These networks are known to mimic the functions of the human brain and can also learn. Edge weights can be changed, new edges can be added or deleted, and existing nodes can be deactivated or new ones added. This creates an overall system that dynamically learns.
[0005] However, there are at least four problems with using KNN in the above manner:
[0006] 1. For an object for which KNN has not been trained, KNN will not provide any information to call the routine assigned to that object.
[0007] 2. The programmed properties or behaviors of detected and identified objects may be incorrect, have changed in the meantime, or remain unknown.
[0008] 3. One or more scientists, experts, etc. select the training data and specify the outcomes to be achieved. Even in unsupervised learning, the training data and hyperparameters are still specified by humans, which does not eliminate the possibility of error.
[0009] 4. ANNs cannot correct errors. Every single piece of information is distributed across all the parameters of the ANN, similar to a hologram, where each data point contains information about the entire image. Therefore, the ANN must always be completely erased and retrained.
[0010] This means that when the ego encounters a previously unknown object, it will not fail to recognize it (unlike KNN), such as the situation in the video where the self-driving car suddenly sees two boxes in the lane. KNN usually stands for Artificial Neural Network.
[0011] In the prior art, it is desirable to enable autonomous control of devices so that time-consuming preparations, such as providing training data and teaching, can be avoided or minimized. Therefore, there is a need for self-learning systems that can collectively learn, or estimate what other participants might or can do. Therefore, there is a need for systems that include multiple participants, where each participant actually understands and further expands upon the others, a process referred to herein as empathy. Summary of the Invention
[0012] Therefore, the present invention aims to provide a method for autonomously controlling a device that operates independently based solely on provided target information and accumulates knowledge of its actions in the process. The proposed method should be able to identify the range of actions of any device's actuators and thus create an improved method or at least an alternative method for autonomously controlling one or more devices. Furthermore, the present invention aims to provide a device for performing the method and a system arrangement comprising a plurality of the proposed devices. Furthermore, the present invention aims to provide a computer program product and a computer-readable storage medium containing instructions for performing the method.
[0013] This problem is solved by the features of claim 1. Further preferred embodiments are given in the dependent claims.
[0014] Accordingly, a method for autonomously controlling a device is proposed, the method comprising: initializing by randomly driving an actuator and reading a plurality of internal sensor units configured to measure internal device properties, and reading a plurality of external sensor units configured to measure external environmental properties; storing a correlation between the driving of the actuator and the read device properties and environmental properties as an action instruction, so that the action instruction converts an output triple of the driving, device properties, and environmental properties into a target triple; and driving the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one stored action instruction, the at least one stored action instruction indicating how to set at least one predefined target device property and / or at least one predefined target environmental property based on the output triple and the target triple.
[0015] The proposed method is used to autonomously control a device (usually physical), i.e. to create action instructions. This can be a car, a robot or a production plant. In terms of mobility, no restrictions are envisaged here, which means that the device can move on land, in the air, on water or even underwater. "Autonomous" in this context means that the method ensures that the device automatically learns and automatically recognizes the possible actions available. These actions can be learned, and sensors can be used to learn which actions lead to which results from which starting points. Generally, the device can set its own goals or transmit goals to an external device. For example, the device's own goal can be to maintain operational capabilities. The internal goal can be, for example, to go to a charging station when the battery is low. The external goal can be transmitted by specifying the tasks or activities that the device should perform.
[0016] During the preparatory process step, initialization occurs by randomly actuating actuators. Actuators are typically physical units that have an effect on the real world. For example, actuators can affect the device itself, such as steering, acceleration, braking, etc., and / or perform targeted manipulation of objects, such as directly, using tools, or through instruments with remote effects, such as projectiles or sprays, or fire extinguishing agents such as water or fire foam. In the automotive example, actuators could be brakes, accelerator pedals, etc., as well as headlights, indicator lights, etc. The actuators are actuated, and the effects of these actions are measured using internal and external sensors. This is recorded and can then be used in the next step of the process. In this way, actions are learned and used in a targeted manner in subsequent process steps.
[0017] Because the proposed device or method can be implemented or controlled without any initial knowledge, the preparatory steps are initiated randomly—that is, arbitrarily—because it must first learn which actions lead to which outcomes. In subsequent iterations of this process, the learned actions are then initiated in a targeted manner. During the random activation process, the system parameters of the actuators are checked, for example, by moving a robotic arm to all possible positions. Each of these cases is recorded, and the actuator's actions are identified as having what effects on the real world. In the example of a spotlight, it might only have parameters for turning it on and off. It is then turned on and off, and the headlight's impact on the environment is measured using external sensors, while internal sensors measure operating parameters such as temperature changes. For advanced LED headlights, both color and intensity can be adjusted. This is tested by randomly driving the system, and the internal and external results are then recorded.
[0018] According to the present invention, internal and external sensor units are proposed. "Internal" and "external" do not refer to the placement of the sensor units. Rather, according to the present invention, the sensor units measure external conditions relative to the device, or similarly, internal conditions relative to the device. Internal conditions refer to all system parameters of the device itself. External conditions refer to environmental variables, i.e., variables of the objects surrounding the device. This means that internal sensor units measure device properties, while external sensor units measure environmental properties. For example, if a device has a certain battery state, this is an internal parameter. If an object is transferred from A to B using a robotic arm, this is an external state, while the state of the robotic arm is an internal state. Therefore, internal and external parameters often interact. For example, actions always result in a decrease in the battery state, which in turn affects the external properties. Viewed from another perspective, this interaction can occur: an actuator can be used to bring the device to a charging station, initiating the charging process, which in turn affects the internal battery state.
[0019] The recorded interactions are then saved and record how the actuation of the actuator affects internal and external parameters. Thus, device properties are stored together with environmental properties and define action instructions. Therefore, actuating an actuator is an operation that affects both the device itself and the environment. For example, if a device moves from a first geographical point to a second geographical point, this changes the device's environment, which is an external parameter, and also changes the device's internal state, for example through temperature changes or changes in battery status. This means that a triple can be stored that indicates what the internal state is, what the external state is, and what action is subsequently performed. This is converted into a new internal state and a new external state. This means that the actuation of the actuator, the device properties, and the environmental properties form an output triple, which is converted into a target triple. The new internal state, the new external state, and possible actions can be specified in the target triple.
[0020] The concept of triples is only used to illustrate the general conversion of system states. Generally speaking, internal states and external states can also be converted into new internal states and new external states by functions. Therefore, an action consists of a mapping function from a first tuple to a second tuple. Those skilled in the art will recognize that the effects of these two descriptions are equivalent. Either the output tuple is converted into a target tuple by a function, or an input triple is formed, which specifies the internal state, external state and action, and the action leads to a target triple, which contains a new internal state, a new external state and another field. This other field can remain empty, but it is preferred to perform at least one action that can currently be performed in this state. The last field can also be filled in the following way: for example, so that a vector containing another action identifier is introduced here.
[0021] For example, the device can drive an actuator called a motor driver and move forward. What is now recorded is an output triplet: the internal state (battery charge), the external state (image signal from the physical environment), and the action command ("move."). This output vector or triplet is now transformed into a target triplet, which contains the new battery charge, the new image signal from the environment, and the action that can now be performed, such as moving forward or reversing. During the forward motion action, for example, braking or activating the headlights can also be specified.
[0022] In an alternative example, a "move" function is used to transform the output triple of battery state and environmental image signals into a target triple describing the new battery state and new environmental image signals. This records what happened internally and externally during a specific action. This knowledge can be reused in other action steps to achieve a predetermined goal.
[0023] In another process step, target device properties or predefined target environment properties can be provided. Thus, the device can be controlled by setting at least one predefined target device property and / or at least one predefined target environment property. The action instructions are already known from the previous method steps and therefore know how to convert the internal state and the external state into another internal state and another external state. For this purpose, a target specification is used, i.e., the target device properties and / or the target environment properties. Therefore, these target properties determine the target to be achieved, which is achieved by executing the action instructions. Therefore, the device state is read and then a stored action is selected based on the actual state to achieve the target state. Preferably, the action instruction that immediately achieves the target specification is executed. However, in the general case, the initial state will be converted to the target state in sequence through multiple action properties. In other words, multiple action instructions are linked in a manner that ultimately achieves the target.
[0024] For example, the program might recognize that a device can move from A to B, C to D, and D to C. This means the stored instructions are that the car can travel from A to B, B to C, and C to D. It also implicitly memorizes that the car can travel from A to C and A to D. It also stores that the car can travel from C to D. Now, using these options, if the destination specification states that the car should travel from A to C, the device has two options: travel from A to B and then to C, or from A to C. Generally speaking, it's also possible to travel from A to D and then from D to C. The destination specification can also define which option is chosen. For example, a maximum duration for the trip can be defined based on time. Furthermore, a maximum energy consumption can be specified. Thus, the actions are randomized during the preparatory process steps and then available at runtime. Overall, the program can be configured to randomly drive the actuators, but it can also reuse previously learned instructions. This means that a learning phase involving random driving can be performed, followed by an execution phase using stored actions. This can also be performed in any order. Furthermore, during the execution phase, the stored action knowledge is refined using internal and external sensors, generating new action instructions, which may target different goals. Furthermore, goals can change during action execution in the following ways: for example, if the battery condition becomes critical, the vehicle's own maneuverability goal is increased so that the program recognizes that it is now necessary to drive to a charging station. This results in a superimposed goal, at least temporarily, that is, driving to a charging station even if this does not achieve the goal of driving to the original target coordinates. Upon reaching a certain charge level, the vehicle's own maneuverability goal is lowered again, while the goal of driving to the destination is increased to allow the journey to continue.
[0025] According to one aspect of the present invention, the method iterates in a learning manner to continuously identify new action commands that, in each case, convert an output triple into a target triple, and stores these action commands for driving the device. This has the advantage that the method is fully or at least partially designed to continuously identify new situations and create action commands to drive the actuator. The method can be modified so that the actuator actuation is no longer random, but rather existing action commands can be combined or portions of existing action commands can be randomized to create new possible action commands with a clear understanding of which interactions they trigger. Furthermore, the program can be designed to start with stored interactions and continue through to device control. This means that when known commands are executed and new parameters are identified, the interactions are also stored. For example, it can be detected that when the same action command is executed twice, different target parameters are generated. This situation can occur, for example, if a crosswind occurs while the car is driving. This situation is stored, and external sensors can then detect that there is wind in the area. This creates new output triples that the actuator can use to randomly try out how to countersteer in this situation. This means that new interactions have been identified, and if such an output triple is identified again, it is now clear which actions must be performed to achieve the desired target triple.
[0026] According to another aspect of the present invention, the internal and / or external actuators are read by storing individual support values and interpolating and / or extrapolating intermediate values. This has the advantage that support values can be selected based on their availability and estimated without having to measure every possible value. Instead, existing measurements can be used to estimate which values would normally occur. This means that additional support values can be calculated mathematically.
[0027] According to another aspect of the present invention, target device properties and / or target environmental properties are at least temporarily influenced by device properties and / or environmental properties. This has the advantage that target specifications can also be modified or prioritized. For example, if a vehicle is traveling from a first geographical point to a second geographical point, prioritizing travel to a charging station during this time can be considered, as otherwise the vehicle would not be able to reach its final destination. This has the advantage of always ensuring that the target will be achieved and creating a fail-safe procedure.
[0028] According to another aspect of the present invention, multiple action instructions are combined to form an orchestration that transforms a source triple into a target triple. This has the advantage that even complex composite actions can be executed, and as the method progresses or is established, new action instructions are continually created and combined to form increasingly efficient orchestrations. Thus, the proposed procedure is iteratively improved.
[0029] According to another aspect of the present invention, external sensors are used to detect objects and the parameters detected by the external sensors are compared with device properties and / or device commands. This has the advantage that a device can virtually clone itself or infer the properties of an object based on its own properties. This can also be referred to as empathy. For example, if an external sensor identifies a device with similar characteristics to an actuating device, the newly identified device is assumed to have similar functionality. To illustrate, let's consider a vehicle implementing the present invention. A first vehicle executes the proposed method and uses external sensors, in this case, an imaging unit, to identify another device, namely a second device. The first device, or first vehicle, has learned that it can only brake to a limited extent at a certain speed and can only perform lateral steering to a limited extent. The first vehicle now recognizes that the second vehicle is of similar size and traveling at a similar speed. The commands are then conceptually projected onto the second vehicle, which recognizes that the second vehicle can only brake to a limited extent and can only make very limited directional changes. If a passing maneuver is performed on a highway, the first vehicle recognizes that the second vehicle can brake or change lanes. Therefore, based on its own behavior or possible output triples and possible target triples, it can be inferred that the object has similar characteristics. In addition, external sensors can be used to monitor the behavior of the detected second vehicle and then update the vehicle's own data memory used to store interactions. This means that output triples and target triples can also be identified, and conclusions about the first vehicle can be inferred from the second vehicle.
[0030] According to another aspect of the present invention, the detected parameters are used to predict the behavior of the detected object. The advantage of this is that, by recognizing that a user's device executed a specific action instruction in a specific context, it is likely that another detected object will also perform the same action in the same context. Thus, the user's own actions or executed instructions are recorded, and it is then assumed that the detected object is likely to behave in the same or at least similar manner.
[0031] According to another aspect of the present invention, correlations between parameters of the detected object and its actions are used to create device instructions. This has the advantage that externally identified correlations related to the proposed device or the device controlled by the method can also be used. Thus, the device not only creates evaluated correlations but also observes the external world and draws conclusions from this about the transformation of output triples into target triples.
[0032] According to another aspect of the present invention, the actuator is a processor, a memory, a motor, a gripper arm, a motion unit, a drive, an end effector, a wiper, an indicator light, a lamp, a steering gear, and / or an actuator. This has the advantage that all possible hardware units can be controlled using the proposed device or according to the method. This design is only an example and is not exhaustive; therefore, any physical unit can be autonomously controlled according to the proposed method.
[0033] According to another aspect of the present invention, the internal sensor unit detects actuator status, actuator settings, actuator position, size, weight, size, speed, acceleration, deceleration, battery level, range, position, processor utilization, memory utilization, and / or system parameters of the device. This has the advantage that the internal sensor unit can measure all parameters and states of the device in which the method is performed or the device is executed or controlled by the method.
[0034] According to another aspect of the present invention, the external sensor unit detects distances, objects, lidar signals, image signals, positions, dimensions, sizes, acoustic signals, and / or external parameters. This has the advantage of allowing the complete environment of the device to be analyzed and recorded. All sensors capable of detecting the environment in any manner can be used individually or in combination.
[0035] The problem is also solved by a device suitable for implementing the method according to one of the preceding claims, the device comprising: an initialization unit, which is suitable for initializing by randomly driving an actuator and reading a plurality of internal sensor units suitable for measuring internal device properties, and reading a plurality of external sensor units suitable for measuring external environmental properties; a memory unit, which is configured to: store the relationship between the driving of the actuator and the read device properties and environmental properties as action instructions, so that the action instructions convert the output triples of driving, device properties and environmental properties into target triples; and a control unit, which is configured to control the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one stored action instruction, the at least one stored action instruction indicating how to use the output triples and the target triples to set at least one predefined target device property and / or at least one predefined target environmental property.
[0036] The problem is also solved by an arrangement comprising a plurality of devices which are connected via communication technology and which exchange action instructions.
[0037] The object is also achieved by a computer program product having control commands for implementing the proposed method or operating the proposed device.
[0038] According to the present invention, the method can be used particularly advantageously to operate the proposed device and unit. Furthermore, the proposed device and unit are suitable for implementing the method according to the present invention. Thus, in each case, the device implements structural features suitable for carrying out the corresponding method. However, these structural features can also be implemented as method steps. The proposed method also provides steps for implementing the functionality of the structural features. Furthermore, physical components can also be provided virtually or in a virtualized form.
[0039] Further advantages, features, and details of the present invention will become apparent from the following description, which describes various aspects of the invention in detail in conjunction with the accompanying drawings. The features recited in the claims and the specification may be essential to the invention individually or in any combination. Likewise, the features recited above, as well as those described further herein, may be used individually or in any combination. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Functionally similar or identical parts or components are sometimes provided with the same reference symbols. The terms "left", "right", "upper" and "lower" used in the description of the embodiments refer to the drawings with generally legible graphic names or orientations of legible reference symbols. The embodiments shown and described should not be understood as conclusive, but rather have an exemplary nature for explaining the present invention. The detailed description is intended to inform those skilled in the art, and therefore, known circuits, structures and methods are not shown or explained in detail in the description so as not to hinder the understanding of this specification. The drawings show:
[0041] Figure 1 : A schematic flow chart of a method for autonomously controlling a device according to one aspect of the present invention;
[0042] Figure 2 : A schematic diagram that can be used to estimate internal and / or external properties. This allows internal and / or external parameters to be measured and interpolated between sampling points, or extrapolated to other sampling points;
[0043] Figure 3 : Schematic diagram of a real-life situation where a child wants to cross the road, the proposed device recognizes this situation and teaches from it;
[0044] Figure 4 : a real-life situation in a road traffic situation, in which the device or method according to one aspect of the present invention is used; and
[0045] Figure 5 : Another realistic situation in a road traffic situation, wherein the device is practicing a turn according to another aspect of the invention.
[0046] Figure 6A: a schematic flow chart of a method for autonomously controlling a device in a juvenile stage according to one aspect of the present invention, the main purpose of which is to clarify its own actions; and Figure 6B : A schematic flow chart of a method for autonomous control of a device according to one aspect of the present invention in the adult stage, which application stage aims to achieve its own goals while taking into account and predicting the behavior of other objects related to the context. DETAILED DESCRIPTION
[0047] Figure 1 A method for autonomously controlling a device is shown in the form of a schematic flow chart, the method comprising: initializing 100 by randomly driving 101 an actuator, reading 102 a plurality of internal sensor units configured to measure internal device properties, and reading 103 a plurality of external sensor units configured to measure external environmental properties; storing 104 a correlation between the driving 101 of the actuator and the read device properties and environmental properties as an action instruction, such that the action instruction converts an output triple of the driving 101, the device properties, and the environmental properties into a target triple; and driving 105 the device according to at least one predefined target device property and / or at least one predefined target environmental property using at least one action instruction stored 104, the at least one action instruction stored 104 indicating how to use the output triple and the target triple to set at least one predefined target device property and / or at least one predefined target environmental property.
[0048] According to one aspect of the present invention, the proposed model uses a central template that allows AI to navigate an unknown and ever-changing world. This approach enables short-term behavioral predictions of observed subjects, allowing the AI to assess evolving situations and factor them into its own behavior design. It also enables learning through observation and empathy. Key aspects of the present invention include:
[0049] - The central template acts as the core of the artificial intelligence and records all objects in the situation;
[0050] - Data storage that enables patterns to be identified;
[0051] -Predict the behavior of all observed objects in the situation;
[0052] -Empathy capabilities based on a central template;
[0053] -The ability to learn through observation;
[0054] - Identify influencing factors that cannot be directly observed;
[0055] - Identify natural laws and cause-and-effect relationships;
[0056] - a computerized form of theory that allows for creation, storage, reuse and continuous improvement; and / or
[0057] -Self-assessment and evaluation of self-generated theories as a basis for cognition, learning, and selection of the best theories to achieve one's goals.
[0058] The article describes how a car, shaped by these concepts, can interpret its behavior. It describes how it can distinguish between a child and a box, how it can judge overtaking on the highway, and how it can recognize the laws of gravity and crosswinds (as examples of recognizing causal relationships and non-obvious objects). It also describes how empathy can be used for observational learning. All of this suggests that the possibilities of this type of artificial intelligence extend far beyond driving a car. However, the core foundation is a physical robot situated in a physical environment, able to directly or indirectly observe other objects.
[0059] Some aspects of the present invention are presented below, which can exemplarily implement the method, device and / or system arrangement. The following aspects should be understood to be merely exemplary and can be applied alone or in combination.
[0060] Possible hardware for the robot or the proposed device:
[0061] This document describes technical equipment that a robot or car may be equipped with according to one aspect of the present invention.
[0062] Actuator:
[0063] In simple cases, such as cars, actuators can be windshield wipers, indicators, lights, etc. Braking, acceleration, and steering are actions that a robot takes to affect its environment; even a change in its own position is a change in context.
[0064] Internal sensors:
[0065] Internal sensor systems measure and provide feedback on the robot's own characteristics, such as height, weight, speed, and acceleration, based on the actions of the actuators. These are essential for evaluating its own movements and are a prerequisite for learning. This enables the robot to identify its own movements, such as the intensity and effectiveness of braking (the degree of negative acceleration), establish connections, and store these connections in memory, allowing it to independently learn and continuously improve its own movements.
[0066] As a physical entity in a physical environment, a robot is subject to the laws of physics. The measured acceleration—the change in velocity—depends not only on mass but also on how deeply the accelerator pedal is depressed or how hard the brakes are applied. If these laws aren't permanently programmed into the robot, but rather discovered, stored, and used independently, internal sensors are needed to learn and store the results of using actuators. For example, a car's internal sensors record data such as size, velocity, and acceleration.
[0067] External sensors:
[0068] According to one aspect of the present invention, an external sensor system detects the environment (distance, objects, etc.) using, for example, lidar, image processing, etc. Naturally, this environmental information must be provided to the robot in an appropriate manner. The specific requirements are described in more detail below. However, due to the use of a central template, the information requirements are relatively low, making it unnecessary to determine in advance which system is most suitable. Not all theoretically available information needs to be recorded and processed.
[0069] Possible software for the robot:
[0070] This article describes the various parts of the artificial intelligence that implements its intelligence according to one aspect of the present invention.
[0071] Data storage:
[0072] This data store is a highly available data storage for data tuples consisting of internal sensors, actuators, actions, and their motivational values (see below). This storage appropriately combines the actions performed and the associated actual, empirical, and observed results. Therefore, it only stores internal data in tabular form, not, for example, the pixels of one or more images. Therefore, only data whose internal meaning is known exists.
[0073] According to one aspect of the present invention, another part of this module repeatedly checks the data for internal consistency. This means that it cannot contain any contradictions, structures that lead to cycles or loops, etc. The goal is that the data can ultimately lead to only one decision. The routine that checks this should also delete redundant stored data, not only to keep the database as lean as possible but also to prevent overfitting. What data is considered redundant depends on the type of theorization.
[0074] Figure 2 Theorization according to one aspect of the invention is shown:
[0075] Using the simple method described in this paper, a central template can reproduce any functional dependency, which can be saved and used for prediction.
[0076] Braking or accelerating depends functionally on the mass being accelerated and the force being applied. However, this function is initially unknown, since the mathematical function of the speed change dV depends on the preconditions Vn (current speed, weight, etc.) and the use of actuators Em (force on the accelerator pedal, steering, etc.), namely:
[0077] dV=f(Vn, Em) (1)
[0078] This function should be learned by the user, rather than programmed in, because this allows for the already learned content to be expanded, modified, and improved in the same way. The effective acceleration during braking (or acceleration) is related to the (mathematically) independent variables of force F (how hard the pedal is pressed) and mass m according to the law F = m*a. The principle by which the robot itself finds this equation is simple. To illustrate, we model the (different, arbitrary) quadratic function y = -((x-4)²) + 9.
[0079] Assume that initially there are only two measurements, at points 11 and 12. Now suppose a prediction p is needed at point x = 3 (point 2), using the straight line a between 11 and 12 to give the value y = 3.8. However, because the error in the subsequent measurement y = 8.0 is too large, point 2R at x = 3.0 and y = 8.0 is added to memory. Therefore, new predictions are made using the degree b1 between [11, 2R] or b2 between [2R, 12]. Over time, many more points (31, 32, 33, ...) are created in the above example. From these points, the functional relationship between x and y can be modeled with arbitrary precision using the corresponding straight lines along the data tuples (c1, c2, c3, ...). The precision is limited only by unacquired experience and the size of the memory. Of course, this can be extended to any number of variables, if computational and memory capacity permit; using a hyperplane a1x1+...+anxn=const instead of a straight line, predictions are made using piecewise linear interpolation or extrapolation.
[0080] However, in order to minimize the number of stored corner points or data tuples, it is also possible to delete the stored corner points or data tuples again: if a post-analysis shows that a certain data tuple is very close to the line / hyperplane between two adjacent tuples, or even on the same line / hyperplane, then the tuple can be deleted as a redundant tuple. This not only reduces the memory requirements, but also reduces the computational effort, and over time the behavior becomes more and more efficient as it adapts more and more to the actual (although still unknown) mathematical function.
[0081] This process also prevents overfitting, which must be taken into account when selecting and determining the scope of training data for CNNs. Therefore, redundant experience is not repeatedly stored, which would completely mask those rare events, the so-called "black swans", and could have serious consequences.
[0082] Storing functional dependencies on variables has another advantage: all possible inverse functions are also stored, since the stored data tuples do not distinguish which values are "dependent variables" and which are "independent variables" at the time of occurrence. In the function above, x is the independent variable, and the corresponding y value can be determined using the function y = -((x-4)²)+9. The x value must be searched in the database; if the x value is not explicitly stored, the y value can be approximately determined using the next two x values. However, you can also search for the y value directly in memory and then determine the x value using piecewise linear interpolation or extrapolation—this is much easier than, for example, trying to determine the inverse function of the above equation.
[0083] With this very simple approach, any functional environment can be modeled with arbitrary accuracy, as in complex movements, as long as the action is sufficiently trained. Of course, the acceptable error ε can also be changed and adjusted over time, either reducing it to increase the required accuracy or increasing it to reduce memory and computational workload, as this in turn affects the number of stored corner points and data points. It is a continuous optimization process.
[0084] According to one aspect of the present invention, a theory of rules and laws for a given situation is defined as a self-contained, consistent data set from which the robot can calculate the theory's predictions through piecewise linear interpolation or extrapolation. This data set includes its own measurements and may also include parameters of a created alter ego. In this way, the robot can independently determine, save, reuse, and improve each of its theories. As part of its self-assessment (see below), it can also select and apply the best specific theory (data set) for a situation, i.e., the most capable theory (see below).
[0085] action:
[0086] According to one aspect of the present invention, an action is an ordered sequence of actuator operations to achieve a goal. Actions can be designed using functional relationships between actuators, internal parameters, and known effects on the robot. Initially, during the early learning phase (see below), the AI performs simple actions. The AI learns to accelerate and brake, then accelerate, wait, hit an obstacle (external braking, presenting an "injury risk"), accelerate, turn, brake, and so on. These simple actions are combined to form increasingly complex "choreographies," which are then further optimized—for example, by reducing pedal pressure for softer braking and smoother cornering as speed decreases. Optimization is achieved using motivation values (see below), which represent an internal assessment of the actions being performed. From the stored (partial) "choreographies," the one with the highest motivation value is selected. In this way, the AI can continuously improve its use of actuators. Furthermore, it can predict its own situation at future points based on current reference values and existing action experience, thus laying the foundation for its own action planning.
[0087] Motivation value:
[0088] Motivational values reflect a reward. On the one hand, they are internally determined after achieving a self-set goal (success); on the other hand, motivational values can also be externally transmitted (a job well done). For example, if a route is driven with strong steering movements at low speed but high lateral acceleration, the motivational value of this series of movements will be lower than the result of "training"; later, the same route can be driven at a faster speed but with lower lateral acceleration. Of course, this must also be kept in mind.
[0089] Motivational values, therefore, represent unspecified, intrinsic goals. They can be associated with long-term goals, such as the journey's end, or short-term, intermediate goals, such as going to the gas station when the tank is empty. As the tank's fuel level decreases, the motivational value of "refueling" slowly increases until it surpasses the motivational value of the long-term goal. The pursuit of the long-term destination is interrupted, replaced by refueling, and then restarted.
[0090] Target System:
[0091] According to one aspect of the invention, the robot is always "on", although it can of course be turned off. Therefore, it is always performing an activity. However, this also includes doing nothing (charging) or optimizing its own long-term memory of action options in a resting state (sleeping). The selected action is always the action with the highest current motivation value. If an action that is not currently being performed reaches a higher motivation value than the action currently being performed, the action will be canceled, terminated or interrupted in order to perform the action with the highest motivation value. If a car traveling from Munich to Hamburg to refuel or charge has a higher motivation value than arriving in Hamburg, it will interrupt the journey at the next opportunity and go to the gas station.
[0092] In the infancy of such a robot or car (see below), during its learning phase, short-distance driving, braking, accelerating, and turning are assigned high motivational values simply because internal long-term memory signals that it should learn or reduce the error in its actions, thereby making its movements more precise. Later, in the adult phase (see below), when such actions cause no or only minimal changes to the data memory, the error reduction approaches zero, and the behavior becomes "boring," receiving only lower motivational values. Then, for example, energy conservation becomes a priority.
[0093] Thus, according to one aspect of the invention, the target system avoids being obsessed with reality, meaning that the interpretation of reality may prove to be wrong in the future.
[0094] Central template, body:
[0095] According to one aspect of the present invention, a central template represents the core and foundation of this artificial intelligence. This central template is called an ontology because it places itself at the center of data processing and primarily views and evaluates everything from its own perspective. Therefore, the ontology does not need to be given or trained on any data that has meaning (e.g., image recognition); it creates everything.
[0096] Subjective polar coordinate system:
[0097] According to one aspect of the present invention, all observational and empirical data of the artificial intelligence are subjective, or derived from the perspective of the subject. This concept of space also follows. There is no Cartesian coordinate system with an assumed origin at a certain location in the observation environment, in which each identified object is assigned corresponding coordinates. Instead, a polar coordinate system can be used. This automatically determines not only the mathematical origin in the subject, but also the subjective orientation of the subject, including motion vectors, in terms of front-back, up-down, left-right, and so on. The surrounding (empty) space is taken for granted. The observed object is identified only by its distance and angle from its own viewpoint, i.e., the origin of the polar coordinate system. Therefore, physically empty space is not considered an independent entity or dimension. It is just there, and there is only an option to move there, unless another object stands or moves (there), in which case the "position" is occupied, or the space is not empty. Another advantage of subjective polar coordinates is reflected in the concept of "mirror body" (see below).
[0098] Childhood:
[0099] Under the premise of unconditional requirements, such as observing the integrity of the object and the robot itself, maintaining operational readiness (battery charge state), and goals such as smooth acceleration and braking and optimized cornering, the robot should / can autonomously learn the use of its actuators and their consequences. Most importantly, it coordinates its actuators with its internal and external sensors and fine-tunes them through incremental training 100-105, interrupted by occasional rest periods, such as to recharge the battery and maintain the internal database (redundancy, consistency, etc.), until the errors detected in plans and outcomes are sufficiently small. In the case of a car, this primarily involves accelerating, steering, braking, reaching the destination (success) or missing the destination (error), etc. Initially, the agent itself sets simple goals for simple actions, executes these goals, and remembers the results, then combines these results into increasingly complex "choreographies." This training phase ends when the actions for the self-set or externally set goal ("drive there") cannot be further optimized—that is, when the mathematical module fails to reduce the error or only slightly reduces it and / or when the database (the number and location of points in the memory) is difficult to optimize. If the error decreases more and more slowly, the infancy phase is drawing to a close. KNN, on the other hand, requires a large amount of carefully selected training data, while Ontology only requires a real environment, a playroom where it can experiment and train its abilities without causing any damage. It follows an intrinsic goal, rather than an external, pre-defined one like KNN, and therefore does not interpret the environment or the world in any way.
[0100] Adulthood:
[0101] In its infancy, the device has already learned the ordered sequence of actuator applications to achieve a given goal. What is the value of this knowledge and experience acquired in its infancy? All the data that the agent learns about the world in which it moves is its own data, its own observations, and its own experiences. From a higher-level scientific or philosophical perspective, it must be stated that this data is true from the perspective of the agent itself, both as it is now and as it was in the past. Therefore, it represents a natural and fixed starting point for cognition. Of course, the agent must maintain its internal data. In addition to the optimization mentioned above, it must also ensure that it is free of contradictions and errors. Otherwise, the robot manufacturer or operator cannot assume that the robot will achieve the intended goal through its designed movements.
[0102] Self-experienced and self-induced environmental changes, such as position changes caused by acceleration, steering, and braking, are experienced and stored by connecting internal and external data. At the level of meaning, there's no need to delve into its own actuators to understand why, for example, pressing the accelerator pedal causes acceleration, steering causes a turn, or braking causes negative acceleration. The car's self-entity simply needs the thought, "It presents three options," to change its position in a targeted manner. At this level, what matters is "the fact," not "why." A street pigeon certainly doesn't know how cars or cyclists work. However, if such a "road user" passes by at a sufficient distance, it will remain where it is; but if the direction of motion changes towards it, it will fly away. It doesn't know the cause, but it is aware of the potential for action and therefore pays close attention.
[0103] Mirror body:
[0104] When an (adult) entity recognizes an object in its observable environment through external sensors, it simply creates a copy of itself and adjusts the copy's parameters based on the observed properties. Thus, the entity serves as a central template for all observable and unobservable objects and phenomena. Hence the name entity and the resulting mirror image for the copy.
[0105] By using the ontology as a central template, there are no longer any unknown objects. There are only unknown parameters, but these can be measured by external sensors, estimated based on personal experience (keyword: bias), and their possible range can be constrained by further observation.
[0106] Just as the original body uses its methods, stored data, and parameters to calculate its behavior and predict its own situation, the mirror body's behavior can be calculated using the same methods, data, and adjusted parameters. The assumptions about the mirror body are limited to the same laws of physics, and short-term goals can be inferred from the observed motion orientation and location. This doesn't matter whether the observed object is a box on the street, a kangaroo in Australia, or an elderly woman with a walker.
[0107] As far as the program is concerned, the ontology should exist as an object. Then, for each new object that appears, a copy of the ontology can be simply created, and the parameters of the new object copy can be adjusted with the help of external sensors and the user's own experience. There can be a list like the following, which contains the mirrored bodies that have been inserted or created with new copies in the current situation:
[0108] class CEgo{...};
[0109] CEgo Ego(...);
[0110] / / Education and learning begin in childhood ...
[0112] / / Beginning of adulthood
[0113] CEgo AlterEgos[]; ...
[0115] / / Identify new objects
[0116] Alter Egos[i] = Ego.clone(); / / Clone the object's body
[0117] AlterEgos[i].adjust(...); / / Parameter adjustment of new object ...
[0119] Thus, the body creates a mirror body internally as a computable image of the observed object in the environment, and as Figure 6B As shown, it can form an inner loop 311 of object recognition, update parameters of observable objects 312, predict the behavior of objects 313, adjust the actuator operation sequence 314, and enter a second outer loop after reaching the target, enter the rest and (also) loading phase 315, until continuing to initialize for the new target 310.
[0120] It doesn't matter whether the object being observed is familiar or completely new. If it is new, the range of possible feature values is larger, which requires higher attention (scanning frequency). This has the following advantages:
[0121] - Understand basically every observation object!
[0122] -Predict the behavior of all external objects using your own methods and data!
[0123] -Human intervention is no longer required to program in the missing pieces.
[0124] -The ontology continuously learns and improves its behavior.
[0125] - The entity is capable of empathy and situational awareness (see below).
[0126] - The original body can be learned by observing the mirror body (see below).
[0127] -Mirror bodies can embody unknown causal relationships and rules (see below).
[0128] Figure 3 Scenario 1 "Children" according to an aspect of the present invention is shown.
[0129] When an agent, for example an artificial intelligence of a simple car, encounters another road user, in this case a child, it creates a copy of itself and adjusts the parameters (distance d from the agent, orientation, speed vK, ...):
[0130] The mirrored entity can then use the original entity's methods to predict when and where it will be. By comparing this prediction with its own calculations of when and where it will be, the original entity can determine whether a potentially dangerous situation is likely to occur and respond accordingly, such as by slowing down. In this way, it can use its own methods and experience to predict a child's short-term behavior. This eliminates the need for any special programming or training for children, which was not present when the AI first encountered children.
[0131] Figure 4 Scenario 2 "Highway" according to an aspect of the present invention is shown.
[0132] Another example that more clearly illustrates the power of this simple concept of mirror bodies is the following scenario: on a highway, there are two cars behind a truck and a car in the fast lane, where the second car is our agent. Based on its understanding of the braking force law (F = m*a), combined with the observed mass parameters of the car and truck, that is, the assumed mass, the measured speed, and the assumed braking effect, the agent can first use its own methods to predict the braking distance of all road users in a possible accident scenario, thereby determining its own safe distance. But let's first deal with the "ideas" or assumptions that the agent makes about other road users in order to predict their behavior:
[0133] -HGV0 is traveling at a certain speed, and the entity (PKWEgo) already knows (without specific training) that such large road users rarely exceed it. Therefore, the truck's mirror image does not anticipate any changes in speed.
[0134] - The PKWEgo, acting as an observation center, will be able to travel at a faster speed and will overtake the HGV0 while taking other road users into consideration. Overtaking allows the PKWEgo to reach its destination faster, so it has a higher motivation value when preparing to overtake.
[0135] PKW1, the other mirror-body, thinks the same way as PWKEgo, because the entity will act in its place. Therefore, the entity assumes that PKW1 also wants to overtake LKW0. Naturally, the mirror-body of Car 1 "sees" PKWEgo and executes its (the entity's) pre-defined overtaking maneuvers, taking into account PKWEgo's presence, for example, by setting its indicator lights and paying special attention to PKWEgo and other road users in the overtaking lane. This is what PKWEgo's entity "thinks," and how it predicts Car 1's behavior.
[0136] Another possible road user, car 2, is approaching at high speed in the fast lane. The main body also creates a mirror body for this and formulates its action prediction based on the clear lane.
[0137] - As shown above, the program, i.e. the ontology, can anticipate the overall context of all relevant road users and plan its own actuator operation sequence accordingly to avoid dangerous progression.
[0138] Mirror body, consciousness:
[0139] Thus, the entity has a complete internal picture of its observable environment with all recognized objects and is able to predict their short-term behavior. This would be a possible and programmable definition of consciousness, a true definition of consciousness, not a fake or simulated one.
[0140] Short-term behavioral prediction based on the ability to put oneself in the position of the observed object means using all data converted into a mirror body to observe the situation and using the calculated possible behavior of the observed object to calculate the progress of the situation.
[0141] The ontology in the highway scenario above can be applied to all observed objects without writing additional programs:
[0142] - Predict the behavior of all road users by assuming they want to travel as quickly and without accidents;
[0143] - continuously adjusts predictions based on actual observed behavior (braking, accelerating, steering, setting indicators, turning on headlights, etc.); and
[0144] - Overtake the truck at the right moment.
[0145] The ontology can easily evaluate the possible behaviors of all the people involved in this situation by applying its own methods and experience (data) to the possible behaviors of all the people involved in this situation and combining them with appropriate parameters. Isn’t this what humans do too?
[0146] These examples clearly demonstrate that the foundation, the basis, and the starting point of this type of artificial intelligence is always the ontology—the central template. The more differentiated its internal and external perception capabilities, and the better it is trained at an early age, the more intelligent its behavior and the greater its potential to predict the behavior of the objects it observes. If parameters are also set for the ontology's own vulnerability—in the case of a car, one might call this deformability—then the ontology can also assess the risk of strong or weak contact. However, the ontology only needs to be built, programmed, and trained (once), and of course, the training can also be transferred from the trained ontology. The objects do not need to be identified (KNN and its training data), nor does the behavior of identified objects need to be programmed. Of course, some might argue that the ontology alone is insufficient to predict the behavior of other road users. But just think about how humans try to assess the behavior of other road users. Humans also don't know the long-term goals of road users on the highway; they can only guess at their short-term goals (getting ahead as quickly as possible without getting into an accident) and prepare accordingly.
[0147] Robot's capabilities:
[0148] Empathy:
[0149] The above shows how the entity uses the mirror body to put itself in the shoes of the objects in the environment so that it can calculate their behavior from their perspective. If you don’t use the word “empathy” specifically for humans, you can call it “sympathy”.
[0150] Learn from observation:
[0151] Learning from observation means observing behaviors that might be better than your own and adopting them when possible. The basis for this is comparing the behavior of others with your own, a comparison made possible by the empathy defined above.
[0152] The subject's empathy is realized here through the mirror body. Through the mirror body, the subject places itself in the situation of other observers in order to predict their behavior based on their perception of the environment. We assume that there is a discrepancy between the predicted behavior and the observed behavior. If the mirror body now changes the sequence of actuator actions to reproduce the observed behavior, then if the observed behavior is advantageous, the subject has the opportunity to replicate it.
[0153] Figure 5 Different "cornering behaviors" in scenario 3 are shown.
[0154] We imagine that the agent has learned to always turn at the same distance from the right edge of the road, i.e. line LE. Now, it observes a car ahead and cuts the curve by turning earlier but more gradually, which also reduces the lateral acceleration and therefore goes around the curve faster on the dotted line LB.
[0155] From this perspective, the empathy described here is a prerequisite for learning by observing and imitating the behavior of others. (There is no reason not to assume that the same applies to humans, where learning by imitation is based on empathy.)
[0156] If the data from external sensors is translated into the mirror body's situation and position and associated with the original body's stored, simple and more complex action sequences that have been copied to the mirror body, the original body can learn from observation: it can recognize the difference between the self-planned action sequence of its own actuators (always maintaining the same distance from the edge of the road) and the observed sequence (earlier braking and turning, smaller steering movements, higher cornering speed and earlier acceleration) and determine that it can negotiate the curve faster this way. This kind of learning is certainly faster than the usual trial-and-error method.
[0157] Identify cause and effect patterns:
[0158] The theoretical part explains how to replicate any observed functional relationship. Let's apply this ability to identify natural laws, and where necessary and appropriate, use mirror bodies for these laws. Of course, they won't be located on the road or driving on the road, but they will cause changes that can be detected by external and / or internal sensors.
[0159] Let's imagine an experiment where our AI entity observes a falling apple. The entity creates a mirrored body and notices the accelerated motion towards the ground. Now, the entity first assumes that the apple itself causes the acceleration; just like the car entity, it also has its own "gas pedal" to accelerate itself. But then, the entity observes that there is no apple on the ground moving by itself, and other objects also fall to the ground and then stop moving. In addition, the entity itself should have learned the law of universal gravitation. On a slope, even without pressing the gas pedal, it will notice and learn about the acceleration, but only when going "downhill", and it must accelerate more when going "uphill". According to the formula we already know (g = acceleration due to gravity, 9.81 m / s2), the larger the angle α, the greater the force:
[0160] F=m*g*sin(α) (2)
[0161] Remember, the entity certainly does not explicitly know this equation or the free fall equation, namely the law of universal gravitation (sin(α=90°)=1), but it implicitly understands this relationship between the relevant variables through data points and piecewise linear interpolation or extrapolation.
[0162] Thus, the entity behaves as if it knows, as physicists say, that as a subject with heavy mass, it is subject to the aforementioned laws of mass gravity. If it now sees another object and creates its mirror image, the mirror image also experiences this mass gravity. This knowledge thus becomes a natural law that affects all entities without being programmed into them. If we observe the behavior of this entity from the outside, we cannot tell whether it truly understands the laws of universal gravitation as physicists do, or is merely pretending to do so. Even if the entity sees a pigeon perched on the street and it flies away as soon as the entity approaches, it recognizes that the flapping of its wings and the upward acceleration are consistent.
[0163] Another example is a strong crosswind. The entity is a large, empty truck. During normal straight-line driving, the lateral acceleration is zero. But then the truck suddenly shifts sideways, triggering the lateral acceleration sensor. In this case, the entity can simply create a new mirror volume for the unexpected force and unknown cause and associate it with the lateral acceleration. If the entity can also record environmental data, such as the location in a forest or plain, or the movement of tree branches, and add this to its data memory, then the entity not only identifies the previously unknown variable (crosswind) but also the possible causal relationship. If the entity can also see and judge the extent of the branch movement, it can even estimate the strength of the force acting on the side and, therefore, the possible impact on driving behavior. In this way, the entity itself creates a theory that incorporates new variables. It's similar to the physicists who proposed dark matter and dark energy, despite their invisibility and unobservability, or Nobel Prize winner Peter Higgs, who conceived of a particle in the 1960s and first detected it at CERN in 2012. The mirror image, the replica of the self, can of course never in this way identify the essence of phenomena with the aim of seeking truth, but it is sufficient for a full grasp of the situation.
[0164] self assessment:
[0165] The first form of knowledge acquisition has already been described. The example of a crosswind was explained in the section "Understanding Causal Laws." A force (initially) unknown to the agent suddenly produces a lateral acceleration in a straight line, which has a measurable effect on the robot or agent. It creates a mirror body for this purpose. However, initially, no additional information signals to the agent when this phenomenon occurs or how strong it is. Only when the agent observes a temporal correlation between the branch's movement, which sometimes increases and sometimes decreases, and its own lateral acceleration, which sometimes increases and sometimes decreases, can it connect the two and assume that they have the same cause. What happens here? The agent cannot identify the crosswind itself, but it can identify its own reaction to it (the lateral acceleration) and the reaction of the branch. This creates a factor that sometimes affects the agent and sometimes does not, but it affects not only the agent but also other observable objects.
[0166] The second form of knowledge acquisition is based on comparing one's own behavior with that of other subjects. This assumes that the observed subject's behavior is "rational" rather than random, although this can be considered rational, making comparison unnecessary for knowledge. By comparing two different behaviors, the subject can determine whether its own capabilities (see below) are inferior, equal, or superior to those of the observed subject in a given situation. This means that the observed subject can perceive more, the same, or fewer situations. For example, the observed truck suddenly slows down on a straight road. It appears to see something unknown to the observing subject, recognizing less than the subject. If inferior capabilities have been identified, this can be seen as an opportunity to specifically analyze these situations in order to at least raise one's own capabilities to the level of the other subject. The comparison is not about whether the behavior is right or wrong (which only leads to well-known problems with theories of truth), but only about whether and when the behavior changes and whether it repeats. The first behavioral comparison involves different situations: does the observed subject behave differently in situations that are different for the subject? The second behavioral comparison involves the same situation: does the observed subject behave repeatedly or differently in situations that are the same for the subject? The subject can draw conclusions from observing these multiple situations over time:
[0167] - If an object exhibits different behavior when repeatedly observed in situations that are identical to the subject, and also changes its behavior in situations that are different to the subject, the subject is less competent;
[0168] - If objects repeatedly observed in situations that are identical to the entity repeat their behavior but do not change their behavior in situations that are different from the entity, the entity is more competent;
[0169] -If objects that are repeatedly observed in situations that are identical to the entity repeat their behavior but change their behavior in situations that are different to the entity, then the abilities of the entities are equal.
[0170] By making this simple comparison, the entity can realize how powerful its own capabilities are in certain situations compared to other objects. Note that this is not an attempt to compare with "real-world truth."
[0171] Through these comparisons, the ontology can evaluate itself and infer, for example, whether it should act more cautiously in certain situations and try to identify what is (not yet) identifiable. This can be understood as an inherent instruction to conduct targeted research as part of "continuous learning and the continuous expansion of cognition and knowledge" (see above).
[0172] Expertise:
[0173] However, the results of these comparisons can also be used to externally evaluate the capabilities of a robot to determine its possible applications. There are no linguistic issues in comparing and determining the strength of capabilities compared to the truthfulness of the knowledge about the real world acquired by the ontology.
[0174] However, the concept of truth associated with a theory has another important aspect: capability. A theory enables predictions. Therefore, understanding the quality of a theory and its subsequent predictions is crucial. Generally speaking, the attribution of truth to a theory is a criterion for being able to apply that truth, allowing it to be applied. The advantage is that this criterion automatically excludes all competing theories, eliminating the problem of deciding which theory to use. The concept of capability must be of a comparable level.
[0175] Expertise is defined and used here as the ability to understand the relevant factors influencing a situation or similar situations. Therefore, it is initially a limited criterion, but its level can be determined through comparative procedures, unlike the validity of a theory, which assumes infinite validity across time and space. However, this cannot be proven, and the criterion does not allow for the comparison of different theories.
[0176] This means that capability is a better concept for describing robot performance than truth, and unlike truth, it can be determined by the robot itself in a formal self-assessment process.
[0177] Figure 6AA schematic diagram of a procedure for autonomously controlling a device during a juvenile or training phase is shown. The goal and endpoint of this phase is to be able to use and deploy the actuator with sufficient accuracy and minimize errors. Initialization 200 is followed by random actuation 201 of the actuator and reading 202 of a plurality of internal sensor units configured to measure internal device properties, as well as reading 203 of a plurality of external sensor units configured to measure external environmental properties. Correlations between the actuation 201 of the actuator and the read device and environmental properties are stored 204 as action instructions, such that the action instructions are output triples of actuation 201, device properties, and environmental properties, which become target triples. The device is actuated 205 using at least one stored 204 action instruction, based on at least one predefined target device property and / or at least one predefined target environmental property, the action instruction specifying how to set the at least one predefined target device property and / or at least one predefined target environmental property based on the output triple and the target triple.
[0178] Figure 6A Showing the infancy stage:
[0179] 200 Initialization, preparation for learning
[0180] 201 pairs of actuators are randomly driven
[0181] 202 Reading a large number of internal sensor units 202
[0182] 203 Internal device characteristics and readings from multiple external sensor units 203
[0183] 204 Save 204 Mutual Relationship
[0184] 205 immediately determines the motion accuracy error. If the error is too large, continue to 201
[0185] Optional: Rest phase for recharging and post-data preparation:
[0186] Error, redundancy, consistency...
[0187] Figure 6B Showing the adult stage:
[0188] 310 Initialization, preparing for long-term goals
[0189] 311 Identify objects in the current context 311
[0190] 312 Create or update the image or its parameters 312
[0191] 313 Using the current mirror body to predict the behavior of the observed object 313
[0192] 314Calculate the optimal actuator deployment 314 to further your own goals
[0193] 315 Rest phase 315 is used for charging and post-data preparation:
[0194] Error, redundancy, consistency...
[0195] In this document, robot, device, system, and ontology are used synonymously.
Claims
1. A method for autonomously controlling a device, the method comprising: - Initialization (100): initialization (100) is performed by randomly driving (101) the actuator and reading (102) a plurality of internal sensor units configured to measure internal device characteristics, and reading (103) a plurality of external sensor units configured to measure external environment characteristics; - storing (104): storing (104) the correlation between the driving (101) of the actuator and the read device properties and environmental properties as an action instruction, so that the action instruction converts the output triple of the driving (101), the device properties and the environmental properties into a target triple; as well as - controlling (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored (104) action instruction, the at least one stored (104) action instruction indicating how to set the at least one predefined target device property and / or the at least one predefined target environment property using the output triplet and the target triplet.
2. The method according to claim 1, characterized in that The method is iterated in a learning manner so that new action instructions are always identified which in each case transform an output triplet into a target triplet and which are stored for driving (105) the device.
3. The method according to claim 1 or 2, characterized in that The reading of the internal actuator ( 102 ) and / or the reading of the external actuator ( 103 ) is carried out in such a way that the individual support values are stored and intermediate values are interpolated and / or extrapolated.
4. The method according to claim 1, wherein The target device property and / or the target environment property is at least temporarily influenced by a device property and / or an environment property.
5. The method according to one of the preceding claims, characterized in that Combine multiple action instructions to form an orchestration that transforms an output triplet into a target triplet.
6. Method according to one of the preceding claims, characterized in that The object is detected by an external sensor unit, and the parameters detected by the external sensor unit are compared with device properties and / or action instructions of the device.
7. The method according to claim 6, characterized in that The detected parameters are used to predict the behavior of the detected object.
8. The method according to claim 6, wherein: The action instruction of the device is generated by utilizing the relationship between the parameters of the detected object and the action of the detected object.
9. The method according to one of the preceding claims, characterized in that Actuators are: processors, memories, motors, gripper arms, motion units, drives, end effectors, wipers, flashers, lights, steering gears and / or actuators.
10. The method according to one of the preceding claims, characterized in that The internal sensor unit detects: actuator status, actuator settings, actuator position, size, weight, specifications, speed, acceleration, deceleration, battery level, range, position, processor utilization, memory utilization and / or system parameters of the device.
11. The method according to one of the preceding claims, characterized in that The external sensor unit detects: distance, objects, lidar signals, image signals, position, size, dimensions, acoustic signals and / or external parameters.
12. An apparatus suitable for carrying out the method according to any one of the preceding claims, the apparatus comprising: an initialization unit configured to: perform initialization (100) by randomly driving (101) the actuator and reading (102) a plurality of internal sensor units configured to measure internal device properties, and reading (103) a plurality of external sensor units configured to measure external environment properties; a memory unit configured to store the relationship between the driving (101) of the actuator and the read device properties and environmental properties as an action instruction, so that the action instruction converts an output triple of the driving (101), the device properties and the environmental properties into a target triple; as well as - a control unit configured to control (105) the device according to at least one predefined target device property and / or at least one predefined target environment property using at least one stored (104) action instruction, the at least one stored (104) action instruction indicating how to set the at least one predefined target device property and / or the at least one predefined target environment property based on the output triplet and the target triplet.
13. A system arrangement comprising a plurality of devices according to claim 12, said devices being connected via communication technology and exchanging action instructions.
14. A computer program product comprising instructions which, when executed by at least one computer, cause the computer to perform the steps of the method according to any one of claims 1 to 11.
15. A computer-readable storage medium comprising instructions, which, when executed by at least one computer, cause the computer to perform the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and model for performing independent path exploration based on operant conditioning
CN105094124A
Machine learning methods and apparatus related to predicting motion(s) of object(s) in a robot's environment based on image(s) capturing the object(s) and based on parameter(s) for future robot movement in the environment
CN109153123A
Home service robot operation state autonomous cognition method and system
CN109407518A
Motion prediction based on appearance
CN113453970A
Robot control method and device, robot and storage medium
CN114454176A