Autonomous control of the device
By using randomly driven actuators and sensors to measure, record, and store action command triplets, the robot system can learn and control autonomously in unknown environments, solving the problem of dependence on training data in existing technologies and achieving efficient autonomous control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 卡尔·阿尔伯特·施赖伯
- Filing Date
- 2023-06-19
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, robots and artificial intelligence systems require a large amount of training data to be effectively controlled, and they have difficulty learning and correcting errors autonomously in unknown environments, resulting in high costs and a high risk of errors.
By using a random-driven actuator, internal and external sensors measure device attributes and environmental attributes, record and store action commands, form action command triplets, iteratively learn and recognize new action commands, and achieve autonomous control.
It enables autonomous learning and control of devices without the need for initial knowledge, and can identify and correct actions in complex environments, thereby improving the system's autonomy and adaptability.
Smart Images

Figure CN120603683B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for autonomously controlling a device, which typically moves in the real physical world and requires no further specification of its physical configuration. The advantage of this method is that the device autonomously learns and continuously improves its learned knowledge or behavior. Overall, it overcomes the disadvantage of prior art methods that require the creation of training data first, which is often the case with traditional artificial intelligence methods. Generally, the method has universal applicability, and the device automatically learns and continuously modifies its own knowledge base. Furthermore, an apparatus for implementing the method and a system arrangement comprising multiple proposed apparatuses are proposed. Additionally, a computer program product and a computer-readable storage medium are proposed for performing the steps of the method or causing a computer to perform the method. Background Technology
[0002] The paper "Dynamic motion learning for multi-DOF flexible-joint robots using active-passive motor babbling through deep learning" by TAKAHASHI KUNIYUKI et al., *Advanced Robotics*, Vol. 31, No. 18, September 17, 2017, pp. 1002-1015, XP055840673, illustrates a learning strategy for performing dynamic motion tasks on multi-DOF flexible-joint robots. Although robots with flexible joints offer several potential advantages, such as utilizing intrinsic dynamics and passively adapting to environmental changes through mechanical compliance, controlling such robots remains challenging due to their increasing dynamic complexity.
[0003] Various artificial intelligence (AI) methods are known in the prior art. These methods envision selecting and providing training data, allowing the algorithm to identify patterns during the training phase. Here, a specific algorithm can identify patterns through the norms of initial data and target values, and then be put into use. This means that implicit knowledge in large amounts of data can be utilized. After the training phase is complete, the appropriately trained algorithm is applied to real-world data, thereby generating implicit dependencies to solve real-world problems. One drawback of this approach is that the training data must be selected and specified first, which limits the resulting AI to the training data. This can negatively impact the learning phase at this stage, making it prone to errors and costly.
[0004] Swarm intelligence is also known in the prior art, which provides multiple devices that collaborate to solve problems in a distributed manner. In this case, the control and coordination of the individual participants is also complex and sometimes prone to errors. This is especially true if distributed learning requires further coordination.
[0005] Artificial neural networks, as known from existing technologies, are based on graph theory and provide neurons and connections, i.e., nodes and edges. These networks are well-known for mimicking the function of the human brain and can also learn. Edge weights can be changed, new edges can be added or old edges can be deleted, and existing nodes can be deactivated or new nodes can be added. This constitutes the overall system of dynamic learning.
[0006] However, there are at least four problems with using KNN in the above manner:
[0007] 1. For objects that have not yet been trained with KNN, KNN will not provide any information to invoke the routines assigned to that object.
[0008] 2. The programming attributes or behaviors of the detected and identified objects may be incorrect, have changed during this period, or remain unknown.
[0009] 3. One or more scientists, experts, etc., select the training data and specify the desired outcome. Even in unsupervised learning, the training data and hyperparameters are still specified manually, which cannot eliminate the possibility of errors.
[0010] 4. ANNs cannot correct errors. Each individual piece of information is distributed across all the parameters of an ANN, much like a hologram where each data point contains information about the entire image. Therefore, ANNs must always be completely removed and completely retrained.
[0011] This means that when the ego encounters a previously unknown object, it will not fail to recognize it (unlike KNN), such as in the video where an autonomous car suddenly sees two boxes in the lane. KNN usually stands for Artificial Neural Network.
[0012] In existing technologies, there is a need to achieve autonomous control of the device in order to avoid or minimize time-consuming preparation work, such as providing training data and instruction. Therefore, there is a need for self-learning systems capable of collective learning or estimating what other participants might or are capable of doing. Consequently, there is a need for systems involving multiple participants, where each participant actually understands and further elaborates on the others; this is referred to in this paper as empathy. Summary of the Invention
[0013] Therefore, the object of this invention is to provide a method for autonomous control of devices, which operates independently based solely on provided target information and accumulates motion knowledge in the process. The proposed method should be able to identify the range of motion of the actuators of any device and create an improved method or at least an alternative method for autonomous control of one or more devices. Furthermore, the object of this invention is to provide devices for executing the method and a system arrangement comprising multiple proposed devices. Additionally, the object of this invention is to provide a computer program product and a computer-readable storage medium containing instructions for executing the method.
[0014] This problem is solved by the features of the present invention. Other preferred embodiments are also provided in the present invention.
[0015] Accordingly, a method for autonomous control of a device is proposed, the method comprising: initializing by randomly driving an actuator and reading multiple internal sensor units configured to measure internal device attributes and multiple external sensor units configured to measure external environmental attributes; storing the correlation between the actuator drive and the read device and environmental attributes as action instructions, such that the action instructions convert an output triplet of the drive, device attributes, and environmental attributes into a target triplet; and driving the device according to at least one predefined target device attribute and / or at least one predefined target environmental attribute using at least one stored action instruction, the at least one stored action instruction indicating how to set at least one predefined target device attribute and / or at least one predefined target environmental attribute based on the output triplet and the target triplet, wherein the method iterates in a learning manner, continuously identifying new action instructions during the iteration, each action instruction converting an initial triplet into a target triplet, and these action instructions are stored for use in controlling the device.
[0016] The proposed method is used for autonomous control of devices (typically physical), i.e., creating action instructions. These could be cars, robots, or manufacturing plants. No limitations are envisioned regarding mobility, meaning the device can move on land, in the air, on water, or even underwater. "Autonomy" in this paper means that the method ensures the device automatically learns and identifies available possible actions. These actions can be learned, and sensors can be used to learn which actions, from which origin, lead to which result. Typically, the device can set its own goals or transmit goals externally. For example, the device's own goal could be to maintain operational capability. Internal goals could be, for example, going to a charging station when the battery is low. External goals can be transmitted by specifying the tasks or activities the device should perform.
[0017] In the preparation step, initialization is performed by randomly driving actuators. Actuators are typically physical units that have an impact on the real world. For example, actuators can affect the device itself, such as steering, acceleration, braking, etc., and / or perform targeted manipulation of an object directly, such as with a robotic arm, or using tools, or via a remotely controlled instrument, such as a projectile or sprayer, or a fire extinguishing agent, such as water or fire-extinguishing foam. In the example of an automobile, the actuator could be a brake, accelerator pedal, etc., or it could be a headlight, indicator light, etc. The actuators are driven, and the effects of these actions are measured using internal and external sensors. These measurements are recorded and can then be used in the next step of the process. In this way, actions are learned and used in a targeted manner in subsequent process steps.
[0018] Because the proposed device or method can be implemented or controlled without any initial knowledge, the preparation step is initiated randomly, i.e., arbitrarily, as it must first learn which action leads to which result. In subsequent iterations of this process, the learned actions are initiated in a targeted manner. During random initiation, the system parameters of the actuator are examined, for example, moving the robotic arm to all possible positions. These situations are recorded each time, and it is identified which action of the actuator had what impact on the real world. In the example of a spotlight, it might only provide parameters for on and off. It is then turned on and off, and the effect of the headlight on the environment is measured using external sensors, and the operating parameters of the headlight, such as temperature changes, are measured using internal sensors. For advanced LED headlights, both color and intensity can be adjusted. This is tested through random driving and with both internal and external results, and then recorded.
[0019] According to the present invention, internal and external sensor units are proposed. "Internal" and "external" do not refer to the arrangement of the sensor units, but rather, according to the present invention, the sensor units measure external conditions relative to the device or similarly measure internal conditions relative to the device. Internal conditions refer to all system parameters of the device itself. External conditions refer to environmental variables, i.e., variables of objects surrounding the device. This means that internal sensor units are used to measure device properties, while external sensor units are used to measure environmental properties. For example, if the device has a certain battery state, this is an internal parameter. If an object is transferred from A to B by a robotic arm, this is an external state, and the state of the robotic arm is an internal state. Therefore, internal and external parameters typically interact; for example, actions always lead to a decrease in battery state, and these actions, in turn, affect external properties. From another perspective, this is an interaction where an actuator can take the device to a charging station and initiate a charging process, which in turn affects the internal battery state.
[0020] The recorded interactions are then saved, along with how the actuator's drive affects internal and external parameters. Therefore, device attributes are stored along with environment attributes, defining the action instructions. Thus, driving an actuator is an operation that affects both the device itself and the environment. For example, if the device moves from a first geographic point to a second geographic point, this changes the device's environment, which is an external parameter, and also alters the device's internal state, for example, through temperature changes or battery state changes. This means that triples can be stored indicating what the internal state is, what the external state is, and what action is subsequently performed. These are then transformed into new internal and external states. This means that the actuator's drive, device attributes, and environment attributes form an output triple, which is then transformed into a target triple. The new internal state, the new external state, and the possible actions can be specified in the target triple.
[0021] The concept of triples is used only to illustrate general transitions in system states. Generally, internal and external states can also be transformed into new internal and external states via functions. Therefore, actions consist of mapping functions from the first tuple to the second tuple. Those skilled in the art will recognize that these two descriptions are equivalent. Either an output tuple is transformed into a target tuple via a function, or an input triple is formed specifying the internal state, external state, and action, which leads to a target triple containing the new internal state, the new external state, and an additional field. This additional field can remain empty, but preferably at least one action currently executable in that state is performed. The last field can also be filled, for example, by introducing a vector containing identifiers for additional actions.
[0022] For example, the device can drive an actuator, called a motor driver, to move forward. Currently recorded is an output triplet containing the internal state (battery level), the external state (image signals of the physical environment), and the action command ("move"). This output vector or triplet is now converted into a target triplet containing the new battery level, the new image signals from the environment, and the action that can now be performed, such as moving forward or reversing. During forward movement, actions such as braking or activating the headlights can also be specified.
[0023] In an alternative example, the "move" function is used to transform the output triplet of image signals of the battery state and environment into a target triplet that describes the new battery state and the new image signals of the environment. This records what happens internally and externally during a specific action. This knowledge can be reused in other action steps to achieve the predetermined goal.
[0024] In additional process steps, target device attributes or predefined target environment attributes can be provided. Therefore, the device can be controlled by setting at least one predefined target device attribute and / or at least one predefined target environment attribute. Action instructions have been learned from previous method steps, and thus, it is known how to transform internal and external states into other internal and external states. For this purpose, target specifications, i.e., target device attributes and / or target environment attributes, are used. These target attributes thus determine the goal to be achieved, which is realized by executing action instructions. Therefore, the device state is read, and then a stored action is selected based on the actual state to achieve the target state. In a preferred case, the action instructions that immediately achieve the target specification are executed. However, in general, the initial state is sequentially transformed into the target state through multiple action attributes. In other words, multiple action instructions are linked in a way that ultimately achieves the goal.
[0025] For example, the program can identify that the device can move from A to B, to C, and to D. This means the stored instructions are that the car can travel from A to B, from B to C, and from C to D. It also implicitly remembers that the car can travel from A to C and from A to D. It also stores that the car can travel from C to D. Now, using these options, if the destination specification stipulates that the car should travel from A to C, the device has two options: travel from A to B and then to C, or travel from A to C. Generally, it can also travel from A to D and then from D to C. The choice can also be specified in the destination specification. For example, the maximum duration of the trip can be limited based on time. Furthermore, the maximum energy consumption can be limited. Therefore, the actions are randomized in the preparation process steps and then available during runtime. Overall, the program can be set to randomly drive the actuator, but it can also reuse previously learned instructions. This means that a randomly driven learning phase can be executed, as well as an execution phase using saved actions. This can also be alternated in any order. Furthermore, during the execution phase, the stored motion knowledge is refined using internal and external sensors, generating new motion commands that may target different objectives. Additionally, the objective can change during execution; for example, if the battery condition becomes critical, the vehicle's maneuverability objective is raised to ensure the program recognizes the need to reach a charging station. This results in at least a temporary superimposed objective—that is, reaching the charging station even if the original target coordinates cannot be achieved. When a certain battery level is reached, the vehicle's maneuverability objective decreases again, while the objective of reaching the destination is raised to continue the journey.
[0026] According to one aspect of the invention, the method iterates in a learning manner to continuously identify new action commands that, in each case, transform output triples into target triples, and stores these action commands to drive the device. This has the advantage that the method is established, overall or at least partially, in such a way that new situations are continuously identified and action commands are created to drive the actuator. The method can be modified so that the driving of the actuator is no longer random, but can combine existing action commands, or it can partially randomize existing action commands to create new possible action commands and clearly understand which interactions they trigger. Furthermore, the procedure can be established from the storage of interactions until the control of the device. This means that interactions are also saved when known commands are executed and new parameters are identified. For example, it can be identified that different target parameters are produced when the same action command is executed twice. This occurs, for example, if a crosswind occurs while the car is driving. This situation is saved, and then external sensors identify that there must be wind. This will create new output triples, which the actuator can use to randomly try how to perform reverse steering in this situation. This means that a new interaction has been identified, and if such an output triple is identified again, then it is now clear what actions must be performed to achieve the desired target triple.
[0027] According to another aspect of the invention, reading from the internal actuator and / or external actuator is performed by storing individual support values and interpolating and / or extrapolating intermediate values. This has the advantage that support values can be selected based on their availability and that these support values can be estimated without measuring every possible value. Instead, existing measurements can be used to estimate which values would be produced under normal circumstances. This means that additional support values can be calculated mathematically.
[0028] According to another aspect of the invention, the target device attributes and / or target environmental attributes are at least temporarily affected by the device attributes and / or environmental attributes. This has the advantage that the target specifications can also be changed or prioritized. For example, if a vehicle is to travel from a first geographic point to a second geographic point, priority can be given to reaching a charging station during this period, otherwise the final destination cannot be reached. Therefore, this has the advantage of always ensuring that the target will be achieved and creating fail-safe procedures.
[0029] According to another aspect of the invention, multiple action instructions are combined to form an arrangement that transforms a source triple into a target triple. This has the advantage that even complex compound actions can be executed, and as the method progresses or is established, new action instructions are continuously created and combined to form increasingly efficient arrangements. Therefore, the proposed procedure is iteratively improved.
[0030] According to another aspect of the invention, an external sensor is used to detect the object, and the parameters detected by the external sensor are compared with device attributes and / or device instructions. This has the advantage that the device can virtually clone itself or infer the attributes of the object based on its own attributes. This can also be called empathy. For example, if the external sensor identifies a device with similar characteristics to the actuator, it is assumed that the newly identified device has similar functionality to the actuator. For illustration, consider a vehicle implementing the invention. The first vehicle executes the proposed method and identifies another device, i.e., a second device, via an external sensor, in this example, an imaging unit. The first device or the first vehicle is aware that it can only brake to a limited extent at a specific speed and can only perform a limited degree of lateral steering. The first vehicle now identifies a second vehicle that is similar in size to the first vehicle and travels at a similar speed. Subsequently, instructions are conceptually projected onto the second vehicle, and the first vehicle identifies that the second vehicle can only brake to a limited extent and can only perform a very limited degree of directional change. If an overtaking maneuver is now being performed on a highway, the first vehicle identifies that the second vehicle can brake or change lanes. Therefore, based on its own behavior or possible output triples and possible target triples, it can be inferred that the object has similar characteristics. Furthermore, external sensors can be used to monitor the behavior of the detected second vehicle, and then update the vehicle's own data storage used to store interactions. This means that output triples and target triples can also be identified, and conclusions about the first vehicle can be inferred from the second vehicle.
[0031] According to another aspect of the invention, the detected parameters are used to predict the behavior of the detected object. The advantage of this is that, since the device itself has executed a specific action command in a specific context, it is highly likely that another detected object will also execute the same action in the same context. Therefore, the user's own actions or executed commands are recorded, and it is then assumed that the detected object may act in the same or at least similar manner.
[0032] According to another aspect of the invention, device instructions are created using the interrelationship between the parameters of the detected object and its actions. This has the advantage that correlations with externally identified features associated with the proposed device or the device controlled by the method can also be used. Therefore, the device not only creates the evaluated interrelationships but also observes the external world and thereby draws conclusions about the transformation from output triples to target triples.
[0033] According to another aspect of the invention, the actuator comprises: a processor, a memory, a motor, a gripping arm, a motion unit, a driver, an end effector, a windshield wiper, an indicator light, a lamp, a steering mechanism, and / or an execution unit. This has the advantage that all possible hardware units can utilize the proposed device or a device controlled according to the method. This design is merely illustrative and not exhaustive; therefore, any physical unit can be autonomously controlled according to the proposed method.
[0034] According to another aspect of the invention, the internal sensor unit detects actuator status, actuator settings, actuator position, size, weight, specifications, speed, acceleration, deceleration, battery level, range, position, processor utilization, memory utilization, and / or system parameters of the device. This has the advantage that the internal sensor unit can measure all parameters and states of the device executed by the method or the device executed or controlled by the method.
[0035] According to another aspect of the invention, the external sensor unit detects distance, object, lidar signal, image signal, position, size, specifications, acoustic signal, and / or external parameters. This has the advantage of allowing analysis and recording of the complete environment of the device. All sensors capable of identifying the environment in any way can be used individually or in combination.
[0036] The problem is also solved by an apparatus suitable for implementing the method according to the invention, the apparatus comprising: an initialization unit adapted to initialize by randomly driving an actuator and reading a plurality of internal sensor units suitable for measuring internal device attributes and a plurality of external sensor units suitable for measuring external environmental attributes; a memory unit configured to store the interrelationship between the actuator drive and the read device attributes and environmental attributes as action instructions, such that the action instructions convert an output triplet of the drive, device attributes, and environmental attributes into a target triplet; and a control unit configured to control the apparatus using at least one stored action instruction based on at least one predefined target device attribute and / or at least one predefined target environmental attribute, the at least one stored action instruction indicating how to use the output triplet and target triplet to set at least one predefined target device attribute and / or at least one predefined target environmental attribute, wherein the method executed by the apparatus is iterative in a learning manner, continuously identifying new action instructions during the iteration, each action instruction converting an initial triplet into a target triplet, and these action instructions are stored for use in controlling the apparatus.
[0037] The problem is also addressed by an arrangement comprising multiple devices connected via communication technology and exchanging action commands.
[0038] This task is also addressed by a computer program product having control commands to implement the proposed method or operation of the proposed apparatus.
[0039] According to the invention, this method is particularly advantageous for operating the proposed apparatus and units. Furthermore, the proposed apparatus and units are adapted to implement the method according to the invention. Thus, in each case, the apparatus implements structural features suitable for performing the corresponding method. However, these structural features can also be designed as method steps. The proposed method also provides steps for implementing the functions of the structural features. Furthermore, physical components can also be provided virtually or in a virtualized manner.
[0040] Further advantages, features, and details of the invention will become apparent from the following description, in which various aspects of the invention will be described in detail with reference to the accompanying drawings. Features mentioned in the claims and description may be essential to the invention, alone or in any combination. Similarly, the foregoing features and those further described herein may be used alone or in any combination. Attached Figure Description
[0041] Components or parts with similar or identical functions are sometimes provided with the same reference numerals. The terms "left," "right," "top," and "bottom" used in the description of embodiments refer to the orientation of the drawings having generally legible graphic names or legible reference numerals. The illustrated and described embodiments should not be construed as conclusive, but rather have an exemplary nature for explaining the invention. The detailed description is intended to inform those skilled in the art; therefore, known circuits, structures, and methods are not shown or explained in detail in the description so as not to hinder understanding of this specification. The drawings show:
[0042] Figure 1 A schematic flowchart of a method for autonomously controlling a device according to one aspect of the present invention;
[0043] Figure 2 This can be used as a schematic diagram to estimate internal and / or external properties. It allows for the measurement of internal and / or external parameters and interpolation between sampling points, or extrapolation of other sampling points.
[0044] Figure 3 A diagram of a real-world scenario where a child wants to cross the road; the proposed device identifies this scenario and teaches accordingly.
[0045] Figure 4 : A real-world road traffic scenario in which the apparatus or method proposed according to one aspect of the present invention is used; and
[0046] Figure 5Another real-world scenario in road traffic, where the device is practicing turning according to another aspect of the invention.
[0047] Figure 6A A schematic flowchart of a method for autonomous control of a device according to one aspect of the present invention in the juvenile stage, the main purpose of which is to define its own actions; and
[0048] Figure 6B A schematic flowchart of a method for autonomous control of a device according to one aspect of the present invention in the adult stage, the application stage being designed to achieve its own objectives while taking into account and predicting the behavior of other objects in relation to the context. Detailed Implementation
[0049] Figure 1 A method for autonomous control of a device is illustrated in the form of a schematic flowchart. The method includes: initializing 100 by randomly driving actuator 101, reading multiple internal sensor units configured to measure internal device attributes 102, and reading multiple external sensor units configured to measure external environmental attributes 103; storing 104 the correlation between the actuator drive 101 and the read device and environmental attributes as action instructions, such that the action instructions convert an output triplet of the drive 101, device attributes, and environmental attributes into a target triplet; and driving device 105 according to at least one predefined target device attribute and / or at least one predefined target environmental attribute using at least one action instruction stored in 104, the at least one action instruction stored in 104 indicating how to use the output triplet and target triplet to set at least one predefined target device attribute and / or at least one predefined target environmental attribute, wherein the method iterates in a learning manner, continuously identifying new action instructions during learning, each new action instruction converting the initial triplet into a target triplet in each case, and these action instructions are stored for use in controlling the device.
[0050] According to one aspect of the invention, the proposed model uses a central template through which artificial intelligence can find its way in an unknown and ever-changing world. This method can make short-term behavioral predictions of observed objects, thereby assessing the development of the situation and incorporating it into its own behavior design. It can also learn through observation and empathy. Key aspects of the invention include:
[0051] - The central template serves as the core of artificial intelligence and is used to record all objects in a context;
[0052] -Data storage enables patterns to be identified;
[0053] - Predict the behavior of all observed objects in the context;
[0054] - Empathy based on a central template;
[0055] - The ability to learn through observation;
[0056] - Identify influencing factors that cannot be directly observed;
[0057] - Identify natural laws and causal relationships;
[0058] - A computerized form of theory that can be created, stored, reused, and continuously improved; and / or
[0059] - Self-assessment and evaluation of self-generated theories serve as the basis for cognition, learning, and selecting the best theory to achieve one's own goals.
[0060] The paper describes a car that behaves according to these principles. It describes how it distinguishes between a child and a box, how it judges overtaking maneuvers on a highway, and how it recognizes the law of universal gravitation and crosswinds (using causal relationships and indirect objects as examples). The paper also describes how empathy can be used for observational learning. All of this suggests that the possibilities of this artificial intelligence extend far beyond driving cars. However, the core foundation is a robot situated in a physically present environment, capable of directly or indirectly observing the physical existence of other objects.
[0061] The following sections will set forth some aspects of the invention that can exemplarily implement the method, apparatus, and / or system arrangement. These aspects should be understood as merely exemplary and may be applied individually or in combination.
[0062] Possible hardware for the robot or the proposed device:
[0063] This article describes the technical equipment that a robot or automobile may be equipped with according to one aspect of the present invention.
[0064] Actuator:
[0065] In simple cases, such as in a car, actuators can be windshield wipers, indicator lights, lights, etc. Braking, acceleration, and steering are actions by which robots influence their environment; even a change in their own position constitutes a change in context.
[0066] Internal sensors:
[0067] An internal sensor system measures and provides feedback on the robot's own characteristics, such as height, weight, speed, and acceleration, based on the actuator's actions. These are essential for evaluating its own actions and are prerequisites for learning. This enables the robot to recognize its own actions, such as braking intensity and effect (the degree of negative acceleration), establish connections, and store these connections in memory for independent learning—that is, to continuously improve its own actions.
[0068] As a physical entity within its physical environment, a robot is bound by the laws of physics. The measured acceleration, or change in velocity, depends not only on its mass but also on the degree to which the accelerator pedal is depressed or the force of braking. If these laws are not permanently programmed into the robot, but rather discovered, stored, and used autonomously, then internal sensors are needed to learn and store the results of using actuators. For example, a car's internal sensors thus record data such as dimensions, speed, and acceleration.
[0069] External sensors:
[0070] According to one aspect of the invention, an external sensor system detects the environment (distance, objects, etc.) through technologies such as lidar and image processing. Of course, environmental information must be provided to the robot in an appropriate manner. Specific requirements will be described in more detail below. However, due to the use of a central template, the information requirements are low, and therefore it is not necessary to determine in advance which system is most suitable. Not all theoretically available information needs to be recorded and processed.
[0071] Possible software for the robot:
[0072] This article describes, according to one aspect of the present invention, the various parts of the presented artificial intelligence that realize its intelligence.
[0073] Data storage:
[0074] This data storage is a high-availability data repository used to store data tuples consisting of internal sensors, actuators, actions, and their motivational values (see below). The repository appropriately combines the executed actions with their associated actual, empirical, and observed results. Therefore, it stores internal data only in tabular form, not, for example, pixels of one or more images. Thus, only data whose internal meaning is known exists.
[0075] According to one aspect of the invention, another part of this module involves repeatedly checking the internal consistency of the data. This means that the data must not contain any contradictions, structures that lead to loops or cycles, etc. The goal is that the data should ultimately lead to only one decision. The routine for checking this should also remove stored redundant data, not only to keep the database as streamlined as possible, but also to prevent overfitting. Which data is considered redundant depends on the type of theorization.
[0076] Figure 2 This illustrates a theorization of one aspect of the invention:
[0077] Using the simple method described in this paper, a central template can reproduce any functional correlation and save the functional correlation for prediction.
[0078] Braking or acceleration depends functionally on the mass of the acceleration and the applied force. However, this function is initially unknown because the mathematical function of the velocity change dV depends on the preconditions Vn (current speed, weight, etc.) and the use of the actuator Em (force on the accelerator pedal, steering, etc.), i.e.:
[0079] dV=f(Vn, Em) (1)
[0080] This function should be learned by the user, not programmed into it, because this allows for the extension, modification, and improvement of the learned content in the same way. The effective acceleration during braking (or acceleration) is determined by the law F=m The acceleration *a*, measured by the internal sensor system, is mathematically independent of the force *F* (the force applied to the pedal) and the mass *m*. The principle by which the robot finds this equation is straightforward. To illustrate this, we model it using a (different, arbitrary) quadratic function y = -((x-4)²) + 9 as an example.
[0081] Suppose initially there are only two measurements at points 11 and 12. Now suppose a prediction p needs to be made at point x=3 (point 2), yielding the value of y=3.8 through the straight line a between 11 and 12. However, since the error of the subsequent measurement y=8.0 is too large, the point 2R at x=3.0 and y=8.0 is included in memory. Therefore, a new prediction will be made using the degree b1 between [11, 2R] or b2 between [2R, 12]. Over time, many more points (31, 32, 33, ...) are created in the above example, through which the functional relationship between x and y can be modeled with arbitrary precision using the corresponding straight lines along the data tuples (c1, c2, c3, ...). The precision is limited only by the experience not yet obtained and the size of the memory. Of course, this can be extended to any number of variables if computation and memory capacity allow; predictions can be made by piecewise linear interpolation or extrapolation using a hyperplane a1x1+...+anxn=const instead of straight lines.
[0082] However, to minimize the number of stored corner tuples or data tuples, these can be removed again: if post-hoc analysis shows that a data tuple is very close to, or even on, the line / hyperplane between two adjacent tuples, it can be removed as a redundant tuple. This not only reduces memory requirements but also computational workload, and over time, the behavior becomes increasingly efficient by adapting to the actual (though still unknown) mathematical function.
[0083] This process also prevents overfitting, which must be taken into account when selecting and determining the range of training data for a CNN. Therefore, redundant experience is not repeatedly stored, which would otherwise completely mask rare events, the so-called "black swan" events, and could have serious consequences.
[0084] Storing the functional dependence of variables has another advantage: it stores all possible inverse functions because the stored data tuples don't distinguish which values are the "dependent variable" and which are the "independent variable" when they occur. In the function above, x is the independent variable, and the corresponding y value could be determined using the function y = -((x-4)²) + 9. The x value must be searched in the database; if the x value is not explicitly stored, the y value can be approximated using the next two x values. However, you can also directly search for the y value in memory and then determine the x value using piecewise linear interpolation or extrapolation—this is much easier than, for example, trying to determine the inverse function of the equation above.
[0085] This very simple method allows for the modeling of any functional environment with arbitrary precision, provided the motion is sufficiently trained, just as in complex motion. Of course, the acceptable error ε can also vary and be adjusted over time, either by decreasing it to improve desired precision or by increasing it to reduce memory and computational workload, as this in turn affects the number of storage corners and data points—a continuous optimization process.
[0086] According to one aspect of the invention, the rules and laws governing the current situation are defined as a self-contained, consistent dataset from which the robot can calculate the predicted values of the theory using piecewise linear interpolation or extrapolation. This dataset contains its own measurements and may also contain parameters from the created alter ego. In this way, the robot can independently identify, save, reuse, and improve each of its theories. As part of self-evaluation (see below), it is also able to select and apply the best specific theory (dataset) for a given situation, i.e., the most capable theory (see below).
[0087] action:
[0088] According to one aspect of the invention, an action is an ordered sequence of actuator operations for achieving a goal. Actions can be designed using functional relationships between actuators, internal parameters, and known effects of the robot. Initially, in the early learning stage (see below), simple actions are performed. The artificial intelligence learns to accelerate and brake, then accelerate, wait, and collide with obstacles (external braking, with a "risk of injury"), accelerate, turn, brake, and so on. These simple actions are combined to form increasingly complex "choreographies," which are then further optimized, for example, by reducing pedal pressure to make braking smoother as speed decreases, and by making turns more fluid. Optimization is achieved through motivation values (see below), which represent an internal evaluation of the actions performed. The choreographies with the highest motivation values are selected from a stored (partial) library of "choreographies." In this way, the artificial intelligence can continuously improve the use of its actuators. Furthermore, it can predict its own situation at future points in time based on current reference values and existing action experience, thus laying the foundation for its own action planning.
[0089] Motivation value:
[0090] Motivation values reflect a reward. On one hand, they are internally determined after a self-set goal is achieved (success); on the other hand, motivation values can also be externally driven (by doing well). For example, if you drive on a route with a strong steering maneuver, low speed, but high lateral acceleration, then the motivation value gained from this series of actions is lower than the result of "training"; subsequently, you drive the same route at a faster speed but with lower lateral acceleration. Of course, this must also be kept in mind.
[0091] Therefore, motivation values represent an unspecified intrinsic goal. They can be associated with long-term goals, such as the destination of a journey, or short-term intermediate goals, such as going to a gas station when the tank is empty: as the tank level decreases, the motivation value for "refueling" slowly increases until it exceeds the motivation value for the long-term goal. The pursuit of the distant destination is interrupted, replaced by refueling, and then the journey begins again.
[0092] Target system:
[0093] According to one aspect of the invention, the robot is always in an "on" state, although it can, of course, be turned off. Therefore, it is always performing an activity. However, this also includes doing nothing (charging) or optimizing its long-term memory of action options (sleep) in a resting state. The selected action is always the one with the highest current motivation value. If an action not currently being performed reaches a higher motivation value than the action currently being performed, that action will be canceled, terminated, or interrupted in order to perform the action with the highest motivation value. If a car traveling from Munich to Hamburg for refueling or charging has a greater motivation than reaching Hamburg, it will abort its journey and head to the gas station the next time it has the opportunity.
[0094] In the early stages of such robots or cars (see below), the learning phase, the motivation for short-distance driving, braking, acceleration, and turning is assigned a high motivation value, simply because internal long-term memory signals that it should learn or reduce motion errors to make its movements more precise. Later, in adulthood (see below), when such behavior has little or no impact on data memory, error reduction approaches zero, and the behavior becomes "boring," receiving only a lower motivation value. Then, for example, energy conservation takes precedence.
[0095] Therefore, according to one aspect of the invention, the target system avoids an obsession with reality, which means that the interpretation of reality may prove to be wrong in the future.
[0096] Central template, main body:
[0097] According to one aspect of the invention, the central template represents the core and foundation of the artificial intelligence. The central template is called an ontology because it places itself at the center of data processing and primarily views and evaluates everything from its own perspective. Therefore, the ontology does not need to be assigned or trained with any data that has its own meaning (e.g., image recognition); it creates everything.
[0098] Subjective polar coordinate system:
[0099] According to one aspect of the invention, all observation and experience data of this artificial intelligence are subjective data, or derived from the perspective of the ontology. The spatial concept also follows this idea. No Cartesian coordinate system is designed with an assumed origin at a location in the observation environment, and no corresponding coordinates are assigned to each identified object in the Cartesian coordinate system. Instead, a polar coordinate system can be used. This not only automatically determines the mathematical origin in the ontology, but also automatically determines the subjective orientation of the ontology, including the forward / backward, up / down, and left / right orientations of the motion vectors. The surrounding (empty) space is taken for granted. The observed object is identified only by its distance and angle from its own viewpoint, i.e., the origin of the polar coordinate system. Therefore, physically empty space is not considered a separate entity or dimension. It is simply there, and merely an option to move there, unless another object is standing or moving there, in which case the "position" is occupied, or the space is not empty. Another advantage of subjective polar coordinates lies in the concept of a "mirror body" (see below).
[0100] Early childhood:
[0101] Under unconditional requirements, such as the integrity of the observed object and the robot itself, maintaining operational readiness (battery charging state), and objectives such as smooth acceleration and braking, and optimized turning, the robot should / can learn autonomously the use of its actuators and their consequences. Most importantly, to coordinate its actuators with its internal and external sensors, it fine-tunes them through progressively increasing training (100-105), occasionally interrupted by rest phases, such as charging the battery, while maintaining the internal database (redundancy, consistency, etc.), until the errors identified from the planning and results become sufficiently small. Taking a car as an example, this mainly involves acceleration, steering, braking, reaching the destination (success), or missing the destination (error). Initially, the robot sets simple objectives for simple actions, executes these objectives and remembers the results, then combines these results to form increasingly complex "orchestrations." This training phase ends when the actions for self-set or externally set objectives ("drive there") cannot be further optimized, i.e., when the mathematical module cannot reduce errors or can only slightly reduce them and / or the database is difficult to optimize (the number and location of points in memory). If the reduction in error becomes increasingly slower, then the early stages are drawing to a close. On the other hand, KNN requires a large amount of carefully selected training data, while ontology only needs a real environment, such as a playroom, where it can experiment and train its abilities without causing any damage. It follows an intrinsic goal, rather than an externally pre-set goal like KNN, and therefore does not interpret the environment or the world in any way.
[0102] Adulthood:
[0103] In its early stages, the device has learned the ordered sequence of its actuators to achieve a predetermined goal. What is the value of the knowledge and experience acquired in this early stage? All the data about the world that the ontology learns in the world it moves through is its own data, its own observations, and its own experiences. From a higher level of scientific theory or philosophical perspective, it must be pointed out that this data is real from the ontology's perspective, whether "so it is now" or "so it was in the past." Therefore, this represents a natural and fixed starting point for cognition. Of course, the ontology must maintain its internal data. In addition to the optimizations already mentioned above, it must also be ensured that there are no contradictions or errors. Otherwise, robot manufacturers or operators cannot assume that the robot will achieve the intended goals through its designed actions.
[0104] The car's own experience and the environmental changes it triggers, such as positional changes caused by acceleration, steering, and braking, are experienced and stored by connecting internal and external data. At the level of the problem's meaning, there's no need to delve into its own actuators to understand the causes—for example, why pressing the accelerator pedal causes acceleration, why steering causes turning, and why braking causes negative acceleration. The car's ego only needs a "it has given three options" to selectively change its position. At this level, what matters is "this is how it is," not "why." A pigeon on the street certainly doesn't know how cars or cyclists work. However, if such a "road user" passes by far enough away, it will remain still; but if the direction of movement changes towards it, it will fly away. It doesn't know the reason, but it is aware of the possibility of action and therefore pays close attention.
[0105] Mirror body:
[0106] When an (adult) ontology identifies objects in the observable environment through external sensors, it simply creates a copy of itself and adjusts the parameters of the copy based on the observed attributes. Therefore, the ontology serves as a central template for all observable and unobservable objects and phenomena. Hence, it is named the ontology, and as a result, a mirror image for the copy.
[0107] By using an ontology as a central template, there are no longer any unknown objects. Only unknown parameters exist, but these parameters can be measured by external sensors, estimated based on personal experience (keyword: bias), and their possible range can be limited through further observation.
[0108] Just as an entity uses its methods, stored data, and parameters to calculate its behavior and predict its own situation, the behavior of a mirror body can also be calculated using the same methods, data, and adjusted parameters. Assumptions about mirror bodies are limited to the same laws of physics, and short-term goals can be deduced from observed motion orientation and orientation, regardless of whether the object of observation is a box on the street, a kangaroo in Australia, or an elderly woman with a walking frame.
[0109] In terms of the program, the ontology should exist as an object. Then, a copy of the ontology can be easily created for each new object that appears, and the parameters of the new object copy can be adjusted using external sensors and the user's own experience. A list can exist containing mirror bodies with new copies inserted or created in the current context:
[0110] Class CEgo {...};
[0111] CEgo Ego (...);
[0112] / / Education and learning begin in early childhood ...
[0113] / / Start of adulthood
[0114] CEgo AlterEgos[]; ...
[0115] / / New object identified
[0116] Alter Egos[i] = Ego.clone(); / / Clones the original object.
[0117] AlterEgos[i].adjust(...); / / Adjust the parameters of the new object ...
[0118] Thus, the ontology internally creates a mirror image, a computable image of the objects observed in the environment, and as... Figure 6B As shown, it can form an inner loop 311 for object recognition, update the parameters of the observable object 312, predict the behavior of the object 313, adjust the executor operation sequence 314, and enter the second outer loop after the target is reached, enter the rest and (and) loading phase 315, until it continues to initialize for a new target 310.
[0119] Whether the object of observation is familiar or entirely new is not important. If it is new, the possible range of feature values is larger, which requires higher attention (scanning frequency). It has the following advantages:
[0120] - Basically understand each observed object!
[0121] -Predict the behavior of all external objects using your own methods and data!
[0122] -Humans no longer need to intervene to fill in the missing parts.
[0123] - The ontology continuously learns and improves its behavior.
[0124] - The entity is capable of empathy and perceiving the current situation (see below).
[0125] - The ontology can learn by observing the mirror image (see below).
[0126] - Mirror images can reflect unknown causal relationships and rules (see below).
[0127] Figure 3 A scenario 1, "children," is shown according to one aspect of the invention.
[0128] When the AI of the subject, such as a simple car, encounters another road user, in this case a child, it creates a copy of itself and adjusts its parameters (distance d from the subject, orientation, speed vK, etc.):
[0129] In this way, the mirror body can use the ontology's methods to predict when and where it will be. By comparing this prediction with its own calculations of when and where it will be, the ontology can determine whether a potentially dangerous situation might occur and react accordingly, such as slowing down. Thus, it can use its own methods and experience to predict a child's short-term behavior. No special programming or training for the child is required, which was not present when the AI first encountered the child.
[0130] Figure 4 Scenario 2, "Highway," is shown according to one aspect of the invention.
[0131] Another example more clearly illustrates the power of the simple concept of a mirror image in the following scenario: On a highway, there are two cars behind a truck, and one car in the fast lane, where the second car is our own image. Based on our own law of braking force (F=m... Based on the understanding of a), combined with observed mass parameters of cars and trucks—that is, assumed mass, measured speed, and assumed braking effect—the ontology can first use its own methods to predict the braking distance of all road users in potential accident scenarios, thereby determining its own safe distance. But let's first address the "ideas," or the ontology's assumptions about other road users, in order to predict their behavior:
[0132] -HGV0 is traveling at a certain speed, and the main body (PKWEgo) already knows (without special training) that such a large number of road users rarely exceed it. Therefore, the truck's mirror body does not predict any changes in speed.
[0133] - As the center of observation, the PKWEgo will be able to travel at higher speeds and overtake the HGV0 while considering other road users. Overtaking allows the PKWEgo to reach its destination faster, thus giving it a higher motivation value when preparing to overtake.
[0134] - As another mirror image, PKW1, due to system limitations, shares the same thinking as PWKEgo, since the original entity will act from its position. Therefore, the original entity assumes that PKW1 also wants to overtake LKW0. Of course, the mirror image of car 1 "sees" PKW1 and executes its (the original entity's) pre-set overtaking maneuver, while taking into account the presence of PKW1, for example, by setting its indicator lights and paying special attention to PKW1 and other road users in the overtaking lane. This is what the original entity of PKW1 "thinks" about, and how it predicts the behavior of car 1.
[0135] Another potential road user, Car 2, is approaching at high speed in the fast lane. The ontology also creates a mirror image of this car and predicts its actions based on the unobstructed lane.
[0136] As shown above, the program, or ontology, can predict the overall situation of all relevant road users and plan its own sequence of actuator operations accordingly to avoid dangerous developments.
[0137] Mirror image, consciousness:
[0138] Therefore, the ontology possesses a complete internal picture of the observable environment with all identified objects and is able to predict their short-term behavior. This would be a possible and programmable constraint on consciousness—a genuine constraint on consciousness, rather than a forged or imitated one.
[0139] Short-term behavior prediction based on the ability to place oneself in the position of the observed object means using all data converted into a mirror image to observe the situation and using the calculated possible behaviors of the observed object to calculate the progress of the situation.
[0140] The ontology in the highway scenario described above can be applied to all observed objects without writing additional programs:
[0141] - Predict their behavior by assuming that all road users want to drive as quickly and without accidents;
[0142] - Continuously adjust predictions based on observed actual behavior (braking, acceleration, steering, indicator lights, headlights, etc.); and
[0143] - Overtake the truck at the right time.
[0144] An ontology can easily assess the possible behaviors of all participants in a given situation by applying its own methods and experience (data) to the possible behaviors of all participants and combining them with appropriate parameters. Don't humans do the same thing?
[0145] These examples clearly demonstrate that the foundation, root, and starting point of this kind of artificial intelligence is always the ontology, the central template. The more differentiated its internal and external perception capabilities, and the better it is trained in its early stages, the more intelligent its behavior and the greater its potential to predict the behavior of observed objects. If we also set parameters for the ontology's own vulnerability—in the case of a car, one might call it deformability—then the ontology can also assess the risk of strong or weak contact. However, the ontology only needs to be built, programmed, and trained (once), and of course, the training part can be transferred from the trained ontology. The observed object does not need to be identified (KNN and its training data), and the behavior of identified objects does not need to be programmed. Of course, one might argue that an ontology alone is insufficient to predict the behavior of other road users. But you only need to consider how humans try to assess the behavior of other road users. Humans do not know what the long-term goals of road users on a highway are; they can only guess at the short-term goals of other road users (getting ahead as quickly as possible without causing an accident) and prepare accordingly.
[0146] Robot's capabilities:
[0147] Empathy:
[0148] The above demonstrates how an ontology uses a mirror image to put itself in the shoes of observed objects in the environment in order to calculate their behavior from their perspective. You can call it "empathy" if you don't use the term "empathy" specifically for humans.
[0149] Learning from observation:
[0150] Learning from observation means observing behaviors that might be better than one's own and adopting them where possible. This is based on comparing the behavior of others with one's own, and this comparison is achieved through empathy, as defined above.
[0151] Empathy in ontology is achieved here through a mirror body, through which the ontology places itself in the context of other observed objects in order to predict their behavior based on their perception of the environment. We assume that there is a difference between the predicted behavior and the observed behavior. If the mirror body alters the sequence of actions of the actuators so that the observed behavior can be reproduced, then the ontology has the opportunity to replicate the observed behavior if it is advantageous.
[0152] Figure 5 Different “curving behaviors” are shown in Scenario 3.
[0153] We imagine the vehicle has learned to always turn at the same distance from the right edge of the road, i.e., the LE line. Now, it observes a car ahead and cuts off the curve by turning earlier but more gently, which also reduces lateral acceleration, thus allowing it to navigate the curve more quickly on the dashed LB line.
[0154] From this perspective, the empathy described here is a prerequisite for learning through observing and imitating the behavior of others. (There is no reason not to assume that this also applies to humans, i.e., learning through imitation based on empathy).
[0155] If data from external sensors is transformed into the situation and position of a mirror image and correlated with simple and more complex action sequences stored in the original body and copied into the mirror image, then the original body can learn from observation: it can recognize the differences between its own self-planned action sequences (always maintaining the same distance from the road edge) and the observed sequences (earlier braking and turning, smaller steering movements, higher turning speeds, and earlier acceleration), and determine that it can navigate curves faster in this way. This learning method is certainly faster than the usual trial-and-error approach.
[0156] Identifying causal relationships:
[0157] The theorization section explains how to replicate any observed functional relationship. Let us apply this ability to identify laws of nature, which, where necessary and appropriate, can also be represented by mirror images. Of course, these mirror images don't exist on roads or travel on roads, but they cause changes that can be detected by external and / or internal sensors.
[0158] Let's imagine our AI ontology observing an experiment with a falling apple. The ontology creates a mirror image and notices the accelerating motion towards the ground. Now, the ontology initially assumes that the apple itself causes the acceleration; like a car, it has its own "accelerator pedal" to accelerate itself. But then, the ontology observes that the apple itself isn't moving, and other objects also fall to the ground and then stop moving. Furthermore, the ontology should have already learned the law of universal gravitation. On a slope, even without pressing the accelerator pedal, it will notice and learn acceleration, but only when "going downhill," while it must accelerate more when "going uphill." According to our known formula (g = gravitational acceleration, 9.81 m / s²), the larger the angle α, the greater the force:
[0159] F=m g sin(α) (2)
[0160] Remember, the ontology does not explicitly know this equation or the equation of free fall, i.e., the law of universal gravitation (sin(α=90°)=1), but it implicitly understands this relationship between the relevant variables through data points and piecewise linear interpolation or extrapolation.
[0161] Therefore, the entity behaves as if it knows, as physicists say, that as a subject with heavy mass, it is subject to the aforementioned law of mass gravity. If it now sees another object and creates its mirror image, then the mirror image also experiences this mass gravity. In this way, this knowledge becomes a law of nature that affects all entities without being programmed into them. If we observe the behavior of this entity from the outside, we cannot determine whether it truly understands the law of universal gravitation like a physicist, or is merely pretending to understand it. Even if the entity sees a pigeon perched on the street, and the pigeon flies away as soon as the entity approaches, it can still recognize that the wing flapping and upward acceleration are consistent.
[0162] Another example is a strong crosswind. The ontology is a large, empty truck. During normal straight-line driving, the lateral acceleration is zero. But then the truck suddenly veers to the side while driving straight, triggering the lateral acceleration sensor. In this case, the ontology can simply create a new mirror image for the unexpected force and unknown cause and associate it with that lateral acceleration. If the ontology can also record environmental data, such as the movement of branches in a forest or plain, and add it to its data storage, then the ontology can not only identify previously unknown variables (crosswinds) but also potential causal relationships. If the ontology can also see and judge the degree of branch movement, it can even estimate the strength of the force acting on the side, thus estimating the possible impact on driving behavior. In this way, the ontology itself creates a theory that includes new variables. This is similar to those physicists who proposed dark matter and dark energy, even though they are invisible and unobservable, or like Nobel laureate Peter Higgs, who conceived of a particle in the 1960s and first detected it at CERN in 2012. A mirror image, a copy of the self, can never be used to identify the essence of a phenomenon in order to seek the truth, but it is sufficient for a full grasp of the situation.
[0163] self assessment:
[0164] The first form of knowledge acquisition has already been described. The example of crosswinds is illustrated in the section "Understanding Causality." A force unknown to the subject suddenly produces lateral acceleration in a straight line, having a measurable effect on the robot or the subject. It creates a mirror image for this. However, initially, no additional information signals to the subject when this phenomenon occurs and its intensity. Only when the subject observes a temporal correlation between the varying intensity of the tree branch's motion and its own varying intensity of lateral acceleration can it correlate the two and assume they have the same cause. What happens here? The subject cannot recognize the crosswind itself, but it can recognize its own response to the crosswind (lateral acceleration) and the tree branch's response. This creates a factor that sometimes affects the subject and sometimes does not, but it affects not only the subject but also other observable objects.
[0165] The second form of knowledge acquisition is based on comparing one's own behavior with that of other objects. This presupposes that the observed object's behavior is "rational" rather than random, although this can be considered rational, thus comparison is no longer a necessary condition for acquiring knowledge. By comparing two different behaviors, the ontology can determine whether its own capabilities (see below) are lower, equal to, or superior to those of the observed object in a given situation. This means that the observed object can perceive more, the same, or fewer situations. The observed truck suddenly slows down on a straight road. It seems to have seen something the observing ontology doesn't know; the ontology recognizes less than the object. If a weaker capability has been identified, this can be seen as an opportunity to specifically analyze these situations in order to at least improve its capabilities to the level of other objects. The comparison is not about whether the behavior is right or wrong (which only leads to problems with well-known theories of truth), but only about whether and when the behavior changes, and whether the behavior is repeated. The first behavioral comparison involves different situations: in different situations for the ontology, does the observed object also exhibit different behaviors? The second behavioral comparison involves the same situations: in the same situations for the ontology, does the observed object's behavior repeat itself, or is it different? The ontology can draw conclusions through long-term observation of these multiple situations:
[0166] - If an object repeatedly observed in the same context for the ontology exhibits different behaviors, and also changes its behavior in different contexts for the ontology, then the ontology's capabilities are poor.
[0167] - If objects repeatedly observe their behavior in the same context as the ontology, but do not change their behavior in different contexts as the ontology, then the ontology is more powerful.
[0168] -If objects repeatedly observe their behavior in the same context for the ontology, but change their behavior in different contexts for the ontology, then the ontology's capabilities are equal.
[0169] Through this simple comparison, an ontology can recognize how powerful its own capabilities are compared to those of other objects in certain situations. Note that this is not an attempt to compare it to "truths of the real world."
[0170] Through these comparisons, an ontology can assess itself and infer, for example, whether it should act more cautiously in certain situations and attempt to identify what (not yet) is identifiable. This can be understood as an intrinsic instruction to conduct targeted research as part of “continuous learning and the ongoing expansion of cognition and knowledge” (see above).
[0171] Professional knowledge:
[0172] However, the results of these comparisons can also be used to externally assess a robot's capabilities to determine its potential applications. There are no linguistic problems regarding the strength of comparative and determinative abilities compared to the truthfulness of the knowledge about the real world acquired by an ontology.
[0173] However, there is another important aspect to the concept of truth related to theory: capability must also be satisfied. A theory enables prediction. Therefore, understanding the quality of a theory and its subsequent predictions is crucial. Generally, attributing truth to a theory is the standard by which that truth can be applied, allowing for its application. The advantage is that this standard automatically excludes all competing theories, thus eliminating the problem of deciding which theory to use. The concept of capability must reach a level commensurate with this.
[0174] Professional knowledge here is defined and used as the ability to grasp the relevant influencing factors of a given situation or similar situations. Therefore, it is initially a finite standard, but its level can be determined through comparative procedures. This differs from the veracity of a theory, which assumes infinite validity in time and space. However, this cannot be proven, and the standard does not allow for comparison between different theories.
[0175] This means that capability is a better description of a robot's performance than the concept of truth, and unlike truth, it can be determined by the robot itself in a formal self-evaluation process.
[0176] Figure 6AA schematic diagram of a procedure for an autonomous control device during the juvenile or training phase is shown. The goal and endpoint of this phase is to be able to use and deploy the actuator with sufficient accuracy and minimize errors. After initialization 200, there is random driving 201 of the actuator and readings 202 of multiple internal sensor units configured to measure internal device attributes, and readings 203 of multiple external sensor units configured to measure external environmental attributes. The correlation between the actuator driving 201 and the read device and environmental attributes is stored 204 as an action command, such that the action command is an output triplet consisting of the drive 201, device attributes, and environmental attributes. Based on at least one predefined target device attribute and / or at least one predefined target environmental attribute, the action command in at least one stored 204 is used to drive the device 205, specifying how to set at least one predefined target device attribute and / or at least one predefined target environmental attribute based on the output triplet and the target triplet.
[0177] Figure 6A It shows the early childhood stage:
[0178] 200 Initialization, preparing for learning
[0179] 201 Randomly drive the actuator
[0180] 202 Reading a large number of internal sensor units 202
[0181] 203 Internal device characteristics and readouts from multiple external sensor units 203
[0182] 204 Save the relationships between 204 entries
[0183] 205 Immediately determine the motion accuracy error. If the error is too large, continue with step 201.
[0184] Optional: Rest phase for charging and post-event data preparation:
[0185] Error, redundancy, consistency...
[0186] Figure 6B The adult stage is shown:
[0187] 310 Initialization: Preparing for Long-Term Goals
[0188] 311 Identifying objects in the current context 311
[0189] 312 Create or update an image or its parameters 312
[0190] 313 Predicting the behavior of the observed object using the current mirror image.
[0191] 314 Calculate the optimal executor deployment 314 to further pursue your own goals.
[0192] 315 Rest Phase: 315 is used for charging and post-event data preparation.
[0193] Error, redundancy, consistency...
[0194] In this document, robot, device, system, and body are used synonymously.
Claims
1. A method for autonomously controlling a device, the method comprising the following steps: - Initialization (100): Initialization is performed by randomly driving the actuator and reading multiple internal sensor units set to measure internal device properties as internal states, and by multiple external sensor units set to measure external environment properties as external states. Specifically, the external sensor unit detects the object, compares the parameters detected by the external sensor unit with device attributes and device action commands, and uses the detected parameters to predict the behavior of the detected object. The action command, acting as a mapping function, transforms the output triples of the actuator's drive, device attributes, and environmental attributes into target triples. These target triples specify the new internal state, the new external state, and the possible actions that can be performed in the new state. - Storage (104): The correlation between the drive (101) of the actuator and the read device attributes and environmental attributes is stored (104) as an action command; and - Control (105): Using at least one stored (104) action instruction, the device is controlled (105) according to at least one predefined target device attribute and / or at least one predefined target environment attribute, the at least one stored (104) action instruction indicating how to use the output triplet and the target triplet to set the at least one predefined target device attribute and / or the at least one predefined target environment attribute, the method is iterative in a learning manner, in which new action instructions are continuously identified, the new action instructions converting the output triplet into the target triplet in each case, and these action instructions are stored for use in controlling the device.
2. The method according to claim 1, characterized in that, The method iterates in a learning manner, thereby constantly recognizing new action instructions that, in each case, convert the output triplet into a target triplet, and these action instructions are stored for use in driving the device.
3. The method according to claim 1 or 2, characterized in that, The reading of internal actuators and / or the reading of external actuators (103) is performed in such a way that each support value is stored and intermediate values are interpolated and / or extrapolated.
4. The method according to claim 1 or 2, characterized in that, The target device attributes and / or the target environmental attributes are at least temporarily affected by the device attributes and / or environmental attributes.
5. The method according to claim 1 or 2, characterized in that, Multiple action instructions are combined to form an arrangement that transforms the output triplet into the target triplet.
6. The method according to claim 1 or 2, characterized in that, The device generates action commands by utilizing the relationship between the parameters of the object being detected and the actions of the object being detected.
7. The method according to claim 1 or 2, characterized in that, The actuator is: motion unit, drive, wiper, light, steering gear or actuator.
8. The method according to claim 1 or 2, characterized in that, The actuator is a motor.
9. The method according to claim 1 or 2, characterized in that, The actuator is a clamping arm.
10. The method according to claim 1 or 2, characterized in that, Internal sensor unit detects: actuator status, actuator settings, actuator position, size, weight, specifications, speed, acceleration, deceleration, and / or battery level.
11. The method according to claim 1 or 2, characterized in that, External sensor unit detection: distance, lidar signal, image signal, position, size, specifications, and / or acoustic signal.
12. An apparatus suitable for carrying out the method of any one of the preceding claims, the apparatus comprising: - An initialization unit is configured to: perform initialization by randomly driving the actuator and reading multiple internal sensor units that are configured to measure internal device attributes as internal states, and by reading external attributes as external states by multiple external sensor units configured for measurement (100). The device is designed to: detect an object using the external sensor unit, compare the parameters detected by the external sensor unit with device attributes and device action commands, and use the detected parameters to predict the behavior of the detected object. The action command, acting as a mapping function, transforms the output triples of the actuator's drive, device attributes, and environmental attributes into target triples. These target triples specify the new internal state, the new external state, and the possible actions that can be performed in the new state. - A memory unit configured to store (104) the correlation between the drive (101) of the actuator and the read device attributes and environmental attributes as an action instruction; The device is further adapted to use at least one stored action instruction to control, based on at least one predefined target device attribute and / or at least one predefined target environment attribute, wherein the at least one stored action instruction indicates how to use the output triplet and the target triplet to set the at least one predefined target device attribute and / or the at least one predefined target environment attribute, and -The device is further designed such that the method iterates in a learning manner, in which new action instructions are continuously identified, each of the new action instructions converting an output triplet into a target triplet, and these action instructions are stored for use in controlling the device.
Citation Information
Patent Citations
Method and model for performing independent path exploration based on operant conditioning
CN105094124A
Machine learning methods and apparatus related to predicting motion(s) of object(s) in a robot's environment based on image(s) capturing the object(s) and based on parameter(s) for future robot movement in the environment
CN109153123A