Autonomous control device and autonomous control method
Patent Information
- Application Number
- PCT/JP2026/010587
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-18
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026010587_01102026_PF_FP_ABST
Abstract
Description
Autonomous Control Apparatus and Autonomous Control Method
[0001] The present invention relates to an autonomous control apparatus that performs autonomous control of an apparatus and an autonomous control method therefor.
[0002] In order to realize AGI (Artificial General Intelligence) and ASI (Artificial Superintelligence), which are advanced forms of AI (Artificial Intelligence), autonomous control by the artificial intelligence itself is required. Although there are various levels of autonomous control, it is considered that for high-level control that spontaneously performs actions and learning, it is important to generate motivations by oneself and connect them to specific objectives and tasks.
[0003] Conventional autonomous AIs include those that perform control based on pre-implemented motivations and those that perform control based on externally given motivations. For example, Patent Document 1 discloses a learning control apparatus that controls actions according to a sequence planned based on prediction of environmental state changes, and sets a target state by learning input-output relationships in actions.
[0004] FIG. 1 simply shows some functions of the apparatus disclosed in Patent Document 1. In this technique, a plurality of types of motivations such as curiosity motivation, achievement motivation, and manipulation motivation are pre-implemented in the learning control apparatus. An evaluation unit 202 observes a state error corresponding to each motivation and sets a target state to be achieved. A planning unit 204 plans an action sequence from the current state to reaching the target state.
[0005] Patent Document 2 discloses a robot apparatus that generates motivations and objectives based on external influences and internal states, and generates and expresses a corresponding expression mode, that is, modality. FIG. 2 simply shows some functions of the apparatus disclosed in Patent Document 2. In this document, an internal motivation generation unit 210 generates motivation for expressing impressions and emotions based on some model. A modality control unit 212 generates modalities such as gestures, facial expressions, and language based on the generated motivation.
[0006] Japanese Patent Publication No. 4525477, Japanese Unexamined Patent Publication No. 2002-331481
[0007] The technology disclosed in Patent Document 1 utilizes pre-implemented motivations, making adaptive motivation difficult. Furthermore, because it utilizes any of multiple motivations, tasks and actions may diverge due to simultaneous task setting, etc. In the technology of Patent Document 2, multiple input data are integrated to generate a single motivation, so the more diverse the input data and motivations, the more complex the modeling for motivation generation can become. Moreover, it is difficult to assign priorities and constraints to multiple motivations of different types.
[0008] This embodiment has been made in view of these problems, and its purpose is to provide a technology that can easily provide adaptive motivation to a device that performs autonomous operation.
[0009] One aspect of this embodiment relates to an autonomous control device. This autonomous control device is characterized by comprising: an intrinsic motivation generation unit that generates intrinsic motivation based on an internal state representing the internal state based on the thinking of an autonomous device; a functional motivation generation unit that generates functional motivation based on at least one of the hardware state of the autonomous device and information on the external environment; a motivation selection unit that selects a motivation to be adopted based on the priority weights assigned to the intrinsic motivation and the functional motivation by different rules; and a control processing unit that controls the behavior of the autonomous device based on the selected motivation.
[0010] Another aspect of this embodiment relates to an autonomous control method. This autonomous control method is characterized by including the steps of: generating an intrinsic motivation based on an internal state representing the internal state based on the thinking of an autonomous device; generating a functional motivation based on at least one of the hardware state of the autonomous device and information about the external environment; selecting a motivation to be adopted based on the weight of priorities assigned to the intrinsic motivation and the functional motivation by different rules; and controlling the behavior of the autonomous device based on the selected motivation.
[0011] Furthermore, any combination of the above components, or any conversion of the representation of this embodiment between methods, apparatus, systems, computer programs, recording media containing computer programs, etc., are also valid as embodiments of this embodiment.
[0012] According to this embodiment, adaptive motivation can be easily provided to a device that performs autonomous operations.
[0013] This is a diagram illustrating some of the functions of the device disclosed in Patent Document 1. This is a diagram illustrating some of the functions of the device disclosed in Patent Document 2. This is a diagram showing an example configuration of an autonomous device to which this embodiment can be applied. This shows an example of the hardware configuration of the autonomous control device in this embodiment. This is a diagram showing the configuration of the functional blocks of the autonomous control device in this embodiment. This is a diagram illustrating the data flow related to intrinsic motivation in this embodiment. This is a diagram illustrating the data flow related to functional motivation in this embodiment. This is a flowchart showing the processing procedure in which the autonomous control device of this embodiment generates and selects motivation and outputs information related to the action plan.
[0014] Figure 3 shows an example configuration of an autonomous device to which this embodiment can be applied. In this example, the autonomous device 62 comprises an autonomous control device 10, an information processing device 64, a sensor 66, a timer 68, an actuator 70, and an output device 72. The autonomous device 62 is basically a device that sets goals based on motivation, generates and executes corresponding action plans. In this respect, the implementation form of the autonomous device 62 is not particularly limited, and it may be a device that moves in the real world such as a vehicle, robot, or drone, or a device that functions in data or a virtual world such as a data analysis device or a virtual assistant.
[0015] Therefore, among the components shown in the figure, elements other than the autonomous control device 10 can vary in various ways depending on the implementation of the autonomous device 62. For example, the autonomous device 62 may be equipped with only some of the information processing device 64, sensor 66, timer 68, actuator 70, and output device 72, or it may be equipped with another mechanism that acquires some kind of state information. Furthermore, at least some of the information processing device 64, sensor 66, timer 68, actuator 70, and output device 72 may be connected to each other and operate in cooperation.
[0016] Furthermore, the autonomous control device 10 may be part of the information processing device 64, or it may be an external device that is communicated with the autonomous device 62 via wired or wireless means. The information processing device 64 and the autonomous control device 10 may communicate with various servers via a network (not shown) to download data necessary for learning or to perform collaborative processing. As such, it will be understood by those skilled in the art that the autonomous device 62 can take on various configurations.
[0017] The autonomous control device 10 acquires information representing the internal state, physical state, and external environment of the autonomous device 62, and based on this, determines the motivation, sets objectives and tasks, and generates an action plan. Here, "motivation" refers to the reason "why" something is done, and "objective" refers to the final goal "what to do" based on the motivation. "Task" refers to the specific goals and work steps to reach the objective, and "action plan" is a structure that includes strategies and schedules on "how" to carry out the objective and tasks.
[0018] In the illustrated example, the autonomous control device 10 acquires information representing internal state, physical state, external environment, etc., from at least one of the information processing device 64, sensor 66, and timer 68. The information processing device 64 is a processor that performs information processing according to the original purpose of the autonomous device 62. The sensor 66 can be any of the general sensors that measure physical quantities inside or outside the autonomous device 62, such as a temperature sensor, image sensor, motion sensor, microphone, touch sensor, or GPS receiver. The timer 68 may measure the elapsed time since a certain state occurred, or it may acquire standard time.
[0019] Furthermore, "internal state" refers to the internal state based on the thoughts of the autonomous device 62, and examples include the type of emotion, the type of mood, the level of stress, and the level of motivation. This information may be generated by the information processing device 64 in its own processing, or it may be converted from measured values by sensors 66 and timers 68 according to predetermined rules.
[0020] "Physical state" refers to the state of the hardware, including the mechanical components and electrical elements that make up the autonomous device 62. Examples include operating status such as performance and usage rate, health status, temperature, cumulative usage time, and the presence or absence of wear and damage. This information is based on output data from the monitoring program executed by the information processing device 64, or measured values from sensors 66 and timers 68. "External environment" refers to the state of the real world in which the autonomous device 62 exists. Examples include temperature, humidity, illuminance, time of day, and noise level. This information is also based on output data from the monitoring program executed by the information processing device 64, or measured values from sensors 66 and timers 68.
[0021] The autonomous control device 10 determines motivation based on this information and controls the information processing device 64, actuator 70, and output device 72 to operate according to the corresponding action plan. The actuator 70 is a device that is installed in a predetermined part of the autonomous device 62 and drives that part. The output device 72 is a device that outputs data to be output, such as a display or speaker, in a predetermined format.
[0022] As a result, for example, the autonomous device 62 can enter a rest mode on its own when it becomes tired, or perform tasks or learn independently when its motivation increases. Due to such state transitions and the passage of time, the internal state, physical state, and external environment can change in various ways. Therefore, the autonomous control device 10 determines the next motivation in response to these changes and controls the autonomous device 62 with a corresponding action plan.
[0023] Generally, humans, who serve as references for artificial intelligence, can generate multiple motivations in parallel in their daily lives and take action based on the motivations they appropriately prioritize. If the autonomous device 62 can respond flexibly to diverse situations, similar to humans, its flexibility and versatility can be further enhanced. Therefore, in this embodiment, by separating the processing system for determining motivation into intrinsic motivation and functional motivation, it is possible to determine actions with appropriate priorities without complicating the model.
[0024] Figure 4 shows an example of the hardware configuration of the autonomous control device 10 in this embodiment. The autonomous control device 10 includes a CPU (Central Processing Unit) 52, main memory 54, storage device 56, and input / output interface 58. The CPU 52 controls the entire autonomous control device 10 by executing the operating system stored in the storage device 56. The CPU 52 also processes programs read from the storage device 56 and loaded into the main memory 54 to determine motivations based on various data, generate action plans, and control the autonomous device 62.
[0025] The storage device 56 stores programs and various data necessary for processing. This data may include a neural network that generates motivation from input data. The main memory 54 is configured, for example, as RAM (Random Access Memory) and temporarily stores data stored in the storage device 56, as well as intermediate data used for processing. The input / output interface 58 is connected to, for example, a communication interface such as a wired or wireless LAN, an output interface that outputs control signals necessary for task execution, and an input interface that receives state information that forms the basis of motivation.
[0026] Figure 5 shows the configuration of the functional blocks of the autonomous control device 10 in this embodiment. In this figure, each element described as a functional block that performs various processes can be configured in hardware terms with the various hardware shown in Figure 4, and in software terms with a program loaded from the storage device 56 into the main memory 54. Therefore, it will be understood by those skilled in the art that these functional blocks can be implemented in various ways using hardware alone, software alone, or a combination thereof, and are not limited to any one of these.
[0027] The autonomous control device 10 includes an internal state information acquisition unit 12 that acquires information relating to the internal state of the autonomous device 62, a physical state information acquisition unit 14 that acquires information relating to the physical state of the autonomous device 62, an external environment information acquisition unit 16 that acquires information relating to the external environment of the autonomous device 62, an intrinsic motivation generation unit 18 that generates intrinsic motivation, a functional motivation generation unit 20 that generates functional motivation, a motivation selection unit 22 that selects a motivation to adopt, a personality acquisition unit 24 that acquires characteristics unique to the autonomous device 62, and a constraint information acquisition unit 26 that acquires information relating to the constraints of motivation.
[0028] The autonomous control device 10 further includes a purpose setting unit 28 for setting objectives based on selected motivations, a task setting unit 30 for setting tasks according to the objectives, an action plan generation unit 32 for creating an action plan based on the tasks, and an output unit 34 for outputting information related to the action plan. Note that the configuration of the functional blocks shown in the figure is merely an example and is not intended to be mandatory. The autonomous control device 10 may also start or stop autonomous control in response to request signals from outside the autonomous control device 10, such as the information processing device 64 of the autonomous device 62. In this case, depending on the request signal, the autonomous control device 10 may switch between operating all functions or some functions.
[0029] The internal state information acquisition unit 12 acquires parameters and status representing the internal state of the autonomous device 62. The internal state information acquisition unit 12 may acquire the internal state determined by the information processing device 64 or the like of the autonomous device 62, or it may generate internal information itself based on information from sensors 66 or the like. The internal state information acquisition unit 12 acquires or updates information related to one or more internal states at predetermined time intervals or irregularly.
[0030] The intrinsic motivation generation unit 18 generates one or more intrinsic motivations at predetermined time intervals or irregularly, based on the latest information acquired by the internal state information acquisition unit 12. Here, intrinsic motivation refers to motivations stemming from internal interests, concerns, or a desire for self-improvement, such as "I want to improve my Japanese language scores" or "I want to learn more about space." The intrinsic motivation generation unit 18 is implemented, for example, by a machine learning model such as a neural network that generates intrinsic motivations from information related to the internal state.
[0031] The physical state information acquisition unit 14 acquires parameters and status representing the physical state of the autonomous device 62. For example, the physical state information acquisition unit 14 acquires information related to one or more physical states from sensors 66, timers 68, etc., at predetermined time intervals or irregularly. The external environment information acquisition unit 16 acquires parameters and status representing the external environment of the autonomous device 62. For example, the external environment information acquisition unit 16 acquires information related to one or more external environments from sensors 66, timers 68, etc., at predetermined time intervals or irregularly. The physical state information acquisition unit 14 and the external environment information acquisition unit 16 may interpret the acquired parameters and status and convert them into a format suitable as data representing the physical state and external environment.
[0032] The functional motivation generation unit 20 generates one or more functional motivations at predetermined time intervals or irregularly, based on the latest information acquired by the physical state information acquisition unit 14 and the external environment information acquisition unit 16. Here, a functional motivation is a motivation for maintaining the normal functioning of the autonomous device 62, and examples include "because the CPU temperature is high," "because an abnormality has been detected in the memory," and "because the outside temperature is low." The functional motivation generation unit 20 is implemented, for example, by a rule-based program that generates functional motivations from information related to the physical state and the external environment.
[0033] The motivation selection unit 22 selects either an intrinsic motivation generated by the intrinsic motivation generation unit 18 or a functional motivation generated by the functional motivation generation unit 20 as the motivation to actually reflect in the movement of the autonomous device 62. The motivation selection unit 22 selects one of the latest intrinsic motivations and functional motivations at a predetermined timing, such as when the previous task is completed.
[0034] The motivation selection unit 22 may select motivations at predetermined time intervals. Alternatively, the motivation selection unit 22 may sequentially acquire intrinsic motivations and functional motivations from the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 at predetermined time intervals, assign priorities to the acquired intrinsic motivations and functional motivations, store them, and select the motivation with the highest priority at the time of completion of the previous task. The priority may be based on the importance of the motivation itself or its chronological relationship to the previous task. Priorities may also be set individually for intrinsic motivations and extrinsic motivations.
[0035] The motivation selection unit 22 differentiates the behavior of priority weighting between intrinsic motivation and functional motivation in selecting motivations. Qualitatively, it assigns weights according to a rule that tends to prioritize functional motivation over intrinsic motivation. This is because, just as survival is paramount for many organisms, the continuation of function is also important for the autonomous device 62. For example, the motivation selection unit 22 acquires information related to the physical state and external environment obtained by the functional motivation generation unit 20, and estimates the level of failure risk of the autonomous device 62 based on at least one of these.
[0036] The functional motivation generation unit 20 assigns weights to each generated motivation such that the priority of functional motivations increases relatively as the risk of failure of the autonomous device 62 increases, and the priority of intrinsic motivations increases relatively as the risk of failure decreases. However, the perspective from which weights are assigned to motivations is not limited to the risk of failure, and weights may be adjusted from multiple perspectives. The functional motivation generation unit 20 ultimately selects the motivation to which the maximum weight is assigned.
[0037] The personality acquisition unit 24 acquires the personality, which is a characteristic unique to the autonomous device 62 that can influence the selection of motivations. Approaches proposed in the field of psychology can be used for the representation of personality. The personality acquisition unit 24 may internally store and use pre-set personality information as a fixed value, or it may update the personality in accordance with the growth of the autonomous device 62.
[0038] Alternatively, the personality acquisition unit 24 may acquire personality information from the information processing device 64 of the autonomous device 62, etc. The motivation selection unit 22 calculates the weights to be assigned to each motivation based on the personality information acquired from the personality acquisition unit 24, so that personality is reflected in motivation selection.
[0039] The rules for calculating the weights of motivations based on personality information are controlled based on general knowledge about personality and thinking, with reference to humans. For example, if a person has a strong desire for self-improvement, the weight of motivations such as "I want to improve my scores" will be increased, while if a person has a desire to maintain the status quo, they will aim to avoid changing their behavioral patterns as much as possible, so the weight of motivations such as "I work long hours" or "It's a frequently chosen action" will be increased. The motivation selection unit 22 may generate the weights to be assigned to each motivation from the personality information using machine learning models such as neural networks, rule-based programs, or other algorithms.
[0040] The constraint information acquisition unit 26 acquires information related to constraints imposed on motivations. Here, constraint information refers to information that specifies motivations that are inappropriate to select under certain conditions. For example, in situations where the generation of sound or light is undesirable, such as in a bedroom at night, a constraint is imposed to exclude motivations that lead to such tasks from the selection criteria. The basis for the constraints can be from various perspectives, such as physical safety, mental safety, and ethics.
[0041] In other words, the conditions for imposing constraints are not limited to location or time, but may be set on any of the various parameters included in the information of the internal state, physical state, or external environment. The information related to the constraint may specify a motive that is always inappropriate, regardless of the specific conditions. A motive that is always inappropriate is, for example, a motive that leads to direct harm to humans, and can be determined regardless of the parameters included in the information of the internal state, physical state, or external environment.
[0042] The constraint information acquisition unit 26 may internally hold information related to preset constraints and use them fixedly, or may update the information based on a setting update request from a user. Alternatively, the constraint information acquisition unit 26 may acquire information related to constraints from an information processing device 64 of the autonomous device 62 or the like. The motivation selection unit 22 acquires information related to constraints from the constraint information acquisition unit 26, and when the actual situation satisfies the constraint condition, calculates a weight to be applied to the motivation so that the correspondingly specified motivation is not selected.
[0043] As a very simple example, the motivation selection unit 22 calculates a weight W(M) for the motivation M as follows. W(M) = (w1(M) + w2(M)) · w3(M) Here, w1(M) is a weight whose behavior changes depending on whether the motivation M is an intrinsic motivation or a functional motivation. For example, when the motivation M is a functional motivation, w1(M) is increased as the failure risk of the autonomous device 62 increases. When the motivation M is an intrinsic motivation, w1(M) can be set to a constant value regardless of the failure risk, or w1(M) can be decreased as the failure risk increases.
[0044] w2(M) is a weight based on the personality acquired by the personality acquisition unit 24, and is a unique value corresponding to the motivation M itself or the type of the motivation M. w3(M) is a weight based on information related to constraints acquired by the constraint information acquisition unit 26, and is "1" if the motivation M is selectable, and "0" if it is not selectable. Alternatively, as the information related to constraints, the degree of inappropriateness of selection under a certain condition may be set for each motivation, and the weight w3(M) may be adjusted with a value between 0 and 1 in accordance with the degree.
[0045] Note that the above formula merely simplifies the concept of motivation selection, and in practice, the motivation selection unit 22 may select a motivation by any of a machine learning model such as a neural network, a rule-based program, or other algorithms. In any case, in the present embodiment, by generating motivations separately into intrinsic motivations and functional motivations, the policy for motivation selection becomes clear, and even if the motivations themselves are diversified, appropriate autonomous control can be performed with a relatively simple model. In addition, adjustment of the selection policy can also be easily performed from various perspectives, so an appropriate motivation can be flexibly selected according to the situation.
[0046] The objective setting unit 28 sets an objective based on the motivation selected by the motivation selection unit 22. The task setting unit 30 sets a task in accordance with the set objective. The action plan generation unit 32 generates an action plan in accordance with the set task. Numerous proposals have already been made for objective and task setting and action plan generation based on a selected motivation as approaches for autonomous learning and autonomous behavior, or as computer control methods. Any of these may be employed in the present embodiment, so a description thereof is omitted here.
[0047] The output unit 34 outputs information on the action plan generated by the action plan generation unit 32 to the outside. For example, the output unit 34 outputs a control signal for realizing the action plan to the actuator 70 of a robot which is the autonomous device 62, or to the information processing device 64 which is the control mechanism thereof. Those skilled in the art will understand that the format and output destination of information output by the output unit 34 can vary in many ways depending on the aspect of the autonomous device 62 and the action plan itself.
[0048] Note that FIG. 5 shows an example in which the objective setting unit 28 acquires the selected motivation, sets an objective, and performs task setting and action plan generation in accordance therewith, but the present embodiment is not limited to this example. For example, motivation information may be received from the motivation selection unit 22, and only one or two of the objective setting unit 28, the task setting unit 30, and the action plan generation unit 32 may operate.
[0049] Alternatively, the objective setting unit 28, the task setting unit 30, and the action plan generation unit 32 may each implement their respective functions in parallel based on the selected motivation. Naturally, depending on these aspects, the information output by the output unit 34 and the output destination may also vary. The objective setting unit 28, the task setting unit 30, the action plan generation unit 32, and the output unit 34 can also be collectively referred to as a "control processing unit" from the perspective of controlling the actual behavior of an autonomous device.
[0050] Figure 6 illustrates the data flow related to intrinsic motivation. In this example, the intrinsic motivation generation unit 18 is implemented using a neural network 42. In this scenario, the "psychological parameters" 40a, 40b, and "intrinsic motivation categories" 44, shown as dashed rectangles in the figure, are also the training data used to construct the neural network 42. The "psychological parameters" 40a and 40b are psychological parameters corresponding to the internal states described above. The "psychological parameters" 40a and 40b can be data that quantifies various psychologically definable states, such as emotional parameters IA (joy, anger, sadness, etc.) and mood parameters IB (exhilaration, depression, etc.).
[0051] In constructing the neural network 42, supervised learning is performed using a dataset of psychological parameters and intrinsic motivations. For example, pairs of psychological parameters and the intrinsic motivations that arose from them can be automatically extracted from a large amount of text data and used as training data. Alternatively, a questionnaire survey could be conducted asking people about the psychological states that led to certain motivations based on their experiences, and the results could be used as training data.
[0052] By constructing the neural network 42 in this manner, an intrinsic motivation generation unit 18 can be realized that outputs a corresponding intrinsic motivation Mi in response to the input of one or more psychological parameters. Although only one intrinsic motivation Mi is shown in the figure, the intrinsic motivation generation unit 18 may output multiple intrinsic motivation Mi in parallel. The figure also shows that the internal state information Ia, Ib, ... acquired by the internal state information acquisition unit 12 corresponds to the psychological parameters IA, IB, .... In other words, the operation of the autonomous control device 10 is to generate intrinsic motivation Mi by inputting the internal state information Ia, Ib, ... acquired by the internal state information acquisition unit 12 to the intrinsic motivation generation unit 18 instead of the psychological parameters IA, IB, ....
[0053] To accurately set objectives and tasks from a diverse range of intrinsic motivations (Mi) and to appropriately control the autonomous device 62 autonomously, it is desirable to adaptively switch the execution model according to the category of intrinsic motivation (Mi). For example, in the field of reinforcement learning, the optimal approach differs depending on whether the focus is on curiosity (see, for example, Deepak Pathak et al., "Curiosity-driven Exploration by Self-supervised Prediction," Proceedings of the 34th International Conference on Machine Learning, 2017, Vol. 70, pp. 2778-2787) or on risk (see, for example, Yun Shen et al., "Risk-Sensitive Reinforcement Learning," Neural Computation, 2014, Vol. 26, No. 7, pp. 1298-1328).
[0054] Preferably, the neural network 42 is configured to output intrinsic motivation Mi in association with the category to which it belongs. Here, the categories of intrinsic motivation include, for example, "curiosity," "stimulus pursuit," "self-growth," and "desire for achievement." If, for example, an intrinsic motivation classified as "curiosity" is selected through such classification, at least one of the objective setting unit 28, task setting unit 30, and action plan generation unit 32 can perform each process using an approach suitable for curiosity. In this case, for example, the method proposed by Pathak et al. mentioned above can be used.
[0055] When an intrinsic motivation classified as "stimulus-seeking" is generated, at least one of the objective setting unit 28, task setting unit 30, and action plan generation unit 32 can perform each process using an approach suitable for stimulus-seeking. In this case, for example, the risk-seeking type can be selected and used in the method proposed by Shen et al. mentioned above. Similarly, for other categories, by selecting and using an appropriate approach or model for each, an efficient and highly accurate action plan can be created. Note that the categories of intrinsic motivations shown in the diagram are examples and may be changed as appropriate depending on the approach or model available.
[0056] When classifying intrinsic motivations in this way, when constructing the neural network 42, the psychological parameters and intrinsic motivation datasets are further associated with the categories to which each intrinsic motivation belongs, and these are used as training data. As a result, the intrinsic motivation generation unit 18 can take internal states Ia, Ib, ... as input and output intrinsic motivations Mi and their categories. However, as mentioned above, the intrinsic motivation generation unit 18 is not limited to a neural network; it may also be implemented using functions that represent the correlation between inputs and outputs, or rule-based programs.
[0057] Figure 7 illustrates the data flow related to functional motivation. In this example, the functional motivation generation unit 20 is implemented using a rule-based program 44. The rule-based program 44 is a program that takes at least one of the following as input: information on physical states, such as the state of the hardware, or information on the external environment, such as temperature and humidity, and outputs functional motivation Mf. There are no limitations on the types or number of physical state or external environment information that can be input to the rule-based program 44 in parallel, nor are there any limitations on the number of functional motivation Mf that can be output at one time.
[0058] The functional motivation generation unit 20 may also output the category to which the determined functional motivation Mf belongs, based on the rule-based program 44. For example, the functional motivation generation unit 20 classifies the functional motivation Mf into categories such as "risk of malfunction," "stability / safety," and "energy efficiency" based on specific numerical values of information about the physical state and external environment. By classifying the functional motivation Mf in this way, just like with intrinsic motivation Mi, it is possible to create efficient and highly accurate action plans using methods appropriate to each category. That is, at least one of the objective setting unit 28, the task setting unit 30, and the action plan generation unit 32 can perform their respective processes using methods corresponding to the category of the selected functional motivation Mf.
[0059] The reason for implementing the functional motivation generation unit 20 as a rule-based program 44 is that, compared to the intrinsic motivation system, it is relatively easier to obtain a clear solution for the functional motivation system. However, this does not mean that the implementation of the functional motivation generation unit 20 is limited to a rule-based program; it can also be a machine learning model such as a neural network, or a function that represents the correlation between input and output. In any case, according to this embodiment, the implementation forms of the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 can be set independently of each other, taking into account the differences in the characteristics of intrinsic motivation and functional motivation. Furthermore, by further classifying each motivation, the approach to action planning can be adjusted at a finer level of granularity.
[0060] Next, the operation of the autonomous control device that can be realized with the above configuration will be described. Figure 8 is a flowchart showing the processing procedure in which the autonomous control device 10 of this embodiment generates and selects motivation and outputs information related to the action plan. This flowchart is started, for example, when a signal requesting the start of control is received from the autonomous device 62. First, the internal state information acquisition unit 12, the physical state information acquisition unit 14, and the external environment information acquisition unit 16 start acquiring information to be used as the basis for motivation (S10). The internal state information acquisition unit 12, the physical state information acquisition unit 14, and the external environment information acquisition unit 16 each acquire predetermined data that they are responsible for, either periodically or irregularly.
[0061] Next, the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 begin generating intrinsic motivation and functional motivation, respectively (S12). In the example described above, the intrinsic motivation generation unit 18 generates intrinsic motivation using a neural network or the like based on information about the internal state. The functional motivation generation unit 20 generates functional motivation using a rule-based program or the like based on at least one of information about the physical state and the external environment. The generation cycles for these may also be constant or variable. Preferably, the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 output the category to which the generated motivation belongs along with the generated motivation.
[0062] The intrinsic motivation generation unit 18 and the functional motivation generation unit 20 transmit information related to the generated motivations to the motivation selection unit 22 as needed. The motivation selection unit 22 stores the transmitted motivation information in memory (not shown), associating it with a timestamp representing the transmission time. The temporal characteristics of motivations vary depending on their content, such as those that are valid for a relatively long period or those whose underlying parameters change moment by moment. Therefore, the motivation selection unit 22, for example, retains currently valid motivations as selection candidates based on the expiration date corresponding to the motivation.
[0063] The motivation-related information transmitted by the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 may include specific numerical values of the internal state, physical state, and external environment that form the basis of the motivation, such as CPU temperature and cumulative usage time. When the intrinsic motivation generation unit 18 and the functional motivation generation unit 20 transmit this information, the motivation selection unit 22 may be configured to update the numerical values in the motivation-related information it holds internally to the latest values. This allows the motivation selection unit 22 to detect changes in the risk of system failure, and can reflect this, for example, in the calculation of the weight W(M) for the motivation M described above.
[0064] The motivation selection unit 22 selects a motivation to be adopted at that time from the candidate motivations stored in memory (S14). The motivation selection unit 22 basically selects one motivation at each selection timing. However, this embodiment is not limited to this, and multiple motivations may be selected depending on the type of motivation. In any case, as described above, the motivation selection unit 22 selects motivations according to a predetermined policy, such as prioritizing functional motivations over intrinsic motivations. In this case, the motivation selection unit 22 may refer to the personality and constraint information of the autonomous device 62 and reflect it in the selection process.
[0065] Next, the objective setting unit 28 sets an objective based on the selected motivation, and the task setting unit 30 sets tasks according to the objective (S16). Then the action plan generation unit 32 generates an action plan according to the tasks (S18). As described above, these processes are not limited to being performed in series; some may be omitted, or multiple processes may be performed in parallel. The output unit 34 generates and outputs control signals, etc., so that the autonomous device 62 operates according to the action plan (S20).
[0066] The autonomous control device 10 waits for the processing of S14 to S20 to begin while the task corresponding to the output in S20 is being performed by the autonomous device 62 (S22, N). However, during this time, the acquisition of state information and information related to the external environment, which was started in S10, and the generation of motivation, which was started in S12, may continue. When the completion of the task is detected by receiving a signal from the autonomous device 62 (S22, Y), and there is no need to terminate the control at that point (S24, N), the motivation selection unit 22 selects a new motivation to be adopted at that point from the candidate motivations stored in memory (S14).
[0067] The objective setting unit 28, task setting unit 30, and action plan generation unit 32 each set the objective, set the task, and generate the action plan based on the newly selected motivation, and the output unit 34 generates and outputs control signals etc. that reflect this (S16-S20). Thereafter, each time a task is completed in the autonomous device 62, the autonomous control device 10 repeats the processing in S14-S20. When it becomes necessary to terminate control due to a request signal from the autonomous device 62 or other reasons, the autonomous control device 10 terminates all processing (Y in S24).
[0068] According to the embodiment described above, an autonomous control device that controls an autonomous device collects various information and generates corresponding motivations. This allows for adaptive motivation in response to internal and external situations that may change over time. Furthermore, by separating the intrinsic motivation system from the functional motivation system, the models, approaches, and implementation forms used for generating and selecting motivations can be optimized according to their characteristics. As a result, adaptive motivation can be easily achieved.
[0069] Furthermore, the relationships involved in the generation of motivations from various types of information can be simplified relatively. As a result, even if the input information and acceptable motivations become more diverse, the ease, flexibility, and interpretability of modeling the data flow, such as the generation and selection of motivations, the setting of objectives and tasks, and the generation of action plans, can be improved. In other words, autonomous devices can be brought closer to the everyday human situation of selecting important motivations from a variety of motivations and taking corresponding actions.
[0070] The present invention has been described above based on embodiments. The above embodiments are illustrative, and it will be understood by those skilled in the art that various modifications are possible in combinations of their respective components and processing processes, and that such modifications also fall within the scope of the present invention.
[0071] The present invention can be used in an autonomous control device and an autonomous control method for performing autonomous control of a device.
[0072] 10 Autonomous control device, 12 Internal state information acquisition unit, 14 Physical state information acquisition unit, 16 External environment information acquisition unit, 18 Intrinsic motivation generation unit, 20 Functional motivation generation unit, 22 Motivation selection unit, 24 Personality acquisition unit, 26 Constraint information acquisition unit, 28 Objective setting unit, 30 Task setting unit, 32 Action plan generation unit, 34 Output unit, 62 Autonomous device.
Claims
1. An autonomous control device comprising: an intrinsic motivation generation unit that generates intrinsic motivation based on an internal state representing the internal state based on the thinking of an autonomous device; a functional motivation generation unit that generates functional motivation based on at least one of the hardware state of the autonomous device and information on the external environment; a motivation selection unit that selects a motivation to be adopted based on the priority weights assigned to the intrinsic motivation and the functional motivation by different rules; and a control processing unit that controls the behavior of the autonomous device based on the selected motivation.
2. The autonomous control device according to claim 1, characterized in that the motivation selection unit estimates the failure risk of the autonomous device using at least one of the hardware state and information on the external environment, and assigns weights according to a rule that the higher the failure risk, the higher the priority of the functional motivation.
3. The autonomous control device according to claim 1 or 2, characterized in that the motivation selection unit changes the weight according to pre-set information relating to selection constraints.
4. The autonomous control device according to claim 1 or 2, characterized in that the motivation selection unit changes the weight according to personality information representing characteristics unique to the autonomous device.
5. The autonomous control device according to claim 1 or 2, characterized in that the intrinsic motivation generation unit generates the intrinsic motivation using a neural network, and the functional motivation generation unit generates the functional motivation using a rule-based program.
6. The autonomous control device according to claim 5, wherein the neural network is constructed to output the intrinsic motivation and the category to which it belongs by inputting the internal state, and when the intrinsic motivation is selected, the control processing unit determines the content of the control using a method corresponding to the category.
7. An autonomous control method comprising: generating intrinsic motivation based on an internal state representing the internal state based on the thinking of an autonomous device; generating functional motivation based on at least one of the hardware state of the autonomous device and information on the external environment; selecting a motivation to be adopted based on the weight of priority assigned to the intrinsic motivation and the functional motivation by different rules; and controlling the behavior of the autonomous device based on the selected motivation.
8. A computer program that enables a computer to implement: a function to generate intrinsic motivation based on an internal state representing the internal state based on the thinking of an autonomous device; a function to generate functional motivation based on at least one of the hardware state of the autonomous device and information about the external environment; a function to select a motivation to adopt based on the weight of priority assigned to the intrinsic motivation and the functional motivation by different rules; and a function to control the behavior of the autonomous device based on the selected motivation.