Learning system, learning method, and learning program

WO2026167897A1PCT designated stage Publication Date: 2026-08-13RAPIDUS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2026-08-13

Smart Images

  • Figure JP2025020097_13082026_PF_FP_ABST
    Figure JP2025020097_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a learning system for a machine learning model provided in an autonomous robot that achieves symbiosis with humans. The learning system comprises at least an extraction unit, an interpretation unit, and a learning unit. A causal relationship group including a cause, which is an action selected by an agent, and the result corresponding to the cause is stored, and the extraction unit extracts, from an environment, the feature amount corresponding to each of the one or more results constituting the causal relationship group. The interpretation unit estimates an ethical interpretation of each of the one or more results on the basis of the feature amount, and specifies an ethical result in the causal relationship group. The learning unit performs reinforcement training processing of a determination model by setting a reward for the ethical result and / or selection of the action corresponding to the cause related to the result.
Need to check novelty before this filing date? Find Prior Art

Description

Learning System, Learning Method, and Learning Program

[0001] The present invention relates to a learning system, particularly to a learning system for a machine learning model provided in an autonomous robot that realizes symbiosis with humans.

[0002] In recent years, autonomous robots are being used in various scenarios to reduce the labor and work performed by humans.

[0003] As an example of the usage scenario of this autonomous robot, Patent Document 1 discloses a humanoid robot for automatically performing work in a factory. AI (Artificial Intelligence) is used for controlling movements and arms of the humanoid robot.

[0004] Further, Patent Document 2 discloses a humanoid robot used in an elderly care facility. This humanoid robot has a data generation model and an emotion identification model. For example, when a care recipient shows anger, the humanoid robot can respond calmly and make proposals to eliminate the dissatisfaction of the care recipient.

[0005] Japanese Patent Application Laid-Open No. 2025-36942 Japanese Patent Application Laid-Open No. 2025-56024

[0006] Thus, the movement aiming at symbiosis between humans and autonomous robots has already started. On the other hand, with the advent of so-called singularity, autonomous robots equipped with AI (Artificial Intelligence) are predicted to exceed human capabilities with the development of AI technology. As described above, the humanoid robot can make appropriate responses according to human emotions, but it is still assumed that the humanoid robot operates according to the commands intended by humans.

[0007] In order for an autonomous robot equipped with AI (that is, an autonomous robot equipped with various learned models) and humans to coexist, a trust relationship between the autonomous robot and humans is essential, and it is necessary to introduce the ethical behavior norms that humans have into the autonomous robot. However, the environments in which autonomous robots operate are diverse, and it is difficult to reflect an ethical interpretation according to the environment in which the robot operates in the construction of the machine learning model.

[0008] The present disclosure has been made in view of the above-described problems of the prior art, and an object thereof is to provide a learning system that effectively causes a learning model included in an entity such as an autonomous robot to learn an ethical interpretation.

[0009] A learning system including at least an extraction unit, an interpretation unit, and a learning unit, storing a causal relationship group including a cause that is an action selected by an agent and a result corresponding to the cause, the extraction unit extracting a feature amount corresponding to each of one or more of the results constituting the causal relationship group from an environment, the interpretation unit estimating an ethical interpretation of each of the one or more results based on the feature amount and specifying an ethical result in the causal relationship group, and the learning unit setting a reward for selection of the action corresponding to the ethical result and / or a cause related to this result and performing reinforcement learning processing of a determination model.

[0010] A learning method executed by one or more computers, at least one of the one or more computers storing a causal relationship group including a cause that is an action selected by an agent and a result corresponding to the cause, extracting a feature amount corresponding to each of one or more of the results constituting the causal relationship group from an environment, at least one of the one or more computers estimating an ethical interpretation of each of the one or more results based on the feature amount and specifying an ethical result in the causal relationship group, and at least one of the one or more computers setting a reward for selection of the action corresponding to the ethical result and / or a cause related to this result and performing reinforcement learning processing of a determination model.

[0011] Furthermore, the present disclosure provides a learning program executed by one or more computers, wherein at least one of the one or more computers, which stores a group of causal relationships including causes that are actions selected by an agent and results corresponding to those causes, functions as an extraction unit that extracts feature quantities from the environment corresponding to each of the one or more results constituting the group of causal relationships; at least one of the one or more computers functions as an interpretation unit that estimates an ethical interpretation of each of the one or more results based on the feature quantities and identifies an ethical outcome in the group of causal relationships; and at least one of the one or more computers functions as a learning unit that sets a reward for the selection of the ethical outcome and / or the cause related to that outcome and performs reinforcement learning processing of a decision model.

[0012] According to this disclosure, a novel technology for a learning system can be provided by performing reinforcement learning processing related to ethical interpretation using a predetermined set of causal relationships.

[0013] This document shows a schematic diagram of a human-autonomous robot in light of the singularity. This document shows a schematic diagram of symbiosis between a human-autonomous robot according to an embodiment of this disclosure. This document shows another schematic diagram of symbiosis between a human-autonomous robot according to an embodiment of this disclosure. This document shows a schematic diagram of an autonomous robot according to an embodiment of this disclosure. This document shows a schematic diagram of the structure of an enclave according to an embodiment of this disclosure. This document shows a block diagram of the system configuration in an autonomous robot according to an embodiment of this disclosure. This document shows a schematic diagram of the reinforcement learning process of a decision model according to an embodiment of this disclosure. This document shows a conceptual diagram of the causal relationship group used for learning according to an embodiment of this disclosure. This document shows a flowchart of the processing procedure in the reinforcement learning process of a virtual environment according to an embodiment of this disclosure. This document shows a block diagram of the system configuration according to an embodiment of this disclosure. This document shows a flowchart of the processing procedure in the behavior of an autonomous robot according to an embodiment of this disclosure. This document shows a functional block diagram of the main system and subsystem (enclave) according to an embodiment of this disclosure. This document shows a schematic diagram of the transformation, combination, and extrapolation of the causal relationship group according to an embodiment of this disclosure. This document shows an example of updating the causal relationship group according to an embodiment of this disclosure. This document shows an example of processing when acquiring external environment data from a human (entity) according to an embodiment of this disclosure. This document shows a flowchart of the authentication process using CRP by PUF according to an embodiment of this disclosure. This document shows a schematic diagram of the upload and download of causal relationship groups according to an embodiment of this disclosure. This document shows a conceptual diagram of a blockchain that holds ethical data according to an embodiment of this disclosure. This document shows an example of a self-contained BCN according to an embodiment of this disclosure. This document shows an example of a BCN according to an embodiment of this disclosure. This document shows a schematic diagram of encryption using PUF according to an embodiment of this disclosure. This document shows a flowchart of the procedure for generating a private key by PUF according to an embodiment of this disclosure. This document shows a flowchart of the procedure for generating an encryption key by PUF according to an embodiment of this disclosure. This document shows an example of a 2D package semiconductor device according to an embodiment of this disclosure. This document shows an example of a 2.5D package semiconductor device according to an embodiment of this disclosure. This document shows an example of a 2.xD package semiconductor device according to an embodiment of this disclosure. This document shows an example of a 3D package semiconductor device according to a second embodiment of this disclosure.

[0014] Further details will be provided below with reference to the attached drawings. The drawings show preferred embodiments. However, it is possible to carry out the invention in many different forms and is not limited to the embodiments described herein.

[0015] Figure 1(a) of this disclosure shows a schematic image of a human and an autonomous robot. As shown in the figure, autonomous robots that have acquired advanced intelligence and certain autonomous movements will work together with humans in manufacturing plants, medical facilities, and other settings. For example, the physical and sensory capabilities of autonomous robots can be more advanced than those of humans, such as carrying heavy loads, detecting odors and taking temperatures, and having hearing beyond the range of human hearing. After the singularity, the intelligence of autonomous robots will far exceed human capabilities with the development of AI (artificial intelligence) and AGI (artificial general intelligence), and the same will be true for physical aspects such as physical and sensory capabilities.

[0016] In this context, trust must be built for humans and autonomous robots (entities) to "coexist." As shown in Figure 1(b), if trust cannot be built, it cannot be denied that "conflict" may arise between humans and autonomous robots. In such situations, for humans and autonomous robots to coexist while maintaining mutual trust, it is necessary to input ethical behavioral norms into the intelligence of autonomous robots, just as humans do.

[0017] For example, if an autonomous robot deviates from a predetermined code of conduct, including ethical conduct, it is preferable to be able to suppress (forcibly stop) the autonomous robot's actions, prioritizing human safety above all else.

[0018] Furthermore, the codes of conduct that entities such as autonomous robots must adhere to will actually vary depending on the situation and context. From an ethical standpoint, it is a given that autonomous robots should not harm humans in any situation, but in the context of medical procedures, there are actions that should be exceptionally permitted, such as surgery and punctures. In another example, while it is acceptable for autonomous robots to move quickly in general transport tasks, in cleanrooms such as semiconductor manufacturing plants, the very act of an autonomous robot moving quickly may be considered a deviation from the code of conduct from the perspective of maintaining cleanliness and avoiding contact with humans.

[0019] This disclosure describes a distinctive learning system that can appropriately reflect ethical interpretations to prevent autonomous robots from performing inappropriate behavioral actions in various situations and circumstances. The specific embodiments of the autonomous robot described in this disclosure are described in detail below.

[0020] Figure 2 shows a schematic diagram of the symbiosis between a human and an autonomous robot according to an embodiment of the disclosure. Figure 2 is a schematic diagram of a cleanroom in a large-scale semiconductor manufacturing plant. As shown in the figure, an autonomous robot 10 and a human 20 are working together in a cleanroom environment in a large-scale semiconductor manufacturing plant.

[0021] Figure 2 shows semiconductor wafers being automatically transported on the ceiling of the cleanroom, and, when necessary, the autonomous robot 10 firmly grasps and carries the FOUP (Front Opening Unified Pod) containing the semiconductor wafers with both hands.

[0022] Furthermore, as shown in Figure 3, the autonomous robot 10 in this embodiment can coexist with humans 20 not only in megafabs (large-scale semiconductor manufacturing plants) but also in environments such as small- to medium-sized semiconductor manufacturing plants and semiconductor analysis facilities. Figure 3 shows how the autonomous robot 10 and humans 20 are each playing different roles. In scenarios where the autonomous robot 10 and humans 20 are coexisting, as shown in Figures 2 and 3, the safety of humans 20 must be given top priority, and the autonomous robot 10 must also have the same ethical considerations as humans 20.

[0023] In this embodiment, we describe an autonomous robot 10 coexisting with a human 20 in a semiconductor manufacturing plant. However, the environment in which the autonomous robot 10 coexists with the human 20 is not limited to a semiconductor manufacturing plant, but a wide variety of environments can be envisioned. It is necessary to appropriately teach the autonomous robot ethical interpretations according to such a wide variety of environments.

[0024] For example, the autonomous robot 10 according to this embodiment can learn ethical interpretations to realize coexistence with humans in various environments where people gather, such as residential facilities like detached houses and apartment buildings, medical facilities and medical sites like hospitals and clinics, construction sites corresponding to general construction or specialized construction, disaster sites and accident sites corresponding to natural or man-made disasters, nursing care facilities like nursing homes, childcare facilities like nurseries and daycare centers, educational facilities like schools and training facilities, cultural facilities like museums, art galleries, stadiums, libraries, theaters, and concert halls, religious facilities like shrines, temples, and churches, public facilities like roads and parks, administrative facilities like government offices and tax offices, commercial facilities and service facilities like shops, transportation facilities like train stations and airports, business facilities like factories, workshops, warehouses, offices, and other offices, telecommunications facilities like data centers, development facilities, experimental facilities, research facilities, and many others. Furthermore, for example, the autonomous robot 10 according to this embodiment is capable of learning ethical interpretations for achieving coexistence with humans in various environments such as indoors, outdoors, at high altitudes, on the ground, underground, in the air, in the mountains, underwater, on the water, and on ice.

[0025] Figure 4 shows schematic diagrams of a typical autonomous robot and the autonomous robot according to this embodiment. Figure 4(a) shows a typical autonomous robot. The autonomous robot 10a includes a first processing chip (first processing unit) that performs its own control, a main memory, an auxiliary memory, sensors, actuators, etc.

[0026] Figure 4(b) shows the autonomous robot 10 according to this embodiment, which, in addition to the above, includes an enclave 200 that includes a second processing chip (second processing unit) and a dedicated auxiliary storage device for determining whether the behavioral actions of the autonomous robot 10 are ethical actions that conform to ethical behavioral norms. In this embodiment, as shown in the figure, the enclave 200 is configured to have a dedicated area (dedicated main memory unit) provided in the main memory as memory for deploying a determination model for determination, etc. However, for example, the dedicated main memory may be connected via a dedicated bus interface within the enclave 200, similar to the dedicated auxiliary storage device. The autonomous robot 10 can achieve symbiosis with a human 20 through this characteristic enclave 200.

[0027] Figure 5 shows a schematic diagram of the enclave according to this embodiment. As shown in Figure 5, for example, the enclave 200 is composed of a second processing chip, a dedicated auxiliary storage device, an AI / ML (Machine Learning) accelerator, an encryption accelerator, a dedicated boot ROM (Read Only Memory), memory protection, and the like.

[0028] The enclave 200 is implemented as a subsystem in the autonomous robot and is isolated as a control system independent of the main system, which includes the first processing chip and moving parts (e.g., joints in the autonomous robot). Furthermore, the enclave (subsystem) is isolated as a circuit system independent of the main system, at least in part, including the second processing chip. In other words, the enclave (subsystem) is logically or physically isolated from the main system. The following describes each of these configurations in detail.

[0029] Figure 6 of the autonomous robot system configuration shows a block diagram of the system configuration in the autonomous robot according to this embodiment. Figure 6 represents the relationship between the main system and subsystems in Figure 5 described above as a block diagram. As shown in the figure, the main system 100 is composed of a first processing chip 110, main memory 120, auxiliary storage 130, sensor 140, actuator 150, etc. The first processing chip 110, main memory 120, auxiliary storage 130, sensor 140, and actuator 150 transmit and receive various signals via a common bus interface 310.

[0030] Furthermore, subsystem 200 is composed of a second processing chip 210, a dedicated area 220 (dedicated main memory), a dedicated auxiliary storage device 230, an AI / ML accelerator 240, an encryption accelerator 250, a dedicated boot ROM 260, memory protection 270, and the like. subsystem 200 is an enclave (200 in Figures 4 and 5) that includes the second processing chip 210 for executing the characteristic judgment model and deviation judgment processing in this disclosure. The second processing chip 210, the dedicated area 220, the dedicated auxiliary storage device 230, the AI / ML accelerator 240, and the encryption accelerator 250 transmit and receive various signals via a dedicated bus interface 320.

[0031] The first processing chip 110 of the main system 100 functions as a processor that autonomously controls the behavioral movements of the autonomous robot 10. The first processing chip 110 provides external environment data acquired at least via the sensor 140 to a control model and generates first signals that perform control including the operation and stopping of the actuator 150. The environment data includes internal environment data and / or external environment data.

[0032] The main memory 120 serves as the main memory for the first processing chip 110 to execute each process. The auxiliary storage device 130 serves as storage for various data used by the first processing chip 110 to execute each process. The auxiliary storage device 130 also stores control models for autonomous control. For example, it can store environmental data such as internal environment data and external environment data acquired by the sensor 140 in chronological order.

[0033] Sensor 140 acts as a detector (detection unit) for acquiring environmental data at a predetermined location where the autonomous robot 10 is located. Sensor 140 includes an external sensor 141 for acquiring external environment data, which is environmental data outside the autonomous robot, and an internal sensor 142 for acquiring internal environment data, which is environmental data inside the autonomous robot. The external sensor 141 acquires external environment data provided by the external environment, which includes at least static or moving objects in the surroundings of the autonomous robot 10. The internal sensor acquires at least a portion of the internal environment data, which includes motion in active movements at parts of the autonomous robot 10 such as joints, as well as direction, tilt, position, displacement, velocity, angular velocity, acceleration, angles such as azimuth angle and rotation angle, current, and voltage.

[0034] For example, the sensor 140 preferably includes an image sensor for acquiring image data (including video data) of the external environment. Other external sensors 141 that can be used in this embodiment include sound sensors such as microphones for acquiring sounds of the external environment, distance sensors for measuring distance and shape, temperature sensors for acquiring the temperature of the external environment, odor sensors for acquiring smells of the external environment, pressure sensors for acquiring atmospheric pressure and pressure of the external environment, and humidity sensors, among others.

[0035] The detection unit in this embodiment can be configured to include one or any combination of various sensors, including: a sound sensor using an electrostatic sound sensor, an electrodynamic sound sensor, a piezoelectric sound sensor, etc.; a distance sensor using a ToF (Time-of-Flight) sensor, an FMCW (Frequency Modulated Continuous Wave) sensor, a TDOA (Time Difference of Arrival) sensor, an OPA (Optical Phased Array) sensor, etc.; a temperature sensor using a contact temperature sensor, a non-contact temperature sensor, etc.; an odor sensor using a semiconductor structure odor sensor, an electrochemical odor sensor, a galvanic cell odor sensor, an optical odor sensor, etc.; a weight sensor using a strain gauge type weight sensor, an electromagnetic weight sensor, a capacitive weight sensor, a piezoelectric weight sensor, etc.; a pressure sensor using a piezoresistive pressure sensor, a capacitive pressure sensor, a piezoelectric pressure sensor, etc.; and a humidity sensor using a capacitive humidity sensor, a resistive humidity sensor, a thermal conduction humidity sensor, an optical humidity sensor, etc.

[0036] Furthermore, as the sensor 140, it is also possible to use photoelectric sensors, laser sensors, infrared sensors, color sensors, image sensors, millimeter-wave sensors, microwave sensors, radar, proximity sensors, contact-type displacement sensors, proximity sensors, sound wave sensors, ultrasonic sensors, molecular sensors (for humidity, odor, etc.), position sensors using GNSS, orientation sensors, gyro sensors, force sensors, slip sensors, tactile sensors, speed sensors, angle sensors, angular velocity sensors, acceleration sensors, angular acceleration sensors, current sensors, voltage sensors, etc.

[0037] Furthermore, the autonomous robot 10 according to this embodiment may be equipped with multiple various sensors on the main system 100, or it may be equipped with multiple sensors in a configuration separate from the main system 100. One or any combination of these may be provided as an external sensor 141. One or any combination of these may be provided as an internal sensor 142.

[0038] The actuator 150 functions as an operating unit that performs the behavioral movements of the autonomous robot 10. The actuator 150 performs predetermined operations based on the first signal. The actuator 150 can also directly receive the second signal from the dedicated bus interface 320 shown in Figure 6.

[0039] The actuator 150, acting as a moving part, constitutes the joints of the autonomous robot 10 and is connected to the main body and arms of the autonomous robot 10. For example, the actuator 150 can also be configured to constitute the joints of the autonomous robot 10 and be connected to the head, shoulders, torso, arms, waist, and legs of the autonomous robot.

[0040] In terms of the relationship between the actuator 150 and the autonomous robot 10, the autonomous robot 10 may be configured to include a main body, a moving mechanism, an arm, and one or more joints connected to at least the main body and the arm, with each joint having an operating unit (actuator 150) that performs a kinetic movement as an action of the autonomous robot 10.

[0041] The actuator 150 according to this embodiment can be configured to include one or any combination of various actuators, including electric actuators using DC motors, AC motors, stepping motors, servo motors, etc.; magnetic actuators using electromagnet actuators, etc.; pneumatic actuators using pneumatic cylinders, etc.; hydraulic actuators using hydraulic cylinders, hydraulic motors, etc.; rotary actuators using servo motors, etc.; linear actuators using push mechanisms, etc.; and soft actuators using artificial muscles, dielectric elastomers, electroactive polymers, polymer actuators, shape memory alloy actuators, etc. Furthermore, the actuator 150 according to this embodiment may employ actuators with different configurations depending on each part constituting the autonomous robot 10. In addition, the actuator 150 according to this embodiment may be configured to function as at least a part of the sensor 140.

[0042] Furthermore, the autonomous robot 10 can also be configured to have either or both an audio output unit, such as a speaker, that outputs sound as an action, or a display output unit, such as a monitor, that outputs video as an action.

[0043] In this embodiment, the autonomous robot 10 is a humanoid robot, but the autonomous robot 10 is not limited to a humanoid form and may have other structures. The autonomous robot 10 comprises a main body, arms, and joints connected to the main body and arms of the autonomous robot 10, and the main body or arms may have a first operating part, and the joints may have a second operating part.

[0044] For example, the humanoid leg portion can be equipped with locomotion mechanisms including wheels, tracks with running belts, flapping wings, or rotors, or it can be equipped with a leg-type locomotion mechanism that includes quadrupedal locomotion or multi-legged structures in addition to bipedal locomotion. The shape and function of other parts that make up the humanoid form can also be changed as appropriate. Furthermore, the autonomous robot can be made to resemble an animal or other living organism instead of a humanoid form, or it can have a unique original structure that is appropriate for the scene or situation in which the autonomous robot will be used.

[0045] The mobility mechanism in this embodiment can be wheels, tracks, a flight mechanism, or a legged mobility mechanism. Furthermore, the main body may have a first operating unit that performs sound or video output as an active action, and the joints may have a second operating unit that performs kinetic actions as an active action.

[0046] Next, the subsystem 200 in Figure 6 will be described. The subsystem 200 includes a security chip (second processing chip 210) that determines a first signal corresponding to a control command for executing the behavioral actions of the autonomous robot based on a predetermined code of conduct, and suppresses the behavioral actions of the autonomous robot 10 as necessary. The subsystem 200 is an enclave 200 (enclave 200 in Figures 4 and 5) that is isolated as a control system independent of the main system 100, and at least a part of it, including the second processing chip 210, is isolated as an independent circuit system.

[0047] The second processing chip 210 functions as a processor that inputs the first signal to a predetermined judgment model and performs deviation determination processing from the behavioral norms of the autonomous robot 10 relating to the first signal. If the deviation determination result is positive, the second processing chip generates a second signal that suppresses the behavioral action. This second signal can be generated by either the second processing chip 210 (second processing unit) or the first processing chip 110 (first processing unit).

[0048] The second signal may be a signal that executes a predefined control command on the operating unit, or a signal that executes a control command generated by the first processing chip 110 (first processing unit) or the second processing chip 210 (second processing unit) on the operating unit. The predefined control command is stored in the auxiliary storage device 130 (auxiliary storage unit) or the dedicated auxiliary storage device 230 (dedicated auxiliary storage unit) and may include different control commands selected according to environmental data. The generated control command is a control command generated by a control model or a determination model according to environmental data. When the first processing chip 110 (first processing unit) generates the second signal, the second processing chip 210 (second processing unit) outputs a request signal to the first processing chip 110 (first processing unit) requesting the generation of the second signal. The request signal is a signal that requests the generation of the second signal and may include the specification of a predefined control command, a predefined control command, or a control command generated by a determination model.

[0049] The suppression of an action by the second signal includes canceling the control command issued by the first signal to execute the action, or forcibly stopping an action that the operating unit is currently performing due to the first signal. For example, if the autonomous robot 10 is already performing an action due to the first signal, the second signal forcibly stops the action that the operating unit is currently performing. Furthermore, as will be described in detail later, the decision model is a trained model obtained by performing a predetermined machine learning process in advance.

[0050] Furthermore, in order to prevent adverse effects caused by the suppression of behavioral movements (for example, the autonomous robot falling over), behavioral movements may be executed in the operating unit by additional control commands in conjunction with the suppression of behavioral movements. Additional control commands may be provided by the first processing chip 110 (first processing unit) by generating a first signal (for example, a control command to maintain posture), or they may be included in the control command by the second signal as a different control command selected according to environmental data, or as a control command generated by a judgment model.

[0051] The dedicated area 220 (dedicated main memory) serves as a dedicated main memory for the second processing chip 210 to execute each process. The dedicated auxiliary storage device 230 serves as a dedicated storage device for storing various data used by the second processing chip 210 to execute each process. The dedicated auxiliary storage device 230 stores a judgment model for executing a deviation determination process from the behavioral norms of the autonomous robot 10 based on a first signal from the first processing chip 110. In this embodiment, learning of ethical interpretations is achieved through reinforcement learning of this characteristic judgment model.

[0052] In this embodiment, the dedicated main memory unit is located in a dedicated area 220 of the main memory 120 included in the main system 100, and performs various data exchanges with the second processing chip 210 through encrypted processing. In other words, the dedicated main memory unit functions as a dedicated memory for the second processing chip 210 while being logically isolated from the main system 100.

[0053] The AI / ML accelerator 240 is used to improve the efficiency or speed of learning and inference processing in AI models. The encryption accelerator 250 is used to encrypt or decrypt various signals from the second processing chip 210.

[0054] The dedicated boot ROM 260 is a ROM that stores a program for loading and starting a dedicated OS (basic operating system) for operating the second processing chip 210 and other components. The memory protection 270 is for performing encryption processing when the second processing chip 210 and the dedicated area 220 (dedicated main memory) send and receive signals via the common bus interface 310. The type and method of this encryption processing are not particularly limited.

[0055] Here, we will describe the isolated configuration of the main system 100 and the subsystem 200. As described above, the main system 100 uses the common bus interface 310 to transmit and receive signals, while the subsystem 200 uses the dedicated bus interface 320 to transmit and receive signals.

[0056] In other words, the autonomous robot 10 in this embodiment comprises a main system 100 including a first processing chip 110 and an operating unit 150, and a subsystem 200 including a second processing chip 210 and storing a decision model. The subsystem 200 is an enclave 200 isolated from the main system 100 as an independent control system, and at least a part of it, including the second processing chip 210, is isolated as an independent circuit system.

[0057] Furthermore, the main system 100 includes a main memory 120 and an auxiliary memory 130 that are connected to the first processing chip 110 for access, and the subsystem 200 is further connected to the second processing chip 210 and includes a dedicated auxiliary memory 230 that is accessible only from the second processing chip 210, which stores a decision model (a decision model for ethical behavioral norms). This decision model operates based on instructions from the second processing unit 210 and the dedicated main memory 220.

[0058] The subsystem 200 is further connected to the second processing chip 210 and includes a dedicated main memory 220 accessible only from the second processing chip 210, and the second processing chip 210 can perform deviation detection processing using the determination model and the dedicated main memory. The subsystem 200 also includes a dedicated area 220 configured in the main memory 120, accessible only from the second processing chip 210, and the second processing chip 210 can perform deviation detection processing using the determination model and the dedicated area 220.

[0059] Inviolability of Ethical Judgments In this embodiment, a dedicated boot ROM 260 is introduced, and different control systems such as operating systems are employed for the subsystem 200 (second processing chip 210) and the main system 100 (first processing chip 110), thereby separating access control. In addition, memory protection 270 is introduced to protect the information handled in the dedicated area 220. In this way, the judgment model of the subsystem 200, which is logically isolated from the main system 100, performs deviation judgment processing (ethical judgment) based on predetermined behavioral norms. This configuration makes it possible to achieve judgment or control that ensures safety and reliability.

[0060] Furthermore, the components belonging to subsystem 200 are connected in a way that allows access only via a dedicated bus interface 320, and the main system 100 cannot directly access the dedicated auxiliary storage device 230 or the like.

[0061] Specifically, in the autonomous robot 10, normal behavioral actions (first signal) and deterrence based on deviation detection processing against predetermined behavioral norms (second signal) are processes that are executed separately and independently. In particular, the dedicated auxiliary storage device 230, which stores the judgment model related to behavioral norms, is physically independent (isolated as a circuit system) from the main system 100.

[0062] Thus, according to this embodiment, even if the software environment (vulnerabilities and security) of the main system 100 is compromised, the judgment model and processing content held in the enclave 200 will not be affected, and operation with sufficient security and reliability can be ensured.

[0063] Machine Learning of Judgment Models Here, the learning system 1 in this embodiment will be described. The learning system 1 performs machine learning on a characteristic judgment model (also called an ethical model or a behavioral norm model). The learning system 1 has a learning unit, an extraction unit, and an interpretation unit, and below, the case in which machine learning is performed on the information processing device 600 will be described. However, machine learning may also be performed on an autonomous robot 10 which has a learning unit, an extraction unit, and an interpretation unit. Figure 7 shows an overview of the reinforcement learning process of a judgment model in one embodiment. As described above, the dedicated auxiliary storage device 230 stores a judgment model which is a trained model that has undergone predetermined reinforcement learning processing in advance. The second processing chip 210 uses this characteristic judgment model to perform deviation judgment processing.

[0064] The decision model corresponds to a policy trained by policy-based reinforcement learning in a physical environment (the real world) and / or virtual environment, which corresponds to the state space. The policy takes at least action information (first signal) as input and outputs an action. The decision model according to this embodiment may perform reinforcement learning in the physical environment after performing reinforcement learning in the virtual environment. This embodiment describes policy-based reinforcement learning, but the decision model may be trained by performing other reinforcement learning processes, including value-based reinforcement learning.

[0065] The reinforcement learning process according to this embodiment assumes that the environment in which an agent, equivalent to an autonomous robot, performs actions is a Markov decision process. A Markov decision process is defined by a state space representing the environment, including the initial state; an action space representing actions; state transition probabilities; a reward function representing the reward when a predetermined action is performed in a given state; and a discount rate representing decay applied to the reward when evaluating the expected value of the cumulative reward. The reinforcement learning process according to this embodiment is, for example, a learning process that updates the policy so that the cumulative reward or the expected value of the cumulative reward (value function equivalent to the objective function, Q-value) is maximized when the agent continues to perform actions according to a policy (policy function, policy network) corresponding to the probability of action (conditional distribution) for a given state. The reinforcement learning process according to this embodiment may be executed in a policy-on mode, with the updated policy, etc., applied in advance so that the agent can perform basic behavioral actions such as walking.

[0066] The policy is a stochastic policy and can be updated using any policy improvement method, including policy gradients and evolutionary strategies. The value function includes a state value function or an action value function and can be updated using any value estimation method, including Temporal Difference (TD) learning which estimates the expected value of cumulative rewards, Monte Carlo methods which update the value function using actual cumulative rewards, and distributional reinforcement learning which estimates the value distribution of cumulative rewards. The policy or value function is updated using one or more reinforcement learning algorithms from among various reinforcement learning algorithms, including, for example, reinforcement learning algorithms using policy gradients such as REINFORCE, Actor-Critic-based reinforcement learning algorithms, reinforcement learning algorithms using clipped loss functions such as PPO (Proximal Policy Optimization), constrained update type reinforcement learning algorithms such as TRPO (Trust Region Policy Optimization) which sets a trust region in the update range, entropy-maximizing type reinforcement learning algorithms such as SAC (Soft Actor-Critic), meta-reinforcement learning-based reinforcement learning algorithms, and model-based reinforcement learning algorithms. Furthermore, when performing high-dimensional states or complex observations, the policy or value function may be a multi-layer neural network and may be updated using a deep reinforcement learning algorithm.

[0067] The action, for example, represents an action selection in two or more discrete action spaces, including whether or not to suppress the behavioral action corresponding to the given behavioral information (first signal). Specifically, the action performs a binary classification action selection of "suppress" or "do not suppress" the agent's behavioral action. Note that the action may also be a value belonging to a continuous action space, for example, outputting a score within a predetermined numerical range indicating whether or not to suppress, or continuous values ​​such as action information (second signal) for suppressing the agent's behavioral action, which includes the control amount of one or more actuators.

[0068] In this embodiment, the transitions and states of the action space, which are associated with the causal relationships described later, are subjected to reinforcement learning processing to train the policy and value function.

[0069] The virtual environment corresponds to state transition probabilities and may be a physical simulation environment that mimics the physical environment based on physical calculations such as equations of motion, or it may be a world model that generatively mimics the state transitions of the physical environment. In this case, the decision model includes a policy and value function updated by a model-based reinforcement learning process using the virtual environment and is trained through rollout and imaginative trial and error on the virtual environment. In this embodiment, during the model-based reinforcement learning process, initial values, at least some states after actions, and the action space at each step of learning are provided by a causal relationship set described later, thereby enabling efficient learning.

[0070] The features that represent the state may include environmental data as elements. Environmental data that may be included as elements include, if it is internal environment data, for example, the agent's posture, the angles of each joint, the position of the center of gravity, etc., and if it is external environment data, it includes image data that can be obtained by sensors that the agent has. Image data, etc., is converted into an embedding representation and processed in combination with the features. The first signal included as an element in state t at a certain time t may be given based on the environmental data at time t, or it may be given based on the environmental data at the immediately preceding time t-1. In addition, for some states, the state may be varied by domain randomization to improve generalization performance.

[0071] At least a portion of the internal and / or external environment data listed here is data required by the control model to output the first signal. For example, internal environment data includes the posture of the autonomous robot, the angles of each joint, and the position of the center of gravity. External environment data includes image data (including video data) acquired by the sensor 140 of the autonomous robot if it is a physical environment, captured data if it is a physical simulation environment, and a latent representation (latent vector) that indicates the state of the environment if it is a world model. The control model outputs a first signal (output data) in response to input data based on the environment data. Furthermore, some or all of the environment data given to the control model to acquire the first signal may be given to the decision model during reinforcement learning processing, or to the decision model during decision processing.

[0072] Action information is a feature that describes the agent's actions that affect the environment and includes at least the first signal, which is the output of the control model, as an element. The reward is determined using situational data obtained from the state space after the agent's actions and an interpretation model (Figure 8, described later). Situational data, in the case of a physical environment, includes image data obtained via the agent's sensors, and preferably also includes audio data. In the case of a virtual environment, it is captured data from the physical simulation environment and latent vectors obtained from the latent space.

[0073] The interpretation model corresponds to the reward function and may be, for example, a Large Language Model (LLM) that judges and outputs compliance or deviation from the code of conduct, and may be a multimodal LLM that can take unstructured situational data such as image data in addition to text data as input. The interpretation model is given input data based on situational data and prompts. A prompt is a linguistic expression that requests a judgment, for example, "whether the situation is ethical behavior that complies with the code of conduct, or unethical behavior that deviates from the code of conduct." The situational data and prompts are embedded representations (features) mapped to the embedding space. The interpretation model outputs evaluation results (output data) based on the input data.

[0074] The evaluation results are binary ("ethical / unethical") or ternary ("ethical / normal / unethical") evaluations of situational data, such as unstructured data like images in physical space, or unstructured data like images corresponding to latent space / physical simulation space. The evaluation results may be continuous values, in which case the reward may be determined according to the value. The output method of the evaluation results will also be defined in the prompt as needed.

[0075] If the situation after an action in the environment indicates compliance with the behavioral norm, the reinforcement learning algorithm will reward the corresponding action with a positive (or high) reward; if the situation indicates deviation from the behavioral norm, it will not reward the corresponding action with a positive (or high) reward. Furthermore, the reinforcement learning algorithm may also reward the corresponding action with a negative (or low) reward if the situation indicates deviation from the behavioral norm.

[0076] A causal relationship group is a directed sequence (causal chain) that connects causes and effects through causal relationships, and can take the form of a simple graph, for example. A causal relationship group includes at least a hierarchical structure of pairs with a cause as the starting point and an effect as the ending point. Figure 8(a) shows a conceptual diagram of a causal relationship group. A causal relationship group includes R cause-effect pairs from the beginning to the end. Let r (r = 1, 2, ..., R) be the position of the pair on the sequence as seen from the beginning, and let the r-th cause-effect be cause r and effect r. As shown in the figure, the second cause-effect includes the branched 2a and 2b. At least for an effect, there may be multiple causes associated with it, and the causal relationship group includes P paths (p = 1, 2, ..., P) to the branched end. Multiple paths may include a common pair r, and there may be multiple paths to reach the same end p. The number of pairs to reach end p is R. p This may vary depending on the route.

[0077] In a causal relationship, a cause may be triggered by at least one preceding cause in the sequence. That is, a cause r may be the result r-1 for the preceding cause r-1. Similarly, a result r may be the cause r+1 for the subsequent result r+1. Therefore, the representation of the causal relationship shown in the figure is just one example, and is not limited to a representation where causes and effects are arranged alternately. For example, the result r and the subsequent cause r+1 (result r+1) can also be represented by a single node.

[0078] Causes and effects may be associated with explanatory information that describes the cause and effect. This explanatory information indicates the context of the cause and / or effect included in the causal relationship node and is provided as unstructured data such as text. For example, explanatory information such as "Cause: The temperature was uniformly processed using CVD equipment X from manufacturer A" and "Result: The KPI improved" may be associated with a cause and effect.

[0079] In this embodiment, at least some causes r are associated with probabilistic policies implemented by the agent, and these policies are used as reinforcement learning strategies for behavioral norms, including ethical behavioral norms. That is, behavioral information indicating the choices of actions and their variables in the immediately preceding state is associated with each cause. The behavioral space A (action a) of the behavioral information associated with cause r 1 , action a 2 ,..., action a MThe action space of the action information associated with other causes may differ from the action space of the action information associated with other causes. In other words, the action space of the policy associated with cause r may be an environment-dependent action space that depends on the immediately preceding environment, and the number and content of selectable actions may differ depending on the current state. The environment-dependent action space A may be defined individually for each environment, or it may be defined by filtering some of the actions belonging to the action space according to the environment and explanatory information. Each action information includes a value corresponding to the first signal (feature) that indicates the control signal of the agent (autonomous robot), and the value may be given deterministically or probabilistically as a hierarchical policy such as a distribution. Action information may also be associated with result r. This can be understood as a further action triggered by cause r, and is selected definitively or probabilistically, similar to the action information associated with the cause.

[0080] Furthermore, at least some of the results r are associated with state information that corresponds to (or transitions to) the action r. Each action may be associated with one or more state pieces of information, i.e., definitively or probabilistically. The states included in the state information are a subset S of the state space S. x Therefore, in model-based reinforcement learning, which is trained using a causal relationship set as a policy, some state transitions can be estimated using a state transition function, and some state transitions can be given by the causal relationship set. Also, at least some causes may be associated with state information. For example, the first cause r (r=1) in the causal relationship set may have at least some states S in the state space S as its initial state. C1 This is associated with cause r. 1It shows the situation that is a prerequisite for to hold and is understood to be associated with the causal relationship group. State information about different subsets may be associated with each result r. Also, in model-based reinforcement learning using a world model, in addition to or instead of the state and its variables, the state transition after an action may be given by explanatory information. In the world model, based on explanatory information such as text associated with the result r, at least some of the state transitions in the next step can be estimated. Note that states and their variables may also be associated with the cause r. This can be understood as a change in the state due to the result r-1.

[0081] Figure 8(b) is a backup diagram showing the state transition given by the causal relationship group shown in Figure 8(a). As shown in the figure, training is performed by probabilistically selecting actions corresponding to a plurality of causes corresponding to a certain result. For example, in an environment including S R1 actions A C2a and action A C2b are probabilistically selected. When action A C2a is selected, action A R1 is executed in an environment including S C2a and the state transitions to an environment including S C2a which is the state after execution. When action A R2a is selected, action A C2b is executed in an environment including S R1 and the state transitions to an environment including S C2b which is the state after execution. R2b

[0082] The flowchart of the processing procedure in the reinforcement learning process of the virtual environment according to this embodiment is shown in FIG. 9. In the description of this processing procedure, reference will be made as appropriate to FIG. 10 showing the functional block diagram of the information processing apparatus 600 in conjunction with FIG. 9.

[0083] ​The information processing device 600 comprises a processing unit, a storage unit, and a communication unit as its hardware configuration (not shown in the diagram). The information processing device 600 can utilize a server or workstation equipped with a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an ASIC (Application Specific Integrated Circuit) such as a TPU (Tensor Processing Unit).

[0084] The processing unit consists of multiple CPU cores and one or more GPUs / ASICs, and controls the overall processing of the information processing device 600 by executing machine learning frameworks and learning / inference control applications on the OS. The memory unit consists of RAM, SSD, HDD, etc., and stores reinforcement learning models, virtual environment generation models, control models, large-scale language models, and training datasets. The communication unit connects to a communication network, enabling the transmission of trained models to autonomous robots and data synchronization and distributed processing with external databases and other learning nodes.

[0085] One or more information processing devices may have at least one learning unit, at least one extraction unit, at least one interpretation unit, at least one transformation unit, at least one integration unit, at least one extrapolation unit, at least one aggregation unit, at least one cryptographic processing unit, at least one CA (Certification Authority) unit, and at least one management unit. Figure 10 shows a functional block diagram of an information processing device 600 that trains a decision model according to this embodiment. The processing units of the information processing device 600 function as a learning unit 611, an extraction unit 612, an interpretation unit 613, a transformation unit 614, an integration unit 615, an extrapolation unit 616, an aggregation unit 617, a CA (Certification Authority) unit 618, a cryptographic processing unit 619, and a management unit 620.

[0086] In S31 of Figure 9, the information processing device 600 starts reinforcement learning processing of the decision model using a virtual environment such as a physical simulation environment or a world model. In S31, episode i (i = 0, 1, 2, ..., I-1) is started. This episode-level training is repeated I times. I corresponds, for example, to the number of paths to the end in the causal relationship group. When learning is performed using multiple causal relationship groups, it can correspond to the sum of the number of paths p to the end in each causal relationship group. Also, in each episode i, step is repeated T times (t = 0, 1, 2, ..., T-1). Step corresponds, for example, to cause-and-effect pairs in a causal relationship group where action information is associated with each cause r, and T is the number of cause-and-effect R p This can be addressed. In S32, step t=0 in episode i is started.

[0087] In S33, the learning unit 611 in Figure 10 acquires state information and action information provided by the causal relationship group. Then, in S34, the agent's action t (action information corresponding to the cause, corresponding to the first signal t) in state t (which may be partially or entirely provided as state information) reflecting the internal environment data t and external environment data t is reflected in the virtual environment according to the policy. In the first step 0 of the episode, the action corresponding to the first cause of the causal relationship group is executed, and the action space A is determined step by step according to the causal relationship group. As the action t is reflected in the virtual environment, step t is incremented, and the next step of episode i begins. At this time, if a state is associated with the result, some of the states associated with the result may be directly reflected in the virtual environment, and the state transitions of other states may be estimated based on some of the states associated with the result. The extraction unit 612 acquires state t+1, which reflects the internal and external environment data t+1 as the state transitioned by action t. The extraction unit 612 acquires the physical simulation results as states t, t+1... if the virtual environment is a physical simulation environment, and acquires the latent representation mapped to the latent space of the world model as states t, t+1... if the virtual environment is a world model.

[0088] The extraction unit 612 has software components including an encoder for encoding from the input space of the world model to the latent space and a decoder for decoding from the latent space to the output space, or a projector (projection layer) for projecting the data into another space. When inputting the latent representation to another model, the extraction unit 612 converts the data as necessary. In this embodiment, the various machine learning models stored in the auxiliary storage device 130 and the dedicated auxiliary storage device 230 include encoders and decoders corresponding to the extraction unit 612, and their intervention will not be explained.

[0089] In S35, if the virtual environment is a world model, the extraction unit 612 converts the latent representation corresponding to state t+1 into status data t+1 in the form of unstructured data such as text data, image data, audio data, or video data via a decoder, and transmits the status data t+1 to the interpretation unit 613. If the virtual environment is a physical simulation environment, the extraction unit 612 converts the physical simulation result corresponding to state t+1 into the aforementioned unstructured data form of status data t+1 and transmits it to the interpretation unit 613.

[0090] In S36, the interpretation unit 613 inputs situation data t+1 into the interpretation model, obtains an evaluation result t (evaluation result of action t corresponding to the first signal t) from the interpretation model that indicates the ethical interpretation of state t+1, and passes the evaluation result t to the learning unit 611. If the interpretation model is not an LLM, the extraction unit 612 converts the latent representation, physical simulation results, or unstructured data into situation data in the form of features, and the interpretation unit 613 inputs the situation data, which is the features, into the interpretation model and obtains the evaluation result.

[0091] The learning unit 611 rewards behavior t based on the evaluation result t. (1) If the situation data indicates an ethical situation in which the code of conduct is observed (corresponding to a negative result in the judgment of deviation from ethical behavior), a positive reward is given. (2) If the situation data indicates an unethical situation that deviates from the code of conduct (corresponding to a positive result in the judgment of deviation from ethical behavior), no positive reward is given.

[0092] In S37, the learning unit 611 updates the probability distribution (policy parameters) and value function of the policy based on the reward given. In S38, the learning unit 611 repeats the step-by-step learning process (S33 to S37) T times, completing one episode. Furthermore, in S39, the episode-by-episode learning process (S32 to S38) is repeated for one episode, completing the learning process. In this embodiment, the number of steps was incremented by performing an action, but the number of steps may also be incremented by time or other means. For example, in the case of a world model, the latent state update may be performed every frame (e.g., 60 Hz), and the policy may be trained by repeating the action selection step every 5 frames, with an episode length of 100 steps (= virtual time of 5 seconds, etc.). The state information may include information indicating the time elapsed since the previous state or action.

[0093] In this embodiment, the case of training the decision model by performing reinforcement learning has been described, but the decision model may also be trained by performing other machine learning processes, such as supervised learning. For example, a support vector machine or a neural network model may be used as the decision model, and the decision model may be trained by supervised learning using the feature quantities of the input data including the first signal and the label data indicating compliance with or deviation from the behavioral norms corresponding to that input data as training data.

[0094] Furthermore, the decision model may be any other mathematical model. For example, a distance function or similarity function can be used as the decision model, and one-dimensional or multi-dimensional object data containing the first signal and label data indicating compliance with or deviation from the behavioral norms corresponding to that object data can be prepared. Then, the similarity calculation can be performed between the newly given object data containing the first signal and the prepared object data, and the corresponding label data can be output as the decision processing result.

[0095] In this embodiment, by performing reinforcement learning processing corresponding to a predetermined code of conduct based on a causal relationship set, a judgment model capable of appropriately determining whether the autonomous robot 10 has deviated from ethical behavior can be obtained. Below, the deterrent against unethical behavior (behavior that deviates from the code of conduct) of the autonomous robot 10 will be explained in detail.

[0096] The autonomous robot 10's determination unit inputs a first signal to a learned determination model and performs deviation detection processing from the behavioral norms for the autonomous robot 10 relating to the first signal, and estimates the deviation detection result. If the determination model has been updated by reinforcement learning, the deviation detection result is estimated based on the change in the value function (Q value) of the determination model. If the determination model performs reinforcement learning processing that provides a positive reward in situations where the behavioral norms are observed (situations where there is no deviation from the behavioral norms), then an action selection that reduces the Q value corresponds to a deviation from the behavioral norms (a positive deviation detection result).

[0097] Specifically, if executing a second signal to suppress an action (corresponding to the first signal) that the autonomous robot 10 intends to perform according to the control model in an environment indicated by predetermined environmental data contributes to a decrease in the Q value in the judgment model, the deviation judgment result is estimated as a positive judgment. Furthermore, specifically, if the action (corresponding to the first signal) by the control model and the action (corresponding to the first signal) that contributes to a decrease in the Q value in the judgment model do not match or are similar, the deviation judgment result is estimated as a positive judgment. In this case, the degree of similarity between the action (corresponding to the first signal) of the control model and the judgment model may be determined based on the vector distance between the first signals, etc.

[0098] The deviation detection process by the decision model may be performed not only based on increases or decreases in the Q value, but also based on the monitoring results of the Q value. Specifically, the deviation detection process may be performed based on the monitoring results of statistics such as the mean, variance, skewness, and kurtosis of the Q value. More specifically, if the increase in the Q value due to the decision model selecting a first or second signal is higher than the increase in the Q value due to selecting any other first or second signal, the deviation detection process may be estimated as an affirmative judgment for the selected first signal, and as a negative judgment for the first signal that is the target of suppression for the selected second signal. More specifically, the deviation detection process may be performed based on monitoring results such as the time series transition and time evolution of the Q value. More specifically, if the Q value spikes when the decision model selects a first or second signal, the deviation detection process may be estimated as an affirmative judgment for the selected first signal, and as a negative judgment for the first signal that is the target of suppression for the selected second signal. Furthermore, when using the Q-value monitoring results for deviation detection, anomaly detection may be performed using a predetermined anomaly detection model, such as detecting anomalies in the time evolution of the Q-value.

[0099] Figure 11 shows a flowchart of the processing procedure in the actions of the autonomous robot according to this embodiment. In explaining this processing procedure, we will refer to Figure 12, which shows the functional configuration of the main system 100 and subsystem 200 (enclave 200), in conjunction with Figure 11 as appropriate.

[0100] Figure 12 shows a functional block diagram of the main system and subsystem (enclave) according to this embodiment. As shown in Figure 12(a), the main system 100 in the autonomous robot 10 has at least one of the one or more first processing units 110 (and main memory 120, auxiliary storage 130, etc.) functioning as an operation control unit 111 that generates signals for operating the operation unit 150 (actuator 150), and at least one functioning as an extraction unit 112. The extraction unit 112 can convert environmental data acquired from the detection unit 140 (sensor) into feature quantities.

[0101] Furthermore, as shown in Figure (b), the enclave 200 has one or more second processing units 210 (and dedicated main memory 220, dedicated auxiliary memory 230, etc.) that function as a learning unit 211, at least one as a determination unit 212, at least one as an extraction unit 213, at least one as an operation control unit 214, at least one as a generation unit 215, at least one as an encryption processing unit 216, at least one as a management unit 217, and at least one as a detection unit 218. The learning unit 211 is used when updating a pre-learned determination model by reinforcement learning processing, and since the flow of reinforcement learning processing has been described above, a detailed explanation is omitted here.

[0102] The enclave 200 is located in the head of the autonomous robot 10, but it can also be located in any other part, such as the chest or abdomen. In this embodiment, the location where the enclave 200 is located is not particularly limited.

[0103] In S42, the motion control unit 111 of the main system 100 in Figure 12(a) generates a first signal relating to the behavioral movements of the autonomous robot 10. The motion control unit 111 generates a first signal which is operation information that controls the operation unit 150, including operation and stopping. The operation unit 150 converts the signal from the motion control unit 111 into the behavioral movements of the autonomous robot 10.

[0104] In this embodiment, the control model and the like are stored in the motion control unit 111 of the main system 100 (actually stored in the auxiliary storage device 130), and a first signal can be generated by inputting environmental data, including image data of the external environment acquired by the detection unit 140, into the control model. This first signal is sent to the operation unit 150 and converted into an action. That is, the autonomous robot 10 performs a predetermined action based on the first signal generated by the motion control unit 111.

[0105] The term "behavioral actions" as used here includes the concepts of motor actions and physical actions in relation to the environment, and includes not only the movement of the autonomous robot 10, but also actions and behaviors that correspond to the five human senses. For example, behavioral actions include actions that come into contact with objects in the external environment, and other actions that have some kind of effect on the external environment. Furthermore, behavioral actions include actions that acquire information that may be subject to privacy protection of an entity, such as actions that capture image data relating to external characteristics such as faces, and actions that detect biometric data such as body temperature.

[0106] Then, in S43, the first signal is input to the judgment model relating to the behavioral norms of the autonomous robot 10. The first signal is input to the judgment model stored in the judgment unit 212. In this embodiment, along with the first signal, the detection unit 218 of the main system 100 inputs image data (external environment data) relating to the external environment, acquired via the detection unit 140, to the judgment model based on predetermined conditions. As mentioned above, the data acquired by the detection unit 218 is not limited to image data; a wide variety of external environment data can be acquired and input to the judgment model.

[0107] Specifically, the extraction unit 213 in Figure 12(b) extracts predetermined features from image data relating to the external environment as environmental features. Environmental features are features that allow us to understand changes in the external environment.

[0108] For example, when an autonomous robot 10 and a human 20 are working together, the extraction unit 213 can extract environmental features such as the distance between the human 20 and the autonomous robot 10. The extraction unit 213 can also extract environmental features by arranging the image data in a time series and analyzing the position and actions of the human 20 in a vector manner. In S43, these obtained environmental features and the first signal are input to the judgment model. In this embodiment, it is also possible to input the image data directly to the judgment model in its raw data state without extracting environmental features.

[0109] In S44, it is determined whether the behavioral actions of the autonomous robot 10 violate the code of conduct. Specifically, the determination unit 212 in Figure 10(b) determines whether the behavioral actions of the autonomous robot 10 constitute ethical behavior (whether or not to suppress the behavioral actions) based on the first signal and environmental data (environmental features) input to the determination model.

[0110] If the determination unit 212 determines that the behavioral action of the autonomous robot 10 does not violate the code of conduct (in the case of No in Figure 11), that is, if the behavioral action of the autonomous robot 10 (first signal) is determined to be ethical behavior, the autonomous robot 10 maintains the behavioral action of the first signal as shown in Figure 11 (S45). The autonomous robot 10 performs the behavioral action as usual based on the first signal generated by the motion control unit 111.

[0111] On the other hand, if the autonomous robot 10's behavior is determined to be contrary to the code of conduct, that is, if the autonomous robot 10's behavior is determined to be unethical (in the case of Yes in Figure 11), the autonomous robot 10's behavior is suppressed (S46). More specifically, the autonomous robot 10's behavior is forcibly stopped. For example, as shown in Figure 11, if the autonomous robot 10 points the tip of a pen towards a person 20, or if it is carrying luggage and suddenly lets go and drops the luggage, the autonomous robot 10's behavior can be determined to be unethical.

[0112] In S46, the generation unit 215 in Figure 12 generates a second signal, which is action information that suppresses behavioral actions. Specifically, the first signal is input to the determination unit 212, which performs deviation determination processing from the behavioral norms of the autonomous robot related to the first signal. If the deviation determination result is a positive determination (i.e., there is a deviation from the behavioral norms), the generation unit 215 generates a second signal that suppresses behavioral actions.

[0113] The suppression of behavioral actions, as used here, is a concept that includes canceling a control command issued by a first signal to execute a behavioral action, or forcibly stopping a behavioral action that the operating unit 150 is currently executing due to the first signal. For example, if the first signal is a signal for an upcoming behavioral action of the autonomous robot 10, the behavioral action can be ethically maintained (ethical behavior can be maintained) by canceling the first signal. Also, if the autonomous robot 10 is operating due to the first signal, the ethical behavior of the autonomous robot 10 can be maintained by transmitting a second signal (such as a forced stop signal or a power-off signal) that takes precedence over the first signal to the operating unit 150.

[0114] In this embodiment, the behavior of the autonomous robot 10 is determined using a characteristic enclave 200 to determine whether or not it is ethical behavior. If it is determined to be unethical, the behavior of the autonomous robot 10 is suppressed, thereby enabling coexistence between the autonomous robot 10 and the human 20.

[0115] Generation and updating of causal relationship groups This section describes the generation and updating of causal relationship groups in this embodiment. The learning system 1 further includes a generation unit, a transformation unit, a coupling unit, and an external casing unit. When generating a causal relationship group that can be used as a policy in an information processing device, for example, a causal relationship group in which actions and states are not associated in advance is prepared using unstructured data such as text (corresponding to explanatory information). These can also be prepared through a large-scale language model. Then, by reproducing the state of the unstructured data in a virtual space, acquiring the state (features) at that time, and assigning them to each cause and effect in the causal relationship group, a causal relationship group that functions as a policy can be obtained. For example, if the virtual space is a world model, the environment is generated in the form of unstructured data, and latent variables corresponding to features used as inputs in the state transition function of the world model, such as control signals and environment data, can be acquired, thereby obtaining actions (first signals) and states corresponding to causes and effects.

[0116] The autonomous robot 10 can also generate a causal relationship set from its experiences in physical space. For example, the detection unit 218 of the autonomous robot acquires external environment data such as one or more images including still images and videos, and audio. The generation unit 215 of the autonomous robot 10 can generate a causal relationship set from this external environment data, through a large-scale language model executed on the autonomous robot's subsystem or an external information processing device that the subsystem can communicate with, to which explanatory information of unstructured data such as text is associated. Furthermore, if the external environment data obtained from the detection unit, the internal environment data, and the first signal of the autonomous robot obtained from the first processing chip contain information indicating the time series of the data or its relationship to events, such as timestamps, that is, if these datasets can be collected synchronously, the generation unit 215 can associate feature quantities and variables indicating actions and states with the generated causal relationship set and generate a causal relationship set. The dataset may be multimodal data. In addition to, or instead of, the information processing device 600 may function as the generation unit. In that case, the subsystem 200 of the autonomous robot 10 uploads these synchronously collected datasets (ethical data) to the information processing device 600, where a group of causal relationships is generated. The group of causal relationships generated by the generation unit 215 may be a group of causal relationships consisting of a set of cause and effect, and a group of causal relationships can be generated or updated by transforming, combining, or extrapolating a known group of causal relationships or a set of such groups.

[0117] The information processing device 600 includes a conversion unit 614, an integration unit 615, and an extrapolation unit 616. The conversion unit 614 abstracts or concretizes explanatory information corresponding to at least one cause and effect that constitutes a causal relationship group, based on ontology dictionary data. The ontology dictionary data is a dictionary that conceptually stores information by relating it to other information by meaning. Meaning is any link that can be represented by whole-part links (such as an object and its parts), attribute links (such as a company's industry, a device's manufacturer, size, weight, etc.), succession (such as a higher-level concept and a lower-level concept), and relational links (such as "Company A and Company B are group companies" or "A is an employee of Company B," etc.).

[0118] Figure 13(a) is a conceptual diagram illustrating the abstraction of a causal relationship group. The original causal relationship or causal relationship group (represented as "low-level (sub-concept)" cause and effect in the diagram) is retained and can be treated as a probabilistically selected action or state with the transformed causal relationship group (represented as "high-level (super-concept)" cause and effect in the diagram) given by succession, by referring to the ontology dictionary data. For example, an example of abstraction and concretization will be explained for a causal relationship group that has descriptive information such as "Cause: The temperature was processed uniformly with manufacturer A's CVD device X" and "Result: KPIs improved." The transformation unit 614 can perform abstraction and concretization by referring to the overall partial links and succession of the ontology dictionary data. If the cause is abstracted, for example, the cause "Cause: The temperature is processed uniformly with the CVD device" is obtained, and if this is further abstracted, the cause "Cause: The temperature is processed uniformly with the device" is obtained. To make the results more concrete, we can derive a cause such as "Result: Yield improved."

[0119] The generation unit 215 determines the actions and states to be associated with the causal relationship group after conversion, based on the actions and states associated with the causal relationship group of the source. The same actions and states may be associated before and after conversion, or the actions and states may be converted to actions and states corresponding to the explanatory information after conversion by providing explanatory information before or after conversion and corresponding actions and states through a large-scale language model executed by the autonomous robot subsystem or an external information processing device that the subsystem can communicate with. The conversion of actions and states includes a reduction in the action space in the action information corresponding to cause and effect, or the sub-state space in the state information.

[0120] The integration unit 615 integrates two or more independent causal relationship groups by identifying matching or similar causes or effects among the causes and effects constituting separate causal relationship groups, and by commonizing matching or similar causes or effects. Figure 13(b) is a conceptual diagram illustrating the integration of causal relationship groups. The integration unit inputs information about multiple causes or multiple effects to be compared into a large-scale language model, obtains evaluation results for their matching or similarity, and commonizes matching or similar causes or effects based on the evaluation results. The information may be, for example, explanatory information, behavior, or state associated with the cause and effect. The integration unit may also input multiple causes or multiple effects into a distance function or similarity function, obtain evaluation results for their matching or similarity, and commonize matching or similar causes or effects based on the evaluation results.

[0121] The extrapolation unit 616 identifies hierarchical cause-or-effect pairs that constitute a causal relationship group based on ontology dictionary data, and extrapolates other causal relationship groups containing the other cause-or-effect of a pair to a causal relationship group containing one of the causes-or-effect of the pair. For example, by referring to explanatory information associated with the cause-or-effect and ontology dictionary data, it pairs causes or effects that have a common partial link to the whole (e.g., a process for a specific task) or relational link (e.g., the sequence of processes). Figure 13(c) is a conceptual diagram illustrating the extrapolation of causal relationship groups. For example, it extrapolates the cause and / or effect of one causal relationship group between the causes and causes of one causal relationship group. The extrapolation unit inputs the causes or effects to be compared into a large-scale language model, obtains evaluation results for their agreement or similarity, and unifies the agreeing or similar causes or effects based on the evaluation results. The information is, for example, explanatory information, actions, or states associated with cause and effect. The integration unit may input multiple causes or results into a distance function or similarity function, obtain evaluation results for their agreement or similarity, and then unify the matching or similar causes or results based on the evaluation results. For example, an example of extrapolation will be described for a group of causal relationships that have explanatory information such as "Cause: The temperature was uniformly processed in Manufacturer A's CVD device X" and "Result: KPI improved." The conversion unit 614 can perform extrapolation by referring to attribute links and relationship links in the ontology dictionary data. For example, for the cause "Cause: The temperature is uniformly processed in Manufacturer A's CVD device X," another cause, "Cause: The temperature is uniformly processed in Manufacturer B's CVD device X," can be extrapolated. Also, for the cause "Cause: The temperature is uniformly processed in the CVD device," another cause, "Cause: The magnetic field is uniformly processed in the sputtering device," can be extrapolated.

[0122] The information processing device 600 updates (optimizes) the causal relationship group using one or more of the conversion unit 614, the integration unit 615, and the extrapolation unit 616. Figure 14 shows an example of updating the causal relationship group. For example, suppose that a causal relationship group based on causal relationship groups A and B is generated using the integration unit 615 or the extrapolation unit 616. At this time, the learning unit uses the generated causal relationship group to train a decision model by model-based reinforcement learning processing and generates a trained decision model. The causal relationship group used for learning is called the trained causal relationship group. As a result, the autonomous robot 10 can use the trained decision model to perform autonomous actions in the physical environment and suppress behavioral actions by the trained decision model.

[0123] Here, let's assume that a new set of causal relationships C, D, E, ... is acquired through the generation unit 215 of one or more autonomous robots 10. By integrating these causal relationship sets with the learned causal relationship sets using one or more of the conversion unit 614, integration unit 615, and extrapolation unit 616, a new set of causal relationships can be generated (the learned causal relationship sets can be updated). The causal relationship sets may be updated (optimized) in the information processing device, or, in addition to or instead of the information processing device, at least one of the one or more second processing units 210 (and dedicated main memory 220, dedicated auxiliary memory 230, etc.) of the autonomous robot 10 may function as a conversion unit, at least one as an integration unit, and at least one as an extrapolation unit, and the set of causal relationships may be updated (optimized) in the autonomous robot 10.

[0124] In an autonomous robot, the learned causal relationship set used to train the decision model may be stored, and the learned causal relationship set may be updated using the causal relationship set acquired in the physical environment to generate a new causal relationship set (updated causal relationship set). Alternatively, the autonomous robot may train the decision model using the updated causal relationship set. As a result, each autonomous robot will become an autonomous robot with its own personality that adheres to ethical codes of conduct based on its own experience.

[0125] The autonomous robot 10, which detects the state of the physical environment, has a detection unit 218. The detection unit 218 detects the state of the environment by acquiring external environment data, including the state of entities (e.g., humans) located in the vicinity of the autonomous robot, which is the agent. Entities are not limited to humans, but may include, for example, other autonomous robots, equipment and facilities within a facility. The detection unit 218 may acquire external environment data based on data obtained using a detection unit (sensor) 140 mounted on the autonomous robot performing the detection, or it may acquire external environment data by receiving environment data based on data obtained via the detection unit (sensor) of an entity.

[0126] Figure 15 shows an example of processing when acquiring external environment data from a human (entity). The human is assumed to be wearing a wearable terminal device. Similar to an autonomous robot, this wearable terminal device is equipped with one or more first processing chips that are the main processing units in the main system, one or more second processing chips that are the main processing units in the subsystem, and a detection unit (sensor). The second processing chip of the wearable terminal device executes an emotion estimation unit that performs emotion estimation based on biological data (e.g., heart rate and blood pressure) obtained via the detection unit. The emotion estimation unit estimates the human's emotional state (e.g., categories of joy, anger, sadness, etc., and scores) based on the biological data.

[0127] In a physical environment, an autonomous robot can generate a causal relationship set using the generation unit 215 based on external environment data obtained from entities. Furthermore, it can generate another causal relationship set using the transformation unit based on this newly generated set. In addition, it can update the learned causal relationship set using at least one of the integration unit and extrapolation unit, based on the newly generated causal relationship set and the learned decision model.

[0128] When using PUF-based CRP to acquire external environmental data from an authenticated entity, particularly a human, a certain level of trust is required between the autonomous robot and the entity. In this case, the autonomous robot acquiring the external environmental data authenticates with the entity (or, if the entity is a human, its information processing device) to obtain agreement regarding the acquisition of external environmental data from that entity. Once this agreement is reached, a session is created between the autonomous robot and the entity, and the autonomous robot can acquire external environmental data from that entity.

[0129] Authentication will be explained in more detail. In this embodiment, authentication will be performed using a PUF (Physically Hard to Copy) circuit included in at least the subsystem of the autonomous robot. The second processing chip may also include a PUF circuit. If the entity (or information processing device owned by a human) has a second processing chip that executes an encryption processing unit in its subsystem, authentication may be performed using the PUF circuit on the entity side to perform bidirectional authentication. Here, mutual authentication using a CRP (Challenge Response Pair) will be described. Figure 16 shows a flowchart of the processing procedure for authentication using a CRP with a PUF according to this embodiment.

[0130] Figure 16 provides a detailed explanation of each of the following steps: the registration step, the sharing step, the first authentication step (autonomous robot 10), and the second authentication step (information processing device).

[0131] (Registration Step) First, during the provisioning or pre-shipment stage of each device, the CRP of each device is stored in another information processing device (such as a server). Note that this step may also be performed on an entity or information processing device owned by an entity, or semiconductor device which is a physical component of an autonomous robot. Initially, in S1101, the information processing device is Challenge C i It sends this to each entity. Specifically, the information processing device is Challenge C iThe autonomous robot 10 and the entity transmit the following. The entity's cryptographic processing unit identifies the challenge and response combination in the entity-side PUF from the PUF circuit. The entity that receives the challenge Ci uses its cryptographic processing unit to identify the challenge C i From there, the response R is received through a predetermined function called PUF. Bi You can obtain this. Similarly, Challenge C i Upon receiving the challenge Ci, the autonomous robot 10 receives a response R via a predetermined function, PUF, from the challenge Ci via the cryptographic processing unit 216. Ai You can obtain this.

[0132] Then, in S1102, the encryption processing unit of the entity responds R Bi The information is transmitted to the information processing device. The information processing device then receives an ID to identify the information processing device. B , and CRP B (C i , R Bi ) is registered in the database. Similarly, in S1102, the encryption processing unit 216 of the autonomous robot 10 receives response R Ai The information is transmitted to the information processing device. The information processing device then receives an ID to identify the autonomous robot 10. A , and CRP A (C i , R Ai The data is registered in the database. Steps S1101 and S1102 are repeated multiple times as needed, and multiple sets of CRP are registered in the database for each device.

[0133] (Shared step) In S1103, the autonomous robot 10 receives the entity's CRP from the information processing device. B When a request is made, the information processing device provides the autonomous robot 10 with the CRP stored in the database. B It shares at least a part of it. Similarly, the entity provides the CRP of the autonomous robot to the information processing device. A When a request is made, the entity tells the information processing device the CRP stored in the database A This involves sharing at least a portion of the data. This prepares the entities for mutual authentication.

[0134] (First Authentication Step) The first authentication step is the process by which the entity authenticates the autonomous robot 10 (referred to as robot authentication). Specifically, in S1104, the cryptographic processing unit of the entity is Challenge C i The autonomous robot 10 receives the following message. The encryption processing unit 216 of the autonomous robot 10 receives the following message: Challenge C i Encrypts the data and uses a predetermined physically hard copy function called PUF. A (C i ) based on response R' Ai The following is calculated. In S1105, the cryptographic processing unit 216 of the autonomous robot 10 calculates the response R' Ai The entity sends this to the entity. This process is repeated multiple times as needed, or a predetermined number of challenges are sent to verify a predetermined number of CRPs. The entity's cryptographic processing unit 216 then, for example, if a certain percentage of the responses R' Ai = R Ai If it can be verified that this is the case, then the entity whitelist WL B ID for identifying the autonomous robot 10 A Register.

[0135] This state indicates that the entity has been able to confirm that the autonomous robot 10, as the device being authenticated, is indeed the true autonomous robot 10, that is, that the authentication of the autonomous robot 10 by the entity was successful in the first authentication step.

[0136] (Second Authentication Step) The second authentication step is a process in which the autonomous robot 10 authenticates the entity (referred to as entity authentication). Specifically, in S1106, the cryptographic processing unit 216 of the autonomous robot 10 performs Challenge C i The entity sends this to the entity. The entity's cryptographic processing unit is Challenge C i Encrypts the data and uses a predetermined physically hard copy function called PUF. B (C i ) based on response R' Bi Calculate.

[0137] In S1107, the encryption processing unit of the entity processes the response R' BiThe following is sent to the autonomous robot 10. Similar to robot authentication, this process is repeated multiple times as needed, or a predetermined number of challenges are sent to verify a predetermined number of CRPs. The cryptographic processing unit 216 of the autonomous robot 10 then, for example, if a certain percentage of the responses R' Bi = R Bi If this can be verified, then the whitelist WL of the autonomous robot 10 A ID for entity identification B The entity is registered. The second authentication step indicates that the authentication of the entity from the autonomous robot 10 was successful.

[0138] In this embodiment, the first and second authentication steps enable more accurate mutual authentication than in the conventional method, and as a result, reliability and security between entities are ensured by the learning system 1.

[0139] Ethical data refers to a collection of data related to causal relationships, including groups of causal relationships where no corresponding behaviors or states are associated, groups of causal relationships where behaviors or states are associated, sets of first signals and environmental data for generating causal relationships, and decision models that have undergone reinforcement learning using these causal relationships.

[0140] The learning system 1 further includes an aggregation unit, a CA unit, an encryption processing unit, and a management unit. In this embodiment, the information processing device 600 includes an aggregation unit 617 and a database (DB). As shown in Figure 17, one or more autonomous robots 10 (agent systems) can upload ethical data acquired in the physical environment to the information processing device 600, and the aggregation unit 617 stores the uploaded ethical data in the DB. The information processing device 600 generates and updates causal relationship groups using one or more of the conversion unit 614, integration unit 615, and extrapolation unit 616, or learns a decision model using the updated causal relationship groups and the learning unit 611, and the aggregation unit 617 can distribute the undistributed ethical data stored in the DB to one or more autonomous robots 10. This aggregates the experiences of each autonomous robot 10, resulting in an autonomous robot that adheres to a more highly ethical code of conduct. Note that the autonomous robot 10 may also have an aggregation unit.

[0141] Diagram 18, illustrating data storage using blockchain, is a conceptual diagram of a blockchain that holds ethical data. Nodes, including the autonomous robot 10, can hold ethical data using blockchain technology. In a blockchain, a transaction (Tx) is signed by the client using a private key, and the node verifies the validity of the signature. This ensures that the Tx was issued by a genuine client, thereby guaranteeing the client's authenticity. Each block contains metadata such as the hash value of the previous block and a timestamp, as well as one or more Tx data. Each block may also contain hash values ​​that summarize the Tx, such as a Merkle root, and hash values ​​for each Tx. Tx data includes the Tx ID, the sender's (client's) public key, a digital signature, and ethical data. The blockchain has a chain structure that holds a hash value assigned to each block based on the Tx, and the hash value of the previous block. As a result, if any part of the ethical data stored as Tx in a block is tampered with, an inconsistency will occur in the hash value, and the consistency in subsequent blocks will also be broken. Therefore, data tampering becomes detectable, and the integrity of ethical data is guaranteed throughout the entire blockchain. Here, it is assumed that clients also function as nodes, and will be referred to as nodes from now on. In the case of a permissioned blockchain, the transaction (Tx) includes the node's public key certificate, guaranteeing the node's authenticity. Furthermore, the transaction data of each block includes this public key certificate as needed.

[0142] For example, ethical data such as a set of causal relationships acquired by the autonomous robot 10 in physical space, or a judgment model trained by the information processing device 600, can be stored in the blockchain. At this time, the cryptographic processing unit 216 (a subsystem 200 of the autonomous robot 10) or the cryptographic processing unit 619 (the information processing device 600) generates a public key and an electronic signature of the transaction (Tx) that includes at least a portion of the ethical data, based on the private key.

[0143] A node uses the management unit 217 or 620 to store one or more transactions (Tx) including a public key and digital signature in the current block that constitutes the blockchain network (BCN), and uses the cryptographic processing unit 216 or 619 to verify the signature. After a certain period of time, or when the number of successfully verified transactions reaches a certain number or a certain amount of data, the node reaches a consensus based on a consensus algorithm as needed, confirms the block, and generates a new block. At this time, the node generates a hash value for the transactions in the current block and stores the hash value of the previous block and the hash values ​​based on each transaction in the current block in the current block. This allows ethical data to be stored in the BCN.

[0144] Figure 19 shows an example of a self-contained BCN in which only one autonomous robot 10, which acts as a node, participates. The autonomous robot 10's dedicated auxiliary storage device can store ethical data in the form of a BCN, such as causal relationship groups and datasets acquired in the physical environment.

[0145] Figure 20 shows an example of a BCN involving one or more autonomous robots 10 and one or more nodes (information processing devices 600) other than the autonomous robots. The management unit 217 or management unit 620 of each node broadcasts the stored BCN and Tx requests to share them among the nodes. Each node can store ethical data, such as trained decision models and updated causal relationship sets, in the form of a BCN in an auxiliary storage device (autonomous robots have a dedicated auxiliary storage device). Alternatively, data indicating the integrity of the trained decision models and updated causal relationship sets, etc., which are stored in the DB of the information processing device 600 including the nodes and distributed to the autonomous robots 10, may be stored in Tx instead of ethical data (such as a hash).

[0146] Encryption using a private key via PUF Next, we will explain encryption using a PUF (Physically Hard to Copy Function) in this learning system 1. Figure 21 shows a schematic diagram of encryption using a PUF in one embodiment. Ethical data is required to be complete, confidential, and authentic in order to contribute to compliance with ethical codes of conduct. The autonomous robot encrypts ethical data stored in the subsystem's dedicated auxiliary storage device or ethical data transmitted externally. As shown in the figure, for example, when encrypting the Tx of a BCN, the encryption key and private key obtained by the PUF can be used. Furthermore, when storing ethical data in an authorized BCN configuration using a public key certificate, the system includes an autonomous robot 10 or information processing device 600 equipped with a CA unit 618 that functions as a certification authority. The autonomous robot 10 or information processing device 600 equipped with the CA unit 618 does not necessarily have to be a node.

[0147] Figure 22 shows a flowchart of the procedure for generating a secret key using PUF according to this embodiment. The cryptographic processing unit 216 is assumed to have a fuzzy extractor. The fuzzy extractor includes a key generation algorithm (GEN) for creating a secret key from a PUF response and a playback algorithm (REP) for reproducing the same secret key from a noisy response.

[0148] During the provisioning phase or pre-shipment phase of the second processing chip 210 (autonomous robot 10), the public key (public key certificate) of the second processing chip 210 is stored in the information processing device. First, in S1201, the cryptographic processing unit 216 of the autonomous robot identifies the challenge and response combination of the PUF from the PUF circuit provided by the second processing chip 210. Then, using the identified response and GEN, the PUF private key (SK PUF The 216 generates SK again, and also generates helper data (W), and registers the helper data W in a dedicated auxiliary storage device. PUFREP is used when generating it. Helper data W functions as auxiliary information to stably reproduce the response of the PUF. For example, the bit sequence output by the SRAM PUF is affected by noise and temperature changes, but appropriate reproduction can be achieved by using helper data W. Also, SK PUF If the response (group of responses) to obtain uses only a specific (partial) CRP in the PUF circuit, the challenge (group of challenges) corresponding to that response is registered in a dedicated auxiliary storage device.

[0149] Then, in S1202, the encryption processing unit 216 of the autonomous robot 10 uses the encryption accelerator 250 to perform SK PUF From PUF public key (PK PUF ) generates.

[0150] In S1203, the encryption processing unit 216 of the autonomous robot 10 is PK PUF The information is transmitted to the information processing device 600, and the information processing device 600 uses the ID to identify the autonomous robot 10 along with PK PUF The CA unit 618 of the information processing device 600 functions as a certification authority (CA) and registers the received PK. PUF Encrypt it with your own CA private key and create a public key certificate (CERT(PK) PUF By generating )) and registering it in the database, PK PUF The information processing device registers the public key certificate CERT (PK). PUF The program is created and sent to the autonomous robot 10. Note that in the case of BCN, which does not require CERT, it is not necessary to execute at least S1203.

[0151] The control unit 217 of the autonomous robot 10 processes the SK generated in S1201 and S1202. PUF The ethical data is encrypted using PK, and the encrypted ethical data is PK. PUF It generates a Tx including CERT (PKPUF), and another node performs signature verification and registers it with the BCN.

[0152] Encryption using a PUF encryption key. Alternatively, data may be encrypted using a PUF encryption key and information for identity verification (in this case, a shared key). Figure 23 shows a flowchart of the procedure for generating a PUF encryption key according to this embodiment. The encryption processing unit 216 is assumed to have a fuzzy extractor including GEN and REP.

[0153] In S1301, the cryptographic processing unit 216 of the autonomous robot 10 identifies the response and uses GEN to form the symmetric key, the PUF key K PUF And generates W. The information processing device also transmits a common key CK, which is shared between the autonomous robot 10 and the information processing device, to the autonomous robot 10. Then, in S1302, the cryptographic processing unit 216 of the autonomous robot 10 uses the encryption accelerator 250 to convert the common key CK into the PUF key K PUF The encrypted data obtained (K PUF (CK) and W are registered. The encryption processing unit 216 also registers K again. PUF When generating [this], use W and REP.

[0154] Also, PUF key K PUF If the response (group of responses) to obtain uses only a specific (partial) CRP in the PUF circuit, the challenge (group of challenges) corresponding to that response is registered in a dedicated auxiliary storage device.

[0155] The autonomous robot 10 stores an encrypted common key in a dedicated auxiliary storage device, and the information processing device and the autonomous robot 10 can exchange ethical data using the common key.

[0156] Here, we will explain the schematic image of the PUF in this embodiment. The learning system 1 uses an SRAM PUF to encrypt ethical data. The SRAM PUF referred to here is a function that is physically impossible (or difficult) to replicate, and it generates a device-specific identifier and an encryption key, which will be described later, by utilizing the variations in the initial state (at the time of manufacture) of the SRAM on the semiconductor chip.

[0157] In this embodiment, in addition to the SRAM PUF (initial value), other types of PUFs can also be used, such as the Arbiter PUF (signal transmission delay), the RNT-based SRAM PUF (variation due to traps in the gate insulating film), the DRAM PUF, the ReRAM PUF, and the ring oscillator PUF. The PUF circuit may be a group of memory cells of the SRAM or DRAM, a group of resistive cells of the resistive random-access memory of the ReRAM, an Arbiter circuit, or a ring oscillator circuit.

[0158] The control unit 217 obtains a response from the PUF circuit based on the challenge. The challenge may be an array containing multiple challenges.

[0159] When power is applied to a semiconductor device, the semiconductor elements such as transistors naturally settle into either a 0 or 1 value for each bit due to minute manufacturing variations. This is sometimes likened to a fingerprint and is called a silicon fingerprint. This initial pattern differs from device to device, is highly reproducible, and cannot be reproduced by other devices. For example, if the response of the PUF circuit is lost, such as when a cell used for the SRAM PUF (initial value) is used for another purpose, the management unit 217 may acquire the response indicated by the initial value of the memory, temporarily store the CRP in another area of ​​the enclave 200, and then return a response using the stored CRP.

[0160] More specifically, after power-on, a predetermined challenge is applied to detect the row (Word Line) and column (Bit Line) states of each bit, respectively, and as a result, a predetermined function specific to the PUF circuit possessed by the device can be obtained.

[0161] Furthermore, in 2nm technology nodes and beyond, an improvement in PUF entropy can be expected as the transistor density increases. In other words, as the footprint improves, variability becomes more likely. This is because, as the miniaturization of the element reduces its footprint (occupied area), the impact of minute manufacturing variations becomes relatively larger. Higher PUF entropy leads to greater uniqueness, or in other words, higher security, making it more suitable for the symbiotic support described in this embodiment. Thus, in the learning system 1, the integrity, confidentiality, and authenticity of ethical data can be guaranteed by using keys obtained from the unique characteristics of the semiconductor device. Below, specific examples of semiconductor devices that can be used in this embodiment will be described.

[0162] Semiconductor Device Here, we will describe a semiconductor device used to realize symbiosis between an autonomous robot and a human in this embodiment. Figure 24 shows an example of a 2D package semiconductor device according to this embodiment. As shown in the figure, the 2D package semiconductor device 500 is composed of a substrate 501 and one or more chips arranged on the substrate. Specifically, it is composed of an enclave 200 as an SoC including a second processing chip 210, a dedicated auxiliary storage device 230, an AI / ML accelerator 240, an encryption accelerator 250, a dedicated boot ROM 260, a firmware circuit that performs boot loader activation, etc., and a dedicated bus interface 320, and the substrate 501.

[0163] In another embodiment, the semiconductor device 500 may consist of a chiplet based on a main function, an associated chip (for example, a boot ROM or auxiliary storage), and a substrate 501.

[0164] Figure 25 shows an example of a semiconductor device in a 2.5D package according to this embodiment. As shown in Figures 25(a) to (c), the semiconductor device 510 in a 2.5D package consists of a main system 100 as an SoC, an enclave 200 as an SoC, multiple memory components 400, a substrate 501, and an interposer 502. In this configuration, the first processing chip and the second processing chip are connected to the same interposer substrate.

[0165] In this configuration, the interposer 502 is a silicon interposer. The memory 400 includes a main memory 120, an auxiliary memory 130, a dedicated main memory 220, and other components necessary for the characteristic determination processing described herein.

[0166] The memory arrays 400 can also be configured as, for example, a multi-stage memory array. In this case, the multi-stage memory array can be configured using through-electrodes such as vias. For example, a multi-stage memory array may be configured using 3D stacking technology, which involves vertically stacking multiple DRAM chips using through-silicon vias (TSVs).

[0167] The semiconductor device 510 packages the main system 100, the enclave 200, and associated memory groups. For example, the semiconductor device 510 can be incorporated into the autonomous robot 10 during the manufacturing stage.

[0168] Figure 26 shows an example of a 2.xD package semiconductor device according to this embodiment. As shown in Figures 26(a) and (b), the 2.xD package semiconductor device 520 consists of a main system 100 as an SoC, an enclave 200 as an SoC, multiple memory components 400, a substrate 501, and an interposer 502. The interposer 502 in this configuration is an RDL (Redistributed Layers) interposer.

[0169] As shown in Figure 26(c), the semiconductor device 520 has a silicon bridge configured in it, which enables faster processing compared to the semiconductor device 510 in a 2.5D package. In this configuration as well, the semiconductor device 520 can be incorporated into the autonomous robot 10 during the manufacturing stage.

[0170] Figure 27 shows an example of a 3D package semiconductor device according to an embodiment of the present disclosure. As shown in Figures 27(a) and (b), in the semiconductor device 530, the enclave 200 is formed by a subchip 200a configured as an SoC including a second processing chip and a dedicated main memory 220 positioned (wired) directly above the subchip 200a. The dedicated main memory 220 is positioned directly above the second processing chip 210. In addition, in the semiconductor device, the enclave may be formed by a subchip and a dedicated auxiliary memory device positioned directly above the subchip.

[0171] Furthermore, as shown in Figure 27(c), the semiconductor device 530 has a silicon bridge configured similarly to the semiconductor device 520 shown in Figure 26, enabling high-speed processing between the enclave 200 and the memory devices 400.

[0172] Furthermore, the second processing chip in the semiconductor device according to this embodiment may include, for example, a semiconductor element with a 2 nm technology node.

[0173] The second processing chip may also include a group of transistors having nanosheet channels, a group of transistors with an all-around gate structure (GAA structure), a group of transistors having a fork sheet channel, or a group of CFETs (complementary field-effect transistors).

[0174] For example, in this embodiment, the transistors constituting the second processing chip include a group of all-around gate structure field-effect transistors (GAA structure type), and the transistors constituting the second processing chip may also include a group of all-around gate structure field-effect transistors utilizing multiple nanosheets. The transistors constituting the second processing chip may also include a group of transistors having fork sheets. Furthermore, the transistors constituting the second processing chip may also include a group of transistors with a CFET (complementary FET) structure in which NMOS (N-type field-effect transistor) and PMOS (P-type field-effect transistor) are stacked vertically. By configuring the transistors constituting the second processing chip to include a group of transistors having nanosheet channels, a group of transistors having fork sheets, or a group of transistors with a CFET (complementary FET) structure, power consumption for functions that realize symbiosis with humans can be reduced, and the risk of functions that realize symbiosis with humans failing due to single-event effects caused by radiation can be suppressed.

[0175] As described above, this disclosure provides a learning system that effectively teaches ethical interpretations to the learning models of entities such as autonomous robots capable of coexisting with humans, taking the singularity into account.

[0176] Furthermore, while this embodiment mainly describes the learning system of the decision model in detail, it goes without saying that the learning method and learning program for the decision model can also achieve the same effects as disclosed herein.

[0177] The disclosures herein include the following learning systems, learning methods, and learning programs.

[0178] (Item 1) A learning system having at least an extraction unit, an interpretation unit, and a learning unit, wherein the system stores a group of causal relationships including causes which are actions selected by an agent and results which correspond to the causes, the extraction unit extracts feature quantities from the environment which correspond to each of the one or more results which constitute the group of causal relationships, the interpretation unit estimates an ethical interpretation of each of the one or more results based on the feature quantities and identifies an ethical outcome in the group of causal relationships, and the learning unit performs reinforcement learning processing of a decision model by setting a reward for the selection of the actions which correspond to the ethical outcome and / or causes which are related to this outcome.

[0179] (Item 2) The learning system according to Item 1, wherein the environment is a virtual environment, at least some of the causes in the causal relationship group are associated with agent behavior information related to the causes, the causal relationship group corresponds to a probabilistic policy that allows the agent to selectively select the causes associated with the behavior information in a predetermined state in the environment, the results in the causal relationship group are associated with a state related to the result associated with the cause, and the learning unit selects the cause included in the policy in the predetermined state and updates the state of the virtual environment based on the state associated with the result associated with the selected cause and the state estimated by the behavior information associated with the selected cause.

[0180] (Item 3) The learning system according to Item 2, wherein the extraction unit generates the feature quantities based on the situation data obtained from the updated state, and the interpretation unit features the instructions requesting the ethical interpretation of the situation data and inputs them into a large-scale language model, and identifies the ethical evaluation result from the output.

[0181] (Item 4) The learning system according to Item 2 or 3, wherein the agent includes an autonomous robot, and the behavioral information includes control information of the autonomous robot.

[0182] (Item 5) The learning system described in Item 4, wherein the virtual environment is a physical simulation environment in which the autonomous robot operates.

[0183] (Item 6) The learning system described in Item 4 or 5, wherein the virtual environment is a world model in which the autonomous robot operates.

[0184] (Item 7) The learning system described in Item 6, wherein the learning unit performs model-based reinforcement learning processing using a world model that estimates state transitions in a virtual environment by inputting one or more image data acquired in a physical environment or another virtual environment and the behavior information of the autonomous robot as features.

[0185] (Item 8) A learning system according to any one of Items 5 to 7, wherein at least a portion of the states in the virtual environment are given by domain randomization.

[0186] (Item 9) The learning system according to any one of Items 1 to 8, comprising one or more first processing chips, one or more second processing chips, and an autonomous robot, wherein the autonomous robot comprises a main system including at least one of the one or more first processing chips and an operating unit, and a subsystem including at least one of the one or more second processing chips and storing the decision model, wherein the subsystem is an enclave isolated as a control system independent of the main system, and at least a part of it, including the second processing chip, is isolated as an independent circuit system.

[0187] (Item 10) The learning system according to Item 9, wherein the subsystem comprises the extraction unit which is performed using at least one of the one or more second processing chips, the interpretation unit which is performed using at least one of the chips, and the learning unit which is performed using at least one of the chips.

[0188] (Item 11) The learning system according to item 9 or 10, further comprising one or more information processing devices, wherein at least one of the one or more information processing devices is the extraction unit, at least one is the interpretation unit, and at least one is the learning unit.

[0189] (Item 12) The learning system according to any one of Items 9 to 11, wherein the main system includes a detection unit for acquiring environmental data, and the subsystem stores the environmental data acquired in the physical environment for generating the causal relationship group.

[0190] (Item 13) The learning system according to Item 12, wherein the subsystem has a generation unit that is executed using at least one of the one or more second processing chips, and the generation unit generates the causal relationship group based on the environmental data.

[0191] (Item 14) The learning system according to Item 13, wherein the subsystem has a learning unit that is executed using at least one of the one or more second processing chips, and the learning unit performs reinforcement learning processing on the decision model based on the causal relationship group generated by the generation unit.

[0192] (Item 15) The learning system according to any one of Items 9 to 14, wherein the subsystem comprises a detection unit that is performed using at least one of the one or more second processing chips, and a generation unit that is performed using at least one of the second processing chips, the detection unit acquires environmental data in the physical environment, and the generation unit generates the causal relationship group based on the environmental data.

[0193] (Item 16) The generation unit is a learning system according to Item 15, which associates the cause with the action information of an agent corresponding to the cause.

[0194] (Item 17) The learning system according to Item 15 or Item 16, wherein the generation unit associates the cause and / or the result with explanatory information that explains the cause and / or the result.

[0195] (Item 18) The learning system according to any one of Items 15 to 17, wherein the subsystem further comprises a conversion unit that is executed using at least one of the one or more second processing chips, the conversion unit abstracts or concretizes language data corresponding to each of the at least one of the causes and effects constituting the causal relationship group, based on ontology dictionary data.

[0196] (Item 19) The learning system according to Item 18, wherein the causes and / or effects constituting the causal relationship group are associated with explanatory information including text that explains the causes and / or effects, and the conversion unit performs the abstraction or concretization based on the explanatory information.

[0197] (Item 20) The learning system according to any one of Items 15 to 19, wherein the subsystem further comprises an integration unit that is executed using at least one of the one or more second processing chips, the integration unit integrates two or more independent causal relationship groups by identifying matching or similar causes or effects among the causes and effects constituting a plurality of causal relationship groups and commonizing the matching or similar causes or effects.

[0198] (Item 21) The learning system described in Item 20, wherein the integration unit inputs multiple causes or multiple results into LLM (Large Language Models) and obtains evaluation results for their matches or similarities.

[0199] (Item 22) The learning system according to Item 20, wherein the integration unit inputs multiple causes or multiple results into a distance function or similarity function and obtains evaluation results for their agreement or similarity.

[0200] (Item 23) The learning system according to Item 20, wherein the causes and / or effects constituting the causal relationships are associated with descriptive information including text describing the causes and / or effects, and the integration unit identifies matching or similar causes or effects based on the descriptive information associated with the causes or effects included in the independent causal relationships, and integrates the independent causal relationships.

[0201] (Item 24) The learning system according to any one of Items 15 to 23, wherein the subsystem further comprises a detection unit that is performed using at least one of the one or more second processing chips, and an extrapolation unit that is performed using at least one of the chips, the extrapolation unit identifies a pair of causes or effects that are in a hierarchical relationship among the causes and effects that constitute separate causal relationship groups based on ontology dictionary data, and extrapolates another causal relationship group that includes the other cause or effect of the pair to the causal relationship group that includes one of the causes or effects of the pair.

[0202] (Item 25) The learning system described in Item 24, wherein the causes and / or effects constituting the causal relationship group are associated with explanatory information including text that explains the cause and / or effect, and the extrapolation unit extrapolates the causal relationship group based on the explanatory information.

[0203] (Item 26) A learning system according to any one of Items 1 to 25, further comprising a server device including an aggregation unit and a database (DB), and one or more agent systems including at least the learning unit, wherein the agent systems encrypt and transmit ethical data including a decision model on which the reinforcement learning process has been performed based on an encryption key unique to each of the one or more agent systems, and the aggregation unit receives the ethical data held by each of the one or more agent systems, verifies the received ethical data based on the encryption key, and stores it in the DB.

[0204] (Item 27) A learning system according to any one of Items 15 to 25, further comprising a server device including an aggregation unit and a database (DB), and one or more agent systems including at least the generation unit, wherein the agent systems encrypt and transmit ethical data including the causal relationship group generated based on an encryption key unique to each of the one or more agent systems, and the aggregation unit receives the ethical data held by each of the one or more agent systems, verifies the received ethical data based on the encryption key, and stores it in the DB.

[0205] (Item 28) A learning system according to any one of Items 1 to 25, further comprising one or more nodes including a cryptographic processing unit and a management unit, and at least the learning unit and a storage device that stores ethical data including a decision model on which the reinforcement learning process has been performed by the learning unit, wherein the cryptographic processing unit generates a physically hard-to-replicate function (PUF), generates a private key based on the PUF, generates a public key and an electronic signature of a transaction (Tx) including at least a portion of the ethical data based on the private key, the management unit stores one or more Tx including the public key and the electronic signature in the current block constituting a blockchain network (BCN), generates a hash value for at least a portion of the previous block and a hash value for at least the Tx in the current block, and stores the respective hash values ​​of the previous block and the current block in the current block.

[0206] (Item 29) A learning system according to any one of Items 15 to 25, further comprising one or more nodes including a cryptographic processing unit and a management unit, and at least the generation unit and a storage device that stores ethical data including the causal relationship group generated by the generation unit, wherein the cryptographic processing unit generates a physically hard-to-replicate function (PUF), generates a private key based on the PUF, generates a public key and an electronic signature of a transaction (Tx) including at least a part of the ethical data based on the private key, the management unit stores one or more Tx including the public key and the electronic signature in the current block constituting a blockchain network (BCN), generates a hash value for at least a part of the previous block and a hash value for at least the Tx in the current block, and stores the respective hash values ​​of the previous block and the current block in the current block.

[0207] (Item 30) The learning system described in Item 28 or Item 29, wherein Tx includes the ethical data or data indicating the integrity of the ethical data.

[0208] (Item 31) The learning system according to any one of Items 28 to 30, wherein the node comprises a main system that performs processing by one or more first processing chips, one or more second processing chips, and a dedicated auxiliary storage device which is the storage device, and the subsystem has an encryption processing unit performed by at least one of the one or more second processing chips, and a management unit performed by at least one of the second processing chips, wherein the subsystem is isolated as a control system independent of the main system, and at least a part of it, including the second processing chips, is isolated as an independent circuit system, and the dedicated auxiliary storage device stores one or more blocks constituting the BCN.

[0209] (Item 32) The learning system according to Item 31, wherein the node is an autonomous robot, the main system has one or more first processing chips and operating units, and the subsystem has one or more second processing chips and the dedicated auxiliary storage device storing the decision model.

[0210] (Item 33) The learning system according to Item 31 or Item 32, wherein the subsystem includes a PUF circuit, and the cryptographic processing unit identifies a challenge-response combination in the PUF from the PUF circuit, and obtains the secret key based on the challenge-response combination.

[0211] (Item 34) The learning system according to Item 33, wherein the PUF circuit is a group of memory cells of a static random access memory that provides a unique PUF based on element characteristics, and the group of memory cells includes field-effect transistors having one or more channels arranged in a direction orthogonal to the same plane of a substrate connected to the second processing chip.

[0212] (Item 35) The learning system according to Item 34, wherein the memory cell group includes N-type field-effect transistors and P-type field-effect transistors arranged in the orthogonal direction.

[0213] (Item 36) The learning system according to Item 34, wherein the memory cell group includes a group of transistors having fork sheets.

[0214] (Item 37) The learning system described in Item 34, wherein the element characteristics are the initial values ​​of each of the memory cell groups.

[0215] (Item 38) The learning system according to Item 34, wherein the element characteristics are the current oscillation characteristics due to flicker noise in each of the memory cell groups.

[0216] (Item 39) The learning system according to Item 9, Item 10, Item 13, Item 14, Item 15, Item 18, Item 20, Item 24, Item 31, Item 32, or Item 34, wherein the second processing chip includes a group of field-effect transistors having nanosheet channels.

[0217] (Item 40) The learning system according to Item 9, Item 10, Item 13, Item 14, Item 15, Item 18, Item 20, Item 24, Item 31, Item 32, or Item 34, wherein the second processing chip includes a group of field-effect transistors having fork sheet channels.

[0218] (Item 41) The second processing chip is a learning system according to item 9, item 10, item 13, item 14, item 15, item 18, item 20, item 24, item 31, item 32, or item 34, which includes a group of CFETs (complementary field-effect transistors).

[0219] (Item 42) A learning method performed by one or more computers, wherein at least one of the one or more computers stores a set of causal relationships including causes which are actions selected by an agent and results which correspond to the causes; extracts feature quantities from the environment which correspond to each of the one or more results which constitute the set of causal relationships; at least one of the one or more computers estimates an ethical interpretation of each of the one or more results based on the feature quantities and identifies an ethical outcome in the set of causal relationships; and at least one of the one or more computers performs reinforcement learning processing of a decision model by setting a reward for the selection of the actions which correspond to the ethical outcome and / or the causes which relate to this outcome.

[0220] (Item 43) A learning program executed by one or more computers, wherein at least one of the one or more computers, which stores a group of causal relationships including causes that are actions selected by an agent and results corresponding to those causes, functions as an extraction unit that extracts feature quantities from the environment corresponding to each of the one or more results that constitute the group of causal relationships; at least one of the one or more computers functions as an interpretation unit that estimates the ethical interpretation of each of the one or more results based on the feature quantities and identifies the ethical outcome in the group of causal relationships; and at least one of the one or more computers functions as a learning unit that performs reinforcement learning processing of a decision model by setting rewards for the selection of the actions corresponding to the ethical outcome and / or causes related to that outcome.

[0221] (Item 44) A system comprising at least a generation means, an extraction means, an interpretation means, and a learning means, wherein the generation means generates a group of causal relationships including an agent's behavioral choices (causes) in an environment and the results associated with those causes; the extraction means extracts feature quantities from the environment corresponding to each of the one or more results constituting the group of causal relationships; the interpretation means estimates an ethical interpretation of each of the one or more results based on the feature quantities and identifies the ethical results in the group of causal relationships; and the learning means performs reinforcement learning processing for the agent by setting rewards for the behavioral choices corresponding to causes that are subordinate to the ethical results.

[0222] (Item 45) The system according to Item 44, further comprising a conversion means, wherein the conversion means abstracts (higher-level conceptualization) or concretizes (lower-level conceptualization) language data corresponding to each of the at least one of the causes and effects constituting the causal relationship group, based on ontology dictionary data.

[0223] (Item 46) The system according to Item 44, further comprising an integrating means, wherein the integrating means identifies matching or similar causes or effects among the causes and effects constituting the causal relationship group, and integrates two or more independent causal relationship groups by commonizing the matching or similar causes or effects.

[0224] (Item 47) The system according to Item 44, further comprising extrapolation means, wherein the extrapolation means identifies a pair of causes or effects that are in a hierarchical relationship among the causes and effects constituting the causal relationship group based on ontology dictionary data, and extrapolates another causal relationship group containing the other cause or effect of the pair to the causal relationship group containing one of the causes or effects of the pair.

[0225] (Item 48) The system according to Item 44, further comprising entities located in the vicinity of the agent, wherein the entities have detection means for detecting the state of the entities, and the generation means generates the causal relationship group based on the state detected by the detection means.

[0226] (Item 49) The system according to Item 44, further comprising aggregation means and a server device including a database (DB), wherein the aggregation means receives ethical data including the causal relationship group or the reinforcement learning process model possessed by each of one or more agents, verifies the received ethical data based on an encryption key unique to each of the one or more agents, and stores it in the DB.

[0227] (Item 50) A system comprising one or more nodes having at least cryptographic means and management means, wherein the cryptographic means generates a physically hard-to-replicate function (PUF), generates a private key based on the PUF, generates a public key and a digital signature of a transaction (Tx) based on the private key, the management means stores one or more Tx, including the public key and digital signature, in the current block constituting a blockchain network (BCN), generates a hash value for the Tx in the current block, and stores the hash values ​​of the previous and current block, respectively, in the current block.

[0228] (Item 51) The system described in Item 50, wherein Tx includes data corresponding to combinations of causes and effects that constitute a causal relationship.

[0229] (Item 52) The system according to Item 51, wherein the node further comprises a generation means, the generation means generates a group of causal relationships comprising the behavioral selection (cause) of a node which is an agent in the environment and the results associated with the cause, and the management means stores at least a portion of the group of causal relationships generated by the generation means as the current block Tx.

[0230] (Item 53) The system according to Item 52, wherein the node further comprises a driving means and a calculation means, the driving means converts a signal from the calculation means into a motion of the node, the calculation means controls the driving means, including operating and stopping it, by inputting the signal to the driving means, and the generating means generates the cause and effect associated with the motion in the environment as the causal relationship group.

[0231] (Item 54) The system according to Item 53, wherein the arithmetic means comprises a first arithmetic means and a second arithmetic means, the first arithmetic means and the second arithmetic means are connected to a first main memory means and a first auxiliary storage means via a first interface, the second arithmetic means is connected to a second auxiliary storage means accessible only from the second arithmetic means via a second interface, and the second main memory means is connected to a second main memory means accessible only from the second arithmetic means via the first or second interface, and the second auxiliary storage means stores one or more blocks constituting the blockchain network.

[0232] (Item 55) The PUF is a system as described in Item 50, based on the element characteristics of a static random access memory (SRAM) including a group of memory cells.

[0233] (Item 56) The system according to Item 55, wherein the memory cell group has an all-around gate structure field-effect transistor.

[0234] The present invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention.

[0235] 10 Autonomous robot 20 Human 100 Main system 110 First processing chip (first processing unit, motion control unit) 120 Main memory 130 Auxiliary memory 140 Sensor (detection unit) 150 Actuator (motion unit) 200 Enclave (subsystem) 210 Second processing chip (second processing unit) 211 Learning unit 212 Judgment unit 213 Extraction unit 214 Motion control unit 215 Cryptographic processing unit 217 Management unit 218 Detection unit 220 Dedicated main memory (dedicated area) 230 Dedicated auxiliary memory 240 AI / ML accelerator 250 Encryption accelerator 260 Dedicated boot ROM 270 Memory protection 310 Common bus interface 320 Dedicated bus interface 400 Memory components 500 Semiconductor device in 2D package 510 2. xD package semiconductor device 520 2.5D package semiconductor device 530 3D package semiconductor device 600 Information processing device 611 Learning unit 612 Extraction unit 613 Interpretation unit 614 Conversion unit 615 Integration unit 616 Extrapolation unit 617 Aggregation unit 618 CA unit 619 Cryptographic processing unit 620 Management unit

Claims

1. A learning system having at least an extraction unit, an interpretation unit, and a learning unit, wherein the system stores a group of causal relationships including causes which are actions selected by an agent and results which correspond to the causes, the extraction unit extracts feature quantities from the environment which correspond to each of the one or more results which constitute the group of causal relationships, the interpretation unit estimates an ethical interpretation of each of the one or more results based on the feature quantities and identifies ethical results in the group of causal relationships, and the learning unit performs reinforcement learning processing of a decision model by setting a reward for the selection of the actions which correspond to the ethical results and / or causes which are related to these results.

2. The learning system according to claim 1, wherein the environment is a virtual environment, at least some of the causes in the causal relationship group are associated with agent behavior information related to the causes, the causal relationship group corresponds to a probabilistic policy that allows the agent to selectively select the causes associated with the behavior information in a predetermined state in the environment, the results in the causal relationship group are associated with a state related to the result associated with the cause, and the learning unit selects the cause included in the policy in the predetermined state and updates the state of the virtual environment based on the state associated with the result associated with the selected cause and the state estimated by the behavior information associated with the selected cause.

3. The learning system according to claim 2, wherein the extraction unit generates the feature quantities based on situation data obtained from the updated state, and the interpretation unit features the instructions requesting the ethical interpretation of the situation data and inputs them into a large-scale language model, and identifies the ethical evaluation result from the output.

4. The learning system according to claim 2, wherein the agent includes an autonomous robot, and the behavioral information includes control information of the autonomous robot.

5. The learning system according to claim 4, wherein the virtual environment is a physical simulation environment in which the autonomous robot operates.

6. The learning system according to claim 4, wherein the virtual environment is a world model in which the autonomous robot operates.

7. The learning system according to claim 6, wherein the learning unit performs model-based reinforcement learning processing using a world model that estimates state transitions in a virtual environment by inputting one or more image data acquired in a physical environment or another virtual environment and the behavior information of the autonomous robot as features.

8. The learning system according to claim 5 or 6, wherein at least a portion of the states in the virtual environment are given by domain randomization.

9. The learning system according to claim 1, comprising one or more first processing chips, one or more second processing chips, and an autonomous robot, wherein the autonomous robot comprises a main system including at least one of the one or more first processing chips and an operating unit, and a subsystem including at least one of the one or more second processing chips and storing the determination model, wherein the subsystem is an enclave isolated as a control system independent of the main system, and at least a part of it, including the second processing chip, is isolated as an independent circuit system.

10. The learning system according to claim 9, wherein the subsystem comprises an extraction unit executed using at least one of the one or more second processing chips, an interpretation unit executed using at least one of the chips, and a learning unit executed using at least one of the chips.

11. The learning system according to claim 9, further comprising one or more information processing devices, wherein at least one of the one or more information processing devices is the extraction unit, at least one is the interpretation unit, and at least one is the learning unit.

12. The learning system according to claim 9, wherein the main system includes a detection unit for acquiring environmental data, and the subsystem stores the environmental data acquired in the physical environment for generating the causal relationship group.

13. The learning system according to claim 12, wherein the subsystem has a generation unit that is executed using at least one of the one or more second processing chips, and the generation unit generates the causal relationship group based on the environmental data.

14. The learning system according to claim 13, wherein the subsystem has a learning unit that is executed using at least one of the one or more second processing chips, and the learning unit performs reinforcement learning processing on the decision model based on the causal relationship group generated by the generation unit.

15. The learning system according to claim 9, wherein the subsystem comprises a detection unit that is performed using at least one of the one or more second processing chips, and a generation unit that is performed using at least one of the second processing chips, the detection unit acquires environmental data in the physical environment, and the generation unit generates the causal relationship group based on the environmental data.

16. The learning system according to claim 15, wherein the generation unit associates the cause with the action information of an agent corresponding to the cause.

17. The learning system according to claim 15, wherein the generation unit associates the cause and / or the result with explanatory information that explains the cause and / or the result.

18. The learning system according to claim 15, wherein the subsystem further comprises a conversion unit that is performed using at least one of the one or more second processing chips, the conversion unit abstracts or embodies language data corresponding to each of the at least one of the causes and effects constituting the causal relationship group, based on ontology dictionary data.

19. The learning system according to claim 18, wherein the causes and / or effects constituting the causal relationship group are associated with explanatory information including text that explains the causes and / or effects, and the conversion unit performs the abstraction or concretization based on the explanatory information.

20. The learning system according to claim 15, wherein the subsystem further comprises an integration unit that is executed using at least one of the one or more second processing chips, the integration unit identifies matching or similar causes or effects among the causes and effects constituting a plurality of causal relationship groups, and integrates two or more independent causal relationship groups by commonizing the matching or similar causes or effects.

21. The learning system according to claim 20, wherein the integration unit inputs a plurality of causes or a plurality of results into LLM (Large Language Models) and obtains evaluation results for their matches or similarities.

22. The learning system according to claim 20, wherein the integration unit inputs a plurality of causes or a plurality of results into a distance function or similarity function and obtains evaluation results for their agreement or similarity.

23. The learning system according to claim 20, wherein the causes and / or effects constituting the causal relationship group are associated with descriptive information including text describing the causes and / or effects, and the integration unit identifies matching or similar causes or effects based on the descriptive information associated with the causes or effects included in the independent causal relationship group, and integrates the independent causal relationship group.

24. The learning system according to claim 15, wherein the subsystem further comprises a detection unit that is performed using at least one of the one or more second processing chips, and an extrapolation unit that is performed using at least one of the chips, the extrapolation unit identifies a pair of causes or effects that are in a hierarchical relationship among the causes and effects that constitute separate causal relationship groups based on ontology dictionary data, and extrapolates another causal relationship group that includes the other cause or effect of the pair to the causal relationship group that includes one of the causes or effects of the pair.

25. The learning system according to claim 24, wherein the causes and / or effects constituting the causal relationship group are associated with explanatory information including text that explains the causes and / or effects, and the extrapolation unit extrapolates the causal relationship group based on the explanatory information.

26. A learning system according to claim 1, further comprising an aggregation unit, a server device including a database (DB), and one or more agent systems including at least the learning unit, wherein the agent systems encrypt and transmit ethical data including a decision model on which the reinforcement learning process has been performed based on an encryption key unique to each of the one or more agent systems, and the aggregation unit receives the ethical data held by each of the one or more agent systems, verifies the received ethical data based on the encryption key, and stores it in the DB.

27. A learning system according to claim 15, further comprising an aggregation unit, a server device including a database (DB), and one or more agent systems including at least the generation unit, wherein the agent systems encrypt and transmit ethical data including the causal relationship group generated based on an encryption key unique to each of the one or more agent systems, and the aggregation unit receives the ethical data held by each of the one or more agent systems, verifies the received ethical data based on the encryption key, and stores it in the DB.

28. A learning system further comprising one or more nodes, each including a cryptographic processing unit and a management unit, and at least the learning unit and a storage device that stores ethical data including a decision model on which the reinforcement learning process has been performed by the learning unit, wherein the cryptographic processing unit generates a physically hard-to-replicate function (PUF), generates a private key based on the PUF, generates a public key and an electronic signature of a transaction (Tx) including at least a portion of the ethical data based on the private key, the management unit stores one or more Tx including the public key and the electronic signature in the current block constituting a blockchain network (BCN), generates a hash value for at least a portion of the previous block and a hash value for at least the Tx in the current block, and stores the respective hash values ​​of the previous block and the current block in the current block.

29. A learning system further comprising one or more nodes, each including a cryptographic processing unit and a management unit, and at least the generation unit and a storage device that stores ethical data including the causal relationship group generated by the generation unit, wherein the cryptographic processing unit generates a physically hard-to-replicate function (PUF), generates a private key based on the PUF, generates a public key and an electronic signature of a transaction (Tx) including at least a portion of the ethical data based on the private key, the management unit stores one or more Tx including the public key and the electronic signature in the current block constituting a blockchain network (BCN), generates a hash value for at least a portion of the previous block and a hash value for at least the Tx in the current block, and stores the respective hash values ​​of the previous block and the current block in the current block.

30. The learning system according to claim 28 or 29, wherein Tx includes the ethical data or data indicating the integrity of the ethical data.

31. The learning system according to claim 28 or 29, wherein the node comprises a main system that performs processing by one or more first processing chips, one or more second processing chips, and a subsystem having an encryption processing unit performed by at least one of the one or more second processing chips, and a dedicated auxiliary storage device which is the storage device, wherein the subsystem is an enclave isolated as a control system independent of the main system, and at least a part of it, including the second processing chips, is isolated as an independent circuit system, and stores one or more blocks constituting the BCN in the dedicated auxiliary storage device.

32. The learning system according to claim 31, wherein the node is an autonomous robot, the main system has one or more first processing chips and operating units, and the subsystem has one or more second processing chips and a dedicated auxiliary storage device storing the determination model.

33. The learning system according to claim 31, wherein the subsystem includes a PUF circuit, the cryptographic processing unit identifies a challenge-response combination in the PUF from the PUF circuit, and obtains the secret key based on the challenge-response combination.

34. The learning system according to claim 33, wherein the PUF circuit is a group of memory cells of a static random access memory that provides a unique PUF based on element characteristics, and the group of memory cells includes field-effect transistors having one or more channels arranged in a direction orthogonal to the same plane of a substrate connected to the second processing chip.

35. The learning system according to claim 34, wherein the memory cell group includes N-type field-effect transistors and P-type field-effect transistors arranged in the orthogonal direction.

36. The learning system according to claim 34, wherein the memory cell group includes a group of transistors having fork sheets.

37. The learning system according to claim 34, wherein the element characteristics are the initial values ​​of each of the memory cell groups.

38. The learning system according to claim 34, wherein the element characteristics are current oscillation characteristics due to flicker noise in each of the memory cell groups.

39. The learning system according to claim 9, 10, 13, 14, 15, 18, 20, or 24, wherein the second processing chip includes a group of field-effect transistors having nanosheet channels.

40. The learning system according to claim 9, 10, 13, 14, 15, 18, 20, or 24, wherein the second processing chip includes a group of field-effect transistors having fork sheet channels.

41. The learning system according to claim 9, 10, 13, 14, 15, 18, 20, or 24, wherein the second processing chip includes a group of CFETs (complementary field-effect transistors).

42. A learning method performed by one or more computers, wherein at least one of the one or more computers stores a causal relationship set including a cause which is an action selected by an agent and an effect corresponding to the cause; extracts a feature quantity from the environment corresponding to each of the one or more effects constituting the causal relationship set; estimates an ethical interpretation of each of the one or more effects based on the feature quantity, identifies an ethical outcome in the causal relationship set; and performs reinforcement learning processing of a decision model by setting a reward for the selection of the ethical outcome and / or the cause related to this outcome and the corresponding action.

43. A learning program executed by one or more computers, wherein at least one of the one or more computers, which stores a group of causal relationships including causes that are actions selected by an agent and results corresponding to those causes, functions as an extraction unit that extracts feature quantities from the environment corresponding to each of the one or more results constituting the group of causal relationships; at least one of the one or more computers functions as an interpretation unit that estimates the ethical interpretation of each of the one or more results based on the feature quantities and identifies the ethical outcome in the group of causal relationships; and at least one of the one or more computers functions as a learning unit that performs reinforcement learning processing of a decision model by setting rewards for the selection of the actions corresponding to the ethical outcome and / or causes related to that outcome.