Information processing system, information processing method, and information processing device

The information processing system adjusts learning efficiency and behavior based on stress evaluation to enhance adaptability to environmental changes, addressing suboptimal performance in machine learning systems.

JP7740327B2Active Publication Date: 2025-09-17SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023506799
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-17
Filing Date
2022-01-19
Publication Date
2025-09-17
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

Machine learning systems, such as robots and agents, often fail to adapt to changes in the external environment or receive less reward than expected due to incorrect past learning or environmental changes, leading to suboptimal behavior.

Method used

An information processing system that includes a learning unit, an input information evaluation unit, and a first parameter calculation unit to evaluate risk and stress levels, adjusting learning efficiency based on cortisol-like parameters to adapt to changes in the environment.

Benefits of technology

The system effectively adapts to environmental changes by adjusting learning efficiency and behavior selection, preventing overlearning and ensuring optimal performance in dynamic situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740327000001
    Figure 0007740327000001
  • Figure 0007740327000002
    Figure 0007740327000002
  • Figure 0007740327000003
    Figure 0007740327000003
Patent Text Reader

Abstract

The present disclosure relates to an information processing system, an information processing method, and an information processing device that make it possible to perform learning adapted to changes in the external environment and conditions. A learning unit learns system behavior selection results pertaining to input information, an input information evaluation unit evaluates an input information danger level for the system, and a first parameter calculation unit calculates, on the basis of an evaluation value of the input information, a first parameter representing stress on the system. In addition, the learning unit changes learning efficiency in accordance with the first parameter. The feature according to the present disclosure can be applied to an information processing system that performs machine learning, for example.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing device, and more particularly to an information processing system, an information processing method, and an information processing device that enable learning to be adapted to changes in the external environment or situation. [Background technology]

[0002] In recent years, machine learning has been used to enable optimal data processing and autonomous control in information processing devices such as robots and agents. One such machine learning technique is reinforcement learning, which learns a policy that maximizes rewards through actions according to a given state.

[0003] Patent Document 1 discloses an information processing device that performs efficient reinforcement learning by using annotations input by a user as rewards. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-64759 Summary of the Invention [Problem to be solved by the invention]

[0005] In machine learning, there are cases where the robot does not behave as expected or receives less reward than expected. This may be due to changes in the external environment or incorrect past learning.

[0006] The present disclosure has been made in light of such circumstances, and aims to enable learning to be performed in a manner that is adaptable to changes in the external environment and circumstances. [Means for solving the problem]

[0007] The information processing system of the present disclosure includes a learning unit that learns the system's behavioral selection results for input information, an input information evaluation unit that evaluates the risk of the input information to the system, and a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information, and the learning unit changes learning efficiency according to the first parameter.

[0008] In the information processing method disclosed herein, an information processing system learns the system's behavioral selection results for input information, evaluates the risk of the input information to the system, calculates a first parameter representing stress in the system based on the evaluation value of the input information, and changes learning efficiency according to the first parameter.

[0009] The information processing device disclosed herein includes a learning unit that learns the system's behavioral selection results for input information, an input information evaluation unit that evaluates the risk of the input information to the system, and a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information, and the learning unit changes learning efficiency according to the first parameter.

[0010] In the present disclosure, the system's behavioral selection results for input information are learned, the risk of the input information to the system is evaluated, a first parameter representing stress in the system is calculated based on the evaluation value of the input information, and the learning efficiency is changed according to the first parameter. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram explaining human stress responses in neuroscience. [Figure 2] FIG. 1 is a block diagram illustrating an example of the configuration of an information processing system. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of the information processing system. [Figure 4] 10 is a flowchart illustrating the flow of a learning process. [Figure 5] 10 is a flowchart illustrating the flow of an input information evaluation process. [Figure 6] FIG. 1 is a diagram showing an example of a time series of cortisol levels. [Figure 7] FIG. 1 is a graph showing the relationship between cortisol levels and hippocampal volume values. [Figure 8] 1 shows an example of study duration depending on cortisol levels. [Figure 9] FIG. 1 is a graph showing the relationship between cortisol levels and learning efficiency. [Figure 10] FIG. 10 is a diagram showing an example of learning efficiency relative to cortisol levels and hippocampal volume values. [Figure 11] FIG. 1 is a block diagram illustrating an example of the configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0012] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below in the following order.

[0013] 1. Human stress response in neuroscience 2. Information Processing System Configuration 3. Learning process flow 4. Application Examples 5. Computer configuration example

[0014] <1. Human stress response in neuroscience> In machine learning, there are cases where the robot does not behave as expected or receives less reward than expected. This may be due to changes in the external environment or incorrect past learning.

[0015] People also feel stressed when they are unable to act as expected or when they receive less reward than expected.

[0016] Figure 1 is a diagram explaining the human stress response in neuroscience.

[0017] External stimuli received from society and the environment as stress affect the medial prefrontal cortex in the frontal lobe of the brain. The medial prefrontal cortex makes context-based fear judgments and judges external stimuli as threats to the social self. Specifically, the medial prefrontal cortex performs error monitoring and control of one's own behavior, and refers to the temporal lobe for past behavioral history. The temporal lobe is the area responsible for memory storage.

[0018] The amygdala is a part of the limbic system and is controlled by the medial prefrontal cortex. The amygdala judges fear based on external stimuli and physical and biological threats to the self. The amygdala also integrates information processed in the amygdala with input from the medial prefrontal cortex. Activation of the amygdala activates the hypothalamic-pituitary-adrenal axis (HPA axis), which secretes the stress hormone cortisol. In other words, the amygdala triggers stress responses such as anxiety and fear.

[0019] Furthermore, the amygdala, which controls emotions, influences the hippocampus, which is also part of the limbic system and controls memory. The hippocampus receives input from the HPA axis, which releases cortisol, and inhibits the HPA axis. In other words, the hippocampus has a negative feedback function that inhibits the release of cortisol from the HPA axis.

[0020] The HPA axis regulates the amount of cortisol secreted by integrating activation (positive input) from the amygdala and inhibition (negative input) from the hippocampus. The HPA axis also has a feedback function that regulates the amount of cortisol secreted by the HPA axis itself so that it does not secrete too much.

[0021] Due to this mechanism, it is known that the amount of cortisol secreted (hereinafter referred to as cortisol amount) increases after a short delay following psychological stress, and then gradually decreases to its original level after the psychological stress is relieved.

[0022] However, the increase or decrease in cortisol levels varies depending on social stress and the situation one is in. For example, the degree of increase in cortisol levels is greater during public speaking or cognitive tasks. Furthermore, when social evaluation is carried out under circumstances that one cannot control, the degree of decrease in cortisol levels is smaller.

[0023] Cortisol secreted from the HPA axis acts on the frontal lobe, which is responsible for decision-making. Normally, the frontal lobe, which is responsible for cognition and reason, makes decisions by suppressing the secretion of cortisol. However, it is known that prolonged cortisol secretion leads to risk-averse decisions, and a sudden increase in cortisol levels leads to immediate reward selection.

[0024] In this way, cortisol secretion dynamically influences decision-making, such as risk judgment and reward prediction.

[0025] The hippocampus also controls the formation and recall of memories. Specifically, the hippocampus forms and reconsolidates memories stored in the temporal lobe. The hippocampus also recalls memories stored in the temporal lobe. These functions of the hippocampus enable people to learn and remember.

[0026] On the other hand, cortisol secreted from the HPA axis reduces hippocampal volume depending on the amount and duration of input to the hippocampus, resulting in a decrease in the hippocampal ability to suppress cortisol.

[0027] For example, a certain level of stress activates the nervous system, improving learning and memory efficiency. However, increasing the intensity or duration of stress can lead to decreased activity or damage to the nervous system. Furthermore, while mild stress improves hippocampal function, excessive stress can decrease hippocampal function or even damage the hippocampus.

[0028] In this way, moderate stress improves learning ability by improving the function of the nervous system and hippocampus, but excessive stress can cause damage to the nervous system and hippocampus.

[0029] Thanks to the brain mechanisms described above, people can learn to adapt to changes in the external environment or situation when they are unable to act as expected or receive less reward than expected (stress).

[0030] The technology disclosed herein enables machine learning to adapt to changes in the external environment and circumstances in response to situations such as failure to behave as expected or receiving less reward than expected (stress).

[0031] <2. Information Processing System Configuration> FIG. 2 is a block diagram illustrating a configuration example of an information processing system according to an embodiment of the present disclosure.

[0032] The information processing system 1 shown in FIG. 2 is configured as a robot, an agent, or the like, and includes an information processing device 10 and a storage device 20.

[0033] The information processing device 10 is configured as a computer such as a PC (Personal Computer) that performs machine learning.

[0034] The storage device 20 is configured as a semiconductor memory, a magnetic storage device, an optical storage device, etc. The storage device 20 may be built into the information processing device 10, may be detachably attached, or may be connected to the information processing device 10 via a network such as the Internet. The storage device 20 stores the results of machine learning (learning results) performed by the information processing device 10.

[0035] The information processing device 10 includes an input unit 31 , a control unit 32 , and an output unit 33 .

[0036] The input unit 31 is configured to include individual sensors capable of acquiring sensor information such as image, sound, temperature, pressure, and tactile information, keys or a touch panel capable of inputting text information, and an external interface capable of inputting external information from external devices or services. The various types of input information input to the input unit 31 are supplied to the control unit 32.

[0037] The control unit 32 is configured with a processor such as a CPU (Central Processing Unit), and controls each unit of the information processing device 10. The control unit 32 performs machine learning based on input information from the input unit 31 and the learning results stored in the storage device 20. The results of the machine learning (learning results) are supplied to the output unit 33 and also stored in the storage device 20.

[0038] The output unit 33 is configured as a display capable of displaying images and characters, a speaker capable of outputting sound, a light-emitting unit that emits light, etc. When the information processing device 10 is configured as an autonomously mobile robot or agent, the output unit 33 is configured as a drive mechanism for moving the information processing device 10.

[0039] In the information processing system 1 shown in FIG. 2, for example, only the storage device 20 may be configured on the cloud, or both the control unit 32 of the information processing device 10 and the storage device 20 may be configured on the cloud.

[0040] FIG. 3 is a block diagram showing an example of the functional configuration of the information processing system 1. As shown in FIG.

[0041] The information processing system 1 in FIG. 3 is made up of an input information evaluation unit 110, a storage device 120, a first parameter calculation unit 130, an action selection unit 140, and an inhibition and storage processing unit 150.

[0042] 3, an input information evaluation unit 110, a first parameter calculation unit 130, an action selection unit 140, and an inhibition / storage processing unit 150 are realized by the control unit 32 in Fig. 2. Also, in Fig. 3, a storage device 120 corresponds to the storage device 20 in Fig. 2.

[0043] The input information evaluation unit 110 has functions corresponding to the medial prefrontal cortex and amygdala in FIG. 1, and evaluates the risk of input information to the information processing system 1 (hereinafter also simply referred to as the system) (possibility of causing damage to the system).

[0044] The input information evaluation unit 110 includes a social risk evaluation unit 111 , a physical risk evaluation unit 112 , and a learning result reference / risk integration unit 113 .

[0045] The social risk assessment unit 111 assesses the social risk that the input information from the input unit 31 poses to the system based on the learning results stored in the storage device 120. The social risk here refers to, for example, the evaluation of the system such as its credibility, influence, and performance, its social status such as whether or not it is possible to control the situation in which the system is placed, and the possibility that the system itself will be restricted by laws and regulations.

[0046] The evaluation value representing the social risk level is supplied to the learning result reference / risk level integration unit 113 .

[0047] The physical risk assessment unit 112 assesses the physical risk that the input information from the input unit 31 poses to the system based on the learning results stored in the storage device 120. The physical risk here refers to, for example, the possibility of coming into contact with something, the possibility of physical damage, the possibility of injuring someone, etc.

[0048] The evaluation value representing the physical risk is supplied to the learning result reference / risk integration unit 113 .

[0049] The learning result reference / risk level integration unit 113 integrates the evaluation values ​​from the social risk level evaluation unit 111 and the physical risk level evaluation unit 112 to calculate an evaluation value x1 that represents the overall risk level that the input information from the input unit 31 poses to the information processing system 1. In addition, the learning result reference / risk level integration unit 113 refers to the learning result that corresponds to the input information from the input unit 31, among the learning results stored in the storage device 120.

[0050] The storage device 120 stores input information 121, behavioral information 122, and evaluation value 123 as past learning results.

[0051] The input information 121 is input information that has been determined to cause damage to the system through past learning, and also includes the risk level previously evaluated for that input information.

[0052] Therefore, the social risk assessment unit 111 and the physical risk assessment unit 112 use the input information 121 to assess the social risk and the physical risk of the input information from the input unit 31, respectively.

[0053] The behavior information 122 is a selection result (countermeasure) of the behavior of the system obtained through past learning in response to the input information 121. Note that there may be multiple pieces of behavior information 122 (countermeasures) for the same input information 121.

[0054] The evaluation value 123 is an evaluation value of a countermeasure obtained by past learning, and is associated with the behavioral information 122. The evaluation value of a countermeasure is derived using an existing learning model based on the amount of reduction in a first parameter (amount of cortisol) described below, the cost and time required for calculation and output, and the like.

[0055] Therefore, the learning result reference / risk level integration unit 113 refers to the behavior information 122 and evaluation value 123 for input information 121 that is the same as or similar to the input information from the input unit 31, from the learning results stored in the storage device 120.

[0056] The evaluation value x1 calculated by the learning result reference / risk level integration unit 113 is supplied to the first parameter calculation unit 130. In addition, the past behavior information 122 and the evaluation value 123 referred to by the learning result reference / risk level integration unit 113 are supplied to the behavior selection unit 140.

[0057] The first parameter calculation unit 130 has a function corresponding to the HPA axis in FIG. 1 and calculates a first parameter representing stress in the system based on three input values ​​x1, x2, and x3. Stress here refers to situations in machine learning, such as not being able to perform expected behavior or receiving less reward than expected. In other words, the first parameter is an internal variable specific to the system that corresponds to the amount of cortisol in the HPA axis. Hereinafter, the first parameter is also referred to as the amount of cortisol.

[0058] The input value x1 is the evaluation value x1 from the above-mentioned input information evaluation unit 110. The input value x2 is a negative feedback value x2 that suppresses the first parameter (cortisol level) and is calculated by the inhibition / storage processing unit 150, which will be described later.

[0059] The calculated cortisol amount is supplied to the behavior selection unit 140 and the inhibition / storage processing unit 150, and is also fed back to the first parameter calculation unit 130 as an input value x3.

[0060] The behavior selection unit 140 has a function corresponding to the frontal lobe in FIG. 1, and selects a behavior for the system based on the cortisol amount from the first parameter calculation unit 130 and the past behavior information 122 and evaluation value 123 from the input information evaluation unit 110. The behavior selection result (selected behavior) of the system is supplied to the inhibition / storage processing unit 150 together with the corresponding input information.

[0061] The inhibition / memory processing unit 150 has a function corresponding to the hippocampus in FIG. 1, and suppresses the amount of cortisol calculated by the first parameter calculation unit 130, and memorizes and learns the selected action of the action selection unit 140.

[0062] The suppression and storage processing unit 150 is made up of a second parameter calculation unit 151 , a negative feedback value calculation unit 152 , and a storage and learning unit 153 .

[0063] The second parameter calculation unit 151 calculates a second parameter that decreases depending on the input amount of cortisol and time, based on the amount of cortisol from the first parameter calculation unit 130. The calculated second parameter is supplied to the negative feedback value calculation unit 152 and the memory / learning unit 153.

[0064] The second parameter is an internal variable specific to the system that corresponds to a hippocampal volume value that indicates the volume of the hippocampus. Hereinafter, the second parameter will also be referred to as the hippocampal volume value.

[0065] The negative feedback value calculation unit 152 calculates a negative feedback value x2 that has a negative correlation with the cortisol amount, based on the cortisol amount from the first parameter calculation unit 130 and the hippocampal volume value from the second parameter calculation unit 151. The calculated negative feedback value x2 is supplied to the first parameter calculation unit 130.

[0066] The memory / learning unit 153 learns the selected behavior of the behavior selection unit 140 and derives an evaluation value for the selected behavior based on the cortisol amount from the first parameter calculation unit 130 and the hippocampal volume value from the second parameter calculation unit 151. The learning result of the selected behavior and the evaluation value are associated with input information and stored in the storage device 120 as input information 121, behavior information 122, and evaluation value 123.

[0067] <3. Learning process flow> Next, the flow of the learning process executed by the information processing system 1 in FIG. 3 will be described with reference to the flowchart in FIG.

[0068] In step S1, the input information evaluation unit 110 executes an input information evaluation process for evaluating the risk of input information to the information processing system 1.

[0069] FIG. 5 is a flowchart illustrating the flow of the input information evaluation process.

[0070] In step S11, the social risk assessment unit 111 evaluates the social risk of the input information from the input unit 31 using an existing learning model based on the input information 121 stored in the memory device 120, thereby obtaining an evaluation value of the social risk of the input information.

[0071] In step S12, the physical risk assessment unit 112 evaluates the physical risk of the input information from the input unit 31 using an existing learning model based on the input information 121 stored in the memory device 120, thereby obtaining an evaluation value of the physical risk of the input information.

[0072] Here, if there is no input information 121 in the storage device 120 that is the same as or similar to the input information from the input unit 31, the social and physical risks of the input information are evaluated as being high.

[0073] In step S13, the learning result reference / risk level integration unit 113 integrates the evaluation value of the social risk level of the input information and the evaluation value of the physical risk level of the input information to obtain an evaluation value x1 of the overall risk level of the input information. The evaluation value x1 of the overall risk level of the input information is supplied to the first parameter calculation unit 130.

[0074] In step S14, the learning result reference / risk level integration unit 113 refers to the past behavior information 122 and the evaluation value 123 for the input information 121 that is the same as or similar to the input information from the input unit 31 in the storage device 120, using an existing learning model. The referred past behavior information 122 and the evaluation value 123 are supplied to the behavior selection unit 140.

[0075] In addition, if there is no input information 121 in the storage device 120 that is the same as or similar to the input information from the input unit 31 (if there is no behavioral information 122 and evaluation value 123 to refer to), the behavioral information 122 and evaluation value 123 set as initial values ​​may be referenced.

[0076] Returning to the flowchart of FIG. 4, in step S2, the first parameter calculation unit 130 calculates the cortisol amount (first parameter) based on the input information evaluation value x1, the negative feedback value x2, and the past cortisol amount x3.

[0077] The cortisol level f is calculated using a function expressed as f(x1, x2, x3)=ax1-bx2-cx3, where coefficients a, b, and c≧0 and x1, x2, and x3≧0.

[0078] That is, the amount of cortisol has a positive correlation with the evaluation value of the input information x1, a negative correlation with the negative feedback value x2, and a negative correlation with the past amount of cortisol x3. In other words, input information from the external environment increases the amount of cortisol, while the feedback value looping within the system acts to decrease the amount of cortisol.

[0079] In addition, by calculating the amount of cortisol based on past input values ​​x1, x2, and x3, the amount of cortisol may fluctuate with a time lag relative to the input information (stress load), similar to the actual brain mechanism.

[0080] In step S3, the behavior selection section 140 selects a behavior of the system based on the amount of cortisol calculated by the first parameter calculation section .

[0081] Specifically, the behavior selection unit 140 selects a behavior for the system based on the evaluation values ​​123 corresponding to the past behavior information 122 as candidates. At this time, the behavior selection unit 140 stores the amount of cortisol over time (time series of cortisol amount) as shown in Fig. 6, and selects a behavior that will reduce the amount of cortisol from the present onwards.

[0082] Furthermore, the behavior selection unit 140 biases the behavior to be selected depending on the level and fluctuation tendency of the cortisol amount, and selects the behavior of the system.

[0083] For example, when cortisol levels are high, the system biases the behaviors it selects to favor those that result in a rapid and significant decrease in cortisol levels, even if this increases the computational cost of the system. On the other hand, when cortisol levels are low, the system biases the behaviors it selects to favor those that result in a long-term decrease in cortisol levels.

[0084] Furthermore, when the cortisol level shows an increase over time as a trend in the amount of cortisol, a different behavior from the previous behavior is selected, whereas when the cortisol level shows a decrease over time as a trend in the amount of cortisol, a behavior that maintains the previous behavior is selected.

[0085] Here, behaviors that reduce cortisol levels include avoidance, taking measures to address the cause, and behaviors that provide other rewards (deceptive behaviors).

[0086] In addition, if the input information from the input unit 31 is new and there is no candidate past behavioral information 122, the system may randomly select a behavior, and the appropriateness of the randomly selected behavior may be determined based on the subsequent change in cortisol levels.

[0087] Returning to the flowchart of FIG. 4, in step S4, the second parameter calculation section 151 calculates a hippocampal volume value (second parameter) based on the amount of cortisol calculated by the first parameter calculation section .

[0088] FIG. 7 is a graph showing the relationship between cortisol levels and hippocampal volume values.

[0089] 7, the hippocampal volume value monotonically decreases over time according to the amount of cortisol input to the second parameter calculation unit 151. However, in the relationship shown in FIG. 7, the hippocampal volume value only decreases over time, so in an actual system, an initial value for the hippocampal volume value is set, and when the amount of cortisol is equal to or less than a predetermined threshold, the hippocampal volume value may be increased over time within a range that does not exceed the initial value.

[0090] Similar to the actual brain mechanisms, hippocampal volume plays a role in regulating the balance between cortisol suppression and memory / learning. Specifically, a large hippocampal volume has the effect of improving learning efficiency, but also of reducing cortisol levels, and a decrease in cortisol levels reduces learning efficiency.

[0091] Returning to the flowchart of Figure 4, in step S5, the negative feedback value calculation unit 152 calculates a negative feedback value x2 that suppresses an increase in the amount of cortisol based on the amount of cortisol calculated by the first parameter calculation unit 130 and the hippocampal volume value calculated by the second parameter calculation unit 151.

[0092] The negative feedback value x2 has a positive correlation with both cortisol levels and hippocampal volume values.

[0093] In step S6, the memory / learning unit 153 learns the selected behavior of the behavior selection unit 140 based on the amount of cortisol calculated by the first parameter calculation unit 130 and the hippocampal volume value calculated by the second parameter calculation unit 151.

[0094] Furthermore, the memory / learning unit 153 derives an evaluation value indicating whether the selected behavior is an action that relieves stress in response to the input information. Specifically, an evaluation value of the selected behavior is derived from the amount of reduction in cortisol levels, the cost and time required for calculation and output, and the selected behavior and the evaluation value are stored in the storage device 120 using an existing learning model.

[0095] As shown in Figure 8, the memory / learning unit 153 learns the relationship between input information and output selection behavior during the period (learning period) from when the cortisol level, which fluctuates over time, rises above a first threshold Th1 to when it falls below the first threshold Th1. Furthermore, the memory / learning unit 153 suppresses learning during the period (learning suppression period) from when the cortisol level rises above the first threshold Th1 to when it exceeds a second threshold that is greater than the first threshold Th1. In the figure, the density of the vertical lines on the curve representing the fluctuations in cortisol level represents the frequency of learning and memory and the level (magnitude) of weighting.

[0096] The memory / learning unit 153 changes the frequency and weighting of learning and memory, in other words, learning efficiency, according to the cortisol level and hippocampal volume value. Learning efficiency has a positive correlation with the hippocampal volume value. Also, as shown in Figure 9, the frequency and weighting of learning and memory (learning efficiency) have an upwardly convex curve relationship with the cortisol level. According to Figure 9, learning efficiency is maximized when the cortisol level is high to a certain extent. In other words, the memory / learning unit 153 temporarily increases learning efficiency during the learning period and decreases it over time. The frequency and weighting of learning and memory (learning efficiency) during the learning period increase according to the cortisol level.

[0097] Thus, learning efficiency varies depending on the relationship between cortisol levels and hippocampal volume values.

[0098] FIG. 10 is a diagram showing an example of learning efficiency relative to cortisol levels and hippocampal volume values.

[0099] In the early stages of increased risk to the system (before the risk becomes excessively high), the cortisol level is likely to be higher (more) than a certain threshold, and the hippocampal volume value is likely to be high. In this state, the system can improve its learning efficiency, learning various coping strategies and learning behaviors that appear to have detected danger in advance so that the risk does not become excessively high. As a result, it can take avoidance behaviors that appear to have detected danger before the risk becomes excessively high.

[0100] In this way, when the risk level rapidly decreases from a state of high learning efficiency, the hippocampal volume value remains high, and the cortisol level falls below the threshold (as shown by arrow #1 in the figure). In this case, further learning is unnecessary, so the system lowers the learning efficiency.

[0101] In addition, when the hippocampal volume value is low and the cortisol level is lower (less than) the threshold, the risk is low or the already learned coping methods are working well, and there is no need for new learning, so the system reduces the learning efficiency.

[0102] Furthermore, if the risk level does not decrease easily from a state of high learning efficiency while measures are being taken to prevent it from increasing, as shown by arrow #2 in the figure, the cortisol level will remain constant and the hippocampal volume will decrease. In this case, the system will continue to learn, but by setting the learning efficiency to a medium level, the frequency and weighting of learning and memory will be slightly lowered. This will prevent overlearning.

[0103] If the cortisol level exceeds a certain level (exceeding the second threshold in Figure 8), the risk is extremely high, and the system's output (behavior) is likely to be limited to a limited range of behaviors, such as not moving or reacting, or repeating the same behavior. This is also likely to be an extreme situation.

[0104] In such cases, the system can prevent overlearning and an increase in computational costs by lowering the learning efficiency and suppressing learning as described above.

[0105] According to the above configuration and processing, the information processing system 1 is able to learn in a way that adapts to changes in the external environment and situation in machine learning, in response to situations such as when the robot is unable to behave as expected or when the reward obtained is less than expected (stress).

[0106] In other words, the system can learn anew regardless of past learning results by increasing the amount of cortisol (first parameter), an internal variable that represents stress, and changing the learning efficiency according to that amount of cortisol, in other words, the stress level.

[0107] This allows the system to re-learn by increasing the cortisol level for input information corresponding to previously learned behavioral selection results and select a new behavior. Also, for new input information that has not been learned before, the system can perform new learning according to the cortisol level and select an appropriate behavior.

[0108] For example, by training the system with weighting relative to past learning results, it is possible to obtain negative learning results that negate past learning results. Here, rather than learning that gradually changes learning results, learning is performed to obtain learning results that are completely different from past learning results. This allows for fine-tuning of the learning model.

[0109] The technology disclosed herein can also allow the system to escape from a local solution when it has fallen into that state. In this case, the fact that the learning accuracy has not improved beyond a certain level is treated as stress, and learning is carried out again based on the possibility that a more optimal solution than the solution found by the current learning results exists.

[0110] Furthermore, in the technology disclosed herein, the system can also perform learning that takes into account time spread and delay. In this case, the correspondence between a certain environmental condition and the amount of reward is not directly learned, but rather, when stress such as damage or a collision occurs, learning is performed that takes into account the surrounding circumstances over time. In a stressful situation, weighting may be expanded along with the time scale to perform learning.

[0111] As described above, in a system to which the technology according to the present disclosure is applied, a negative feedback value x2 that suppresses the amount of cortisol is input to the first parameter calculation unit 130 that calculates the amount of cortisol (first parameter). This prevents excessive learning and memorization, making it possible to prevent overlearning.

[0112] Furthermore, because a system employing the technology disclosed herein mimics the brain mechanisms involved in actual stress responses, it is possible to select output behaviors that are similar to those of real people. Therefore, by applying the technology disclosed herein to systems that communicate with people, such as smart speakers for casual conversation or chatbots, it is possible to provide more natural and lifelike interactions to people.

[0113] <4. Application Examples> Therefore, hereinafter, an application example of a system to which the technology according to the present disclosure is applied will be described.

[0114] (Application example 1: Communication system) A system to which the technology according to the present disclosure is applied can be applied to a communication system.

[0115] The communication system may be configured as, for example, a dialogue system or a text generation system. Dialogue systems include chat systems such as smart speakers and conversational AI, and automated response systems in call centers. Text generation systems include chatbots, article generation systems, and story generation systems. Here, the output format of the behavior selected in the communication system is text or voice without any interaction.

[0116] In a communication system, input information is feedback information for text or voice (dialogue) output by the communication system. The feedback information is composed of at least one of voice information and character information.

[0117] The feedback information may be, for example, a user's reaction, which may include evaluation information such as "likes" on social networking services (SNS), responses to surveys, or estimation results of emotions based on biometric information or facial expression analysis of the user.

[0118] The feedback information may also be the actual results of KPIs (Key Performance Indicators), such as sales contract rate and PV (Page Views), which are expected as the results of communication.

[0119] In the communication system, the input information evaluation unit 110 evaluates the expectation level of such feedback information as the social risk level. If the expectation level is lower than expected, the amount of cortisol, which indicates stress, increases.

[0120] The memory and learning unit 153 learns communication with the user and changes the learning efficiency according to the amount of cortisol. When the amount of cortisol increases, learning is performed again or an existing learning model is fine-tuned.

[0121] In particular, communication trends change over time and cultures. Dialogue and text generation based on existing learning models may no longer be appropriate depending on the region or time.

[0122] For example, sentences generated with the assumption that they will satisfy users may not be well-received, buzzwords, popular phrases, and sentences with specific nuances may become outdated over time, and certain words and nuances may take on discriminatory or negative connotations due to social events.

[0123] In contrast, a communication system that applies the technology disclosed herein allows for re-learning and fine-tuning of existing learning models in response to changes in region and time, thereby enabling appropriate dialogue and sentence generation that is suited to changes in region and time.

[0124] (Application example 2: Claim processing system) A system to which the technology disclosed herein is applied can be applied to a complaint processing system that processes customer complaints.

[0125] In the complaint processing system, the memory / learning unit 153 learns how to deal with customer complaints.

[0126] In the complaint processing system, the input information is a customer complaint, which is composed of at least one of image information, audio information, and text information.

[0127] In the complaint processing system, the input information evaluation unit 110 evaluates the degree of difficulty in dealing with the customer or the content of the complaint as a social risk. Specifically, the input information evaluation unit 110 evaluates the emotional state of the customer from the customer's voice parameters, facial expressions, and gestures. The input information evaluation unit 110 also evaluates the emotional state from words included in the results of speech recognition and text recognition, and recognizes the content of the complaint. Furthermore, the input information evaluation unit 110 may evaluate the customer's social status, such as whether the customer is a VIP customer or an influencer, or the relationship between the customer and the system.

[0128] The action selection unit 140 selects an action that the complaint processing system will actually take from among candidate complaint handling methods.

[0129] For example, if the difficulty of handling is low, it is highly likely that the customer is not very angry or that the complaint is similar to past complaints, so an action such as nodding along to what the customer is saying and suggesting a response may be selected.

[0130] If the difficulty of handling is high, it is likely that the customer is furious or that the complaint is unprecedented, so actions such as apologizing or remaining silent are chosen.

[0131] If no coping method is available, an apology or a similar complaint handling method will be randomly selected, and subsequent changes in cortisol levels will be monitored.

[0132] Furthermore, when there are multiple ways to deal with a problem, the system will bias the selected behavior depending on the cortisol level. Specifically, when the cortisol level is high, the system will select behaviors such as not arguing with the problem, as it does not want to increase the difficulty of the problem. On the other hand, when the cortisol level is low, the system will select behaviors such as arguing with the problem or making suggestions from the system, as the difficulty of the problem is low.

[0133] The memory and learning unit 153 learns the content of the complaint, the customer's emotions, and the system's response when the cortisol level is high from the time the system starts processing the complaint until it finishes. For complaints that have never been processed before, the cortisol level rises, so learning progresses. Furthermore, when the customer is extremely angry or the content of the complaint is incomprehensible, the risk level increases significantly, and the cortisol level also rises significantly. However, since the learning results in such situations are unlikely to be valid in similar situations in the future, learning does not progress in order to avoid overlearning.

[0134] (Application example 3: Recommendation system) A system to which the technology according to the present disclosure is applied can be applied to a recommendation system that makes recommendations to users.

[0135] The recommendation system may be configured as a product recommendation system for books, goods, or other items on an EC (Electronic Commerce) site, or as a content recommendation system for videos, music, or other items. The recommendation system may also be configured as a person or company recommendation system on a matching site, or as various search systems.

[0136] In a recommendation system, input information is feedback information for the recommendation results (search results) output by the recommendation system. The feedback information is composed of at least one of image information, audio information, and text information.

[0137] The feedback information is considered to be the reaction of users and society to the recommendation results (search results).

[0138] In the recommendation system, the input information evaluation unit 110 evaluates the negativity of such feedback information as a social risk. When feedback information with a high negativity is received, the amount of cortisol, which indicates stress, increases.

[0139] Highly negative feedback information can be the reaction of users and society to inappropriate recommendations, such as presenting products that actually require parental control without any restrictions. Also, highly negative feedback information can be the reaction of users and society to socially problematic recommendations, such as presenting humans as image search results for animals, or presenting specific content disproportionately as recommendations for crime-related content.

[0140] The input information evaluation unit 110 may evaluate the degree of negativity of feedback information input to the system, or may evaluate the degree of negativity of feedback information crawled on the Web. Furthermore, the input information evaluation unit 110 may use various sensing results for the user as feedback information and evaluate the degree of negativity thereof.

[0141] The memory and learning unit 153 learns the recommendation results of the recommendation system and changes the learning efficiency according to the cortisol level. When highly negative feedback information is received and the cortisol level rises, learning is performed again or the existing learning model is fine-tuned. This allows more appropriate recommendation results / search results to be presented the next time a recommendation / search is made.

[0142] (Application example 4: Robot) A system to which the technology according to the present disclosure is applied can be applied to a robot.

[0143] The robot here may be configured as, for example, a picking robot or a working robot. The robot may also be configured as a mobile robot such as a guide robot, a baggage carrying robot, an autonomous vehicle, or a drone.

[0144] In the robot, the memory / learning unit 153 learns the robot's own behavior in the surrounding environment.

[0145] In addition, in a robot, input information is sensor information such as image, sound, pressure, and temperature.

[0146] In the robot, the input information evaluation unit 110 evaluates the possibility of the robot itself being damaged by the surrounding environment as a physical risk. Specifically, the input information evaluation unit 110 recognizes objects present around the robot and evaluates the risk. The input information evaluation unit 110 also evaluates physical quantities related to the object, such as the amount of light or volume emitted by the object, and the pressure or temperature of contact with the object, and estimates the distance to the object to evaluate the possibility of a collision. Furthermore, if the object is a person or animal, the input information evaluation unit 110 may estimate the emotion of the person or animal and analyze whether or not the person or animal recognizes the robot.

[0147] The behavior selection unit 140 selects the behavior that the robot will actually take from among candidate behaviors of the robot.

[0148] For example, if there are few objects around the robot and no object that could damage the robot exists, the risk is low, so the robot will select an action such as moving freely.

[0149] If there are many objects around the robot and there are objects that could damage the robot, the risk is high, so the robot may select actions such as stopping until the number of objects decreases or changing its course. Also, if the objects are people or animals, the robot may select actions such as announcing its presence with sound or light.

[0150] If no coping strategy is available, a random action such as stopping movement or moving to an open space is chosen, and subsequent changes in cortisol levels are monitored.

[0151] Furthermore, when there are multiple ways to deal with the situation, the behavior selected is biased depending on the cortisol level. Specifically, when cortisol levels are high (high), the desire is to reduce the danger as quickly as possible, so complex behaviors such as communicating with the target, such as emitting light or sound to alert the target, are selected. On the other hand, when cortisol levels are low (low), the desire is to reduce the danger in the long term, so simple behaviors that place less strain on the surroundings, such as changing the robot's movement speed, are selected.

[0152] The memory and learning unit 153 learns surrounding objects, their actions, and the system's response to them while the cortisol level is elevated. Cortisol levels rise in situations that have never been experienced before or when encountering objects that are recognized for the first time, leading to learning. Furthermore, when there are an extremely large number of surrounding objects or when objects are moving at high speed, the level of danger increases significantly, and the cortisol level also rises significantly. However, since the learning results in such special situations are unlikely to be valid in similar situations in the future, learning does not proceed in order to avoid overlearning.

[0153] With this configuration, when a critical event occurs to the robot, the robot can learn to avoid danger by inputting the conditions before and after the event.

[0154] For example, if a person suddenly jumps into the robot's path and threatens to collide with it, the robot will learn that there is a possibility that a person may suddenly jump out at a certain time and location.The robot will also learn safer avoidance behavior under further or specific conditions.

[0155] Furthermore, not only avoidance actions but also actions according to conditions may be associated. For example, an autonomous vehicle may learn a safety measure such as using high beams to more reliably avoid a collision. Also, in the case of fog, a safety measure such as honking the horn to more reliably avoid a collision may be learned.

[0156] Furthermore, in the above example, the input information evaluation unit 110 evaluates the possibility that the robot itself will be damaged by the surrounding environment, but it may also be configured to evaluate the danger that the robot poses to the surrounding environment.

[0157] This will enable the control of robots and autonomous systems that do not cause stress to people. For example, the distance between a robot and a person, and the speed at which they approach each other, can be automatically set according to the situation. It is also possible to control parameters such as the speed and cornering of autonomous driving.

[0158] Furthermore, the technology disclosed herein makes it possible to realize smooth interactions between robots or between AIs, and to enable robots and AIs to make ethical and value judgments based on stress, as well as to evaluate these judgments.

[0159] <5. Computer configuration example> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers that are built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0160] FIG. 11 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0161] In the computer 500, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.

[0162] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a storage unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0163] The input unit 506 includes a keyboard, a mouse, a microphone, etc. The output unit 507 includes a display, a speaker, etc. The storage unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives removable media 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0164] In the computer 500 configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the memory unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.

[0165] The program executed by the computer 500 (CPU 501) can be provided by being recorded on a removable medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0166] In the computer 500, the program can be installed in the storage unit 508 via the input / output interface 505 by inserting the removable medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the storage unit 508. Alternatively, the program can be installed in the ROM 502 or the storage unit 508 in advance.

[0167] The program executed by computer 500 may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0168] In this specification, the steps describing a program to be recorded on a recording medium include not only processes that are performed chronologically in the order described, but also processes that are not necessarily performed chronologically but are performed in parallel or individually.

[0169] The embodiments of the technology according to the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technology according to the present disclosure.

[0170] For example, the technology according to the present disclosure can be configured as a cloud computing system in which a single function is shared and processed jointly by multiple devices via a network.

[0171] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0172] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0173] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0174] Furthermore, the technology according to the present disclosure can have the following configuration. (1) a learning unit that learns the system's action selection results for input information; an input information evaluation unit that evaluates the risk of the input information to the system; a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information; Equipped with The learning unit changes the learning efficiency in accordance with the first parameter. Information processing system. (2) The risk level includes at least one of a social risk level and a physical risk level to the system. The information processing system according to (1). (3) The system further includes an action selection unit that selects, based on the time series of the first parameter, an action that reduces the first parameter as an action of the system. (2) An information processing system according to the present invention. (4) The action selection unit biases the action to be selected according to the first parameter. (3) An information processing system according to the present invention. (5) Actions that decrease the first parameter include avoidance, countermeasures to the cause, and other actions that result in rewards. An information processing system according to (3) or (4). (6) The learning unit learns the behavior of the system during a period from when the first parameter exceeds a first threshold to when the first parameter falls below the first threshold. An information processing system according to any one of (1) to (5). (7) The learning unit temporarily improves learning efficiency during the period and decreases it over time. (6) An information processing system according to (6). (8) The learning unit increases the frequency and weight of learning according to the magnitude of the first parameter during the period. (7) An information processing system according to (7). (9) The learning unit suppresses learning of the behavior of the system when the first parameter exceeds a second threshold that is higher than the first threshold. An information processing system according to any one of (6) to (8). (10) a second parameter calculation unit that calculates a second parameter that decreases according to an input amount of the first parameter; The learning unit increases the frequency and weight of learning depending on the magnitudes of the first parameter and the second parameter. (3) An information processing system according to the present invention. (11) a feedback value calculation unit that calculates a feedback value that has a negative correlation with the first parameter based on the first parameter and the second parameter; The first parameter calculation unit calculates the first parameter based on the evaluation value of the input information and the feedback value. (10) An information processing system according to (10). (12) Further, a storage device is provided for storing the learning results of the behavior of the system, The input information evaluation unit evaluates the risk level of the input information based on the learning result stored in the storage device. An information processing system according to any one of (3) to (11). (13) The action selection unit selects an action of the system based on the learning result referred to by the input information evaluation unit. (12) An information processing system according to (12). (14) The input information includes at least one of sensor information including images, sounds, temperatures, pressures, and tactile information, character information, and external information input from an external device or service. An information processing system according to any one of (1) to (13). (15) It is structured as a communication system, the learning unit learns communication with a user; The input information evaluation unit evaluates the expectation of feedback information for the text or voice output by the communication system. An information processing system according to any one of (1) to (14). (16) It is configured as a complaint processing system that processes customer complaints. the learning unit learns how to deal with the complaints, The input information evaluation unit evaluates the degree of difficulty in dealing with the customer or the complaint. An information processing system according to any one of (1) to (14). (17) The system is configured as a recommendation system that makes recommendations to users, the learning unit learns the recommendation results of the recommendation system; The input information evaluation unit evaluates the degree of negativity of feedback information regarding the recommendation result. An information processing system according to any one of (1) to (14). (18) It is configured as a robot, the learning unit learns the behavior of the robot in relation to a surrounding environment; The input information evaluation unit evaluates the possibility of the robot being damaged by the surrounding environment. An information processing system according to any one of (1) to (14). (19) The information processing system Learn the system's behavioral choices based on input information, assessing the risk of the input information to the system; Calculating a first parameter representing stress in the system based on the evaluation value of the input information; The learning efficiency is changed according to the first parameter. Information processing methods. (20) a learning unit that learns the system's action selection results for input information; an input information evaluation unit that evaluates the risk of the input information to the system; a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information; Equipped with The learning unit changes the learning efficiency in accordance with the first parameter. Information processing device. [Explanation of symbols]

[0175] 1 Information processing system, 10 Information processing device, 20 Storage device, 31 Input unit, 32 Control unit, 33 Output unit, 110 Input information evaluation unit, 111 Social risk evaluation unit, 112 Physical risk evaluation unit, 113 Learning result reference and risk integration unit, 120 Storage device, 130 First parameter calculation unit, 140 Action selection unit, 150 Inhibition and memory processing unit, 151 Second parameter calculation unit, 152 Negative feedback value calculation unit, 153 Memory and learning unit, 500 Computer, 501 CPU

Claims

1. a learning unit that learns the system's action selection results for input information; an input information evaluation unit that evaluates the risk of the input information to the system; a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information; Equipped with The learning unit changes the learning efficiency in accordance with the first parameter. Information processing system.

2. The risk level includes at least one of a social risk level and a physical risk level to the system. The information processing system according to claim 1 .

3. The system further includes an action selection unit that selects, based on the time series of the first parameter, an action that reduces the first parameter as an action of the system. The information processing system according to claim 2 .

4. The action selection unit biases the action to be selected in accordance with the first parameter. The information processing system according to claim 3 .

5. The behaviors that decrease the first parameter include avoidance, countermeasures against the cause, and other behaviors that can obtain rewards. The information processing system according to claim 3 .

6. The learning unit learns the behavior of the system during a period from when the first parameter exceeds a first threshold to when the first parameter falls below the first threshold. The information processing system according to claim 1 .

7. The learning unit temporarily improves learning efficiency during the period and decreases it over time. The information processing system according to claim 6.

8. The learning unit increases the frequency and weight of learning depending on the magnitude of the first parameter during the period. The information processing system according to claim 7 .

9. The learning unit suppresses learning of the behavior of the system when the first parameter exceeds a second threshold that is higher than the first threshold. The information processing system according to claim 6.

10. a second parameter calculation unit that calculates a second parameter that decreases according to an input amount of the first parameter; The learning unit increases the frequency and weight of learning depending on the magnitudes of the first parameter and the second parameter. The information processing system according to claim 3 .

11. a feedback value calculation unit that calculates a feedback value that has a negative correlation with the first parameter based on the first parameter and the second parameter; The first parameter calculation unit calculates the first parameter based on the evaluation value of the input information and the feedback value. The information processing system according to claim 10.

12. Further, a storage device is provided for storing the learning results of the behavior of the system, The input information evaluation unit evaluates the risk level of the input information based on the learning result stored in the storage device. The information processing system according to claim 3 .

13. The action selection unit selects an action of the system based on the learning result referred to by the input information evaluation unit. The information processing system according to claim 12.

14. The input information includes at least one of sensor information including images, sounds, temperatures, pressures, and tactile information, character information, and external information input from an external device or service. The information processing system according to claim 1 .

15. It is structured as a communication system, the learning unit learns communication with a user; The input information evaluation unit evaluates the expectation of feedback information for the text or voice output by the communication system. The information processing system according to claim 1 .

16. It is configured as a complaint processing system that processes customer complaints. the learning unit learns how to deal with the complaints, The input information evaluation unit evaluates the degree of difficulty in dealing with the customer or the complaint. The information processing system according to claim 1 .

17. The system is configured as a recommendation system that makes recommendations to users, the learning unit learns the recommendation results of the recommendation system; The input information evaluation unit evaluates the degree of negativity of feedback information regarding the recommendation result. The information processing system according to claim 1 .

18. It is configured as a robot, the learning unit learns the behavior of the robot in relation to a surrounding environment; The input information evaluation unit evaluates the possibility of the robot being damaged by the surrounding environment. The information processing system according to claim 1 .

19. The information processing system Learn the system's behavioral choices based on input information, assessing the risk of the input information to the system; calculating a first parameter representing stress in the system based on the evaluation value of the input information; The learning efficiency is changed according to the first parameter. Information processing methods.

20. a learning unit that learns the system's action selection results for input information; an input information evaluation unit that evaluates the risk of the input information to the system; a first parameter calculation unit that calculates a first parameter representing stress in the system based on the evaluation value of the input information; Equipped with The learning unit changes the learning efficiency in accordance with the first parameter. Information processing device.

Citation Information

Patent Citations

  • Multi-microgrid unplanned island control method based on artificial emotion reinforcement learning

    CN112186796A

  • Arousal state determination device, arousal state determination method, and program

    JP2018064759A