Brain-inspired Adaptive Control Electronic Device and Method for Solving Bias-Variance Tradeoff

By combining low variance and low bias intelligent systems, adaptive control based on the forecast error baseline, the deviation variance trade-off problem of the intelligent system under environmental changes is solved, and high performance and high adaptability are achieved.

CN115113726BActive Publication Date: 2025-07-29KOREA ADVANCED INST OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210034666.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-18
Filing Date
2022-01-13
Publication Date
2025-07-29
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

When facing environmental changes, it is difficult for existing intelligent systems to maintain low deviation error and low variance error at the same time, resulting in significantly lower performance when environmental changes. The existing compromise method cannot flexibly respond to context changes in the actual environment.

Method used

By combining low-variance intelligent systems and low-bias intelligent systems, the environmental forecast error baseline is realized, and the forecast error baseline is dynamically updated to track environmental changes and maintain low forecast error.

Benefits of technology

It realizes the maintenance of low variance error and low deviation error while environmental changes, improves the performance and adaptability of the intelligent system and reduces the total error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115113726B_ABST
    Figure CN115113726B_ABST
Patent Text Reader

Abstract

Multiple embodiments of the present invention relate to an electronic device and a method for brain-inspired adaptive control for solving the bias-variance trade-off, which estimate a prediction error baseline for an environment based on a first prediction error of a low-variance intelligent system and a second prediction error of a low-bias intelligent system for the environment, and implement an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the estimated prediction error baseline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Multiple embodiments of the present invention relate to an electronic device and a method for brain-inspired adaptive control for solving the bias-variance tradeoff. Background Art

[0002] In the design of engineering control and learning systems, the bias-variance tradeoff belongs to one of the most fundamental issues that have not been solved yet. Although intelligent systems with high complexity are beneficial for solving specific problem situations (low bias error), they will show a significant performance degradation (high variance error) even when there are slight changes in the environment. On the other hand, although intelligent systems with low complexity have a small performance difference (low variance error) with the change of the environment, they will generally show low performance (high bias error).

[0003] In order to develop an optimal intelligent system, preferably, the existing mainstream methodology compromises and selects a sub-optimal system with the smallest sum of the bias error and the variance error. However, this compromising methodology has a potential risk of performance degradation because it cannot flexibly respond to the context (also known as the context) changes of the actual environment. On the contrary, humans can show a learning mode that can flexibly adapt to context changes. Summary of the Invention

[0004] Technical Problem

[0005] An object of the present invention is to provide an adaptive control system as described below, that is, a new form of adaptive control system is derived from a brain computing algorithm that always maintains a low error value through appropriate movement control between an intelligent system with a low bias error and an intelligent system with a low variance error.

[0006] Technical Solution

[0007] Multiple embodiments of the present invention provide an electronic device and a method for brain-inspired adaptive control for solving the bias-variance tradeoff.

[0008] According to multiple embodiments, the method of the electronic device may include: an estimation step of estimating a prediction error baseline for the environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system; and an adaptive control system implementation step of implementing an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the estimated prediction error baseline.

[0009] According to multiple embodiments, an electronic device may include: a memory; and a processor connected to the memory and configured to execute at least one instruction stored in the memory. The processor estimates a prediction error baseline for the environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system, and implements an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the estimated prediction error baseline.

[0010] According to multiple embodiments, the present invention may store one or more programs for executing a method. The method includes: an estimation step of estimating a prediction error baseline for the environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system; and an adaptive control system implementation step of implementing an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the estimated prediction error baseline.

[0011] Effects of the Invention

[0012] Multiple embodiments of the present invention may implement a brain-inspired adaptive control system for solving the bias-variance trade-off. That is, the electronic device moves and combines a low-variance intelligent system (which refers to an intelligent control or learning system that generates low-variance errors) and a low-bias intelligent system (which refers to an intelligent control or learning system that generates low-bias errors) based on a prediction error baseline of the environment. Thus, the bias-variance trade-off can be solved based on the characteristics of the natural intelligent system of the human brain. In this case, the electronic device may update the prediction error baseline based on changes in the environment to track the total prediction error and maintain a low prediction error. Therefore, the adaptive control system may simultaneously have low-variance errors and low-bias errors and maintain a low prediction error. Brief Description of the Drawings

[0013] Figure 1 Diagrams showing an electronic device according to multiple embodiments.

[0014] Figure 2 , Figure 3 , Figure 4 and Figure 5 Diagrams for explaining the features of an electronic device according to multiple embodiments.

[0015] Figure 6 Diagrams showing the method of an electronic device according to multiple embodiments.

[0016] Figure 7 , Figure 8 , Figure 9a and Figure 9b Diagrams for explaining the performance of an electronic device according to multiple embodiments. Detailed Description of the Embodiments

[0017] Hereinafter, a plurality of embodiments of the present specification will be described with reference to the accompanying drawings.

[0018] Developing an intelligent system for solving problem situations necessarily involves a bias-variance tradeoff. Since an intelligent system with high complexity overfits the problem situations encountered during the learning process, even if there are minor changes in the environment, it cannot operate properly, resulting in a significant reduction in performance (low bias error and high variance error). On the contrary, although an intelligent system with too low complexity has low performance due to insufficient learning (underfitting), it has the flexibility to cope with minor changes (high bias error and low variance error). To select the best intelligent system in this bias-variance tradeoff, a commonly used method is to use an intelligent system with a lower total error (the sum of the bias error and the variance error), that is, to use an intelligent system with appropriate complexity. However, this method is only a compromise solution, and since there are still problems with high errors, it cannot be regarded as a solution to the bias-variance tradeoff. In particular, if the environment undergoes large changes like actual situations and becomes diverse, even higher errors will be generated.

[0019] To solve the above problems, it is necessary to develop an intelligent system that satisfies the following conditions: First, even under various environmental changes, the total error can be minimized; second, low error can be always maintained through appropriate and effective adaptive control between a system with low variance error and a system with low bias error.

[0020] To minimize the total error even in various environmental changes, an intelligent system that can accurately respond to the error distribution that changes with the environment is required. Since low-level environmental changes do not cause large changes in the corresponding distribution, a low-complexity intelligent system can be used to cope with them. However, if the environmental changes result in an error distribution completely different from the previous situation, it cannot be accurately responded to. To cope with this, the intelligent system itself needs to track the error that changes with the environment. If the changed error is tracked by updating the prediction error baseline, the deviation between the environmental error distribution and the error distribution predicted by the intelligent system will be reduced, and even if the environment changes significantly, the total error can be minimized.

[0021] That is, as a system that can implement multiple intelligent systems to which an application belongs and flexibly control the intelligent systems in a manner with low error according to the situation, a natural intelligent system including a human can respond to environmental changes through this flexible control. In particular, it is well known that two reinforcement learning algorithms classified as model-based and model-free organisms respectively have low bias error and high complexity, and low variance error and low complexity. Therefore, appropriate control between these two reinforcement learning algorithms can be a method to solve the bias-variance error by reducing the total error.

[0022] Figure 1 FIG. for showing the electronic device 100 of multiple embodiments Figure 2 、 Figure 3 、 Figure 4 and Figure 5 FIG. for explaining the features of the electronic device 100 of multiple embodiments

[0023] Referring to Figure 1 In multiple embodiments, the electronic device 100 may include at least one of an input module 110, an output module 120, a memory 130, or a processor 140. In an embodiment, at least one of the multiple structural elements of the electronic device 100 may be omitted, and at least one other structural element may be added. In some embodiments, at least two of the multiple structural elements of the electronic device 100 may form an integrated circuit.

[0024] The input module 110 may input a signal to be used to at least one structural element of the electronic device 100. The input module 110 may include at least one of an input device, a sensing device, and a receiving device. The input device may enable a user to directly input a signal to the electronic device 100. The sensing device generates a signal by detecting surrounding changes. The receiving device is used to receive a signal from an external device. For example, the sensing device may include an inertial measurement unit (IMU). The inertial measurement unit may include a gyroscope, an acceleration sensor, and a geomagnetic sensor. The acceleration sensor can detect the roll angle, yaw angle, pitch angle, etc. For example, the input device may include at least one of a microphone, a mouse, or a keyboard. In some embodiments, the input device may include at least one of a touch circuit and a sensing circuit. The touch circuit is used to detect a touch, and the sensing circuit is used to measure the intensity of the force generated according to the touch.

[0025] The output module 120 can output information to the outside of the electronic device 100. The output module 120 may include at least one of a display device, an audio output device, and a transmitting device. The display device can visually output information, the audio output device can output information as an audio signal, and the transmitting device can wirelessly transmit information. For example, the display device may include at least one of a display, a holographic projection device, and a projector. As an example, the display device can be a touch screen, which is implemented by assembling at least one of the touch circuit or the sensing circuit of the input module 110. For example, the audio output device may include at least one of a speaker or a receiver.

[0026] According to an embodiment, the receiving device and the transmitting device can be implemented through a communication module. In the electronic device 100, the communication module can communicate with an external device. Since the communication module establishes a communication channel between the electronic device 100 and the external device, communication can be performed with the external device through the communication channel. Among them, the external device may include at least one of a vehicle, a satellite, a ground station, a server, or other electronic devices. The communication module may include at least one of a wired communication module or a wireless communication module. The wired communication module can be connected to an external device in a wired manner to implement wired communication. The wireless communication module may include at least one of a short-range communication module or a long-range communication module. The short-range communication module can communicate with an external device through a short-range communication method. For example, the short-range communication method may include at least one of Bluetooth, WiFi direct, and infrared data association (IrDA). The long-range communication module can communicate with an external device through a long-range communication method. Among them, the long-range communication module can communicate with an external device through a network. For example, the network may include at least one of computer networks such as a cellular network, the Internet, a local area network (LAN), or a wide area network (WAN).

[0027] The memory 130 can store various data used by at least one structural element of the electronic device 100. For example, the memory 130 may include at least one of a volatile memory or a non-volatile memory. The data may include at least one program and related input data or output data. The program can be stored in the memory 130 as software including at least one instruction, and may include at least one of an operating system, middleware, and an application program.

[0028] The processor 140 can control at least one structural element of the electronic device 100 by running the program in the memory 130. Thus, the processor 140 can perform data processing or operations. In this case, the processor 140 can run the instructions stored in the memory 130.

[0029] According to multiple embodiments, the processor 140 may implement an optimal intelligent system by combining a low-variance intelligent system and a low-bias intelligent system. Among them, the low-variance intelligent system represents an intelligent system with low variance error, and the low-bias intelligent system may represent an intelligent system with low bias error. In this case, the processor 140 may flexibly combine the low-variance intelligent system and the low-bias intelligent system based on the changes in the environment to maintain a low prediction error (PE, prediction error), thereby implementing an adaptive control system as the optimal intelligent system. Thus, the information processing process of the human brain that solves the bias-variance trade-off can be implemented as an intelligent system migrated to the model to achieve an adaptive control system.

[0030] According to one embodiment, the low-variance intelligent system may include a model-free (MF, model-free) reinforcement learning (hereinafter, may be referred to as MF) algorithm, and the low-bias intelligent system may include a model-based (MB, model-based) reinforcement learning (hereinafter, may be referred to as MB) algorithm. The MF algorithm, the MB algorithm, and the MF+MB algorithm that combines the MF algorithm and the MB algorithm according to multiple embodiments may have the following Figure 2 characteristics as shown. That is, the MF algorithm has a high bias error and a low variance error, and the MB algorithm may have a high bias error or a low bias error and a high variance error. In contrast, the MF+MB algorithm combined according to multiple embodiments may have a low bias error and a low variance error.

[0031] To this end, the processor 140 may estimate a prediction error baseline (PE baseline, prediction error baseline) for the environment based on the first prediction error of the low-variance intelligent system of the environment and the second prediction error of the low-bias intelligent system. Among them, the prediction error baseline may change corresponding to the changes in the environment in a dynamic environment. Thus, the processor 140 may update the prediction error baseline based on the changes in the environment. In this case, the prediction error baseline may represent the lowest value within the range of the prediction error achieved by the combination of the low-variance intelligent system and the low-bias intelligent system. According to one embodiment, the first prediction error is the reward prediction error (RPE, reward prediction error), and the second prediction error is the state prediction error (SPE, state prediction error). Among them, as Figure 3As shown, the processor 140 estimates the prediction error for the environment through value learning based on the first prediction error and the second prediction error. Meanwhile, the prediction error baseline for minimizing the prediction error can be estimated through strategy control.

[0032] Moreover, the processor 140 can implement an adaptive control system by combining a low-variance intelligent system and a low-bias intelligent system based on the prediction error baseline. In this case, the processor 140 can control the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline. Among them, as Figure 3 shown, the processor 140 can determine the combination ratio of the low-variance intelligent system and the low-bias intelligent system for achieving the prediction error baseline through adaptive control, and combine the low-variance intelligent system and the low-bias intelligent system through the combination ratio. Thus, an adaptive control system that is the optimal intelligent system for the adaptive environment can be implemented. And the adaptive control system can simultaneously have low-variance error and low-bias error and maintain a low prediction error.

[0033] According to an embodiment, the processor 140 can implement an adaptive control system by combining a model-free reinforcement learning algorithm and a model-based reinforcement learning algorithm. In this case, as Figure 4 shown, the processor 140 can estimate the prediction error for the environment through value learning based on the reward prediction error (RPE) of the model-free reinforcement learning algorithm and the state prediction error (SPE) of the model-based reinforcement learning algorithm for the environment, and simultaneously estimate the prediction error baseline for minimizing the prediction error through strategy control. Moreover, as Figure 5 shown, the processor 140 can determine the combination ratio of the model-free reinforcement learning algorithm and the model-based reinforcement learning algorithm for achieving the prediction error baseline, and combine the model-free reinforcement learning algorithm and the model-based reinforcement learning algorithm based on the combination ratio. Among them, as Figure 5 shown in part (a) of Figure 5 when the prediction error baseline is fixedly applied (fixed prediction error baseline) in the change of the environment, the combination ratio of the model-free reinforcement learning algorithm and the model-based reinforcement learning algorithm can also be maintained in the change of the environment. In contrast, as Figure 5As shown in part (b), the processor 140 can adaptively determine the combination ratio of the model-free reinforcement learning algorithm and the model-based reinforcement learning algorithm for implementing the variable prediction error baseline through adaptive control. Thus, an adaptive control system that is the optimal intelligent system for the adaptive environment can be realized.

[0034] Figure 6 A diagram of a method for an electronic device 100 showing multiple embodiments.

[0035] Referring to Figure 6 , in step 610, the electronic device 100 can estimate a prediction error baseline for the environment based on the first prediction error of the low-variance intelligent system of the environment and the second prediction error of the low-bias intelligent system. In this case, the prediction error baseline can represent the lowest value within the prediction error range achieved through the combination of the low-variance intelligent system and the low-bias intelligent system. According to an embodiment, the low-variance intelligent system includes a model-free reinforcement learning algorithm, and the low-bias intelligent system can include a model-based reinforcement learning algorithm. In this case, the first prediction error can be a reward prediction error (RPE), and the second prediction error can be a state prediction error (SPE). Among them, as Figure 3 or Figure 4 shown, the processor 140 can estimate the prediction error for the environment through value learning based on the first prediction error and the second prediction error, and at the same time estimate the prediction error baseline for minimizing the prediction error through policy control.

[0036] Next, in step 620, the electronic device 100 can combine the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline. In this case, the processor 140 can control the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline. Thus, an adaptive control system that is the optimal intelligent system for the adaptive environment can be realized. According to an embodiment, the processor 140 can combine a model-free reinforcement learning algorithm and a model-based reinforcement learning algorithm to implement the adaptive control system. Among them, as Figure 3 or Figure 5 shown in part (b), the processor 140 can determine the combination ratio of the low-variance intelligent system and the low-bias intelligent system for implementing the prediction error baseline through adaptive control, and combine the low-variance intelligent system and the low-bias intelligent system according to the combination ratio.

[0037] According to multiple embodiments, the electronic device 100 may repeatedly execute step 610 and step 620. Among them, the prediction error baseline may change corresponding to the changes in the environment in a dynamic environment. Thus, the processor 140 may update the prediction error baseline based on the changes in the environment in step 610. Moreover, the processor 140 may adaptively determine the combination ratio of the low-variance intelligent system and the low-bias intelligent system for implementing the variable prediction error baseline. Thus, an adaptive control system as the optimal intelligent system for the adaptive environment can be realized. Therefore, the adaptive control system may simultaneously have low-variance error and low-bias error and maintain a low prediction error.

[0038] Figure 7 and Figure 8 FIG. is a diagram for explaining the performance of the electronic device 100 according to multiple embodiments.

[0039] Referring to Figure 7 and Figure 8 , compared with other adaptive control systems, the adaptive control system based on the variable prediction error baseline has excellent performance. As shown in part (a) of Figure 7 , the MB algorithm, the MF algorithm, the fixed combination (fixed MB+MF) algorithm, and the variable combination (variable MB+MF) algorithm are compared, and the variable combination algorithm is an adaptive combination algorithm. The results show that only the exceedance probability of the adaptive combination (variable MB+MF) algorithm is greater than 0.99, which indicates that the performance of the adaptive combination (variable MB+MF) algorithm is better than the performance of the remaining algorithms. Also, the fitness indices of the adaptive combination (variable MB+MF) algorithm such as the Bayesian information criterion (BIC) and the Akaike information criterion (AIC) are less than the fitness indices of the remaining algorithms, which indicates that the performance of the adaptive combination (variable MB+MF) algorithm is better than the performance of the remaining algorithms. On the other hand, as shown in part (b) of Figure 7 , the influence of the context of the environment on the selection operation of the adaptive combination (variable MB+MF) algorithm is evaluated. The results show that the adaptive combination (variable MB+MF) algorithm performs an appropriate selection operation for the context of the environment. On the other hand, as shown in Figure 8As shown, a parameter recovery analysis was performed to detect overfitting to the problem situation during the learning process. In this case, the influence of the context of the environment on the selection of the MB algorithm, MF algorithm, fixed combination (fixed MB+MF) algorithm, and variable combination algorithm, where the variable combination algorithm is the adaptive combination (variable MB+MF) algorithm, was compared. The results showed that, compared with the remaining algorithms, the adaptive combination (variable MB+MF) algorithm can perform appropriate selection according to the context of the environment.

[0040] Figure 9a and Figure 9b FIGS. are for illustrating the performance of the electronic device 100 according to multiple embodiments.

[0041] Referring to Figure 9a and Figure 9b , the predicted error baseline estimated according to multiple embodiments reflects the information processing process of the human brain that solves the bias-variance trade-off. That is, in the human brain region, the reliability of the MB algorithm, MF algorithm, and the predicted error baseline estimated therefrom is very high compared with the neural activation pattern of the medial prefrontal cortex. This indicates that the electronic device 100 can solve the brain-like adaptive control of the bias-variance trade-off through appropriate flexible control between an intelligent system with low bias error and an intelligent system with low variance error.

[0042] According to multiple embodiments, the electronic device 100 can implement a brain-like adaptive control system that solves the bias-variance trade-off. That is, the electronic device 100 flexibly combines a low-variance intelligent system and a low-bias intelligent system based on the predicted error baseline of the environment, and thus can solve the bias-variance trade-off based on the characteristics of the natural intelligent system of the human brain. In this case, the electronic device 100 can update the predicted error baseline corresponding to the change of the environment to track the total predicted error and maintain a low predicted error. Therefore, the adaptive control system can simultaneously have low variance error and low bias error and maintain a low predicted error.

[0043] Since the currently developed intelligent systems belong to a conservative method of reducing the failure risk, the improvement of effective performance is limited due to low complexity. However, according to multiple embodiments, since the adaptive control system can solve the bias-variance trade-off, the performance of existing intelligent systems can be significantly improved at the performance level. Thus, multiple embodiments can be applied or adapted to multiple fields. For example, these fields include the field of sensor-based control systems, the field of human-robot / computer interaction, the field of intelligent Internet of Things (IoT), the field of professional analysis and intelligent education, the field of user-targeted advertising (AD), etc., but are not limited thereto.

[0044] The first field is the field of sensor-based control systems. Taking an automobile as an example, the control systems that were previously controlled by mechanical devices have recently been gradually replaced by electronic devices. This process of electronicization is difficult to obtain additional gains due to the characteristics of electronicization only in the way of simulating mechanical processes. This is because, in the case of functional failures or errors, the losses are large, but the generation of errors can be significantly reduced through the control of an intelligent system that can minimize the deviation variance error. Therefore, it can be applied to the development of low-cost and high-performance control systems.

[0045] The second field is the field of human-robot / computer interaction. Since all the working bases of natural intelligence are caused by high-dimensional recognition functions based on minimizing the deviation variance error, the behavior of humans can be predicted more accurately through such a system. As a representative example, in the field of affective computing, the purpose is to assist human behavior in a manner consistent with the state by reading emotions, which is one of the types of human recognition states. According to multiple embodiments, in addition to simply reading emotions, a system that can effectively assist human behavior can be constructed through the prediction of computer-recognizable emotions and environmentally similar changes in context (such as arousal and non-arousal), so as to assist humans in achieving excellent results.

[0046] The third field is the field of intelligent Internet of Things. In particular, in the field of the Internet of Things, since it is necessary to control a variety of devices, a variety of recognition functions applicable to the control of each device can be formed. In this case, in terms of the generality of multiple embodiments, in the process of controlling each device, the function that a person wants to use more effectively can be predicted in a way that does not cause overfitting by recognizing environmental changes, so as to assist the person.

[0047] The fourth field is the field of professional analysis and intelligent education. The solution related to the deviation variance trade-off achieved by recognizing environmental changes refers to the situation where optimal learning is achieved in the learning process of humans. Multiple embodiments can clarify the learning process of humans. First, the recognition of environmental changes is immature, and second, which part of the deviation variance error is immature. Therefore, it is particularly important to construct an educational system through multiple embodiments that can effectively and efficiently cultivate the working execution abilities of judges, doctors, financial experts, military operation commanders, etc., who are important for decision-making ability. And for such a customized system for intelligent education, which part of the immaturity is crucial can be analyzed in advance.

[0048] The fifth field is the field of user-targeted advertising. Currently, advertising push technology recommends new advertisements based on a person's historical search records. However, this advertising push technology does not fully consider the environmental changes that a person experiences at each moment. This will result in the system pushing advertisements that are completely inconsistent with the user's interests and poor performance. If multiple embodiments are used, more accurate user-targeted advertising can be achieved by constructing the following-described system in the environment that a person experiences, that is, pushing advertisements with a lower probability (error) of not making a choice.

[0049] Although intelligent systems are very useful, when a functional failure occurs, they need to be used with caution in fields that can lead to fatal results. However, recently, artificial intelligence based on introducing intelligent systems into control systems has gradually become a market trend. In this trend, between the existing control systems (low complexity and low performance) and the latest control systems based on artificial intelligence (high complexity and high performance), intelligent systems that solve the bias-variance trade-off can be used as efficient and effective control systems (low complexity and high performance). All intelligent systems necessarily produce a bias-variance trade-off. Multiple embodiments can not only solve the bias-variance trade-off, but also successfully implement functions even in high-variation environments. Since they are less complex than the latest intelligent systems developed based on deep learning and more high-performance than previous intelligent systems, they can be applied to all industries that mainly engage in intelligent systems.

[0050] The method of the electronic device 100 according to multiple embodiments may include: a presumption step 610 of presuming a forecast error baseline for the environment based on the first forecast error of the low-variance intelligent system of the environment and the second forecast error of the low-bias intelligent system; and an adaptive control system implementation step 620 of implementing an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the presumed forecast error baseline.

[0051] According to multiple embodiments, the adaptive control system implementation step 620 may include a control step of controlling the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the forecast error baseline.

[0052] According to multiple embodiments, the forecast error baseline may change corresponding to the change of the environment.

[0053] According to multiple embodiments, the forecast error baseline may represent the lowest value within the forecast error range achieved by the combination of the low-variance intelligent system and the low-bias intelligent system.

[0054] According to multiple embodiments, the low-variance intelligent system may include a model-free (MF) reinforcement learning algorithm, and the low-bias intelligent system may include a model-based (MB) reinforcement learning algorithm.

[0055] According to multiple embodiments, the first prediction error may be a reward prediction error (RPE), and the second prediction error may be a state prediction error (SPE).

[0056] According to multiple embodiments, the above method may be repeatedly executed as the environment changes.

[0057] According to multiple embodiments, the electronic device 100 may include a memory 130; and a processor 140 connected to the memory 130 and configured to execute at least one instruction stored in the memory 130.

[0058] According to multiple embodiments, the processor 140 may estimate a prediction error baseline for the environment based on the first prediction error of the low-variance intelligent system and the second prediction error of the low-bias intelligent system of the environment, and implement an adaptive control system by combining the low-variance intelligent system and the low-bias intelligent system based on the estimated prediction error baseline.

[0059] According to multiple embodiments, the processor 140 may control the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline.

[0060] According to multiple embodiments, the prediction error baseline may change corresponding to the change of the environment.

[0061] According to multiple embodiments, the prediction error baseline may represent the lowest value within the prediction error range achieved by the combination of the low-variance intelligent system and the low-bias intelligent system.

[0062] According to multiple embodiments, the low-variance intelligent system may include a model-free (MF) reinforcement learning algorithm, and the low-bias intelligent system may include a model-based (MB) reinforcement learning algorithm.

[0063] According to multiple embodiments, the first prediction error may be a reward prediction error (RPE), and the second prediction error may be a state prediction error (SPE).

[0064] The devices described above can be implemented by hardware structure elements, software structure elements, and / or a combination of hardware structure elements and software structure elements. For example, the devices and structure elements described in the embodiments can be implemented using one or more general-purpose computers or special-purpose computers such as processors, controllers, arithmetic logic units (ALUs), digital signal processors, microcomputers, field programmable gate arrays (FPGAs), programmable logic units (PLUs), microprocessors, or other devices that can execute and respond to instructions. The processing device can execute an operating system (OS) and one or more software applications executed on the above operating system. Also, the processing device can access, store, operate, process, and generate data in response to the execution of software. For ease of understanding, the case of using only one processing device is illustrated, and those of ordinary skill in the art to which the present invention pertains will appreciate that the processing device can include multiple processing elements and / or multiple types of processing elements. For example, the processing device can include multiple processors or one processor and one controller. Also, other processing configurations such as parallel processors can be included.

[0065] The software can include a computer program, code, instruction, or a combination of one or more of them, and can configure the processing device in a manner that works as needed or issue instructions to the processing device independently or collectively. The software and / or data can be embodied as any type of mechanical, structural element, physical device, computer storage medium, or device for the purpose of being interpreted by the processing device or providing instructions or data to the processing device. The software is distributed over computer systems connected by a network, so that it can be stored or executed by a distributed method. The software and data can be stored in one or more computer-readable recording media.

[0066] The methods of multiple embodiments can be recorded in a computer-readable medium in the form of program instructions implemented by various computer units. In this case, the medium can continuously store computer-executable programs, or can temporarily store them for running or downloading. Moreover, the medium can be various recording units or storage units combined by single or multiple hardware, and is not limited to the medium directly accessible to the computer system, and can also be dispersed on the network. As an example, the medium can include magnetic media such as hard disks, floppy disks and magnetic disks, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and read-only memories (ROMs), random access memories (RAMs), flash memories, etc. to store program instructions. And, as another example of the medium, it can also be an application store that distributes applications or a recording medium or storage medium managed by websites, servers, etc. that provide and distribute various other software.

[0067] The multiple embodiments of this specification and the terms used herein do not limit the technology described in this specification to a specific embodiment, and should be understood to include various changes, equivalent technical solutions and / or alternative technical solutions of the corresponding embodiments. In the process of describing the drawings, similar reference numerals may be used for similar structural elements. Unless clearly indicated in the context, singular expressions may also include plural expressions. In this specification, expressions such as "A or B", "at least one of A and / or B", "A, B or C" or "at least one of A, B and / or C" can include all combinations of the listed items. Expressions such as "first", "second", "the first" or "the second" are only used to modify the corresponding structural elements, have nothing to do with the order or importance, and are only used to distinguish one structural element from other structural elements, and do not limit the corresponding structural elements. At the functional or communication level, when a certain (for example, the first) structural element is "combined" or "coupled" with another (for example, the second) structural element, the above-mentioned certain structural element can be directly connected or coupled to the other structural element, or the above-mentioned certain structural element can be connected to the other structural element through other structural elements (for example, the third structural element).

[0068] The term "module" used in this specification includes a unit composed of hardware, software or firmware, and can be used interchangeably with terms such as logic, logic block, component or circuit, for example. A module can be an integrated component or the smallest unit or a part thereof that executes one or more functions. For example, a module can be composed of an application-specific integrated circuit (ASIC).

[0069] According to multiple embodiments, among the illustrated multiple structural elements, each structural element (e.g., module or program) may include singular or plural individuals. According to multiple embodiments, one or more structural elements or steps may be omitted from the corresponding structural element or one or more other structural elements or steps may be added. Instead of or additionally, multiple structural elements (e.g., modules or programs) may be integrated into one structural element. In this case, the integrated structural element may perform one or more functions of each of the multiple structural elements in the same or similar manner as the corresponding structural element among the multiple structural elements before integration. According to multiple embodiments, the steps performed by a module, program, or other structural element may be executed sequentially, in parallel, repeatedly, or in a matching manner, or one or more of the multiple steps may be executed in a different order or omitted, or one or more other steps may be added.

Claims

1. A method for an electronic device, characterized in that, Comprising: A presumption step of presuming a prediction error baseline for the above environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system, wherein the prediction error baseline represents the lowest value within the range of prediction errors achieved through the combination of the low-variance intelligent system and the low-bias intelligent system; And An adaptive control system implementation step of implementing an adaptive control system by combining the above low-variance intelligent system and the above low-bias intelligent system based on the presumed prediction error baseline; Wherein the adaptive control system implementation step includes controlling the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline.

2. The method of an electronic device according to claim 1, wherein The above prediction error baseline changes corresponding to the change of the environment.

3. The method of the electronic device according to claim 1, wherein, The above low-variance intelligent system includes a model-free reinforcement learning algorithm, and the above low-bias intelligent system includes a model-based reinforcement learning algorithm.

4. The method of an electronic device according to claim 1, characterized in that, The above first prediction error is a reward prediction error, and the above second prediction error is a state prediction error.

5. The method of an electronic device according to claim 2, characterized in that, The above method is repeatedly executed as the above environment changes.

6. An electronic device, characterized in that Comprising: A memory; and A processor connected to the above memory for running at least one instruction stored in the above memory, The above processor presumes a prediction error baseline for the above environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system, wherein the prediction error baseline represents the lowest value within the range of prediction errors achieved through the combination of the low-variance intelligent system and the low-bias intelligent system; Based on the presumed prediction error baseline, an adaptive control system is implemented by combining the above low-variance intelligent system and the above low-bias intelligent system; wherein the implementation of the adaptive control system includes controlling the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline.

7. The electronic device according to claim 6, wherein The above prediction error baseline changes corresponding to the change of the environment.

8. The electronic device according to claim 6, wherein The above low-variance intelligent system includes a model-free reinforcement learning algorithm, and the above low-bias intelligent system includes a model-based reinforcement learning algorithm.

9. The electronic device according to claim 6, wherein, The above first prediction error is a reward prediction error, and the above second prediction error is a state prediction error.

10. A non-transitory computer-readable storage medium, characterized in that, Storing one or more programs for executing the method, The above method includes: A presumption step of presuming a prediction error baseline for the above environment based on a first prediction error of a low-variance intelligent system of the environment and a second prediction error of a low-bias intelligent system, wherein the prediction error baseline represents the lowest value within the range of prediction errors achieved through the combination of the low-variance intelligent system and the low-bias intelligent system; and An adaptive control system implementation step of implementing an adaptive control system by combining the above low-variance intelligent system and the above low-bias intelligent system based on the presumed prediction error baseline, wherein the adaptive control system implementation step includes controlling the combination ratio of the low-variance intelligent system and the low-bias intelligent system based on the prediction error baseline.

11. The non-transitory computer-readable storage medium according to claim 10, wherein The above prediction error baseline changes corresponding to the change of the environment.

12. The non-transitory computer-readable storage medium according to claim 10, wherein The above low-variance intelligent system includes a model-free reinforcement learning algorithm, and the above low-bias intelligent system includes a model-based reinforcement learning algorithm.

13. The non-transitory computer-readable storage medium according to claim 10, wherein The above first prediction error is a reward prediction error, and the above second prediction error is a state prediction error.

14. The non-transitory computer-readable storage medium according to claim 11, wherein The above method is repeatedly executed as the above environment changes.