Reinforcement learning based scheme for tuning memory interfaces

By tuning the parameters of DRAM devices using a reinforcement learning-based machine learning model, the problem of low efficiency in existing technologies has been solved, and the stability and reliability of the devices have been improved, while power consumption has been reduced.

CN116457796BActive Publication Date: 2025-11-28QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180072783.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-29
Filing Date
2021-10-14
Publication Date
2025-11-28
Estimated Expiration
2041-10-14

AI Technical Summary

Technical Problem

In the prior art, the parameter tuning methods for DRAM devices are inefficient and cannot effectively adapt to changes in the manufacturing, packaging, and deployment environments, resulting in suboptimal device performance.

Method used

By employing a reinforcement learning-based machine learning model, the parameters of DRAM devices are tuned to improve stability and reliability and reduce power consumption by generating a set of reward values ​​and a reward function.

Benefits of technology

It enables parameter tuning of DRAM devices, improves device stability and reliability, reduces power consumption, adapts to different environmental changes, and meets actual deployment needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116457796B_ABST
    Figure CN116457796B_ABST
Patent Text Reader

Abstract

A method performed by a machine learning system includes generating a set of reward values based on a set of parameter values selected by the machine learning system, each reward value in the set of reward values corresponding to a parameter value in the set of parameter values programmed at a device. The method also includes determining a reward function for maximizing a reward of a set of parameters corresponding to the device based on the set of reward values. The method further includes tuning a parameter in the set of parameters based on the reward function.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 084,508, filed on October 29, 2020, entitled “REINFORCEMENT LEARNING BASEDSCHEME FOR TUNING MEMORY INTERFACES”, the disclosure of which is expressly incorporated herein by reference in its entirety. Background Technology

[0003] field

[0004] Various aspects of this disclosure generally relate to memory interface tuning based on reinforcement learning.

[0005] background

[0006] High-speed dynamic random access memory (DRAM) can be tuned to operate reliably at speeds exceeding 1 GHz. Tuning can be specified for several components due to variations in, for example, manufacturing, packaging, and / or printed circuit board (PCB). In most cases, DRAM is tuned during manufacturing. Once deployed, one or more parameters or components can be retuned to address variations in operating conditions and / or hardware degradation. Conventional systems perform limited on-chip task mode tuning. Improved on-chip tuning may be desired.

[0007] Overview

[0008] In one aspect of this disclosure, a method performed by a machine learning model is disclosed. The method includes generating a set of reward values ​​based on a set of parameter values ​​selected by the machine learning model. The method further includes determining a reward function based on the set of reward values ​​for maximizing a reward corresponding to a set of parameters of a device. The method further includes tuning the parameters in the set of parameters based on the reward function.

[0009] Another aspect of this disclosure relates to an apparatus including means for generating a set of reward values ​​based on a set of parameter values ​​selected by the machine learning model. The apparatus further includes means for determining a reward function based on the set of reward values ​​to maximize the reward corresponding to the set of parameters of the apparatus. The apparatus further includes means for tuning the parameters in the set of parameters based on the reward function.

[0010] In another aspect of the disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code is executed by a processor and includes program code to generate a set of reward values based on a set of parameter values selected by the machine learning model. The program code also includes program code to determine a reward function for maximizing a reward of a set of parameters corresponding to the device based on the set of reward values. The program code further includes program code to tune a parameter of the set of parameters based on the reward function.

[0011] Another aspect of the disclosure relates to an apparatus. The apparatus has a memory, one or more processors coupled to the memory, and instructions stored in the memory. The instructions, when executed by the processor, are operable to cause the apparatus to generate a set of reward values based on a set of parameter values selected by the machine learning model. The instructions also cause the apparatus to determine a reward function for maximizing a reward of a set of parameters corresponding to the device based on the set of reward values. The instructions additionally cause the apparatus to tune a parameter of the set of parameters based on the reward function.

[0012] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer-readable media, user equipments, base stations, wireless communication devices, and processing systems, as substantially described and as illustrated by the drawings and specification.

[0013] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows can be better understood. Additional features and advantages will be described hereinafter. The disclosed concepts and specific examples can be readily utilized as bases for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions are not to be regarded as a departure from the scope of the appended claims. The novel features of the concepts disclosed will be more particularly described below, by way of example, and compared to prior approaches. The disclosed concepts are not limited in their application to the details of construction and the arrangement of components set forth in the following description or illustrated in the drawings. The concepts are capable of other constructions and of being practiced or carried out in various ways. Features and aspects of the disclosed concepts can be better understood from the following description taken in conjunction with the accompanying drawings. Each of the drawings is provided for the purpose of illustration and description, and not as a definition of the limits of the claims. BRIEF DESCRIPTION OF DRAWINGS

[0015] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings in which like reference characters identify correspondingly throughout and wherein:

[0016] Figure 1 An example implementation of designing a machine learning model using a system on chip (SOC), including a general purpose processor, is illustrated in accordance with certain aspects of the present disclosure.

[0017] Figure 2 is a block diagram illustrating an example of a reinforcement learning model in accordance with aspects of the present disclosure.

[0018] Figure 3 is a diagram illustrating an example of a data eye in accordance with aspects of the present disclosure.

[0019] Figure 4 is a diagram illustrating an example of Bayesian optimization for a machine learning model in a Gaussian process in accordance with aspects of the present disclosure.

[0020] Figure 5 An example of a dynamic random access memory (DRAM) parameter tuning system in accordance with aspects of the present disclosure is illustrated.

[0021] Figure 6 is a flow diagram illustrating an example process performed, for example, by a parameter tuning device in accordance with aspects of the present disclosure.

[0022] DETAILED DESCRIPTION

[0023] The detailed description set forth below, in connection with the appended drawings and description of various configurations, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without

[0024] Based on the teachings herein those skilled in the art will appreciate that the scope of the disclosure is intended to cover all aspects of the present disclosure including those which are presented in the claims below. It is therefore contemplated to cover any and all modifications, variations or equivalents that fall within the scope of the present disclosure. It is intended that the specification and examples be considered as exemplary only, with a true scope of the disclosure being indicated by the following claims.

[0025] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0026] While certain aspects are described, numerous variations and permutations of the aspects are possible. While some benefits and advantages of the preferred aspects have been described, additional benefits and advantages can be realized and are within the scope of the disclosure. Furthermore, some benefits and advantages are not to be construed as sub- stantial features of the preferred aspects. The preferred aspects are not to be construed as limiting of the scope of the disclosure, but rather as representative thereof. The scope of the disclosure is defined by the claims and their equivalents.

[0027] Memory for a system on a chip (SOC) can be tuned to reliably operate at high speeds, such as speeds above 1 GHz. The memory can be a high speed dynamic random access (DRAM) memory or another type of memory. For ease of explanation, aspects of the present disclosure use DRAM as an example. Still, aspects of the present disclosure are not limited to DRAM. The tuning can be specified for several components due to, for example, variations in manufacturing, packaging, and / or printed circuit board (PCB). In most cases, the DRAM is tuned during the manufacturing packaging process. Once deployed, one or more parameters or components can be retuned to address variations in operating conditions and / or hardware degradation.

[0028] For example, factory boot can be performed for data bus tuning to compensate for routing skew on the printed circuit board. The tuning can find a stable operating point for the data eye. The stable operating point can be the center of the data eye in a reference voltage and clock delay circuit shmoo plot. In some cases, periodic tuning is triggered to compensate for variations in operating conditions, such as an increase in operating temperature and voltage.

[0029] As DRAM speeds increase, improved optimization of on-chip parameters can be specified to extract the maximum margin of each platform. The on-chip parameters can include DRAM interface parameters, such as pull-up / pull-down, on-die termination, and / or read / write turnaround time. Some conventional systems do not optimize the DRAM interface parameters. Other conventional systems perform limited optimization. For example, conventional systems can perform partial optimization by modeling and simulating real-time traffic. The partial optimization can be limited to the accuracy of the real-time traffic model, which does not account for device-to-device variations.

[0030] In some conventional systems, additional optimization can be performed on a limited set of interfaces in a tuning environment, such as a lab. The optimization performed in the tuning environment can be referred to as a universally best-fit optimization. The settings of the final optimization can be broadcast to various DRAM devices. Optimizing parameters in the tuning environment can not account for a large number of devices manufactured and shipped across different combinations of part, platform, form factor, and original equipment manufacturer. Additionally, the reliability of the optimization performed in the tuning environment decreases as DRAM speeds increase. That is, the universally best-fit optimization does not account for the maximum stability that any individual system can provide. As a result, the performance of the DRAM devices can be limited by sub-optimal parameters, such as pre-tuned parameters.

[0031] Tuning all parameters for all types of DRAM can be infeasible due to the number of parameters. The parameters can include a combination of dependent, independent, and interdependent parameters. Some parameters can be specified for tuning per bit, which increases the tuning time due to the increased number of data lines in high density DRAM. In other examples, a brute force method tries all combinations to find the best fit parameters for factory boot. The brute force method can not account for periodic training performed while the DRAM is online (e.g., actively in use). Additionally, the brute force method can be time consuming.

[0032] Aspects of the present disclosure relate to sampling a set of parameters (e.g., operating configurations) from an individual DRAM and tuning the parameters of that DRAM based on the sampled set of parameters. The tuned parameters can improve the reliability and stability of a high speed DRAM interface. In one configuration, the parameters are tuned without sampling all of the parameters. Additionally, aspects of the present disclosure improve the stability and reliability of a high speed DRAM interface. In one configuration, performance is improved and power is reduced by optimizing DRAM parameters (e.g., on-chip parameters) with a machine learning model.

[0033] Figure 1 An example implementation of a system on chip (SOC) 100 according to certain aspects of the present disclosure is illustrated, which can include a central processing unit (CPU) 102 or multi-core CPU configured for parameter tuning. Variables, system parameters associated with a computing device, delays, frequency bin information, and task information can be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or can be distributed across multiple blocks. The NPU 108 can execute a machine learning model. Instructions executed at the CPU 102 can be loaded from a program memory associated with the CPU 102 or can be loaded from the memory block 118.

[0034] The SOC 100 can also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which can include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU, DSP, and / or GPU. The SOC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which can include a global positioning system).

[0035] The SOC 100 can be based on an ARM instruction set. In an aspect of the disclosure, instructions loaded into the general purpose processor 102 can include code to generate a set of reward values based on a set of parameter values selected by a machine learning model, determine a reward function for maximizing a reward for a set of parameters corresponding to the device based on the set of reward values, and tune a parameter in the set of parameters based on the reward function.

[0036] Learning functions for devices implementing machine learning models can generally be classified into supervised learning and unsupervised learning. Reinforcement learning can be based on unsupervised learning. As an example, reinforcement learning can learn decisions, classifications, and / or actions based on rewards. That is, reinforcement learning can learn to take appropriate actions in an environment by learning to maximize future rewards. Aspects of the disclosure relate to selecting one or more parameters to improve performance of a device while reducing power consumption. The one or more parameters can be an example of an action.

[0037] A reinforcement learning model can be specified for a decision task in a reinforcement learning framework. An input to an agent of the reinforcement learning model can be a state s. An output of the reinforcement learning model can be a state-action value function Q() (e.g., Q(s, a)) for all available actions a for a given state s. The state-action value function Q() provides a cumulative reward (e.g., quality (Q) value). The Q value can be a sum of an immediate reward for selecting an action a from a state s and a possible highest Q value from a subsequent state (e.g., a state s+1 after taking the action a from the current state s). The true value of the value Q(s, a) can not be initially known with respect to a combination of the state s and the action a. The agent can select one or more actions a in a certain state s and can receive one or more rewards for the action a. In this way, the agent learns to select an action that maximizes a reward (e.g., an optimal action).

[0038] Figure 2 FIG. 2 is a block diagram illustrating an example of a reinforcement learning model 200, in accordance with aspects of the disclosure. One or more components of the reinforcement learning model 200, such as an agent 204, an environment 206, and / or an interpreter 202, can be components of the SOC 100. Alternatively, the reinforcement learning model 200 can be a separate device that works in conjunction with the SOC 100. In Figure 2 In an example, the agent 204 can learn to take actions based on a state s observed by the interpreter 202. Additionally, the agent 204 can interact with the environment 206 to maximize a total reward that can be observed by the interpreter 202 and given as reinforcement feedback to the agent 204. In some examples, the agent 204 and the interpreter 202 can be implemented as the same or separate components.

[0039] In some examples, reinforcement learning is modeled as a Markov Decision Process (MDP). An MDP is a discrete, time-stochastic control process. An MDP provides a mathematical framework for modeling decision making in situations where outcomes can be partially random while also being under the control of a decision maker. In an MDP, at each time step, the process is in a state s from a finite set of states S, and the decision maker can choose any action a from a finite set of available actions A in that state. The process responds by stochastically moving to a new state and giving the decision maker a corresponding reward at the next time step. In other examples, a partially observable MDP (POMDP) is used. A POMDP can be used when the state can be unknown when an action is taken, and thus the probabilities and / or rewards can be unknown.

[0040] As described above, aspects of the present disclosure relate to increasing the stability and reliability of high-speed interfaces (e.g., dynamic access memory (DRAM), peripheral component interconnect express (PCIe), etc.). In one implementation, a machine learning model tunes one or more on-chip parameters to improve performance while reducing power consumption. The on-chip parameters can be tuned at an end-user device (e.g., a real-world deployment). The tuning can be referred to as active adaptive correction of interface parameters.

[0041] A system on chip (SOC), such as the SOC 100 described with reference to Figure 1 The SOC 100 can be implemented in various components, such as a vehicle or a robotic device. The safety and reliability of semiconductor devices (e.g., SOCs) in these components can prevent economic loss and harm to human life. Various standards, such as the automotive safety integrity level (ASIL) standards (ISO 26262), standardize and certify the safety and reliability of such components. Conformance testing can be performed to determine whether a component conforms to the standard body specifications.

[0042] In some implementations, conformance testing can be performed via eye diagram analysis. This testing method provides an analysis of the waveform signal integrity of double data rate (DDR) memory by providing a data eye (e.g., eye diagram). The data eye can represent a voltage and time plot of an electrical signal. For example, the eye can be represented as a plot of a reference voltage (Vref) versus clock delay (CDC). The quality of the data eye, such as the eye opening, symmetry, and / or distortion, can indicate the stability of the signal. The data eye can be inspected by a human or a trained neural network to reveal the amount of jitter on the device, non-monotonic edges, and / or other issues on the device. In conventional systems, an engineer can inspect the data eye each time the parameters are tuned. The parameters can be continuously tuned until a satisfactory data eye is obtained.

[0043] As described, a data eye diagram is a time-domain representation of a signal from which the electrical quality of the signal can be visualized. Figure 3is a diagram illustrating an example of a data eye 300 in accordance with aspects of the present disclosure. In Figure 3 In an example, a reference signal (VREF) is provided to a device, such as a DDR memory, to produce the data eye 300. As shown in Figure 3 In an example, a reference signal (VREF) is provided to a device, such as a DDR memory, to produce the data eye 300. As shown in

[0044] As described, high-speed external interfaces, such as DRAM, have various configurable parameters that correspond to electrical characteristics and circuit metrics. These parameters can be adjusted (e.g., tuned) to improve or degrade the quality of transmitted signals. For example, these parameters can be tuned to balance load / impedance or drive strength to reduce negative effects on data lines, such as signal reflections, noise, and / or crosstalk. Configurable parameters can include, for example, signal impedance, drive strength, clock phase delay, on-die termination, and / or data bus inversion.

[0045] Additionally, as described, these parameters can be tuned by engineers in a lab to improve signal stability of high-speed interfaces. DRAM reliability can be based on signal stability, as signal stability provides resilience to noise and environmental changes. In conventional systems, initial parameter values are selected based on educated guesses. Iterative tuning can be performed based on satisfaction or violation of stability and power criteria. The number of parameters for high-speed interfaces can be in the range of thousands or tens of thousands.

[0046] Most parameters can depend on process voltage temperature (PVT) characteristics. As such, power performance of devices, such as DRAM devices, can be limited by sub-optimal parameters. Performance can be further limited as engineers can not have sufficient time and / or other resources to tune parameters for all devices. As described, in most cases, tuning is performed in a validation lab, and generalized parameters can be propagated to all devices. Due to the number of parameters, including combinations of parameters, conventional systems can not be able to accommodate a wide range of environmental changes in a test environment (e.g., real-world deployment). Conventional systems can perform limited tuning (e.g., data bus tuning and periodic tuning) to partially overcome the described constraints.

[0047] In conventional systems, a brute force function and / or a deep learning model can tune parameters of an interface. For example, a brute force function samples all possible combinations to obtain parameters that satisfy a stability criterion. As described above, the stability criterion can be determined by analyzing data eyes. The brute force function is time consuming and infeasible for real-world deployment. That is, the brute force function can not feasibly tune parameters of a deployed device. In other conventional systems, a deep learning model can collect a majority of samples (e.g., seventy percent of samples) from an entire sample space to predict parameters for increasing stability of a device. Given the number of configurable parameter combinations, conventional tuning methods, such as deep learning models and brute force functions, are infeasible for end-user applications due to time constraints.

[0048] According to aspects of the disclosure, a machine learning based solution, such as a reinforcement learning (RL) model, tunes parameters to improve stability of a device, such as a DRAM device. As described, a stable signal can be resilient to noise and environmental changes. Thus, improved signal stability improves reliability of the signal. The number of parameters sampled by the machine learning based solution is less than the number of parameters sampled by a brute force function and a deep learning model. As such, the machine learning based solution for tuning parameters can be less time consuming compared to the brute force function and the deep learning model.

[0049] In one configuration, the machine learning based solution is a RL model designated to model a non-linear function and improves data collection by predicting a sample point based on one or more previous sample points. In particular, the RL model can simulate a human learning scenario in which the RL model assimilates available data to predict outcomes of new actions, performs the actions to collect new data, and assimilates the new data to predict further actions and outcomes until a goal is achieved.

[0050] In aspects of the disclosure, the RL model requests one or more sample points when a certainty of the modeled function is less than a certainty threshold. Over multiple iterations, the RL model improves an accuracy of the underlying function, thereby enabling prediction of an optimal point of parameters to improve signal stability. In one configuration, each parameter is treated as a constrained dimension in an n-dimensional parameter space. Stability and / or other characteristics can be represented as feedback or a reward. The optimal point of parameters can be a point in the parameter space that maximizes a reward function.

[0051] Figure 4 is a diagram illustrating an example of Bayesian optimization of a machine learning model in a Gaussian process according to aspects of the disclosure. As described, the machine learning model can be a reinforcement learning (RL) model. To facilitate explanation, a system (e.g., a device) receives a single input x and generates a non-linear output based on a function f(x).

[0052] In Figure 4 In the example, as shown in block 400a, the machine learning model selects a first parameter value x 402 and determines a corresponding output of the function f(x). The first parameter value x 402 can be selected randomly, based on historical data from previous parameter adjustments or previous training, in accordance with aspects of the present disclosure. The first parameter value x 402 can be a value of a parameter of a device, such as a DRAM device. The output of the function f(x) can be a reward value based on a reward criterion, such as an eye stability and / or power performance of data eyes corresponding to the first parameter value x 402. That is, the first point 404 can correspond to a data eye quality as a function of f(x) of input x. As shown in block 400a, the output of the function f(x) is a first point 404 on a reward value plot 406 of the system. The reward value plot 406 represents data eye quality (e.g., reward values) based on different parameter values x (x-axis) of the function f(x) (y-axis).

[0053] Based on the first point 404, the machine learning model predicts a function f'(x) 408 of a smooth curve (in block 400b). That is, the machine learning model predicts a reward function f'(x) based on all available sampled points. Additionally, the machine learning model can identify points of the predicted reward function f'(x) that have an uncertainty greater than a threshold value. If multiple reward functions f'(x) can be predicted for a parameter value x corresponding to a point, then the uncertainty of the point can be greater than the threshold value. That is, if multiple reward functions f'(x) 408 of smooth curves correspond to a parameter value x, then the uncertainty can be higher than the threshold value.

[0054] In the current example, at block 400b, the machine learning model determines that the uncertainty is greater than the threshold value at a point corresponding to a second parameter value x 410. The second parameter value x 410 is provided to the system and a corresponding reward value is returned to the machine learning model. That is, at block 400c, the machine learning model determines a second point 412 based on the second parameter value x 410. Additionally, at block 400c, the machine learning model predicts a second smooth curve reward function f'(x) 414 based on the first point 404 and the second point 412. Additionally, at block 400c, the machine learning model determines that the uncertainty is greater than the threshold value for a point on the second smooth curve reward function f'(x) 414 corresponding to a third parameter value x 416.

[0055] In Figure 4In the example of FIG. 4, at block 400d, the machine learning model determines a third point 418 corresponding to a third parameter value x 416. Additionally, at block 400d, the machine learning model predicts a third smooth curve reward function f'(x) 420 based on the first point 404, the second point 412, and the third point 418. Additionally, at block 400d, the machine learning model determines that the uncertainty for a point on the third smooth curve function f'(x) 420 corresponding to a fourth parameter value x 422 is greater than the threshold.

[0056] At block 400e, the machine learning model determines a fourth point 424 corresponding to the fourth parameter value x 422. Additionally, at block 400e, the machine learning model predicts a fourth smooth curve reward function f'(x) 426 based on the first point 404, the second point 412, the third point 418, and the fourth point 424. Additionally, at block 400e, the machine learning model determines that the uncertainty for all points on the fourth smooth curve reward function f'(x) 426 is less than the threshold. As such, at block 400e, the predicted reward function f'(x) can be similar to the actual function f(x), and the training can be complete. After determining the reward function f'(x), the RL model can select a parameter value that maximizes the value returned by the reward function f'(x).

[0057] According to aspects of the disclosure, the number of sampled points (such as the first point 404, the second point 412, the third point 418, and the fourth point 424) for a machine learning based solution (e.g., the RL model) can be less than the number of sampled points for a deep learning based solution.

[0058] Referring to Figure 4 The described RL model can be a component of a parameter tuning system, such as a DRAM parameter tuning system. Figure 5 An example of a DRAM parameter tuning system 500 according to aspects of the disclosure is illustrated. As shown in Figure 5 As shown in FIG. 5, the DRAM parameter tuning system 500 can include an RL model 502, a parameter tuning input interface 504, a device 506, a parameter tuning output interface 508, and a data eye rating module 510. The device 506 can be a type of DRAM, such as a double data rate (DDR) DRAM.

[0059] In one configuration, the RL model 502 tunes parameters of the device 506 by maximizing a reward of a reward function (f'(x)). In one configuration, the RL model 502 generates the reward function based on reward values provided by the data eye rating module 510. Different types of rewards can be considered. In one configuration, the data eye rating module 510 generates the reward values based on reward criteria, which can include stability and / or power performance criteria. In one configuration, the data eye rating module 510 generates the reward values based on a reward function that is a function of the stability and / or power performance criteria. Figure 5In the example, the RL model 502 can determine the reward function by sampling the parameters of the device 506.

[0060] In one configuration, when determining the reward function, the RL model 502 selects one or more parameter values. For example, the RL model 502 can select a first parameter value x 402 as described with respect to Figure 4 The first value x 402 can be randomly selected. The first parameter value x 402 can be received at the parameter tuning input interface 504, which can program the device 506 according to the first parameter value x 402. The parameter tuning output interface 508 samples a data eye based on the current programmed configuration of the device 506. In this example, the parameter tuning output interface 508 samples a data eye based on the programmed first parameter value x 402. The data eye can be a shmoo plot of interface reference voltage versus sampling delay. The data eye rating module 510 receives the data eye and generates a rating corresponding to one or more of stability or power performance of the data path at a given operating frequency. In one configuration, the data eye rating module 510 is a convolutional neural network trained to generate a stability score for the data line of the device 506. The power performance refers to the amount of power used by the data path. The rating can be output to the RL model 502, and the RL model 502 can generate the reward function f'(x) based on the rating. That is, the rating output by the data eye rating module 510 can correspond to a point on the reward value plot 406, such as the first point 404 described with respect to Figure 4

[0061] The RL model 502 can continue to sample parameters from the device 506 until the uncertainty of the predicted function f'(x) is less than a threshold. After generating (e.g., learning) the predicted function f'(x), the RL model 502 can tune the parameters of the device 506. The parameters can be tuned to maximize the reward of the predicted function f'(x). For example, after generating the reward function, the RL model 502 can be the parameter value that yields the highest reward from the reward function. The reward can be the predicted stability score and / or power performance. The parameter value with the highest reward can be selected for tuning the device 506. The device 506 can be programmed with the selected value. That is, the selected value can be applied to one or more parameters of the device 506.

[0062] As described above, in one configuration, the parameters of the device are individually tuned to improve power performance. Each parameter can be considered a constrained dimension in an n-dimensional parameter space. Stability and / or power can be represented as feedback or reward. Thus, the optimal parameter can correspond to the point in the parameter space (e.g., f'(x)) where the reward function is maximized.

[0063] As described with respect to Figure 4 ​As described, a machine learning model (e.g., an RL model) samples parameters (e.g., inputs) to learn a prediction function f'(x). To improve training, parameters (e.g., DRAM interface parameters) can be categorized into groups. These groups can include, for example, limited range integer parameters, global parameters, per-channel parameters, and per-bit parameters. The RL model can be trained on individual groups.

[0064] In one configuration, the RL model can be initialized with limited range integer parameters, which can include parameters with a limited number of values (e.g., less than ten values). These limited range integer parameters can be independent of each other. A subset of the limited range integer parameters can be collected offline before sampling the parameters with the RL model for training. The limited range integer parameters can provide initial data for learning a predicted function f'(x) of the RL model, such that an initial prediction function f'(x) can be learned without sampling all of the limited range integer parameters.

[0065] Global parameters can be designated for coarse tuning. Global parameters can be parameters that apply to all bits in a DRAM interface, and can be general top-level parameters. The RL model can tune global parameters for a single data line, and apply the tuned parameters to other data lines (e.g., bits in the DRAM interface). Global parameters can be used to obtain a semi-converged configuration of the RL model.

[0066] After tuning global parameters, the RL model 502 can tune per-channel parameters. Per-channel parameters can apply to all bits in any given DRAM channel. Bits in a DRAM channel can be physically and logically adjacent. In one configuration, a parameter value for one data line in a channel can be applied to all data lines in the channel. For example, the device 506 can include eight eight-bit channels. In this example, for each channel, the RL model 502 can tune one bit, and apply the tuned parameter to the remaining seven bits.

[0067] After tuning per-channel parameters, the RL model 502 can tune individual bits for a single data line (e.g., per-bit training). Because the RL model 502 tunes global parameters and per-channel parameters before tuning per-bit parameters, the number of per-bit parameters can be limited. Thus, the time for tuning per-bit parameters can be reduced compared to traditional tuning systems.

[0068] Aspects of the disclosure can enable autonomous tuning of devices, such as memory (e.g., DRAM). Additionally, aspects of the disclosure can improve the reliability of SOCs. Reliable SOCs can be designated for mission critical devices, such as automobiles and autonomous industrial devices. In some examples, a parameter tuning method can be triggered for mission critical devices as well as other types of devices in response to health checks, fault detection, and / or data eye analysis. The parameter tuning method can provide fault prevention by updating (e.g., tuning) parameters to mitigate faults.

[0069] Aspects of the disclosure are not limited to DRAM and / or SOCs. Other types of devices are contemplated, such as devices with high speed interfaces (e.g., Peripheral Component Interconnect Express (PCIe)).

[0070] Figure 6 FIG. 6 is a diagram illustrating an example process 600 performed, for example, by a parameter tuning device, in accordance with aspects of the present disclosure. The process 600 can be performed by a machine learning system that implements a machine learning model, such as a reinforcement learning model.

[0071] As Figure 6 shown in FIG. 4, at block 602, the process 600 generates a set of reward values based on a set of parameter values selected by the machine learning system. Each reward value in the set of reward values can correspond to a parameter value of the set of parameter values programmed at a device. The device can be a dynamic random access memory (DRAM), a high speed DRAM, or another type of device with programmable parameters.

[0072] At block 604, the process 600 determines a reward function for maximizing a reward corresponding to the set of parameters of the device based on the set of reward values. In one configuration, the process 600 can determine the reward function based on the process described with reference to Figure 4 For example, the process 400 can randomly select a first parameter value of the set of parameter values, program the device with the first parameter value, sample a first data eye in response to programming the device with the first parameter value, generate a first reward value of the set of reward values based on the first data eye, and determine the reward function based on the first reward value. The first reward value can be generated by determining reward criteria including data path stability and / or power consumption of the device based on the first data eye, and generating the first reward value based on the determined reward criteria.

[0073] In one configuration, to determine the reward function, the process 600 can determine that an uncertainty of the reward function is greater than a threshold and select a second parameter value of the set of parameter values based on the uncertainty. In response to selecting the second parameter value, the process 600 can program the device with the second parameter value and sample a second data eye in response to programming the device with the second parameter value. Additionally, the process 600 can generate a second reward value of the set of reward values based on the second data eye and update the reward function based on the second reward value.

[0074] At block 606, the process 600 tunes a parameter of the set of parameters based on the reward function. In one configuration, the parameters can include global parameters corresponding to all bits of an interface of the device, channel parameters corresponding to bits of a channel of the device, or per-bit parameters corresponding to a single data line of the device. The tuning can include identifying a number of reward values from a number of values for the parameter based on the reward function. Each reward value of the number of reward values can correspond to one of the number of values. Additionally, the tuning can include selecting a value of the number of values corresponding to a maximum reward value of the number of reward values and applying the selected value to the parameter. In one configuration, the parameter can be tuned in response to a health check of the device, a failure of the device, or a data eye analysis of the device.

[0075] The various operations of methods described above can be performed by any suitable entity or set of entities. These entities can include various hardware and / or software components and / or modules, including, but not limited to circuitry, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in the figures, those operations can have corresponding counterpart means-plus-function components with similar numbering.

[0076] As used, the term "determining" encompasses a wide variety of actions. For example, "determining" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure), ascertaining and the like. Additionally, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and the like. Furthermore, "determining" can include resolving, selecting, choosing, establishing and the like.

[0077] As used, the phrase "at least one of" a set including one or more items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c.

[0078] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any commercially available processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0079] The steps of a method or algorithm described in connection with the disclosure can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in any form of storage medium that is known in the art. Some examples of storage media that can be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, and so forth. A software module can comprise a single instruction, or many instructions, and can be distributed over several different code segments, in different programs, and across multiple storage media. A storage medium can be coupled to a processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The disclosure, in an embodiment, comprises a computer program product. The computer program product comprises a computer readable storage medium having stored thereon instructions that, when executed by a processor of a computer system, cause the computer system to perform steps or actions as described herein.

[0080] The disclosed methods comprise one or more steps or actions for accomplishing a described process. The method steps and / or actions can be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.

[0081] The functions described can be implemented in hardware, software, firmware or any combination thereof. If implemented in hardware, an example hardware configuration can include a processing system in a device. The processing system can be implemented with a bus architecture. The bus can include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus can link together various circuits such as a processor, machine-readable medium, and buses interface. The bus interface can be used to connect a network adapter to the processing system via the bus. The network adapter can be used to implement signal processing functionality. For certain aspects, a user interface (e.g., keypad, displays, mouse, joystick, etc.) can also be connected to the bus. The bus can also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art, and therefore, will not be further described.

[0082] The processor can be responsible for managing the bus and general processing, including the execution of software stored on the machine-readable medium. The processor can be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry that can execute software. Software shall be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Machine-readable media can include a non-transitory machine-readable medium, a machine-readable storage medium, a tangible machine-readable medium, a storage device, and / or a memory device. Examples of machine-readable media include random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs), magnetic disks, other optical and non-optical data storage devices, and / or any other machine-readable media suitable for storing the desired information and / or instructions and combinations thereof. Machine-readable media can be embodied in a computer program product.

[0083] In a hardware implementation, the machine-readable media can be part of the processing system separate from the processor. However, as those skilled in the art will readily appreciate, the machine-readable media, or any portion thereof, can be external to the processing system. As examples, the machine-readable media can include a transmission line, a carrier wave modulated by a data signal, and / or a computer product, all

[0084] The processing system can be configured as a general- purpose processing system with one or more microprocessors providing processor functionality and external memory providing at least a portion of the machine-readable media all linked together with other supporting circuitry through an external bus architecture. Alternatively, the processing system can include one or more neuromorphic processors for implementing the described models. As another alternative, the processing system can be implemented with an application specific integrated circuit (ASIC) with the processor, bus interface, user interface, supporting circuitry, and at least a portion of the machine-readable media integrated into a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry that can perform the various functionality described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality for the processing system, depending on the particular application and the overall design constraints imposed on the overall system.

[0085] The machine-readable media can comprise a number of software modules. The software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. The software modules can include a transmission module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. By way of example, a software module can be loaded into RAM from a hard drive when a triggering event occurs. During execution of the software module, the processor can load some of the instructions into cache to increase access speed. One or more cache lines can then be loaded into a general register file for execution by the processor. When referring to the functionality of a software module below, it will be understood that such functionality is implemented by the processor when executing instructions from that software module. Furthermore, it should be appreciated that aspects of the present disclosure result in improvements to the functioning of the processor, computer, machine, or other system implementing such aspects.

[0086] If implemented in software, the functions can be stored or transmitted as one or more instructions or codes on or through a computer-readable medium. Computer-readable media includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Additionally, any connection is also legitimately referred to as computer-readable media. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared (IR), radio, and microwave), then that coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. The disks and discs used include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs. Disks, where disks often magnetically reproduce data, and discs optically reproduce data using lasers. Therefore, in some aspects, computer-readable media may include non-transient computer-readable media (e.g., tangible media). Additionally, in other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0087] Therefore, some aspects may include a computer program product for performing the given operations. For example, such a computer program product may include a computer-readable medium on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the described operations. In some aspects, the computer program product may include packaging material.

[0088] Furthermore, it should be understood that modules and / or other suitable means for performing the described methods and techniques may be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the described methods. Alternatively, the various methods described can be provided via a storage device (e.g., RAM, ROM, physical storage media such as CDs or floppy disks) so that the device can acquire the various methods once the storage device is coupled to or provided to the user terminal and / or base station. In addition, any other suitable techniques suitable for providing the described methods and techniques to the device may be utilized.

[0089] It will be understood that the claims are not limited to the precise configuration and components illustrated above. Various modifications, changes, and adaptations will be apparent to others skilled in the art with the benefit of this disclosure.

Claims

1. A method performed by a machine learning system, comprising: A set of reward values ​​is generated based on a set of memory parameter values ​​of the device memory, which is selected by the machine learning system. Each reward value in the set of reward values ​​corresponds to a performance metric associated with a corresponding memory parameter value in the set of memory parameter values ​​that is programmed at the device memory. A reward function is determined based on the set of reward values ​​to maximize the reward for the set of memory parameters corresponding to the device memory; as well as The memory parameters in the memory parameter set are tuned based on the reward function.

2. The method of claim 1, further comprising: Randomly select a first memory parameter value from the set of memory parameter values; The device memory is programmed using the first memory parameter values; A first data eye is sampled in response to programming the device memory with the first memory parameter value; A first reward value is generated based on the first data eye to form the reward value set; as well as The reward function is determined based on the first reward value.

3. The method of claim 2, further comprising: The uncertainty of the reward function is determined to be greater than a threshold; The second memory parameter value in the set of memory parameter values ​​is selected based on the aforementioned uncertainty. The device memory is programmed using the second memory parameter value; The second data eye is sampled in response to programming the device memory with the second memory parameter value; A second reward value is generated based on the second data eye to form the reward value set; as well as The reward function is updated based on the second reward value.

4. The method of claim 2, wherein generating the first reward value comprises: Based on the first data eye, a reward criterion is determined for at least one of the data path stability or power consumption of the device memory. as well as The first reward value is generated based on the determined reward criteria.

5. The method of claim 1, wherein tuning the memory parameters includes: Based on the reward function, multiple reward values ​​are identified from multiple values ​​used for the memory parameters, each of the multiple reward values ​​corresponding to one of the multiple values; Select the value that corresponds to the maximum reward value among the plurality of reward values; as well as The selected value is applied to the memory parameter.

6. The method of claim 1, wherein the memory parameters include global parameters corresponding to all bits of the interface of the device memory, channel parameters corresponding to the bits of the channel of the device memory, or per-bit parameters corresponding to a single data line of the device memory.

7. The method of claim 1, wherein the device memory includes dynamic random access memory (DRAM).

8. The method of claim 1, further comprising tuning the memory parameters in response to a health check of the device memory, a failure of the device memory, or a data eye analysis of the device memory.

9. An apparatus for a machine learning system, comprising: At least one processor; Device memory coupled to the at least one processor; as well as Instructions, which are stored in the device memory and, when executed by the processor, are operable to cause the device to: A set of reward values ​​is generated based on a set of memory parameter values ​​of the device memory, which is selected by the machine learning system. Each reward value in the set of reward values ​​corresponds to a performance metric associated with a corresponding memory parameter value in the set of memory parameter values ​​that is programmed at the device memory. A reward function is determined based on the set of reward values ​​to maximize the reward for the set of memory parameters corresponding to the device memory; as well as The memory parameters in the memory parameter set are tuned based on the reward function.

10. The apparatus of claim 9, wherein the instructions further cause the apparatus to: Randomly select a first memory parameter value from the set of memory parameter values; The device memory is programmed using the first memory parameter values; A first data eye is sampled in response to programming the device memory with the first memory parameter value; A first reward value is generated based on the first data eye to form the reward value set; as well as The reward function is determined based on the first reward value.

11. The apparatus of claim 10, wherein the instructions further cause the apparatus to: The uncertainty of the reward function is determined to be greater than a threshold; The second memory parameter value in the set of memory parameter values ​​is selected based on the aforementioned uncertainty. The device memory is programmed using the second memory parameter value; The second data eye is sampled in response to programming the device memory with the second memory parameter value; A second reward value is generated based on the second data eye to form the reward value set; as well as The reward function is updated based on the second reward value.

12. The apparatus of claim 10, wherein the instructions cause the apparatus to generate the first reward value by: Based on the first data eye, a reward criterion is determined for at least one of the data path stability or power consumption of the device memory; and The first reward value is generated based on the determined reward criteria.

13. The apparatus of claim 9, wherein the instructions cause the apparatus to tune the memory parameters by: Based on the reward function, multiple reward values ​​are identified from multiple values ​​used for the memory parameters, each of the multiple reward values ​​corresponding to one of the multiple values; Select the value that corresponds to the maximum reward value among the plurality of reward values; as well as The selected value is applied to the memory parameter.

14. The apparatus of claim 9, wherein the memory parameters include global parameters corresponding to all bits of an interface of the device memory, channel parameters corresponding to bits of a channel of the device memory, or per-bit parameters corresponding to a single data line of the device memory.

15. The apparatus of claim 9, wherein the device memory includes dynamic random access memory (DRAM).

16. The apparatus of claim 9, wherein the instructions further cause the apparatus to tune the memory parameters in response to a health check of the device memory, a failure of the device memory, or a data eye analysis of the device memory.

17. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by at least one processor and comprising: Program code for generating a set of reward values ​​based on a set of memory parameter values ​​of the device memory, the set of memory parameter values ​​being selected by a machine learning system, wherein each reward value in the set of reward values ​​corresponds to a performance metric associated with a corresponding memory parameter value in the set of memory parameter values ​​that is programmed at the device memory; Program code for determining a reward function based on the set of reward values ​​to maximize the reward for the set of memory parameters corresponding to the device memory; as well as Program code for tuning memory parameters in the set of memory parameters based on the reward function.

18. The non-transient computer-readable medium of claim 17, wherein the program code further comprises: Program code for randomly selecting a first memory parameter value from the set of memory parameter values; Program code for programming the device memory with the first memory parameter values; Program code for sampling a first data eye in response to programming the device memory with the first memory parameter value; Program code for generating a first reward value for the reward value set based on the first data eye; as well as Program code for determining the reward function based on the first reward value.

19. The non-transient computer-readable medium of claim 18, wherein the program code further comprises: Program code used to determine that the uncertainty of the reward function is greater than a threshold; Program code for selecting a second memory parameter value from the set of memory parameter values ​​based on the uncertainty; Program code for programming the device memory with the second memory parameter value; Program code for sampling a second data eye in response to programming the device memory with the second memory parameter value; Program code for generating a second reward value for the reward value set based on the second data eye; as well as Program code for updating the reward function based on the second reward value.

20. The non-transient computer-readable medium of claim 18, wherein the program code for generating the first reward value comprises: Program code for determining a reward criterion based on the first data eye, either data path stability or power consumption of the device memory; as well as Program code used to generate the first reward value based on the determined reward criteria.

21. The non-transient computer-readable medium of claim 17, wherein the program code for tuning the memory parameters comprises: Program code for identifying multiple reward values ​​from multiple values ​​for the memory parameters based on the reward function, each of the multiple reward values ​​corresponding to one of the multiple values; Program code for selecting the value corresponding to the maximum reward value among the plurality of reward values; as well as Program code used to apply the selected value to the memory parameter.

22. The non-transient computer-readable medium of claim 17, wherein the memory parameters include global parameters corresponding to all bits of an interface of the device memory, channel parameters corresponding to bits of a channel of the device memory, or per-bit parameters corresponding to a single data line of the device memory.

23. The non-transient computer-readable medium of claim 17, wherein the device memory includes dynamic random access memory (DRAM).

24. The non-transient computer-readable medium of claim 17, wherein the program code further includes program code for tuning the memory parameters in response to a health check of the device memory, a failure of the device memory, or a data eye analysis of the device memory.

25. An apparatus for implementing a machine learning system, comprising: Apparatus for generating a set of reward values ​​based on a set of memory parameter values ​​of a device memory, the set of memory parameter values ​​being selected by the machine learning system, each reward value in the set of reward values ​​corresponding to a performance metric associated with a corresponding memory parameter value programmed at the device memory in the set of memory parameter values; A means for determining, based on the set of reward values, a reward function for maximizing the reward corresponding to a set of memory parameters of the device memory; as well as A means for tuning memory parameters in the set of memory parameters based on the reward function.

26. The apparatus of claim 25, further comprising: A means for randomly selecting a first memory parameter value from the set of memory parameter values; A means for programming the device memory with the first memory parameter values; A means for sampling a first data eye in response to programming the device memory with the first memory parameter value; A means for generating a first reward value for the reward value set based on the first data eye; as well as A means for determining the reward function based on the first reward value.

27. The apparatus of claim 26, further comprising: A means for determining that the uncertainty of the reward function is greater than a threshold; A means for selecting a second memory parameter value from the set of memory parameter values ​​based on the uncertainty; A means for programming the device memory with the second memory parameter value; A means for sampling a second data eye in response to programming the device memory with the second memory parameter value; A means for generating a second reward value for the reward value set based on the second data eye; as well as A means for updating the reward function based on the second reward value.

28. The apparatus of claim 26, wherein the means for generating the first reward value comprises: A means for determining a reward criterion based on the first data eye, either the stability of the data path including the device memory or the power consumption; as well as A means for generating the first reward value based on the determined reward criteria.

29. The apparatus of claim 25, wherein the means for tuning the memory parameters comprises: A means for identifying a plurality of reward values ​​from a plurality of values ​​for the memory parameters based on the reward function, each of the plurality of reward values ​​corresponding to one of the plurality of values; A means for selecting, from the plurality of values, the value corresponding to the maximum reward value among the plurality of reward values; as well as A means for applying the selected value to the memory parameter.

30. The device of claim 25, wherein the memory parameters include global parameters corresponding to all bits of an interface of the device memory, channel parameters corresponding to bits of a channel of the device memory, or per-bit parameters corresponding to a single data line of the device memory.

Citation Information

Patent Citations

  • Electric power communication network routing method based on deep reinforcement learning

    CN111010294A

  • Plant control supporting apparatus, plant control supporting method, plant control supporting program, and recording medium

    EP3428744A1