Fast tuning of reinforcement learning policies for improved evaluation of multiple design options

Fast-tuned reinforcement learning policy training techniques enhance CAD/AI generative processes by clustering architectures, transferring policies, and performing meta-learning to expedite the evaluation of design options, addressing inefficiencies in existing systems.

WO2025221348A1PCT designated stage Publication Date: 2025-10-23SIEMENS IND SOFTWARE NV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014899
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-02-07
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing CAD/AI generative processes for system/architecture design are slow and inefficient, hindering the evaluation of numerous design options and the identification of optimal or surprising designs due to time and resource constraints.

Method used

Implementing fast-tuned reinforcement learning (FT-RL) policy training techniques that cluster architectures, transfer policies across varying inputs and outputs, and perform serial or parallel meta-learning to generate final tuned policies for rapid evaluation of design options.

Benefits of technology

Enables faster evaluation of architectural designs, allowing quick identification of optimal or surprising designs by improving the speed and efficiency of CAD/AI generative processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014899_23102025_PF_FP_ABST
    Figure US2025014899_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the invention are directed to a computer-implemented method of evaluating a system-under-design (SUD). The computer-implemented method includes applying, using a processor system, first reinforcement learning (RL) operations to an electronic representation of a first design of the SUD to generate a first set of performance indicators. The first RL operations include a first final tuned RL policy. The first final tuned policy is based at least in part on multiple RL policies associated with multiple different designs of the SUD.
Need to check novelty before this filing date? Find Prior Art

Description

FAST TUNING OF REINFORCEMENT LEARNING POLICIES FOR IMPROVEDEVALUATION OF MULTIPLE DESIGN OPTIONSTECHNICAL FIELD

[0001] The present embodiments relate to computing systems, computer- implemented methods, and computer program products configured and arranged to perform fast tuning of reinforcement learning (RL) policies using unsupervised-learning- based architecture groupings of multiple design points, thereby improving the speed and efficiency at which a relatively large number of computer-generated design options can be evaluated.BACKGROUND

[0002] Computer-aided design (CAD) is the use of computer-based software to aid in design processes. CAD software can be used to create two-dimensional (2-D) drawings or three-dimensional (3-D) models of a system-under-design (SUD). CAD software generally includes a variety of tools that enable a designer to optimize and streamline workflow; increase productivity; improve the quality and level of detail in the design; improve documentation communications; and often contribute toward a manufacturing design database. CAD software outputs come in the form of electronic files, which can be used in tandem with computer-aided manufacturing (CAM) software to control manufacturing and / or fabrication processes. CAD / CAM is software routinely used to design variety of products such as electronic circuit boards in computers and other devices.

[0003] CAD software can include CAD simulation functionality that allows virtual experiments to be performed on a SUD model instead of a physical prototype of the SUD. Conventional CAD simulations determine electronically (or virtually) how relatively predictable stressors (e.g., high heat conditions, load stress, physical pressure,etc.) in a relatively simple and predictable virtual environment impact a SUD model. CAD simulation functionality can also include artificial intelligence (Al) functionality that generates the SUD model and further assists with the system design process. In its simplest form, Al is a field that combines computer science and robust datasets to enable problem-solving. Al also encompasses sub-fields of machine learning and deep learning. Machine learning and deep learning are implemented as neural networks having input layers, hidden layers and output layers. Machine learning neural networks differ from deep learning neural networks in that deep learning has more hidden layers than machine learning. Al systems can be implemented as Al algorithms that seek to create expert systems operable to perform a variety of tasks, such as predictions or classifications, based on input data.

[0004] CAD / AI simulation functionality can also include generative functionality that provides simulation, modeling and generative support to facilitate so-called “generative engineering” activities. Every design process begins with the creation of a functional architecture and system-level design. Traditionally, these early design processes have been an expert-centric engineering process. However, CAD / AI functionality can be configured and arranged to include generative functionality operable to generate and evaluate system / architecture design options during the early concept phase. For example, a CAD / AI / Generative software tool commercially available from Siemens AG under the tradename Simcenter™ (referred to herein as “Simcenter”) combines system simulation, optimal control methods, and reinforcement learning with a machine learning and scientific computing stack to generate, simulate and assess hundreds or even thousands of system / architecture design options based on the requirements. The result is a selection of system / architectural design options from which the engineer can choose. In most use cases, such CAD / AI / Generative functionality generates system / architecture design options faster and in larger numbers that could be done relying on heavy manual and expert-centric processes.

[0005] As the capability of the above-described C D / AI / Generative functionality to generate system / architecture design options for complex real-world systems continues to grow, there is a need to improve the speed and efficiency of these CAD / AI / Generative processes. Improving the speed and efficiency of CAD / AI / Generative processes would enable the evaluation of various system / architectural design options faster, which results in an improved ability to quickly identify “optimal” and / or or “surprising” system / architecture designs that satisfy design requirements / constraints. Improving the speed and efficiency of CAD / AI / Generative processes would further enable the ability of such processes to identify improved and sophisticated system / architecture designs that may have gone unrealized due to time and / or budget constraints associated computer- implemented processes that consume a large amount of computing resources.BRIEF SUMMARY

[0006] Embodiments of the invention are directed to a computer-implemented method of evaluating a system-under-design (SUD). The computer-implemented method includes applying, using a processor system, first reinforcement learning (RL) operations to an electronic representation of a first design of the SUD to generate a first set of performance indicators. The first RL operations include a first final tuned RL policy. The first final tuned RL policy is based at least in part on multiple RL policies associated with multiple different designs of the SUD.

[0007] In addition to any one or more of the features described herein, the computer- implemented method further includes applying, using the processor system, second reinforcement learning (RL) operations to an electronic representation of a second design of the SUD to generate a second set of performance indicators.

[0008] In addition to any one or more of the features described herein, the second RL operations include a second final tuned RL policy.

[0009] In addition to any one or more of the features described herein, the second final tuned RL policy is based at least in part on the multiple policies associated with the multiple different designs of the SUD.

[0010] In addition to any one or more of the features described herein, the computer- implemented method further includes comparing the first set of performance indicators with the second set of performance indicators.

[0011] In addition to any one or more of the features described herein, the computer- implemented method further includes selecting the first design of the SUD or the second design of the SUD based at least in part on a result of comparing the first set of performance indicators with the second set of performance indicators.

[0012] In addition to any one or more of the features described herein, the first final tuned RD policy is generated based at least in part on performing fast-tuning reinforcement-learning (FT-RD) policy generation operations.

[0013] In addition to any one or more of the features described herein, the FT-RD policy generation operations include clustering the multiple different designs of the SUD to generate SUD design clusters; generating initial polices based at least in part on the SUD design clusters; and generating the first final tuned RD policy based at least in part on tuning one of the initial policies.

[0014] In addition to any one or more of the features described herein, the FT-RD policy generation operations include clustering SUD responses of the multiple different designs of the SUD to generate SUD design response clusters; and transferring policies across different number of inputs and outputs of the multiple different designs of the SUD to identify which parts of NNs representing the multiple different designs of the SUD can be transferred.

[0015] In addition to any one or more of the features described herein, the FT-RL policy generation operations further include performing serial-type transfer learning or parallel-type meta-learning within and across the SUD design response clusters to generate a set of final tuned RL policies including the first final tuned RL policy.

[0016] Embodiments of the invention are also directed to computer systems and computer program products having substantially the same features as the computer- implemented method described above.

[0017] The present invention is defined by the following claims, and nothing in this section should be taken as a limitation on those claims. Further aspects and advantages of the invention are discussed below in conjunction with the disclosed embodiments and can be later claimed independently or in combination. Additional features and advantages are realized through techniques described herein. Other embodiments and aspects are described in detail herein. For a better understanding, refer to the description and to the drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The subject matter which is regarded as embodiments is particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other features and advantages of the embodiments are apparent from the following detailed description taken in conjunction with the accompanying drawings in which:

[0019] FIG. 1 depicts a simplified block diagram of a system design and its subelements that can be utilized in implementing embodiments of the invention;

[0020] FIG. 2 depicts a simplified block diagram of an example simulator and / or simulation system that can be utilized in implementing embodiments of the invention;

[0021] FIG. 3 depicts a simplified block diagram illustrating non-limiting examples of feasible system architecture designs that can be generated using the simulator shown in FIG. 2;

[0022] FIG. 4 depicts a simplified block diagram illustrating a non-limiting example of a reinforcement learning (RL) system that can be utilized in embodiments of the invention;

[0023] FIG. 5A depicts a simplified block diagram illustrating a model of a biological neuron operable to be utilized in neural network (NN) architectures in accordance with aspects of the invention;

[0024] FIG. 5B depicts a simplified block diagram illustrating a deep learning NN architecture in accordance with aspects of the invention;

[0025] FIG. 6 depicts a fast-tuning reinforcement-learning (FT-RL) policy generation system in accordance with embodiments of the invention;

[0026] FIG. 7 depicts another FT-RL policy generation system in accordance with embodiments of the invention;

[0027] FIG. 8 depicts a flow diagram illustrating a computer-implemented methodology in accordance embodiments of the invention;

[0028] FIG. 9 depicts a non-limiting example of the FT-RL policy generation system shown in FIGS. 6 and / or 7 combined with the simulator shown in FIG. 2 to generate a suite of final tuned policies for use in evaluating each proposed system architecture generated by the simulator shown in FIG.2;

[0029] FIG. 10 depicts a simplified block diagram illustrating a non-limiting example of a NN architecture operable to implement transfer learning between a System-A NN and a System-B NN in accordance with aspects of the invention;

[0030] FIG. 11 depicts a simplified block diagram illustrating a non-limiting example of a meta-learning technique that can be used to implement aspects of the invention;

[0031] FIG. 12 depicts a simplified block diagram illustrating another non-limiting example of an RL system that can be utilized in embodiments of the invention;

[0032] FIG. 13 depicts a machine learning system that can be utilized to implement aspects of the invention;

[0033] FIG. 14 depicts a learning phase that can be implemented by the machine learning system shown in FIG. 13; and

[0034] FIG. 15 depicts details of an exemplary computing system capable of implementing various aspects of the invention.

[0035] In the accompanying figures and following detailed description of the disclosed embodiments, the various elements illustrated in the figures are provided with three digit reference numbers. In some instances, the leftmost digits of each reference number corresponds to the figure in which its element is first illustrated.DETAILED DESCRIPTION

[0036] Embodiments of the invention provide fast policy training using reinforcement learning for generative engineering designs represented as a relatively large number of system architectures. In generative engineering, there are many components that can be connected in different ways resulting in different architectures in which the number and type of components (i.e., the “configuration) can differ. In a simple scenario, the same architecture can have different parameter values due to having the same components but different configuration parameters. This is a simpler case due to the simple number of inputs and outputs across configurations. Ideally, a more general case can be applied where each policy can be trained across the architecture. However,this more general case is difficult to achieve because the number of inputs and outputs differ.

[0037] Embodiments of the invention address the shortcomings in known techniques for leveraging transfer learning, meta-learning, and similar learning-acceleration techniques by providing novel fast-tuned RL policy training techniques that can be utilized where the number of inputs and outputs differ for the environments to which the RL policies will be applied. The fast-tuned RL policy training techniques disclosed herein result in faster training across architectures / configurations. The number of policies that can train quickly using embodiments of the invention allows the evaluation of various architectural design choices faster, which results in quick identification of “optimal” or “surprising” designs.

[0038] Lor example, if a team of engineers / designers is tasked with using CAD / AI / Generative software tools to design a hybrid electric vehicle (HEV), and multiple designs are generated where multiple electric motors, batteries and other components are be used and connected in different ways, it is not immediately clear what the performance of each design is in terms of speed performance and fuel consumption. Lor each design, reinforcement learning (RL) policies (one per architecture) are generated and tuned in accordance with aspects of the invention then used to evaluate each architecture across different key performance indicators (KPIs) to choose the best one using, for example, a Pareto-type analysis. Because the CAD / AI / Generative tool generate many possible architecture designs, any improvement on training the RL policies multiplies with the number of architectures / configurations. This translates into finding the better architecture faster.

[0039] In some embodiments of the invention, a fast-tuning reinforcement-learning (LT-RL) policy generation system is provided and used to cluster architectures; analyze the clustered architectures to generate initial policies; and perform tuning and / or fine- tuning on the initial policies to generate a final tuned RL policy for each architecture. Asimulator simulates the architectures using the associated final tuned RL policy to generate KP results for each architecture. The FT-RL policy generation system compares the KPI results to select one or more best or desired system architectures.

[0040] In some embodiments of the invention, an FT-RL policy generation system is provided and used to cluster architectures based on their system responses; transfer policies across different number of inputs and outputs of the architectures to identify which parts of NNs representing the architectures can be transferred; and perform serialtype transfer learning or parallel-type meta-learning within and across clusters to generate a final tuned RL policy for each architecture. A simulator simulates the architectures using the associated final tuned RL policy to generate KPI results for each architecture. The FT-RL policy generation system compares the KPI results to select one or more best or desired system architectures.

[0041] Turning now to a more detailed description of aspects of the invention, the terms “system design” and “system architecture” are related concepts, but they refer to different aspects of the overall processes used to design a system. As illustrated in FIG.1, a System Design 100 is the process of defining the System Design Sub-elements, which include the System Architecture 110, Components, Modules, Interfaces, and Data for a system (e.g., a hybrid electric vehicle (HEV)) to satisfy specified requirements. More specifically, System Design 100 involves translating the requirements of a system into a detailed design that can be implemented. System Design 100 focuses on the internal structure and behavior of the system, considering factors such as Performance, Scalability, Reliability, Maintainability, and Security. System Design 100 involves making decisions about the System Design Sub-elements, their interactions, and their relationships to external systems or modules.

[0042] As further illustrated in FIG. 1, the System Architecture 110 refers to the overall structure and organization of a system (e.g., an HEV), including its Components, their relationships, and the principles and guidelines governing their design andevolution. The System Architecture 110 provides a high-level view of the system, defining its major Components, their responsibilities, and how they interact with each other and with external systems. System architecture 110 is concerned with the system’s overall design philosophy, the allocation of functionality to different Components, and the coordination and integration of those Components.

[0043] In summary, the System design 100 is a more detailed and specific activity focused on designing the internal components and behavior of a system (e.g., an HEV), while the System Architecture 110 provides a broader perspective and focuses on the overall structure and organization of the system. System architecture 100 sets the foundation for System Design 100 by defining the high-level structure and principles that guide the design process.

[0044] FIG. 2 depicts a simplified block diagram of an example simulator 200 (and / or simulation system) that can be utilized in implementing embodiments of the invention. The simulator 200 can be any suite of computer-implemented CAD / CAM / AI technologies that can be used to aid a designer in creating a variety of systems, products, and / or devices within a relevant domain. Systems / products / devices that are moving through their design process are referred to herein as a system-under-design (SUD), a product-under-design (PUD), and / or a device-under-design (DUD). In this detailed description, operations described as being applied to a SUD apply equally to a PUD, a DUD, or any other thing that can designed and created. In embodiments of the invention, the simulator 200 can be implemented to include CAD simulation functionality that allows virtual experiments to be performed on a SUD model instead of a physical prototype of the SUD. Conventional CAD simulations determine electronically (or virtually) how relatively predictable stressors (e.g., high heat conditions, load stress, physical pressure, etc.) in a relatively simple and predictable virtual environment impact a SUD model.

[0045] CAD simulation functionality of the simulator 200 can also include Al functionality that generates the PUD model and further assists with the system design process. In its simplest form, Al is a field that combines computer science and robust datasets to enable problem-solving. Al also encompasses sub-fields of machine learning and deep learning. Machine learning and deep learning are implemented as neural networks (NNs) having input layers, hidden layers and output layers. Machine learning NNs differ from deep learning NNs in that deep learning has more hidden layers than machine learning. Al systems can be implemented as Al algorithms that seek to create expert systems operable to make predictions or classifications based on input data.

[0046] In embodiments of the invention, the CAD simulation functionality of the simulator 200 further includes Al-driven generative engineering functionality operable to further assist a designer (e.g., an engineer) with the various tasks associated with working on a SUD. As shown in FIG. 2, the generative engineering functionality of the simulator 200 is operable to receive a set of initial SUD constraints 210 associated with a particular application (e.g., Application A, which can be a new engine design for an HEV) and generate therefrom a relatively large number of feasible designs 220 for the SUD. Each of the feasible designs 220 can include a relatively large number of design choices 222, and each of the design choices 222 can include a relatively large number of tunable parameters 224. In some embodiments of the invention, the initial SUD constraints 210 are generated by the designer. In some embodiments of the invention, the initial SUD constraints 210 are also computer / AI generated based on minimal design goals generated by the designer. In some embodiments of the invention, the relatively large number of feasible designs 220 and their associated design choices 222 and tunable parameters 224 are in the form of electronic files configured and arranged to facilitate further electronic and computer simulated processing and analysis of the relatively large number of feasible designs 220 and their associated design choices 222 and tunable parameters 224.

[0047] In some embodiments of the invention, the CAD simulation functionality of the simulator 200 can include substantially the same features and functionality as a suite of software tools sold commercially under the tradename Simcenter™ Studio (referred to herein as “Simcenter”). Simcenter combines system simulation, optimal control methods, and reinforcement learning with a machine learning and scientific computing stack to simulate and assess a large number (e.g., hundreds or many thousands) of architectures based on the requirements (e.g., the initial SUD constraints 210). The result is a selection of architectural options (e.g., the large number of feasible designs 220, design choices 222, and tunable parameters 224) from which the designer can choose. In most use cases, Simcenter provides a much larger number of design options much faster than could feasibly be generated by relying on the individual expertise of engineers and known rules, which only allow the evaluation of a limited number of system architecture variants before choosing the “right” concept for time sensitive projects. Simcenter allows for the evaluation of the most important KPIs, which improves the reliability of business and engineering decisions.

[0048] Non-limiting examples of two (2) of the relatively large number of feasible designs 220 (shown in FIG. 2) are depicted in FIG. 3 as System Architecture-A and System Architecture-B. System Architecture-A represents one design option, generated by the simulator 200 (shown in FIG. 2), for satisfying the initial SUD constraints 210 (shown in FIG. 2); and System Architecture-B represents another design option, generated by the simulator 200, for satisfying the initial SUD constraints 210. In embodiments of the invention, the initial SUD constraints 210 can be at a sufficiently high level that many different designs can satisfy the initial SUD constraints 210. For example, the initial SUD constraints 210 can prompt the simulator 200 to generate design options for an HEV with a first constraint that sets a maximum size and voltage rating of the HEV battery, and with a second constraint that sets the mile per gallon (MPG) rating of the vehicle within a predetermined MPG range. There are many HEV system designs that could satisfy these constraints, and each HEV design generated by the simulator 200will have its own distinct configurations of the system design sub-elements (e.g., the System Design Sub-elements of the System Design 100 shown in FIG. 1). For ease of illustration and explanation, System Architecture-A and System Architecture-B only show the components (e.g., Component-A, Component-B, Component-C, Component-D, and Component-E) of the associated system architecture. However, it is understood that System Architecture-A and System Architecture-B generated by the Simulator 200 will each include any combination of the various System Design Sub-elements of the System Design 100.

[0049] Referring still to FIG. 3, Configuration- A of System Architecture-A includes a configuration of components that includes Component-A, Component-B, and Component-C. Component-A, Component-B, and Component-C are configured and arranged to process Inputs-A to the System Architecture-A and generate Outputs-A. As previously noted, Configuration-A of System Architecture-A represents one system design option for satisfying the initial set of SUD constraints 210 (shown in FIG. 2). The details of how Component-A, Component-B, and Component-C are connected to one another is one example of the design choices (e.g., design choices 222 shown in FIG. 2) associated with System Architecture-A; and the various settings and performance ratings of Component-A, Component-B, and Component-C are examples of the tunable parameters (e.g., tunable parameters 224 shown in FIG. 2) associated with the design choices for System Architecture-A. Similarly, Configuration-B of System Architecture-B includes a configuration of components that includes Component-A, Component-D, and Component-E. Component-A, Component-D, and Component-E are configured and arranged to process Inputs-B to the System Architecture-B and generate Outputs-B. As previously noted, Configuration-B of System Architecture-B represents another system design option for satisfying the initial set of SUD constraints 210. The details of how Component-A, Component-D, and Component-E are connected to one another is one example of the design choices (e.g., design choices 222) associated with System Architecture-B; and the various settings and performance ratings of Component-A,Component-D, and Component-E are exampling the tunable parameters (e.g., tunable parameters 224) associated with the design choices for System Architecture-B.

[0050] FIG. 3 also illustrates that Configuration-A of System Architecture-A includes generic features 302 and non-generic features 304. In general, the generic features 302 represent features of System Architecture-A that are common with other instances of the relatively large number of feasible design options 220 (shown in FIG. 2) generated by the simulator 200 (shown in FIG. 2); and the non-generic features 304 represent features of System Architecture-A that are unique to System Architecture-A and not shared with other instances of the relatively large number of feasible design options 220 generated by the simulator 200. Similarly, the generic features 312 represent features of System Architecture-B that are common with other instances of the relatively large number of feasible design options 220 generated by the simulator 200; and the non-generic features 314 represent features of System Architecture-B that are unique to System Architecture- B and not shared with other instances of the relatively large number of feasible design options 220 generated by the simulator 200.

[0051] FIG. 4 illustrates an example of how reinforcement learning can be used to compare and evaluate the relatively large number of feasible designs 220 (shown in FIG. 2). As shown in FIG. 4, a given instance of the relatively large number of feasible designs 220 can be represented as a simulation model 410, and the various design options of a given instance of the relatively large number of feasible designs 220, design choices 222 (shown in FIG. 2), and tunable parameters 224 (show in FIG. 2) can be represented as a decisions-to-be-made module 412.

[0052] The simulation model 410 is analyzed by a reinforcement learning (RL) “brain” 420. The RL brain 420 implements a RL machine learning technique in which a learning agent of the RL brain 420 learns what actions (Action 422) to take to maximize a numerical reward signal (e.g., State and Reward 424). The RL operations performed by the RL brain 420 are based on the idea of framing problems as a Markov decision process(e.g., decisions-to-be-made 412) where the agent of the RL brain 420 learns a control policy to always pick the best possible action (e.g., Action 422) for a given state (e.g., State and Reward 424) of the system (e.g., System Architecture- A shown in FIG. 3) represented by the simulation model 410. The agent of the RL brain 420 can be implemented using known RL algorithms, and the policy of the RL brain 420 can be implemented as one or more NNs (e.g., deep learning NN architecture (or model) 520 shown in FIG. 5B) each having input layers, hidden layers and output layers. Ideally, the RL system depicted in FIG. 4 is somewhat random and dynamic, making a reward-based learning approach superior in comparison to other traditional control theories. Successful application of NNs in conjunction with reinforcement learning (hence the name “Deep Reinforcement Learning”) enabled the analysis of complex scenarios (e.g., complex System Designs 100 shown in FIG. 1) that were previously deemed impossible. Deep reinforcement learning (DRL) follows this method, using a deep NN to represent the policy.

[0053] In general, RL systems use a large amount of “trial and error” episodes or interactions with environments to learn a good policy. The simulation model 410 essentially becomes the environment in which the RL agent of the RL brain 420 will learn. Because the simulation model 410 learns by a continuous process of receiving rewards (e.g., State and Reward 424) for every action (e.g., Action 422) taken, it can learn to respond to unforeseen environments. For systems having relatively high complexity level, it would be extremely difficult, if not impossible, to write a computer program that could effectively manage every possible combination of circumstances occurring in everyday scenarios. As a result, the system used to evaluate the simulation model 410 must be highly adaptive. RL techniques are especially useful in evaluating complex systems because they do not require lots of pre-existing knowledge or data to provide useful recommendations and solutions.

[0054] There are various training techniques that can be used to improve the speed and efficiency of the various training operations that need to be performed when generating NNs such as the NN elements used to implement policy of the RL brain 420. For example, a model’s ability to learn a new task can be accelerated through the transfer of knowledge from a related task that has already been learned. In such knowledge transfers, a base network is first trained on a base dataset and task, and then the learned features are repurposed or transferred to a second target network to be trained on a target dataset and task. This process will tend to work if the features are general, meaning suitable to both the base task and the target task. For example, referring back to FIGS. 3 and 4, the training operations required by the RL brain 420 and simulation model 410 can be improved if the generic features 302 shared by System Architecture- A and other instances of the relatively large number of feasible designs 220 can be used as a starting point then fine-tuned to incorporate the specific non-generic features (e.g., non-generic features 304) of the associated system architecture (e.g., System Architecture- A). However, known techniques for efficiently and effectively extracting and repurposing generic features of system models (e.g., generic features 302) are largely ineffective when the system models do not have sufficient overlap. For example, although System Architecture-A and System Architecture-B are both able to satisfy the initial set of SUD constraints 210, System Architecture-A and System Architecture-B sufficiently different in the configurations, connections, inputs and outputs that known techniques for efficiently and effectively extracting generic features of system models (e.g., generic features 302) in order to create a foundation model are largely ineffective.

[0055] Non-limiting examples of the above-described known techniques for efficiently and effectively extracting and repurposing generic features of system models (e.g., generic features 302) include so-called “transfer learning” techniques (e.g., transfer learning system 1000 shown in FIG. 10 and described in greater detail subsequently herein) and so-called “meta-learning” techniques (e.g., meta-learning system 1100 shown in FIG. 11 and described in greater detail subsequently herein). Transfer learning is amachine learning method where a machine learning model developed for a first task is reused as the starting point for a model on a second, different but related task. For example, in a deep learning application, pre-trained models are used as the starting point on a variety of computer vision and natural language processing tasks. Transfer learning leverages through reuse the vast knowledge, skills, computer, and time resources required to develop NN models. Transfer learning techniques have been developed that leverage training data from a different but related domain in an attempt to avoid the significant amount of time it takes to develop labeled training data for a given domain. The domain associated with the to-be-learned (TBL) task is referred to as the target domain (TD), and the domain of the different but related task is referred to as the source domain (SD). Transfer learning is possible because NNs tend to learn different patterns in the data. For example, in computer vision, initial layers learn low-level features such as lines, dots, and curves. Top layers learn high-level features built on top of low-level features. Most of the patterns, especially low-level ones, are common to many different computer image data sets. Instead of learning them every time from scratch, transfer learning uses existing parts of the network and reuse them with a bit of tweaking for other purposes. However, as previously noted, transfer learning techniques are largely ineffective unless there is sufficient overlap between the SD and the TD.

[0056] Meta-learning in machine learning refers to learning algorithms that learn from other learning algorithms. Most commonly, this means the use of machine learning algorithms that learn how to best combine the predictions from other machine learning algorithms in the field of ensemble learning. Meta-learning algorithms typically refer to ensemble learning algorithms like stacking that learn how to combine the predictions from ensemble members. Meta-learning also refers to algorithms that learn how to learn across a suite of related prediction tasks, referred to as multi-task learning.

[0057] Commonly, machine learning techniques attempt to find out what algorithms work best with the data that will be used. These algorithms learn from historical data toproduce models, and those models can be used later to predict outputs for a given task. Meta-learning algorithms do not use directly that kind of historic data, but they instead learn from the outputs of machine-learning models. This means that meta-learning algorithms require the presence of other models that have already been trained on data. For example, if the goal is to classify images, machine learning models take images as input and predict classes while meta-learning models take predictions of those machine learning models as input and based on that, predict classes of the images. In that sense, meta-learning occurs one level above machine learning. However, as previously noted, meta-learning techniques are largely ineffective unless there is sufficient overlap between the models and predictions that are used to transfer information / learning.

[0058] Embodiments of the invention address the shortcomings in known techniques for leveraging transfer learning, meta-learning, and similar techniques by providing novel fast-tuned RL policy training techniques that can be utilized where the number of inputs and outputs differ for the environments to which the RL policies will be applied. The fast-tuned RL policy training techniques disclosed herein result in faster training across architectures / configurations. The number of policies that can train quickly using embodiments of the invention allows the evaluation of various architectural design choices faster, which results in quick identification of “optimal” or “surprising” designs. The fast-tuned RL policy training techniques can be implemented in accordance with aspects of the invention as fast-tuned RL (FT-RL) policy training system 600, 600A, examples of which are shown in FIGS. 6, 7, 10, and 11. The FT-RL policy training systems 600, 600A shown in FIGS. 6, 7, 10, and 11 provide systematic techniques configured and arranged to cluster architectures based on their system responses; transfer RL policy across different number of inputs and outputs and identify which parts of NNs need to be transferred; and perform either transfer learning within clusters or meta- learning and across clusters to generate a suite of final tuned policies 650, 650A (shown in FIGS. 6, 7, and 9) for each of the relatively large number of feasible designs 220. The FT-RL policy training system 600, 600 A, in accordance with aspects of the invention,find reinforcement learning methodologies resulting in faster training across architectures / configurations. The number of policies (e.g., final tuned policies 650, 650A) that can train quickly using the FT-RL policy training system 600, 600A allows the evaluation of various architectural designs (e.g., designs 220), design choices (e.g., design choices 222), and tunable parameters (e.g., tunable parameters 224) faster which results in quick identification of “optimal” or “surprising” designs.

[0059] FIGS. 5A and 5B illustrate additional details of NNs that can be used in implementing aspects of the invention, including the specifically the final tuned policies 650, 650A (shown in FIGS. 6 and 7, respectively). NNs are a specific category of machines that can mimic human cognitive skills. In general, a NN is a network of artificial neurons or nodes inspired by the biological neural networks of the human brain. In FIG. 5 A, the biological neuron is modeled as a node 502 having a mathematical function, f(x), depicted by the equation shown in FIG. 5A. Node 502 receives electrical signals from inputs 512, 514, multiplies each input 512, 514 by the strength of its respective connection pathway 504, 506, takes a sum of the inputs, passes the sum through a function, f(x), and generates a result 516, which may be a final output or an input to another node, or both. In the present specification, an asterisk (*) is used to represent a multiplication. Weak input signals are multiplied by a very small connection strength number, so the impact of a weak input signal on the function is very low.Similarly, strong input signals are multiplied by a higher connection strength number, so the impact of a strong input signal on the function is larger. The function f(x) is a design choice, and a variety of functions can be used. A suitable design choice for f(x) is the hyperbolic tangent function, which takes the function of the previous sum and outputs a number between minus one and plus one.

[0060] FIG. 5B depicts a simplified example of a deep learning NN architecture (or model) 520. In general, NNs can be implemented as a set of algorithms running on a programmable computer (e.g., processor 1504 of the computing system 1500 shown inFIG. 15). In some instances, NNs are implemented on an electronic neuromorphic machine that attempts to create connections between processing elements that are substantially the functional equivalent of the synapse connections between brain neurons. In either implementation, NNs incorporate knowledge from a variety of disciplines, including neurophysiology, cognitive science / psychology, physics (statistical mechanics), control theory, computer science, artificial intelligence, statistics / mathematics, pattern recognition, computer vision, parallel processing and hardware (e.g., digital / analog / VLSI / optical). The basic function of a NN is to recognize patterns by interpreting sensory data through a kind of machine perception. Real-world data in its native form (e.g., images, sound, text, or time series data) is converted to a numerical form (e.g., a vector having magnitude and direction) that can be understood and manipulated by a computer. The NN can be “trained” by performing multiple iterations of learning-based analysis on the real-world data vectors until patterns (or relationships) contained in the real-world data vectors are uncovered and learned.

[0061] NNs use feature extraction techniques to reduce the number of resources required to describe a large set of data. The analysis on complex data can increase in difficulty as the number of variables involved increases. Analyzing a large number of variables generally requires a large amount of memory and computation power. Additionally, having a large number of variables can also cause a classification algorithm to over-fit to training samples and generalize poorly to new samples. Feature extraction is a general term for methods of constructing combinations of the variables in order to work around these problems while still describing the data with sufficient accuracy.

[0062] Although the patterns uncovered / learned by a NN can be used to perform a variety of tasks, two of the more common tasks are labeling (or classification) of real- world data and determining the similarity between segments of real-world data. Classification tasks often depend on the use of labeled datasets to train the NN to recognize the correlation between labels and data. This is known as supervised learning.Examples of classification tasks include identifying objects in images (e.g., stop signs, pedestrians, lane markers, etc.), recognizing gestures in video, detecting voices, detecting voices in audio, identifying particular speakers, transcribing speech into text, and the like. Similarity tasks apply similarity techniques and (optionally) confidence levels (CLs) to determine a numerical representation of the similarity between a pair of items.

[0063] Returning still to FIG. 5B, the simplified NN architecture / model 520 is organized as a weighted directed graph, where the artificial neurons are nodes (e.g., Nl- N13), and where weighted directed edges (i.e., directional arrows) connect the nodes. The NN architecture / model 520 is organized such that nodes Nl, N2, N3 are input layer nodes, nodes N4, N5, N6, N7 are first hidden layer nodes, nodes N8, N9, N10, Nl 1 are second hidden layer nodes, and nodes N12, N13 are output layer nodes. Having multiple hidden layers indicates that the NN architecture / model 520 is a deep learning NN architecture / model. Each node is connected to every node in the adjacent layer by connection pathways, which are depicted in FIG. 5B as directional arrows each having its own connection strength. For ease of illustration and explanation, one input layer, two hidden layers, and one output layer are shown in FIG. 5B. However, in practice, multiple input layers, multiple hidden layers, and multiple output layers can be provided. When multiple hidden layers are provided, the NN model 520 can perform unsupervised deeplearning for executing classification / similarity type tasks.

[0064] Similar to the functionality of a human brain, each input layer node Nl, N2, N3 of the NN 520 receives Inputs directly from a source (not shown) with no connection strength adjustments and no node summations. Each of the input layer nodes Nl, N2, N3 applies its own internal f(x). Each of the first hidden layer nodes N4, N5, N6, N7 receives its inputs from all input layer nodes Nl, N2, N3 according to the connection strengths associated with the relevant connection pathways. Thus, in the first hidden layer node N4, its function is a weighted sum of the functions applied at input layer nodes Nl, N2, N3, where the weight is the connection strength of the associated pathway intothe first hidden layer node N4. A similar connection strength multiplication and node summation is performed for the remaining first hidden layer nodes N5, N6, N7, the second hidden layer nodes N8, N9, N10, N11, and the output layer nodes N12, N13.

[0065] The NN architecture / model 520 can be implemented as a feedforward NN or a recurrent NN. A feedforward NN is characterized by the direction of the flow of information between its layers. In a feedforward NN, information flow is unidirectional, which means the information in the model flows in only one direction — forward — from the input nodes, through the hidden nodes (if any) and to the output nodes, without any cycles or loops. In contrast to recurrent NNs, which have a bi-directional information flow, feedforward NNs are trained using the backpropagation method.

[0066] FIG. 6 depicts an FT-RL policy training system 600 in accordance with embodiments of the invention. A cloud computing system 50 is in wired or wireless electronic communication with the FT-RL policy training system 600. The cloud computing system 50 can supplement, support or replace some or all of the functionality (in any combination) of the FT-RL policy training system 600. Additionally, some or all of the functionality of the FT-RL policy training system 600 can be implemented as a node of the cloud computing system 50.

[0067] In the FT-RL policy training system 600, the simulator 200 (also shown in FIG. 2) is utilized. The simulator 200 shown in FIG. 6 has substantially the same features and functionality as the simulator 200 shown in FIG. 2 and described in greater detail previously herein. In the system 600, the simulator 200 generates the relatively large number of feasible designs 220 (shown in FIG. 2), which are implemented in FIG. 6 to include system architectures represented by System Architecture- 1 through System Architecture-N, where N is a whole number. In embodiments of the invention, N can be any whole number. Based on the complexity of the SUD, N can be ten thousand (10,000) or more. Although the relatively large number of feasible designs 220 are provided in FIG. 6 as system architectures, embodiments of the invention can represent the relativelylarge number of feasible designs 220 at any level of the System Design 100 (shown in FIG. 1). An FT-RL policy training module 610 is operable to receive and analyze the system architectures represented by System Architecture- 1 through System Architecture- N in order to generate a suite of final tuned policies 650 that provide one policy per instance of the system architectures represented by System Architecture- 1 through System Architecture-N.

[0068] The operations performed by the FT-RL policy training module 610 include, in accordance with embodiments of the invention, architecture clustering 620, transfer policy generation 630, and per architecture policy tuning 640. The architecture clustering 620 can be implemented using known machine-learning-based clustering techniques to organize the system architectures represented by System Architecture- 1 through System Architecture-N into clusters. In general, clustering is the process of arranging a group of objects in such a manner that the objects in the same group (which is referred to as a cluster) are more similar to each other than to the objects in any other group. Because clustering is unsupervised machine learning, it does not require a labeled dataset. A result of the architecture clustering 620 is that architectures having the most features in common will end up in the same cluster, which facilitates the identification of the various instances of system architectures within that cluster that share generic features (e.g., generic features 302 shown in FIG. 3).

[0069] The generic features in a given cluster provide a basis for using the transfer policy module 630 to pre-train a suite of initial policies that can be used as starting point for generating the policies in the suite of final tuned policies 650 for the system architectures that share generic features (e.g., generic features 302 shown in FIG. 3). The transfer policy module 630 uses RL to, in effect, generate initial policies and transfer them from one architecture to another one (Block 806 shown in FIG. 8). When the transfer policy module 630 transfers policy from one architecture to another one, the first and last layer of the actor are changed, and the first layer of the critic is changed, therebyproviding them with a consistent number of input and outputs. In RL, actor-critic is a temporal difference (TD) version of a policy gradient. The actor is represented by one NN; and the critic is represented by another NN. The actor NN decides which action should be taken, and the critic NN informs the actor NN how well the action performed and how it should be adjusted. The learning of the actor NN can be based on a policy gradient approach. In comparison, the critic NN evaluates the action produced by the actor NN by computing the value function. If there are similar types of observations and actions, the channel weights are replicated across different inputs and outputs. Once the actor NN and critic NN are initialized, a short training is performed to determine which parameters are most affected. In general, NN parameters are weights learned by the model during training. The rest of the parameters having less impact are fixed, and the chosen parameters are trained. At a last stage of the policy transfer operations, all parameters are tuned.

[0070] The process (e.g., changing the first and last layers of the actor, and changing the first layer of the critic, thereby making the actor / critic have a consistent number of input and outputs) can be performed in accordance with embodiments of the invention in the following manner. RL policies (actors and critics of each architecture) are functions with tunable parameters, and in general RL policies are neural networks. The input dimension of the actor is equal to the number of measurements of the architecture. The output of actor is equal to the number of actuators of the architecture. Similarly, the input of the critic is equal to the number of measurements and actuators of the architecture. The transfer policy module 630 requires transferring actor and critic across different architectures. However, the number of measurements and actuators could be different across architectures. Therefore, for the actor, the first and last layers of the actor and the first layer of the critic may need to be changed to make it consistent with the number of measurements and actuators of the new architecture. Without this change, RL policies cannot be transferred because the inputs and output would not be compatible. The reason the in-between layers are maintained the same and transferred to the next RL policy is tokeep the useful information from the previous architecture and transfer this useful information to the new architecture without having to learn it from scratch and by fine- tuning.

[0071] As a non-limiting example, consider a speed tracking problem for a car where the car’s speed is kept as close as possible to the desired speed profile. Our measurements and actuators in Architecture- 1 (e.g., Architecture- 1 shown in FIG. 6) could be position (x) and speed measurements (v) and diesel engine torque actuator (Ta). Similarly, in Architecture-2 (e.g., Architecture-2 shown in FIG. 6), we may only measure speed (v), not measure the position, and use the diesel (Ta) and electric engine torques as actuators (Te). As shown by this example, the number of measurements in Architecture- 1 is two (2), i.e., inputs (x,v), and the number of actuators in Architecture- 1 is one (1), i.e., output (Ta). In Architecture-2, the number of measurements is one (1), i.e., input (v) and the number of actuators is two (2), i.e., outputs (Ta,Te). Thus, we cannot directly transfer an actor policy of Architecture- 1 whose number of inputs is two (2), i.e., (x,v), and whose number of outputs is one (1), i.e., (Ta) to the actor policy of Architecture-2 whose number of inputs is one (1), i.e., (x) and whose number of outputs is two (2), i.e., (Ta,Te).Therefore, input and output layers (i.e., first and last layers of the actor) have to be changed when transferring the policy. For example, the first layer of the actor in Architecture- 1 could be a multi-layer-perceptron (MLP) layer with an input size of two (2) (number of measurements, equivalently inputs) and an output size of sixteen (16) hidden dimension, i.e., MLP(2,16). A MLP is a name for a modern feedforward artificial neural network that includes fully connected neurons with a nonlinear kind of activation function, organized in at least three layers, notable for being able to distinguish data that is not linearly separable. We cannot use the MLP layer with input size two (2) in the first layer of actor in Architecture-2 because the number of actuators, equivalently the number of outputs, is one (1). Therefore, the first layer in the actor of Architecture- 1, MLP(2,16) has to be changed to MLP(1,16) in the first layer in the actor of Architecture-2. A similaranalogy can be made for the last layer of the actors of Architecture- 1 and Architecture-2 based on the number of measurements, equivalently the outputs of the actors.

[0072] Similarly, the critic of Architecture- 1 could have three (3) inputs (two measurements + one actuator) and one output and the critic of Architecture-2 can have three (3) inputs (one measurement + two actuators) and one (1) output. Although, they are input-output wise compatible, the physical meaning of the inputs are different for critic and it would make sense to change the input layer for critic as well. Based on the example set forth in the immediately preceding paragraph, the inputs of the critic of Architecture- 1 are two measurements (x,v) and one actuator (Ta) and the corresponding first layer for the critic will be MLP(3,16) where we assume sixteen (16) is the hidden dimension size. When we transfer the critic to the Architecture-2, the first layer’s inputs of the critic of Architecture-2 are the measurement (x) and two actuators (Ta,Te). Although, the first layer of the critic of Architecture- 1 can be transferred to that of Architecture-2, doing so will have an adverse impact because the inputs to the first layer of the critic of Architecture- 1 are (x,v,Ta) and the inputs to the first layer of Architecture- 2 are (x,Ta,Te). Because they are physically different variables, it makes sense to reinitialize the first layer and not transfer from Architecture- 1 as is.

[0073] Each of the initial policies generated using the operations at the transfer policy generation module 630 is associated with one or more of the system architectures represented by System Architecture- 1 through System Architecture-N. For example, the architecture clustering module 620 can identify sufficient overlap between System Architecture- 1 and System Architecture-4 that the transfer policy generation module 630 generates an initial Policy- A that can be used as a starting point for generating (e.g., through using per architecture tuning module 640) a final tuned policy (e.g., Final Tuned Policy- 1 shown in FIG. 9) associated with System Architecture- 1. Similarly, based on the sufficient overlap between System Architecture- 1 and System Architecture-4, Policy- A can also be used as a starting point for generating (e.g., through using per architecturetuning module 640) a final tuned policy (e.g., Final Tuned Policy-4 shown in FIG. 9) associated with System Architecture-4. The per architecture policy tuning module 640 is operable to applying tuning and / or fine-tuning training operations to an initial policy using non-generic features (e.g., non-generic features 304 shown in FIG. 3) of the system architecture associated with initial policy. Fine-tuning is a technique for adapting a pretrained machine learning model (e.g., an initial policy generated by the transfer policy generation module 630) to new data (e.g., non-generic features 304 shown in FIG. 3) or tasks. Fine-tuning adjusts a model’s internal weights to bias it towards new data, without overwriting everything it has learned. This allows the model to retain its general skills while gaining new specialized skills.

[0074] Accordingly, the FT-RL policy training system 600 generates and uses RL policies that are “fast-tuned” in a manner that improves the speed and efficiency of the CAD / AI / Generative processes used in the simulator 200. Improving the speed and efficiency of CAD / AI / Generative processes used in the simulator 200 enables the evaluation of the large number of feasible designs 220 faster, which results in an improved ability to quickly identify “optimal” and / or or “surprising” system / architecture designs that satisfy the initial SUD constraints 210. Improving the speed and efficiency of CAD / AI / Generative processes used in the simulator 200 further enable the ability of such processes to identify improved and sophisticated system / architecture designs that may have gone unrealized due to time and / or budget constraints associated computer- implemented processes that consume a large amount of computing resources.

[0075] FIG. 7 depicts another FT-RL policy training system 600A in accordance with embodiments of the invention. The FT-RL policy training system 600A has substantially the same features and functionality as the FT-RL policy training system 600 (shown in FIG. 6) except the FT-RL policy training system 600A provides additional details of how clustering, initial policy generation, and policy tuning can be performed in accordance with aspects of the invention. FIG. 8 depicts a computer-implemented methodology 800in accordance with aspects of the invention; and FIG. 9 depicts a non-limiting example of the FT-RL policy training system 600, 600A shown in FIGS. 6 and / or 7 combined with the simulator 200 shown in FIG. 2 to generate the suite of final tuned policies 650 for use in evaluating each proposed system architecture (e.g., system architectures represented by System Architecture- 1 through System Architecture-N) generated by the simulator 200 shown in FIG.2. FIG. 9 further illustrates the final tuned policies 650, 650A, where each of the final tuned policies 650, 650A is represented by Final Tuned Policy-1 through Final Tuned Policy-N, where each of the final tuned policies represented by Final Tuned Policy- 1 through Final Tuned Policy-N is associated with one of the system architectures represented by System Architecture- 1 through System Architecture-N. The computer- implemented methodology 800 can be performed by the FT-RL policy training systems 600, 600A shown in FIGS. 6 and 7, as well as by the configuration of systems shown in FIG. 9. The following description of the operations of the FT-RL policy training system 600A will be provided with reference to systems and computer-implemented methods shown in FIGS. 7, 8, and 9.

[0076] Referring initially to FIG. 7, the FT-RL policy training system 600A includes a suite of simulated architecture responses 710 and an FT-RL policy training module 610A configured and arranged to generate a suite of final tuned policies 650A (one policy per architecture). A cloud computing system 50 is in wired or wireless electronic communication with the FT-RL policy training system 600A. The cloud computing system 50 can supplement, support or replace some or all of the functionality (in any combination) of the FT-RL policy training system 600 A. Additionally, some or all of the functionality of the FT-RL policy training system 600A can be implemented as a node of the cloud computing system 50.

[0077] The FT-RL policy training module 610A includes an architecture clustering segment 620A and an architecture tuning segment 640 A. The architecture tuning segment 640 A includes a transfer policy module 630 A, a transfer learning module 742,and a meta-learning module 744, configured and arranged as shown. In the system 600A, the simulator 200 shown in FIG. 2 is utilized to generate the simulated architecture responses 710 (Block 802 shown in FIG. 8; simulator 200 shown in FIG. 9) and provide the same to the architecture clustering module 620A, which performs clustering operations (Block 804 shown in FIG. 8) on the simulated architecture responses 710. In some embodiments of the invention, the clustering operations include, for each architecture, injecting a time and discrete signals to identify the time and frequency domain based dynamical properties. Based on these properties, clustering approaches such as K-means++ or hierarchical clustering can be used to determine dynamically close architectures for clustering and the minimal group number.

[0078] The transfer policy module 630 A uses RL to, in effect, generate initial policies and transfer them from one architecture to another one (Block 806 shown in FIG. 8). When the transfer policy module 630A transfers policy from one architecture to another one, the first and last layer of the actor are changed, and the first layer of the critic is changed, thereby making it consistent with the number of input and outputs. In RL, actor-critic is a temporal difference (TD) version of policy gradient. The actor is represented by one NN; and the critic is represented by another NN. The actor NN decides which action should be taken, and the critic NN informs the actor NN how well the action performed and how it should be adjusted. The learning of the actor NN is based on a policy gradient approach. In comparison, the critic NN evaluates the action produced by the actor NN by computing the value function. If there are similar type of observations and actions, the channel weights are replicated across different inputs and outputs. Once the actor NN and critic NN are initialized, a short training is performed to determine which parameters are most affected. In general, NN parameters are weights learned by the model during training. The rest of the parameters having less impact are fixed, and the chosen parameters are trained. At a last stage of the policy transfer operations, all parameters are tuned.

[0079] The FT-RL policy training system 600A tunes and / or fine-tunes the initial policies generated by the transfer policy module 630 A to generate the fine-tuned policies 650A. In accordance with embodiments of the invention, any set of suitable NN tuning operations can be utilized. In accordance with some embodiments of the invention, the NN tuning operations can include serial NN tuning operations applied per architecture, along with parallel NN tuning operations applied per cluster. In the FT-RL policy training system 600A depicted in FIG. 7, the serial NN tuning operations applied per architecture are implemented as a transfer learning module 742; and the parallel tuning operations are implemented as a meta-learning module 744. In accordance with embodiments of the invention, an evaluation (Decision Block 808 shown in FIG. 8) is performed on the initial policies generated by the transfer policy 630A to determine whether to tune the initial policy using the transfer learning module 742 (Block 810 shown in FIG. 12) or the metalearning module 744 (Block 812 shown in FIG. 12). The evaluation to determine whether to tune the initial policy using the transfer learning module 742 or the metalearning module 744 is based at least in part on, for example, a design choice rule. A rule can be created that defines a distance metric for architectures within the cluster. If this metric is small (i.e., architectures within the cluster are close to each other, transfer learning can be performed and tuned sequentially (even parallel if computation is available). Because the architectures are close, there is not much need to incur the overhead of meta-learning. However, if architectures are not close, each sequential training may take longer time. In this case, it may be beneficial to train a meta-policy resulting in somewhat good performance within the cluster, and each architecture can be fine-tuned individually using the meta-policy.

[0080] The transfer learning module 742 can be implemented by, for each architecture, initializing the actor NN and critic NN according to transfer policy rules (e.g., transfer policy 630, 630A). The initial actor NN and critic NN are chosen with respect to the cluster distance metrics within or across clusters. This initialized actor NN and critic NN are trained as described in the transfer policy module 630A.

[0081] FIG. 10 depicts a simplified block diagram illustrating a non-limiting example of a NN architecture of a transfer learning system 1000 operable to implement transfer learning operations performed by the transfer learning module 742 and portions of the transfer policy module 630A. More specifically, the transfer learning system 1000 illustrates transfer learning between a Source NN-A representing System Architecture-A (shown in FIG. 3) and a Target NN-B representing a system architecture that shares the generic features 302 of System Architecture-B. Transfer learning is a machine learning method where a machine learning model developed for a first task is reused as the starting point for a model on a second, different but related task. For example, in a deep learning application, pre-trained models are used as the starting point on a variety of computer vision and natural language processing tasks. Transfer learning leverages through reuse the vast knowledge, skills, computer, and time resources required to develop NN models. Transfer learning techniques have been developed that leverage training data from a different but related domain in an attempt to avoid the significant amount of time it takes to develop labeled training data for a given domain. The domain associated with the to-be-learned (TBL) task is referred to as the target domain (TD) (represented by the Source NN-A shown in FIG. 10, which represents System Architecture-A shown in FIG. 3), and the domain of the different but related task is referred to as the source domain (SD) (represented by the Target NN-B shown in FIG.10, which represents a system architecture that shares the generic features 302 of System Architecture-A shown in FIG. 3). Transfer learning is possible because NNs tend to learn different patterns in the data. For example, in computer vision, initial layers learn low- level features such as lines, dots, and curves. Top layers learn high-level features built on top of low-level features. Most of the patterns, especially low-level ones, are common to many different computer image data sets. Instead of learning them every time from scratch, transfer learning uses existing parts of the network and reuse them with a bit of tweaking for other purposes. However, as previously noted, transfer learning techniques are largely ineffective unless there is sufficient overlap between the SD and the TD.

[0082] The meta learning module 744 (shown in FIG. 7) can be implemented by, within each architecture cluster generated by the architecture clustering module 620A, training a meta-policy for the given cluster. This meta policy is fine-tuned for each architecture within the cluster. When tuning is finished for one cluster, the next closest cluster is tuned with respect to the cluster distance metrics. The meta-policy for the next cluster is initialized with the previously closest cluster’s meta-policy. Embodiments of the invention use a meta-learning approach that interacts with multiple architectures within the same cluster and tunes one meta policy such as using so-called “first-order meta-learning algorithms.” The underlying principle is to define a cost function aggregated over the performance of multiple architectures and learn the policy improving all of them and in result learning features important to all of architectures. The key aspect is to tune one policy to find a good initial policy for the next fine tuning step. When measurements and actions of architectures are different, separate first and last layers are needed with shared in-between layers. Once a meta-policy is learned, each architecture’s policy is initialized with the meta-policy and fine-tuned separately to learn additional features specific to the architecture considered.

[0083] FIG. 11 depicts a simplified block diagram illustrating a non-limiting example of a meta-learning system 1100 that can be used to implement the meta learning module 744. The Meta Model (MM) collects measurements, i.e., mi, ... ,mn, from architectures within the same cluster, i.e., Architecture- 1 , ... , Architecture -n and applies actions to architectures by computing outputs, ai, ... , an, of the neural network function, MM, for given inputs, mi, ... , mn. Besides this input-output mapping, MM also collects measurements and actions and optimizes the meta-model neural network, MM, by minimizing the aggregated cost function defined over the architectures. This results in a meta-model policy, MM, giving good results for all architectures within the cluster. Each architecture box (Architecture- 1 through Architecture-n) represents the system within a cluster and can be thought as a model of architecture. For given actions, a;, its output (measurements) can be computed by a simulation or the real hardware. Thus, for theMeta-Learning Phase structure, each architecture box simulates or computes the measurements given actions, and the MM box computes the corresponding actions using its MM policy NN. MM minimizes the aggregated cost function over architectures with the parameters of meta-model policy.

[0084] Referring still to FIG. 11 , in the Per Architecture Tuning Phase after MetaLearning, this diagram is a standard Reinforcement Learning (RL) workflow. The only difference is that the RL policy is initialized by the MM policy as indicated on the MM box. In the RL Policy initialized with MM, for given actions, each “RL Policy initialized with MM” computes the actions. By collecting measurements and actions, it minimizes the cost function for its specific architecture (not the aggregated one). In the Architecture Box (Architecture- 1 through Architecture-n), each box represents the system within a cluster and can be thought as a model of architecture. For given actions, ai, its output (measurements) can be computed by a simulation or the real hardware. Thus, for the Per Architecture Tuning Phase after Meta-Learning Structure, each architecture box simulates or computes the measurements given actions, and the “RL Policy initialized with MM” box computes the corresponding actions using its RL policy. The RL Policy box minimizes the its cost function for its specific architecture with the parameters of RL policy.

[0085] The fine-tuned policies output from the transfer learning module 742 and the meta-learning module 744 are combined to generate the final tuned policies 650A (Block 814 shown in FIG. 8). At this stage of the operation of the simulator 200 and the FT-RL policy training system 600, 600A of FIG. 9, one final tuned policy (Final Tuned Policy-1 through Final Tuned Policy-N) is provided for each system architecture (System Architecture- 1 through System Architecture-N). At block 816 shown in FIG. 8, the simulator 200 again simulates the system architectures (System Architecture- 1 through System Architecture-N) using the associated final tuned policy (Final Tuned Policy- 1 through Final Tuned Policy-N) to generate KPI results for each system architecture(System Architecture- 1 through System Architecture-N). At block 818 shown in FIG. 8, the FT-RL policy training system 600, 600A (e.g., using the simulator 200) compares the KPI results generated at block 816 to select one or more best or desired system architectures (System Architecture- 1 through System Architecture-N shown in FIGS. 6 and 9). In embodiments of the invention, the simulator 200 can include functionality perform automated analysis of the KPI results. For example, given a tuned final RL policy, the simulator 200 can compute KPIs and plots the Pareto front with respect to all KPIs and allow the designer to choose, or the simulator 200 can choose itself provided that additional preferences are defined.

[0086] FIG. 12 depicts a combined block / flow diagram of a reinforcement learning system 1200 that can be used to implement aspects of the invention. The reinforcement learning system 1200 includes an agent 1210 configured and arranged to interact with an environment 1220. The agent 1210 includes a policy 1212 (i.e., a mapping from states to (preferably optimal) actions) and a reinforcement algorithm 1214. The agent 1210 receives observations (or states) 1208 and reward signals 1204 as inputs signals. The observation 1208 indicate the current state of the environment 1220, while the reward input signal 1204 indicates a reward associated with a prior action of the agent 1210 (e.g., for an immediately preceding action 1202). Based on the observations / states 1208 and the reward signals 1204, the agent 1210 chooses an action 1202 (location of the displayed region 126), which is applied to the environment 1220. Responsive to the action 1202, a new observation / state 1208 and reward 1204 for the environment 1220 are determined. The reinforcement learning algorithm 1214 of the agent 1210 seeks to learn values of observations / states 1208 (or state histories) and tries to maximize utility of the outcomes. The values of observations / states 1208 can be defined by the reward function described by the boxed text shown in FIG. 12. Thus, the reinforcement learning algorithm 1214 constructs a model of the environment 1220, wherein the model attempts to learn what operations are possible in each observation / state 1208, and what observation / state 1208will result from performing an operation in a given state. The cycle repeats as the inputs 1208, 1204 are continuously provided to the agent 1210.

[0087] The observations / states 1208 can be defined as a signal conveying to the agent 1210 some sense of “how the environment is” at a particular time. The observations / states 1208 can be whatever information is available to the agent 1210 about the environment 1220. The observation / state signal 1208 can be produced by any suitable preprocessing system (including sensors and sensor analysis circuitry) capable of evaluating the state of the environment 1220.

[0088] The policy 1212 defines how the learning agent 1210 behaves at a given time. Roughly speaking, the policy 1212 is a mapping from perceived observations / states 1208 of the environment 1220 to the actions 1202 to be taken when in those states. In some cases, the policy 1212 can be a simple function or lookup table, whereas in other cases it can involve extensive computation such as a search process. The policy 1212 is the core of a reinforcement learning agent 1210 in the sense that it alone is sufficient to determine behavior. In general, the policy 1212 can be stochastic.

[0089] The reward signal 1204 defines the goal in the reinforcement learning problem. On each time step, the environment 1220 sends to the reinforcement learning agent 1210 the reward signal 1204. The objective of the agent 1210 is to maximize the total reward 1204 it receives over the long run. The reward signal 1204 thus defines what are good and bad events for the agent 1210. The reward signal 1204 is the primary basis for altering the policy 1212. If an action 1202 selected by the policy 1212 is followed by low reward 1204, the policy 1212 may be changed to select some other action 1202 in that situation in the future. In general, the reward signal 1204 can be stochastic functions of the state of the environment 1220 and the actions 1202 that were taken.

[0090] Whereas the reward signal 1204 indicates what is good in an immediate sense, the reward function (shown in the text box at the bottom of FIG. 12) specifies what isgood (or valuable) in the long run. The reward function provides an estimate the value of associated with the inputs 1206, 1204 and is executed by the reinforcement learning algorithm 1214 of the agent 1210. Roughly speaking, the value of an observation / state 1208 is the total amount of reward 1204 the agent 1210 can expect to accumulate over the future, starting from that observation / state 1208. Accordingly, the reward signal 1204 determines the immediate, intrinsic desirability of environmental observations / states 1208, values (defined by the reward function) indicate the long-term desirability of observations / states 1208 after taking into account the observations / states that are likely to follow, and the rewards 1204 available in those observations / states 1208. For example, an observation / state 1208 might always yield a low immediate reward 1204 but still have a high value because it is regularly followed by other observations / states 1208 that yield high rewards 1204. The actions 1202 are chosen based on value judgments. The agent 1210 seeks actions 1202 that bring about observations / states 1208 of highest value (as measured by the reward function) not the highest reward signal 1204 because such actions 1202 obtain the greatest amount of reward signals 1204 over the long run. The reward signals 1204 are essentially given directly by the environment 1220, but values (as defined by the reward function) must be estimated and re-estimated from the sequences of observations / states 1208 the agent 1210 makes over its entire lifetime.

[0091] Reinforcement learning performed by the system 1200 is different from supervised machine learning. Supervised machine learning is learning from a training set of labeled examples provided by a knowledgeable external supervisor. Each example is a description of a situation together with a specification — the label — of the correct action the system should take to that situation, which is often to identify a category to which the situation belongs. The object of this kind of machine learning is for the system to extrapolate, or generalize, its responses so that it acts correctly in situations not present in the training set. Supervised machine learning alone is not adequate for learning from interaction. In interactive problems it is often impractical to obtain examples of desired behavior that are both correct and representative of all the situations in which the agenthas to act. In uncharted territory, which is where learning is expected to be most beneficial, an agent must be able to learn from its own experience.

[0092] Reinforcement learning is also different from unsupervised machine learning, which is typically attempting to find structure hidden in collections of unlabeled data. The terms supervised machine learning and unsupervised machine learning would seem to exhaustively classify machine learning paradigms, but they do not. Although both reinforcement learning and unsupervised learning do not rely on examples of correct behavior, they differ in that reinforcement learning is attempting to maximize a reward signal instead of attempting to find hidden structure. Uncovering structure in an agent’s experience can be useful in reinforcement learning but by itself does not address the reinforcement learning problem of maximizing a reward signal. Accordingly, it is appropriate to consider reinforcement learning to be a third machine learning paradigm, alongside supervised learning and unsupervised learning.

[0093] An example of machine learning techniques that can be used to implement aspects of the invention (e.g., the clustering operations 620, 620A shown in FIGS. 6 and 7) will be described with reference to FIGS. 13 and 14. Machine learning models configured and arranged according to embodiments of the invention will be described with reference to FIG. 13. Detailed descriptions of an example computing system and network architecture capable of implementing one or more of the embodiments of the invention described herein will be provided with reference to FIG. 15.

[0094] FIG. 13 depicts a block diagram showing a machine learning or classifier system 1300 capable of implementing various aspects of the invention described herein. More specifically, the functionality of the system 1300 is used in embodiments of the invention to generate various models and sub-models that can be used to implement computer functionality in embodiments of the invention. The system 1300 includes multiple data sources 1302 in communication through a network 1304 with a classifier 1310. In some aspects of the invention, the data sources 1302 can bypass the network1304 and feed directly into the classifier 1310. The data sources 1302 provide data / information inputs that will be evaluated by the classifier 1310 in accordance with embodiments of the invention. The data sources 1302 also provide data / information inputs that can be used by the classifier 1310 to train and / or update model(s) 1316 created by the classifier 1310. The data sources 1302 can be implemented as a wide variety of data sources, including but not limited to, sensors configured to gather real time data, data repositories (including training data repositories), and outputs from other classifiers. The network 1304 can be any type of communications network, including but not limited to local networks, wide area networks, private networks, the Internet, and the like.

[0095] The classifier 1310 can be implemented as algorithms executed by a programmable computer such as a processing system 1500 (shown in FIG. 15). As shown in FIG. 13, the classifier 1310 includes a suite of machine learning (ML) algorithms 1312; natural language processing (NLP) algorithms 1314; and model(s) 1316 that are relationship (or prediction) algorithms generated (or learned) by the ML algorithms 1312. The algorithms 1312, 1314, 1316 of the classifier 1310 are depicted separately for ease of illustration and explanation. In embodiments of the invention, the functions performed by the various algorithms 1312, 1314, 1316 of the classifier 1310 can be distributed differently than shown. For example, where the classifier 1310 is configured to perform an overall task having sub-tasks, the suite of ML algorithms 1312 can be segmented such that a portion of the ML algorithms 1312 executes each sub-task and a portion of the ML algorithms 1312 executes the overall task. Additionally, in some embodiments of the invention, the NLP algorithms 1314 can be integrated within the ML algorithms 1312.

[0096] The NLP algorithms 1314 include speech recognition functionality that allows the classifier 1310, and more specifically the ML algorithms 1312, to receive natural language data (text and audio) and apply elements of language processing, information retrieval, and machine learning to derive meaning from the natural language inputs andpotentially take action based on the derived meaning. The NLP algorithms 1314 used in accordance with aspects of the invention can also include speech synthesis functionality that allows the classifier 1310 to translate the result(s) 1320 into natural language (text and audio) to communicate aspects of the result(s) 1320 as natural language communications.

[0097] The NLP and ML algorithms 1314, 1312 receive and evaluate input data (i.e., training data and data -under-analysis) from the data sources 1302. The ML algorithms 1312 includes functionality that is necessary to interpret and utilize the input data’s format. For example, where the data sources 1302 include image data, the ML algorithms 1312 can include visual recognition software configured to interpret image data. The ML algorithms 1312 apply machine learning techniques to received training data (e.g., data received from one or more of the data sources 1302) in order to, over time, create / train / update one or more models 1316 that model the overall task and the sub-tasks that the classifier 1310 is designed to complete.

[0098] Referring now to FIGS. 13 and 14 collectively, FIG. 14 depicts an example of a learning phase 1400 performed by the ML algorithms 1312 to generate the abovedescribed models 1316. In the learning phase 1400, the classifier 1310 extracts features from the training data and coverts the features to vector representations that can be recognized and analyzed by the ML algorithms 1312. The features vectors are analyzed by the ML algorithm 1312 to “classify” the training data against the target model (or the model’s task) and uncover relationships between and among the classified training data. Examples of suitable implementations of the ML algorithms 1312 include but are not limited to neural networks, support vector machines (SVMs), logistic regression, decision trees, hidden Markov Models (HMMs), etc. The learning or training performed by the ML algorithms 1312 can be supervised, unsupervised, or a hybrid that includes aspects of supervised and unsupervised learning. Supervised learning is when training data is already available and classified / labeled. Unsupervised learning is when training data isnot classified / labeled so must be developed through iterations of the classifier 1310 and the ML algorithms 1312. Unsupervised learning can utilize additional learning / training methods including, for example, clustering, anomaly detection, neural networks, deep learning, and the like.

[0099] When the models 1316 are sufficiently trained by the ML algorithms 1312, the data sources 1302 that generate “real world” data are accessed, and the “real world” data is applied to the models 1316 to generate usable versions of the results 1320. In some embodiments of the invention, the results 1320 can be fed back to the classifier 1310 and used by the ML algorithms 1312 as additional training data for updating and / or refining the models 1316.

[0100] In aspects of the invention, the ML algorithms 1312 and the models 1316 can be configured to apply confidence levels (CLs) to various ones of their results / determinations (including the results 1320) in order to improve the overall accuracy of the particular result / determination. When the ML algorithms 1312 and / or the models 1316 make a determination or generate a result for which the value of CL is below a predetermined threshold (TH) (i.e., CL < TH), the result / determination can be classified as having sufficiently low “confidence” to justify a conclusion that the determination / result is not valid, and this conclusion can be used to determine when, how, and / or if the determinations / results are handled in downstream processing. If CL > TH, the determination / result can be considered valid, and this conclusion can be used to determine when, how, and / or if the determinations / results are handled in downstream processing. Many different predetermined TH levels can be provided. The determinations / results with CL>TH can be ranked from the highest CL>TH to the lowest CL>TH in order to prioritize when, how, and / or if the determinations / results are handled in downstream processing.

[0101] In aspects of the invention, the classifier 1310 can be configured to apply confidence levels (CLs) to the results 1320. When the classifier 1310 determines that aCL in the results 1320 is below a predetermined threshold (TH) (i.e., CL < TH), the results 1320 can be classified as sufficiently low to justify a classification of “no confidence” in the results 1320. If CL > TH, the results 1320 can be classified as sufficiently high to justify a determination that the results 1320 are valid. Many different predetermined TH levels can be provided such that the results 1320 with CL>TH can be ranked from the highest CL>TH to the lowest CL>TH.

[0102] The functions performed by the classifier 1310, and more specifically by the ML algorithm 1312, can be organized as a weighted directed graph, wherein the nodes are artificial neurons (e.g. modeled after neurons of the human brain), and wherein weighted directed edges connect the nodes. The directed graph of the classifier 1310 can be organized such that certain nodes form input layer nodes, certain nodes form hidden layer nodes, and certain nodes form output layer nodes. The input layer nodes couple to the hidden layer nodes, which couple to the output layer nodes. Each node is connected to every node in the adjacent layer by connection pathways, which can be depicted as directional arrows that each has a connection strength. Multiple input layers, multiple hidden layers, and multiple output layers can be provided. When multiple hidden layers are provided, the classifier 1310 can perform unsupervised deep-learning for executing the assigned task(s) of the classifier 1310.

[0103] Similar to the functionality of a human brain, each input layer node receives inputs with no connection strength adjustments and no node summations. Each hidden layer node receives its inputs from all input layer nodes according to the connection strengths associated with the relevant connection pathways. A similar connection strength multiplication and node summation is performed for the hidden layer nodes and the output layer nodes.

[0104] The weighted directed graph of the classifier 1310 processes data records (e.g., outputs from the data sources 1302) one at a time, and it “learns” by comparing an initially arbitrary classification of the record with the known actual classification of therecord. Using a training methodology knows as “back-propagation” (i.e., “backward propagation of errors”), the errors from the initial classification of the first record are fed back into the weighted directed graphs of the classifier 1310 and used to modify the weighted directed graph’s weighted connections the second time around, and this feedback process continues for many iterations. In the training phase of a weighted directed graph of the classifier 1310, the correct classification for each record is known, and the output nodes can therefore be assigned “correct” values. For example, a node value of “1” (or 0.9) for the node corresponding to the correct class, and a node value of “0” (or 0.1) for the others. It is thus possible to compare the weighted directed graph’s calculated values for the output nodes to these “correct” values, and to calculate an error term for each node (i.e., the “delta” rule). These error terms are then used to adjust the weights in the hidden layers so that in the next iteration the output values will be closer to the “correct” values.

[0105] FIG. 15 illustrates an example of a computer system 1500 that can be used to implement the various processor-related operations and / or cognitive algorithms described herein. The computer system 1500 includes an exemplary computing device (“computer”) 1502 configured for performing various aspects of the content-based semantic monitoring operations described herein in accordance embodiments of the invention. In addition to computer 1502, exemplary computer system 1500 includes network 1514, which connects computer 1502 to additional systems (not depicted) and can include one or more wide area networks (WANs) and / or local area networks (LANs) such as the Internet, intranet(s), and / or wireless communication network(s). Computer 1502 and additional systems are in communication via network 1514, e.g., to communicate data between them.

[0106] Exemplary computer 1502 includes processor cores 1504, main memory (“memory”) 1510, and input / output component(s) 1512, which are in communication via bus 1503. Processor cores 1504 includes cache memory (“cache”) 1506 and controls1508, which include branch prediction structures and associated search, hit, detect and update logic, which will be described in more detail below. Cache 1506 can include multiple cache levels (not depicted) that are on or off-chip from processor 1504. Memory 1510 can include various data stored therein, e.g., instructions, software, routines, etc., which, e.g., can be transferred to / from cache 1506 by controls 1508 for execution by processor 1504. Input / output component(s) 1512 can include one or more components that facilitate local and / or remote input / output operations to / from computer 1502, such as a display, keyboard, modem, network adapter, etc. (not depicted).

[0107] A cloud computing system 50A is in wired or wireless electronic communication with the computer system 1500. The cloud computing system 50A can supplement, support or replace some or all of the functionality (in any combination) of the computing system 1500. Additionally, some or all of the functionality of the computer system 1500 can be implemented as a node of the cloud computing system 50A.

[0108] For the sake of brevity, conventional techniques related to making and using the disclosed embodiments may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs to implement the various technical features described herein are well known. Accordingly, in the interest of brevity, many conventional implementation details are only mentioned briefly or are omitted entirely without providing the well-known system and / or process details.

[0109] For convenience, some of the technical operations described herein are conveyed using informal expressions. For example, a processor that has data stored in its cache memory can be described as the processor “knowing” the data. Similarly, a user sending a load-data command to a processor can be described as the user “telling” the processor to load data. It is understood that any such informal expressions in this detailed description should be read to cover, and a person skilled in the relevant art wouldunderstand such informal expressions to cover, the formal and technical description represented by the informal expression.

[0110] Many of the functional units of the systems described in this specification have been labeled as modules. Embodiments of the invention apply to a wide variety of module implementations. For example, a module can be implemented as a hardware circuit including custom VLSI circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices or the like. Modules can also be implemented in software for execution by various types of processors. An identified module of executable code can, for instance, include one or more physical or logical blocks of computer instructions which can, for instance, be organized as an object, procedure, or function. Nevertheless, the executables of an identified module need not be physically located together, but can include disparate instructions stored in different locations which, when joined logically together, function as the module and achieve the stated purpose for the module.

[0111] The various components / modules / models of the systems illustrated herein are depicted separately for ease of illustration and explanation. In embodiments of the invention, the functions performed by the various components / modules / models can be distributed differently than shown without departing from the scope of the various embodiments of the invention describe herein unless it is specifically stated otherwise.

[0112] Aspects of the invention can be embodied as a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.

[0113] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0114] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0115] The terms “about,” “substantially,” “substantial,” “approximately,” and equivalents thereof are intended to include the degree of error associated with measurement of the particular quantity based upon the equipment available at the time of filing the application. For example, “about” can include a range of ± 8% or 5%, or 2% of a given value.

[0116] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, and / or groups thereof.

[0117] While the present invention has been described with reference to an exemplary embodiment or embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof without departing from the scope of the present invention. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present invention without departing from the essential scope thereof. Therefore, it is intended that the present invention not be limited to the particular embodiment disclosed as the best mode contemplated for carrying out this present invention, but that the present invention will include all embodiments falling within the scope of the claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method of evaluating a system-under-design (SUD), the computer-implemented method comprising: applying, using a processor system, first reinforcement learning (RL) operations to an electronic representation of a first design of the SUD to generate a first set of performance indicators; wherein the first RL operations comprise a first final tuned RL policy; and wherein the first final tuned RL policy is based at least in part on multiple RL policies associated with multiple different designs of the SUD.

2. The computer-implemented method of claim 1 further comprising applying, using the processor system, second reinforcement learning (RL) operations to an electronic representation of a second design of the SUD to generate a second set of performance indicators.

3. The computer-implemented method of claim 2, wherein the second RL operations comprise a second final tuned RL policy.

4. The computer-implemented method of claim 3, wherein the second final tuned RL policy is based at least in part on the multiple RL policies associated with the multiple different designs of the SUD.

5. The computer-implemented method of claim 4 further comprising comparing the first set of performance indicators with the second set of performance indicators.

6. The computer-implemented method of claim 5 further comprising selecting the first design of the SUD or the second design of the SUD based at least in part on a result of comparing the first set of performance indicators with the second set of performance indicators.

7. The computer-implemented method of claims 1, 5, or 6 wherein the first final tuned RL policy is generated based at least in part on performing fast-tuning reinforcement learning (FT-RL) policy generation operations.

8. The computer-implemented method of claim 7, wherein the FT-RL policy generation operations comprise: clustering the multiple different designs of the SUD to generate SUD design clusters; generating initial polices based at least in part on the SUD design clusters; and generating the first final tuned RL policy based at least in part on tuning one of the initial policies.

9. The computer-implemented method of claim 7, wherein the FT-RL policy generation operations comprise: clustering SUD responses of the multiple different designs of the SUD to generate SUD design response clusters; and transferring policies across different number of inputs and outputs of the multiple different designs of the SUD to identify which parts of NNs representing the multiple different designs of the SUD can be transferred.

10. The computer-implemented method of claim 9, wherein the FT-RL policy generation operations further comprise performing serial-type transfer learning orparallel-type meta-learning within and across the SUD design response clusters to generate a set of final tuned RL policies comprising the first final tuned RL policy.

11. A computer system comprising a processor communicatively coupled to a memory, wherein the processor performs processor operations comprising: applying first reinforcement learning (RL) operations to an electronic representation of a first design of a system-under-design (SUD) to generate a first set of performance indicators; wherein the first RL operations comprise a first final tuned RL policy; and wherein the first final tuned RL policy is based at least in part on multiple RL policies associated with multiple different designs of the SUD.

12. The computer system of claim 11, wherein the processor operations further comprise applying second reinforcement learning (RL) operations to an electronic representation of a second design of the SUD to generate a second set of performance indicators.

13. The computer system of claim 12, wherein the second RL operations comprise a second final tuned RL policy.

14. The computer system claim 13, wherein the second final tuned RL policy is based at least in part on the multiple RL policies associated with the multiple different designs of the SUD.

15. The computer system of claim 14, wherein the processor operations further comprise comparing the first set of performance indicators with the second set of performance indicators.

16. The computer system of claim 15, wherein the processor operations further comprise selecting the first design of the SUD or the second design of the SUD based at least in part on a result of comparing the first set of performance indicators with the second set of performance indicators.

17. The computer system of claims 11, 15, or 16, wherein the first final tuned RL policy is generated based at least in part on performing fast-tuning reinforcementlearning (FT-RL) policy generation operations.

18. The computer system of claim 17, wherein the FT-RL policy generation operations comprise: clustering the multiple different designs of the SUD to generate SUD design clusters; generating initial polices based at least in part on the SUD design clusters; and generating the first final tuned RL policy based at least in part on tuning one of the initial policies.

19. The computer system of claim 17, wherein the FT-RL policy generation operations comprise: clustering SUD responses of the multiple different designs of the SUD to generate SUD design response clusters; and transferring policies across different number of inputs and outputs of the multiple different designs of the SUD to identify which parts of NNs representing the multiple different designs of the SUD can be transferred.

20. The computer system of claim 19, wherein the FT-RL policy generation operations further comprise performing serial-type transfer learning or parallel-typemeta-learning within and across the SUD design response clusters to generate a set of final tuned RL policies comprising the first final tuned RL policy.