Manufacturing biologic pharmaceutical using causal models

The system addresses inefficiencies in biopharmaceutical manufacturing control by using a causal model to iteratively select and measure control settings, enhancing production efficiency and adaptability to environmental factors.

JP2025112242AInactive Publication Date: 2025-07-31SOLVENTUM INTELLECTUAL PROPERTIES CO
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024075651
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2024-05-08
Publication Date
2025-07-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing techniques for controlling biopharmaceutical manufacturing environments rely on modeling-based or active control methods that are inefficient in determining optimal control settings, failing to account for environmental responses and changes effectively.

Method used

A system that repeatedly selects and measures the impact of control settings on biopharmaceutical quality, using a causal model to update and refine the relationship between input settings and quality metrics, including environmental characteristics that cannot be controlled.

Benefits of technology

This approach rapidly and accurately models the causal relationship between control settings and biopharmaceutical quality, improving production efficiency and resilience to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025112242000001_ABST
    Figure 2025112242000001_ABST
Patent Text Reader

Abstract

To provide methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for optimizing a process of manufacturing a biologic pharmaceutical.SOLUTION: In a control system, the method comprises repeatedly performing the following: selecting a configuration of input settings for manufacturing a batch of a biologic pharmaceutical based on a causal model that measures current causal relationships between input settings and a measure of a quality of batches of the biological pharmaceutical; determining a measure of the quality of a batch of the biological pharmaceutical manufactured using the configuration of input settings; and adjusting, based on the measure of the quality of the batch of the biological pharmaceutical, the causal model.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to the control of one or more biopharmaceutical sites and the determination of a causal relationship between control settings for manufacturing biopharmaceuticals at the site and environmental responses received from the biopharmaceutical site.

[0002] Existing techniques for determining which control settings to use to control the environment generally employ either modeling-based techniques or rely on active control of the system.

[0003] In modeling-based techniques, the system passively observes data, i.e., a historical mapping of control settings to environmental responses, discovers patterns in the data, and attempts to learn a model that can be used to control the environment. Examples of modeling-based techniques include decision forests, logistic regression, support vector machines, neural networks, kernel machines, and Bayesian classifiers.

[0004] In active control techniques, the system relies on active control of the environment for the generation and application of knowledge. Examples of active control techniques include randomized control experiments, e.g., bandit experiments.

Summary of the Invention

[0005] This specification describes a system and method implemented as a computer program on one or more computers in one or more locations that selects control settings for manufacturing biopharmaceuticals.

[0006] According to one aspect, there is provided a method comprising repeatedly performing: i) selecting a configuration of input settings for manufacturing a batch of a biological pharmaceutical based on a causal model that measures the current causal relationship between the input settings and the measured values of the quality of the batch of the biological pharmaceutical; ii) determining the measured values of the quality of the batch of the biological pharmaceutical manufactured using the configuration of the input settings; and iii) adjusting the causal model based on the measured values of the quality of the batch of the biological pharmaceutical.

[0007] In some implementations, the method further comprises selecting a configuration of input settings based on a set of internal control parameters and adjusting the internal control parameters based on the measured values of the quality of the batch.

[0008] In some implementations, the measured values of the quality of the batch include one or more of the batch yield, the measured value of the batch activity, the measured value of the batch effectiveness, the measured value of the batch impurities, the measured value of the viability of the transfected cells, the measured value of the transfection efficiency, the measured value of the purity of the biological pharmaceutical, the measured value of the monodispersity, or the measured value of the cell density. In some implementations, the purity of the biological pharmaceutical includes one or more of the measured value of the host cell protein impurities, the measured value of the DNA impurities, the measured value of the RNA impurities, or the measured value of the host cell metabolite impurities.

[0009] In some implementations, the input settings include one or more of the following: settings related to processing steps, settings related to the chemical composition of a substance, settings related to physical controls during manufacturing, or settings related to the purification of a batch of a biological pharmaceutical. In some implementations, the settings related to processing steps include one or more of the following: the order of processing steps, the time spent in each processing step, the time spent in each sub-process of a given processing step, or the selection of parameters for each processing step. In some implementations, the settings related to chemical composition include the selection of one or more of the following: a nutrient medium, an ion exchange resin, an affinity chromatography matrix, a polymer, a wash buffer, or an elution buffer. In some implementations, the settings related to physical controls during manufacturing include one or more of the following: manufacturing time, purification time, temperature during manufacturing, temperature during purification, humidity during manufacturing, humidity during purification, flow rate used during manufacturing, flow rate used during purification, G-force used during manufacturing, measurement of magnetism used during manufacturing, measurement of electricity used during manufacturing, or measurement of agitation used during manufacturing. In some implementations, the settings related to the purification of a batch of a biological pharmaceutical include purifying the batch using one or more of the following: electrostatic interactions, hydrophobic interactions, hydrophilic interactions, affinity chromatography, or size exclusion chromatography.

[0010] In some implementations, the method further includes selecting a configuration based on a causal model and a predetermined set of external variables, and the method further includes adjusting control parameters that parameterize the influence of the predetermined set of external variables. In some implementations, the predetermined set of external variables includes one or more external variables related to the environmental climate during manufacturing. In some such implementations, the external variables related to the environmental climate include one or more of the following: humidity during manufacturing, external temperature during manufacturing, or level of light during manufacturing.

[0011] Particular embodiments of the subject matter described herein can be implemented so as to realize one or more of the following advantages.

[0012] Using the methods described herein, the production of biological pharmaceuticals can be rapidly improved. By repeatedly selecting various control settings and measuring the impact of the control settings on the quality of the biological pharmaceutical, the control system can generate a causal model that more rapidly and accurately models the causal relationship between the control settings and the quality of the biological pharmaceutical than other prior art control systems.

[0013] The control system can also take into account the characteristics of the environment that cannot be controlled but affect the quality of the biological pharmaceutical. Thus, the causal model can independently model the relationship between the control settings and the quality of the biological pharmaceutical for various configurations of the environmental characteristics, making the quality of the biological pharmaceutical less vulnerable to changes in these characteristics.

[0014] In some implementations, the control system can continue to operate and use the causal model to select control settings for manufacturing the biological pharmaceutical. Thus, the system can continuously update the causal model while optimizing the quality of the biological pharmaceutical using the causal model.

Brief Description of the Drawings

[0015]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

[0016] Like reference numerals and designations in the various drawings indicate like elements.

DETAILED DESCRIPTION OF THE INVENTION

[0017] This specification generally describes a control system for controlling an environment when the state of the environment changes. In particular, the system controls the environment to determine the causal relationship between the control settings of the environment and the environmental response to the control settings. In particular, the environment is one or more biological pharmaceutical sites. A biological pharmaceutical site is a location where a biological pharmaceutical is manufactured, such as a biological pharmaceutical manufacturing plant or purification plant. Here, a biological pharmaceutical is defined as a pharmaceutical manufactured, extracted, or semi-synthesized from a biological source, such as a vaccine, blood component, gene therapy agent, or living cell used in cell therapy. The system selects manufacturing control settings, including the purification of batches of biological pharmaceuticals at the site. The environmental response is a measure of the quality of the biological pharmaceutical produced using the input settings selected by the system.

[0018] For example, the measurements of the quality of the biological pharmaceutical for which the causal relationship has been determined may include (i) sensor readings or other environmental measurements that reflect the state of the manufactured biological pharmaceutical, (ii) performance metrics that measure the performance of the control system based on the environmental measurements, such as a merit index or objective function, or (iii) both.

[0019] In particular, the control system repeatedly selects control settings that each include respective settings for each of a set of controllable elements of the biological pharmaceutical site. Generally, when different control settings are selected, there are differences in system performance. That is, the values of the environmental response are different.

[0020] More specifically, by repeatedly selecting control settings and measuring the impact of the control settings on the measured values of the quality of the biological pharmaceutical, the control system updates a causal model that models the causal relationship between the control settings and the environmental response, i.e., updates the maintained data that identifies the causal relationship between the control settings and the system performance.

[0021] Although the causal model is called a "causal model", in some implementations, the model can consist of multiple causal models, each corresponding to a different segment of the environment, i.e., a segment of the environment that shares specific characteristics.

[0022] In some implementations, the control system can continue to operate and use the causal model to select control settings for the biological pharmaceutical site. In other implementations, when certain criteria are met, the control system can provide the causal model to an external system or provide the user with data that displays the causal relationships identified by the causal model for use in controlling the biological pharmaceutical site. For example, the criteria can be met after the system has controlled the biological pharmaceutical site for a specific period of time or after the settings have been selected a specific number of times. As another example, the criteria can be met when the causal relationships identified in the maintained data meet certain criteria, e.g., when there are non-overlapping confidence intervals.

[0023] While updating the causal model, the system repeatedly selects different control settings and measures the impact of each possible control setting on the measured values of the quality of the biological pharmaceutical produced using the control settings, based on the internal parameters of the control system and the characteristics of the biological pharmaceutical site.

[0024] In other words, the internal parameters of the control system define both (i) how the system updates the causal model and (ii) how the system determines the control settings selected based on the current causal model. While updating the causal model, the control system repeatedly adjusts at least some of the internal parameters as more environmental responses become available to assist in identifying causal relationships.

[0025] FIG. 1A shows a control system 100 that selects a control setting 104 applied to a biopharmaceutical site 102. Each control setting 104 defines the setting of each of a plurality of controllable elements of the biopharmaceutical site 102. Generally, the controllable elements of the biopharmaceutical site 102 are elements that can be controlled by the system 100 and can take on a plurality of different possible settings.

[0026] The control setting 104 can include settings related to the steps for processing the biopharmaceutical. For example, the control setting 104 can include the order of the processing steps, the time spent in a given processing step, and / or the selection of parameters for a given processing step. The control system 100 can also direct the time spent on one or more sub-processes of a given processing step.

[0027] The control setting 104 can also include settings related to the chemical composition of a given substance. The substance can be the biopharmaceutical itself or any other substance used during the manufacture of the biopharmaceutical. For example, the control setting 104 can include the selection of a nutrient medium, an ion exchange resin, an affinity chromatography matrix, a polymer, a wash buffer, and / or an elution buffer.

[0028] Control setting 104 can also include physical controls during manufacturing, i.e., settings related to controlling the environment and parameters of the manufacturing process. For example, control setting 104 can include the manufacturing time and / or purification time of a biological pharmaceutical. The control system 100 can also direct the temperature during manufacturing and / or purification, the humidity during manufacturing and / or purification, or the flow rate used during manufacturing and / or purification. Control setting 104 can also include the G-forces used during manufacturing, the magnetic measurements used during manufacturing, the electrical measurements used during manufacturing, and / or the agitation measurements used during manufacturing.

[0029] Control setting 104 can also include settings related to the purification of biological pharmaceuticals. In these cases, the purification process can occur after a separate manufacturing process, during the manufacturing process, or by itself. That is, a given biological pharmaceutical site purifies only biological pharmaceuticals. The control system 100 can select how the biological pharmaceutical is purified. For example, the control settings can include the selection of electrostatic interactions, hydrophobic interactions, hydrophilic interactions, affinity chromatography, and / or size exclusion chromatography.

[0030] During operation, the control system 100 repeatedly selects control setting 104 and monitors the environmental response 130 to control setting 104. The environmental response 130 can be measured using one or more measurements of the quality of a batch of biological pharmaceuticals manufactured using control setting 104. For example, the quality measurements can include the batch yield, a measurement of the batch activity, a measurement of the batch effectiveness, and / or a measurement of the impurities in the batch. If transfection is part of the manufacturing process of a biological pharmaceutical, the response also includes a measurement of the viability of the transfected cells or a measurement of the transfection efficiency. The quality measurements can also include measurements of monodispersity, e.g., a measurement of the dispersion of non-aggregated proteins, and / or a measurement of cell density.

[0031] As a specific example, the measured values of the quality of a batch of biological pharmaceuticals may include the measured value of the purity of the biological pharmaceutical. Here, the purity of the biological pharmaceutical may include the measured value of host cell protein impurities, the measured value of DNA impurities, the measured value of RNA impurities, and / or the measured value of host cell metabolite impurities.

[0032] The system can calculate a performance metric for the environmental response 130, that is, calculate a single value representing the performance of the system when controlling the environment to optimize the quality of the biological pharmaceutical. An exemplary performance metric that combines all the measured values of quality used in the system is the weighted sum of the selected measured values of quality.

[0033] As another example, the performance metric can be, for each measured value of quality, the weighted sum of the difference between the measured value of quality and the baseline or desired value of the measured value of quality, that is, to minimize the deviation of the system outside the allowable range of each measured value of quality. Another example of such a performance metric is the weighted sum of a function that is zero when the measured value of quality is within the allowable range for each measured value of quality, and when the measured value of quality is outside the allowable range, it is equal to the difference from the measured value to the nearest end point of the allowable range.

[0034] System 100 also monitors the characteristics 140 of the environment 102. Generally, the characteristics 140 can include any data that can change the effect of the control settings 104 on the environmental response 130 but is not considered in the control settings, that is, data that characterizes the environment that cannot be controlled by the control system 100. For example, the environmental characteristics 140 of the biological pharmaceutical site 102 may include the humidity during production, the external temperature during production, and / or the level of light during production if their respective measured values cannot be adjusted by the biological pharmaceutical site 102. In some cases, note that these measured values may be controllable by the biological pharmaceutical site 102, for example, whether site 102 can adjust the lighting within the site. In these cases, the adjustable measured values are included as control settings rather than environmental characteristics 140.

[0035] System 100 uses environmental response 130 to model the causal relationship between control settings and environmental response, i.e., to update causal model 110 that models how different settings of different elements affect the value of the environmental response.

[0036] In particular, for each controllable element of biopharmaceutical site 102 and for each different type of environmental response, causal model 110 measures the causal effect that the various possible settings of the controllable element have on the environmental response, and the current level of uncertainty of the system regarding the causal effect of the possible settings.

[0037] As a specific example, for each different possible setting of a given controllable element and for each different type of environmental response, causal model 110 includes an effect measure that represents the effect of the possible setting on the environmental response compared to the other possible settings of the controllable element, e.g., the average estimated value of the true average effect of the possible setting, and a confidence interval, e.g., the 95% confidence interval of the effect measure that represents the current level of uncertainty of the system regarding the causal effect.

[0038] Thus, the system calculates, for example, a confidence interval that specifies the upper and lower 95% bounds of the effect of the control setting on system performance. Specifically, this enables the system to identify cases where the selection of various control settings results in (clinically) significant or non-significant differences. The system can refrain from testing controllable elements that do not result in significant differences. For example, if the upper and lower bounds of the confidence interval are such that even at the maximum effect they do not result in a clinically meaningful difference, and the cost of continuing the testing / exploration of that controllable element is high, the system can request permission to remove that control setting. As a clinical example, imagine that the confidence intervals for several control settings of systolic blood pressure indicate that the effect of the setting is ±0.02 points. Regarding continuing the experiment with that setting, e.g., in a range where there are health-related or other high costs, the cost of the experiment is such that the system can request permission to stop the experiment because the benefit of a 0.02 pressure drop on the probability of a cardiac event does not exceed the cost.

[0039] Before starting to control the biological pharmaceutical site 102, the control system 100 receives an external input 106. The external input 106 can include data received by the control system 100 from any of various sources. For example, the external input 106 can include data received from a user of the system, data generated by another control system that previously controlled the biological pharmaceutical site 102, data generated by a machine learning model, or some combination of these.

[0040] Generally, the external input 106 specifies at least (i) possible initial values for the settings of controllable elements of the biological pharmaceutical site 102, and (ii) which environmental responses the control system 100 is to track during operation.

[0041] For example, the external input 106 can specify that the control system 100 needs to track the measured values of a particular sensor of the biological pharmaceutical site, performance metrics, i.e., merit indices or other objective functions derived from the particular sensor measurements and optimized by the system 100 while controlling the biological pharmaceutical site, or both.

[0042] The control system 100 uses the external input 106 to generate an initial probability distribution (the "baseline probability distribution") over the initial possible setting values of the controllable elements. By using the external input 106 to initialize these baseline probability distributions, the system 100 ensures that settings that do not violate the constraints imposed by the external data 106 are selected, and, if desired by the user of the system 100, do not deviate from the historical range of control settings that have already been used to control the biological pharmaceutical site 102. For example, if there are certain ranges of control settings known to be unsafe for the biological pharmaceutical site 102, the external data 106 can define those ranges so that the system 100 never selects control settings within the unsafe ranges.

[0043] The control system 100 also uses the external input 106 to initialize a set 120 of internal parameters, i.e., to assign baseline values to the set of internal parameters. Generally, the internal parameters 120 define how the system 100 selects control settings, assuming the current causal model 110, i.e., assuming the current causal relationships determined by the system 100 and the uncertainty of the system with respect to the current causal relationships. The internal parameters 120 also define how the system 100 updates the causal model 110 using the received environmental response 130.

[0044] As will be described in more detail below, the system 100 updates at least some of the internal parameters 120 while updating the causal model 110. That is, some of the internal parameters 120 can be fixed to the baseline values initialized during the operation of the system 100, but the system 100 repeatedly adjusts other internal parameters 120 during operation to enable the system to measure more effectively and, in some cases, utilize causal relationships.

[0045] Specifically, in order to control a biopharmaceutical site, during operation, the system 100 repeatedly identifies procedural instances within the biopharmaceutical site based on the internal parameters 120.

[0046] Each procedural instance is a collection of one or more entities within the biopharmaceutical site associated with a time frame. Entities within the biopharmaceutical site are either a subset of the biopharmaceutical site, i.e., either an appropriate subset or an inappropriate subset. In particular, an entity is a subset of the biopharmaceutical site that can acquire an environmental response and can be affected by the applied control settings.

[0047] For example, if a biological pharmaceutical site includes a plurality of physical entities from which sensor measurements can be obtained, a given procedural instance can include an appropriate subset of the physical entities to which a set of control settings is applied. The number of subsets into which the entities within the biological pharmaceutical site can be divided is defined by internal parameter 120.

[0048] In particular, the way in which system 100 divides entities into subsets at any given point during the operation of the system is defined by an internal parameter that defines the spatial extent of the control settings applied by the system to the instance. The spatial extent of the instance identifies the subset of the environment that is assigned to the instance. That is, so that the environmental response obtained from that subset is associated with the instance.

[0049] For example, a procedural instance can include one or more machines that operate using control settings 104. In such a case, the spatial extent can define the number and type of machines within the procedural instance. System 100 obtains an environmental response to the control settings selected for a given group of machines. For example, as described above, system 100 can select control settings related to the measurement of agitation applied by the machine, and then system 100 can use the selected agitation measurement value to track a selected measurement of the quality of the batch of biological pharmaceuticals processed by the machine.

[0050] The length of the time frame associated with the entities within any given procedural instance is also defined by internal parameter 120. In particular, the time frame that the system assigns to any given procedural instance is defined by an internal parameter that defines the temporal extent of the control settings applied by the system. This time frame, that is, the temporal extent of the instance, defines whether future environmental responses determined by system 100 are caused by the control settings selected for the procedural instance.

[0051] Since the internal parameter 120 changes during the operation of the system 100, the instances generated by the system 100 may also change. That is, the system can change the way in which the procedural instances are identified when the system changes the internal parameter 120.

[0052] The system 100 then selects the settings for each instance based on the internal parameter 120 and, in some cases, based on the environmental characteristics 140.

[0053] In some cases, that is, when the system 100 is exploring the space of possible settings, the system 100 selects the settings for all instances based on the baseline probability distribution.

[0054] In other cases, that is, when the system 100 is using the causal relationships that have already been determined to optimize the objective function, the system 100 continues to select the settings for other instances (the "baseline instances") based on the baseline probability distribution while using the current causal model 110 to select the settings for some instances (the "hybrid instances"). More specifically, at any given point in time during the operation of the system 100, the internal parameter 120 defines the ratio of hybrid instances to the total number of instances.

[0055] The system 100 also determines, for each instance, which environmental responses 130 are associated with the instance, that is, which are used when updating the causal model 110 based on the internal parameter 120.

[0056] Next, the system 100 sets the settings 104 for each instance and monitors the environmental responses 130 for the settings selected for the instance. The system 100 maps the environmental responses 130 to impact measurement values for each instance and uses the impact measurement values to determine the causal model update 150 that is used to update the current causal model 110.

[0057] In particular, the system determines, based on internal parameter 120, which historical procedural type instances (and the environmental responses 130 associated with the instances) should be considered by causal model 110, and determines causal model update 150 based only on these determined historical procedural type instances. Which historical procedural type instances are considered by causal model 110 is determined by a set of internal parameters 120 that define a data inclusion window. The data inclusion window specifies, at any given point in time, one or more historical time frames during which the procedural type instance should occur in order for the result of that procedural type instance, i.e., the environmental response 130 associated with that procedural type instance, to be considered by causal model 110.

[0058] For those internal parameters that have been changed by system 100, system 100 also periodically updates 160 the data maintained by system 100 for those internal parameters based on causal model 110. In other words, when causal model 110 changes during the operation of system 100, system 100 also updates internal parameter 120 to reflect the changes in causal model 110. When system 100 assigns several control settings for using the current causal model 110, system 100 can also determine internal parameter update 160 using the difference in system performance between "hybrid" instances and "baseline" instances.

[0059] FIG. 1B shows data from an exemplary causal model. In particular, in the example of FIG. 1B, the causal model is represented as a chart 180 showing control settings, i.e., different possible settings for different controllable elements on the x-axis, and the causal effects on the control settings on the y-axis. In particular, for each possible setting of each controllable element, the causal model shows an impact measurement value and a confidence interval around that impact measurement value.

[0060] These causal relationships are shown in more detail for a particular controllable element 192 in the element-specific chart 190. The element-specific chart 190 shows that there are five possible settings for the controllable element 192, and the possible settings are referred to as levels in the chart 190. For each of the five settings, the chart includes a bar representing the impact measurement value and a representation of the confidence interval around the bar as an error bar around the impact measurement value. Thus, the information of the causal model for any given setting of the controllable element includes the impact measurement value and the confidence interval around that impact measurement value. For example, in the case of the second setting 192 (shown as IV-LV2 in the figure), the chart 190 shows a bar 194 indicating the impact measurement value of the second setting, an upper bar 196 above the bar 194 indicating the upper limit of the confidence interval of the second setting, and a lower bar 198 below the bar 194 indicating the lower limit of the confidence interval of the second setting.

[0061] FIG. 1B shows a single causal model, and from the following description, it is understood that the system can maintain and update multiple different causal models, one for each cluster of the procedural instances, for any given controllable element.

[0062] FIG. 2 is a flowchart of an exemplary process 200 for controlling an environment. For convenience, the process 200 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as the control system 100 of FIG. 1 appropriately programmed, can execute the process 200.

[0063] The system assigns a baseline value to a set of internal parameters and assigns a baseline probability distribution to each of the controllable elements of the environment (step 202).

[0064] In particular, the system receives external data, for example, from a user of the system or from data obtained from previous control of the system environment by another system, and then uses the external data to assign baseline values and generate a probability distribution. Generally, the external data specifies the initial constraints under which the system operates when controlling the environment.

[0065] In particular, the external data identifies the possible control settings for each controllable element within the environment. That is, the external data identifies, for each controllable element within the environment, which possible settings the system can select for the controllable element when controlling the system.

[0066] In some cases, the external data can specify additional constraints on the possible control settings, for example, that the setting of a particular controllable element depends on the settings of other controllable elements, or that a particular entity can be associated only with a particular subset of the possible control settings of a particular controllable element.

[0067] Thus, the external data defines the search space of the possible combinations of control settings that the system can explore when controlling the environment.

[0068] In some implementations, these constraints may change during the operation of the system.

[0069] For example, the system can receive additional external inputs that change the possible control settings of one or more controllable elements, or the range of values of spatial and temporal ranges.

[0070] As another example, if the system determines that the optimal setting for one of the controllable elements or one of the internal parameters is approaching the boundary of the search space defined by the external constraints, for example, if the impact measure of the causal model indicates that the optimal setting is the one closest to one of the boundaries of the search space, the system can, for example, request permission from the system administrator or other users of the system to expand the space of possible values of the controllable element or internal parameter.

[0071] As yet another example, if the external data specifies that the possible settings of a controllable element may be any value within a continuous range, the system can first discretize the range in one way and change the discretization to prioritize a segment if the confidence interval strongly enough indicates that the optimal value is included in that segment of the continuous range.

[0072] As yet another example, if the system determines that a particular controllable element has no causal effect on the environmental response, for example, if it is likely that the impact measures for all possible settings of the controllable element are all zero, the system can request permission to remove the controllable element from system control.

[0073] Next, the system generates a baseline (or "prior") probability distribution for the possible control settings for each controllable element of the environment. For example, if the external data specifies only the possible values of a given controllable element and does not assign any priority to any of the possible values, the system can generate a uniform probability distribution over all possible values, assigning equal probabilities to each possible value. As another example, if, for example, based on the historical results of controlling the environment, the external data prioritizes a particular setting over other settings for a given controllable element, the system can generate a probability distribution that assigns a higher probability to the prioritized setting.

[0074] The system also assigns baseline values to each internal parameter of the system. In particular, the internal parameters of the system include (i) a set of internal parameters (referred to as "spatial range parameters") that define the spatial range of the procedural instances generated by the system, and (ii) a set of internal parameters (referred to as "temporal range parameters") that define the temporal range of the procedural instances generated by the system.

[0075] If the system includes multiple entities, the system can maintain separate sets of spatial range parameters and temporal range parameters for each of the multiple entities. In other cases where the system includes multiple entities, the system maintains only a single set of spatial and temporal range parameters that apply to all of the multiple entities. In still other cases where the system includes multiple entities, the system initially maintains a single set of spatial and temporal range parameters and, during operation of the system, can switch to maintaining separate sets of spatial or temporal range parameters if system performance improves, i.e., if different entities respond to control settings in different ways than other entities.

[0076] Furthermore, in some implementations, the system maintains separate sets of temporal range parameters for various controllable elements.

[0077] The system also maintains a set of internal parameters (referred to as "data inclusion window parameters") that define the data inclusion window used by the system. In some implementations, the system maintains a single set of data inclusion window parameters that apply to all controllable elements. In some other implementations, the system maintains a separate set of data inclusion window parameters for each controllable element of the environment. That is, when updating the causal model, the system can use different data inclusion windows for different controllable elements. As described in more detail below, when the system clusters procedural instances into multiple clusters, the system can either (a) maintain a separate set of data inclusion window parameters for each cluster or (b) maintain a separate set of data inclusion window parameters for each cluster and each controllable element. That is, different clusters can use different data inclusion windows for the same controllable element.

[0078] In an implementation where the system utilizes a causal model, the internal parameters also include (iv) a set of internal parameters (referred to as "ratio parameters") that define the ratio between the hybrid instance and the baseline instance. In some implementations, the system maintains a single set of ratio parameters that apply to all controllable elements. In some other implementations, the system maintains a separate set of ratio parameters for each controllable element of the environment. That is, when selecting control settings, the system can use different ratios for different controllable elements. As will be explained in more detail below, when the system clusters procedural instances into multiple clusters, the system can either (a) continue to maintain a single set of ratio parameters across all clusters, (b) maintain a separate set of ratio parameters for each cluster, or (c) maintain a separate set of ratio parameters for each cluster and each controllable element. That is, different clusters can use different ratios when selecting control settings for the same controllable element.

[0079] As will be explained below, in an implementation where the system clusters instances into multiple clusters, the internal parameters also include (v) a set of internal parameters (referred to as "clustering parameters") that define the current clustering strategy.

[0080] Generally, the clustering parameters are or define hyperparameters of the clustering technique used by the system. Examples of such hyperparameters include the cluster size of each cluster, i.e., the number of procedural instances within each cluster, and the environmental characteristics used to cluster the procedural instances.

[0081] The system maintains a set of clustering parameters for each controllable element. That is, for each controllable element, the system generates clusters of procedural instances of that controllable element using different hyperparameters when applying clustering techniques.

[0082] The internal parameters can also optionally include any of various other internal parameters that affect the operation of the control system. For example, the internal parameters may include a set of internal parameters that define how to update the causal model (e.g., a set of weights, each representing the relative importance of each environmental characteristic during trend matching between procedural instances that can be used to calculate the d-score as described below).

[0083] As described above, the system changes at least some of these internal parameters during operation.

[0084] For each set of internal parameters that the system changes during environmental control, the system can change the values using (i) a heuristic-based approach, (ii) probabilistically sampling values to optimize a merit index of the internal parameters, or (iii) both.

[0085] For any set of internal parameters that are changed based only on heuristics, the system maintains a single value of the internal parameter and repeatedly adjusts that single value based on the heuristics.

[0086] For any set of internal parameters that are changed by probabilistically sampling, the system maintains parameters that define the range of possible values of the internal parameters and maintains a causal model that identifies the causal relationship between the possible values of the internal parameters and the merit index of the internal parameters. The merit index of the internal parameters may be different from the performance metric used in the causal model of the control settings. The system then selects values from within the range of possible values for at least some of the instances at any given point in time, based on the current causal model.

[0087] When the system uses heuristics to update a set of internal parameters in addition to probabilistic sampling, the system can use the heuristics to update the range of possible values. That is, the range of possible values is updated by a heuristic-based approach, but the causal model of the values within the range at any given point in time is updated by probabilistic sampling.

[0088] For any internal parameters that the system does not change during the control of the environment, the system can maintain either a fixed range of values and a fixed probability distribution over the fixed range of values, or a fixed single value that is always used during the operation of the system.

[0089] Depending on the content included in the external data, the system assigns to each internal parameter either a baseline value derived from the external data or a baseline value of a default value.

[0090] For example, external data generally identifies a range of values for spatial and temporal extents. For example, if the spatial extent is not fixed and is an internal parameter that can be changed by the system, the external data can specify the minimum and maximum values of the spatial extent. Similarly, if the temporal extent is not fixed and is an internal parameter that can be changed by the system, the external data can specify the minimum and maximum values of the temporal extent.

[0091] Next, the system uses external data to assign initial values to the spatial range parameters such that the parameters define the range of values specified by the external data, and also assigns initial values to the temporal range parameters such that the parameters define the range of values specified by the external data.

[0092] For other internal parameters, the system assigns default values. For example, the system can initialize the clustering parameter to indicate that the number of clusters is 1 (i.e., clustering is disabled at the start of environmental control), and initialize the ratio parameter to indicate that there are no hybrid instances (i.e., the system only explores at the start of environmental control). The system can also initialize the data inclusion window parameter to indicate that all historical procedural type instances completed in the data inclusion window are included.

[0093] The system executes the start phase (step 204). In the start phase, the system selects control settings for the procedural type instances based on the baseline probability distribution of the controllable elements and updates the causal model using the environmental response. That is, unless a historical causal model is provided as part of the external data, the system does not consider the current causal model when determining which control settings to assign to the procedural type instances.

[0094] Instead, the system selects control settings using the baseline probability distribution according to an assignment scheme that enables the system to effectively calculate the impact measure, i.e., the d-score, at a later time. In other words, the assignment scheme selects control settings in a way that takes into account the block scheme the system will use to calculate the impact measure, i.e., it assigns control settings to various procedural instances that will later be able to identify blocked groups in order to calculate the impact measure between the blocked groups. The block scheme (and thus the assignment scheme) employed by the system can be any of various schemes that reduce unexplained variation between different control settings. Examples of block schemes that can be used in the system include one or more of double-blind assignment, pairwise assignment, Latin square assignment, propensity matching, etc. Generally, the system can use any suitable block scheme that assigns procedural instances to blocked groups based on the current environmental characteristics of the entities within the procedural instance.

[0095] If the system can change either or both of the spatial extent and the temporal extent, the system changes the spatial extent parameter, the temporal extent parameter, or both during the initialization phase, such that values for the spatial and temporal extents that are more likely to result in sufficiently orthogonal procedural instances are more likely to be selected. A group of instances is considered orthogonal if the control setting applied to one of the instances within the group does not affect the environmental response associated with any of the other instances within the group.

[0096] The selection of control settings and causal model updates during the initialization phase are described in more detail below with reference to FIG. 3. Changing the spatial or temporal extent parameter is described in more detail below with reference to FIG. 11.

[0097] In some implementations, the system continues this initialization phase during system operation. That is, the system continues to explore the space of possible control settings and compiles the results of the exploration into a causal model.

[0098] As an example, the system can continue this initialization phase when updating the causal model with respect to multiple different environmental responses, i.e., when there is no merit index or objective function used by the system when utilizing the causal model, rather than a single merit index or objective function.

[0099] In some of these implementations, the system continues to explore the space of possible control settings while also adjusting a particular one of the set of initial parameters based on the causal model, such as spatial extent parameters, temporal extent parameters, data inclusion window parameters, clustering parameters, etc.

[0100] In some other implementations, when the system determines that a particular criterion is met, the system begins to execute a different phase. In these implementations, during the initialization phase, the system fixes and holds particular parameters of the internal parameters. For example, the system can fix and hold the data inclusion window parameter to indicate that all historical instances need to be incorporated into the causal model. As another example, the system can fix the internal parameters of the clustering to indicate that clustering should not be performed.

[0101] In particular, in these other implementations, when the system determines that the criterion is met, the system can begin to execute an exploitation phase (step 206).

[0102] For example, if the amount of procedural instances for which environmental responses have been collected exceeds a threshold, the system can start executing the exploitation phase. As a specific example, the system can determine that the threshold is met if the total number of such procedural instances exceeds the threshold. As another specific example, the system can determine that the threshold is met if the minimum number of environmental responses related to any one possible setting of any controllable element exceeds the threshold.

[0103] Furthermore, in some cases, the system does not adopt the initialization phase and proceeds directly to the exploitation phase, i.e., it does not execute step 204.

[0104] When a threshold is used, the system can determine the threshold in any of various ways.

[0105] As an example, when the system assigns settings to instances based on a causal model, the system can determine that the threshold is met when sufficient environmental responses of instances have been collected such that the likelihoods of different selected settings vary. The method of assigning likelihoods based on a causal model will be described in more detail below with reference to FIG. 5.

[0106] As another example, the system can determine a threshold that is the number of procedural instances required for a statistical test that the system performs to determine a confidence interval in order to generate an accurate confidence interval, i.e., the number of procedural instances that satisfy the statistical assumptions of the confidence calculation.

[0107] As another example, the system can determine the threshold to be equal to the number of procedural instances required to provide a desired power to the causal model, i.e., as determined by a power analysis.

[0108] During the exploitation phase, the system selects some control settings of the procedural instances based on the current causal model and continues to select the control settings of other procedural instances based on the baseline values of the internal parameters.

[0109] In particular, the system changes a ratio internal parameter to change the ratio between the number of procedural instances that are hybrid instances, i.e., instances for which control settings are assigned based on a causal model, and the number of procedural instances that are baseline instances, i.e., instances for which control settings are assigned based on a baseline probability distribution and for which the number is greater than zero.

[0110] The system begins to designate certain instances as hybrid instances during the exploit phase, so the system begins to use the differences in system performance among the hybrid instances and can explore the instances to adjust the values of internal parameters, such as the ratio internal parameter, the data inclusion window parameter, and the like.

[0111] The selection of control settings, the causal model update, and the update of internal parameters during the exploit phase will be described in more detail below with reference to FIG. 3.

[0112] In some implementations, when the system determines that certain criteria are met, the system begins a clustering phase (step 208). That is, if the system is configured to cluster procedural instances, when the clustering criteria are met, the system begins the clustering phase. If the system is not configured to cluster instances, the system does not cluster procedural instances at any point during the operation of the system.

[0113] Generally, the system creates a sub-population of similar procedural instances, taking clustering into account. In real-world situations, different procedural instances across the population may respond differently to different control settings. The optimal control settings for one procedural instance may not be optimal for another. These differences can affect the distribution of performance metrics seen across instances. If one control setting is chosen for the entire population, it may negatively impact the overall utility, i.e., the overall performance of the system. To maximize the overall utility across the entire population, the system can cluster instances into sub-populations, considering their individual characteristics (modeled by environmental characteristics) and their feedback characteristics (modeled by the performance metrics received for control settings). The system selects control settings at the level of these sub-populations.

[0114] Depending on the implementation and criteria, the system can start the clustering phase during the initialization phase or the exploit phase. That is, although clustering is step 208 and the initialization phase and the exploit phase are steps 204 and 206 respectively in FIG. 2, the clustering phase is shown to overlap with the initialization phase, the exploit phase, or both.

[0115] During the clustering phase, before assigning control settings to procedural instances, the system clusters the procedural instances into clusters based on the current values of the clustering internal parameters and the characteristics of the procedural instances. As described above, the clustering internal parameters for any given controllable element define the hyperparameters of the clustering technique used for clustering that controllable element.

[0116] When the system starts the clustering phase, at any given point in time, the system maintains a separate causal model for each cluster. That is, the system identifies separate causal relationships within each cluster. As described above, the system can also maintain separate sets of internal parameters for at least some of the internal parameters of each cluster.

[0117] In the following description, generally, it is described that separate sets of ratio parameters and data inclusion window parameters are maintained for each cluster and for each controllable element. However, if the system maintains only one set of some types of parameters for each cluster, the calculations described need to be performed only once for each cluster, and the result of a single calculation should be understood to be usable for each controllable element of the cluster. Similarly, if the system maintains only a single set of some types of parameters for all clusters, the calculations described need to be performed only once, and the result of a single calculation is usable for all controllable elements in all clusters.

[0118] During the exploit phase, when clustering is started, within a given cluster, the system selects some control settings for the procedural instances within the cluster based on the current causal model and continues to select control settings for other procedural instances based on the baseline values of the internal parameters.

[0119] The system can determine the point in time to start clustering using any of various criteria. That is, it can determine the point in time when the internal parameters of the clustering begin to change from the baseline value indicating that the total number of clusters needs to be set to 1.

[0120] As an example, one criterion may include that sufficient environmental responses have been collected when, for example, the amount of collected environmental responses exceeds a threshold. As a specific example, the system can determine that the threshold is met when the total number of environmental responses exceeds the threshold. As another specific example, the system can determine that the threshold is met when the minimum number of environmental responses associated with any one possible setting of any controllable element exceeds the threshold.

[0121] As another example, another criterion can specify that the system can start clustering when it determines that for any one of the controllable elements, different environmental characteristics have different effects on the causal effects of different control settings of that controllable element. As a specific example, this criterion can be that the d-score distribution of any controllable element is statistically different between any two procedural instances, that is, the d-score distribution of the causal model based only on the environmental responses of one procedural instance is statistically different from the d-score distribution of the causal model based only on the environmental responses of another procedural instance up to the threshold level of statistical significance, in which case the system can start clustering.

[0122] The selection of control settings, causal model updates, and updates of internal parameters during the clustering phase will be described in more detail below with reference to FIG. 3.

[0123] FIG. 3 is a flowchart of an exemplary process 300 for performing iterations of environmental control. For convenience, process 300 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as the control system 100 of FIG. 1 appropriately programmed, can execute process 300.

[0124] The system can repeatedly execute process 300 to update a causal model that measures the causal relationship between control settings and environmental responses.

[0125] The system determines a set of current procedural instances based on current internal parameters (step 302). As will be described in more detail below with reference to FIG. 4A, the system determines spatial and temporal ranges based on the current internal parameters, for example, based on the possibility of resulting in instances where different spatial and temporal ranges are orthogonal, and then generates current procedural instances based on the spatial and temporal ranges.

[0126] As described above, each procedural instance is a collection of one or more entities within the environment and is associated with a time frame. The time frame associated with a given procedural instance defines, as will be described in more detail below, which environmental responses the system attributes to or associates with the procedural instance.

[0127] In some cases, for each controllable element, the system also determines the period during which the selected setting for the controllable element is applied as part of the time frame associated with the controllable element (e.g., the entire time frame, the first quarter of the time frame, or the first half of the time frame). Generally, the period during which the setting is applied can be fixed to a value independent of the time frame, made a fixed part of the time frame, or made an internal parameter that is changed by the system for a part of the time frame.

[0128] Determining the current set of procedural instances will be described in more detail below with reference to FIG. 4A.

[0129] If the environment contains only a single physical entity, the current set of instances may contain only one instance. Alternatively, the system can identify multiple current instances, each current instance containing a single physical entity but being separated temporally, i.e., at least by the temporal range of the entity.

[0130] The system assigns control settings for each current instance (step 304). The way the system assigns control settings to any given instance depends on the control phase the system is currently executing.

[0131] As described above, at the start of environmental control, i.e., before sufficient information becomes available to determine causal relationships with any degree of confidence, the system operates in an initialization phase. In the initialization phase, the system selects the control settings of the instance without considering the current causal model, i.e., the system explores the space of possible control settings. That is, the system selects the control settings of the instance according to the baseline probability distribution for the possible control settings of each controllable element.

[0132] In some implementations, during the initialization phase, the system changes internal parameters that determine procedural instances of the spatial range, the temporal range, or both, in order to identify the degree to which each possible value of the spatial and temporal ranges is likely to result in instances that are orthogonal to each other.

[0133] As described above, in some implementations, the set of control phases includes only the initialization phase, and the system continues to operate throughout this start-up phase, i.e., continues to explore the space of possible control settings while compiling environmental responses to update the causal model.

[0134] In some other implementations, when a specific criterion is met, the system transitions to an exploitation phase. In the exploitation phase, the system selects some of the control settings of the current instance based on the current causal model, i.e., utilizes the causal relationships reflected in the current causal model and continues to select the control settings of other current instances based on the baseline values of the internal parameters.

[0135] Furthermore, in some implementations, the system starts performing clustering either in the initialization phase or in the exploitation phase.

[0136] When clustering is being performed, the system clusters the procedural instances into clusters. Within each cluster, the system proceeds independently as described above.

[0137] That is, during the initialization phase, the system independently selects settings within each cluster using the baseline distribution, but during the exploitation phase, the system independently assigns control settings to some of the current instances based on the current causal model within each cluster and continues to select the control settings of other current instances based on the baseline values of the internal parameters within each cluster.

[0138] By performing clustering, the system can conditionally assign control settings based on (i) the interaction of factors between the impact of the settings on the environmental response and the environmental characteristics of the instances, such as the attributes of the instances that cannot be manipulated by the control system, (ii) the factor interactions of different independent variables, or (iii) both.

[0139] The selection of control settings in the exploitation phase when clustering is used and when it is not used will be described in more detail below with reference to FIG. 5.

[0140] The system obtains the environmental response for each of the procedural instances (step 306).

[0141] Specifically, the system monitors the environmental response and determines which environmental response to attribute to which current instance based on the time frame associated with each procedural instance.

[0142] More specifically, for each procedural instance, the system associates each environmental response that (i) corresponds to an entity of the procedural instance and (ii) is received during some portion of the time frame associated with the procedural instance, with the procedural instance. As a specific example, to limit the carry-over effect from a previous control setting assignment, the system can associate each environmental response that corresponds to an entity within the instance and that is received beyond a threshold period after the start of the time frame, e.g., the second half of the time frame, the last third of the time frame, or the last quarter of the time frame, with the procedural instance. In some implementations, this threshold period is fixed. In other implementations, the system defines this threshold period and maintains a set of internal parameters that change the period during operation of the system.

[0143] The system updates a causal model based on the acquired environmental responses (step 308). Causal model updates are described in more detail below with reference to FIG. 6.

[0144] The system updates at least some of the internal parameters based on the current performance of the system, i.e., compared to the baseline performance of the system, as reflected in the updated causal model, or both (step 310).

[0145] In particular, the system can update any of a set of various internal parameters based on a heuristic-based approach, by stochastic variation, or both. The heuristic-based approach can include heuristics derived from one or more of the following: the updated causal model, the current performance of the system compared to the baseline performance of the system, or criteria determined using prior statistical analysis.

[0146] In other words, for each set of internal parameters that the system can change, the system uses one or more of the above techniques to update the set of internal parameters so that the system can measure causal relationships more accurately.

[0147] In some cases, even if the system can change internal parameters, the system is restricted to fix a particular set of internal parameters. For example, the system can fix the data inclusion window parameters and the clustering parameters during the initialization phase. As another example, the system can fix the clustering parameters until a particular criterion is met, and can start changing all internal parameters under the control of the system during the exploitation phase after the criterion is met.

[0148] Updating a set of internal parameters will be described in more detail below with reference to FIGS. 8-12.

[0149] Generally, the system can execute steps 302-306 at a frequency different from step 308, and can execute step 310 at a frequency different from both steps 302-306 and step 310. For example, the system can execute multiple iterations of steps 302-306 for each iteration of step 308 executed, that is, it can collect environmental responses for multiple different sets of instances before updating the causal model. Similarly, the system can execute multiple different instances of step 308 before executing step 310, that is, it can execute multiple different causal model updates before updating the internal parameters.

[0150] FIG. 4A is a flowchart of an exemplary process 400 for determining a procedural instance. For convenience, process 400 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 properly programmed, can execute process 400.

[0151] The system selects the spatial extent of each entity in the environment (step 402). The spatial extent of a given entity defines the segment of the environment that affects the environmental response obtained from the given entity when controlled by a given set of control settings. The spatial extent of a given entity is defined by either a set of spatial extent parameters, e.g., a set of spatial extent parameters specific to the given entity, or a set shared among all entities. In some implementations, the internal parameters of the spatial extent are fixed, i.e., kept at the same value or randomly sampled from a fixed range throughout the control of the environment. For example, if the environment contains only a single entity, each procedural instance contains the same single entity. As another example, if the environment contains multiple entities but there is no uncertainty as to which entity is affected by the control settings, the spatial extent parameters can be fixed to values that ensure that the generated instances are orthogonal.

[0152] If the spatial extent is not fixed and a single value is maintained for the spatial extent parameters (i.e., the spatial extent parameters are updated based on heuristics only), the system selects the current value of the spatial extent parameters of each entity as the spatial extent of the entity. If the spatial extent is not fixed and a range of values is defined by the spatial extent parameters, the system samples a value for the spatial extent from the range currently defined by the spatial extent parameters based on the current causal model of the spatial extent parameters of the entity.

[0153] By selecting the spatial extent of the entities, the system defines the number of entities included in each procedural instance and which entities are included in each procedural instance. In particular, the system generates procedural instances such that a procedural instance does not cover a segment of the environment that is further partially within the spatial extent of the entities of another procedural instance.

[0154] Figure 4B shows an example of a map 420 of an environment that includes a plurality of physical entities, each associated with a spatial extent. In particular, Figure 4B shows an environment within a portion of the United States that includes a plurality of physical entities represented as dots in the figure. The spatial extent selected for each entity by the system is represented by a shaded circle. For example, the system may maintain a range of possible radii for each entity and select the radius of the shaded circle for each entity from that range. As can be seen from the example of Figure 4B, different entities can have different spatial extents. For example, entity 412 has a shaded circle of a different size than entity 414.

[0155] As can also be seen from the example of Figure 4B, the system can also, if desired, apply additional criteria to reduce the likelihood that the procedural instances are not orthogonal. In particular, the system selects for each entity a buffer (represented by a dashed circle) that extends beyond the spatial extent of the entity and requires that entities of different instances cannot have spatial extents within that buffer.

[0156] For the spatial extent and for the buffer, certain entities within the environment are not selected as part of the procedural instances in the iteration shown in Figure 4B. These unselected entities, for example, entity 416, are represented as non-shaded dots or dashed circles. In particular, the system did not select these entities because their spatial extents intersected the spatial extent or buffer of another entity selected as part of the procedural instance. For entity 416, the entity was not selected because the spatial extent of entity 416 would intersect the spatial extent or buffer of entity 414. For example, considering the sampled spatial extents and buffers, the system was able to select a spatial extent that maximized the number of procedural instances that could be included in the current set of instances without violating any of the criteria.

[0157] The system selects a time range for each procedural instance or, if different controllable elements have different time ranges, for each controllable element of each procedural instance (step 404). As described above, the time range defines a time frame associated with each procedural instance or a time frame associated with a controllable element within the procedural instance.

[0158] In some cases, the time range can be fixed, i.e., before the control system operates for the user of the system, it is known which environmental responses observed for a given entity in the environment need to belong to the procedural instance containing that entity. In other cases, the time range may be unknown or associated with some degree of uncertainty, i.e., the user of the system does not know or specify exactly how long the effect of the setting can be observed after a set of settings is applied.

[0159] If the time range is not fixed, the system samples a value for the time range from the range currently defined by the time range parameter based on the current causal model of the time range parameter. As described above, different entities (and thus different procedural instances) can have different sets of time range parameters, or all entities can share the same set of time range parameters.

[0160] The system generates a procedural instance based on the selected spatial range and the selected time range (step 406). In other words, the system divides the entities in the environment based on the spatial range, i.e., so that an entity in a procedural instance does not have a spatial range (or buffer if used) that intersects the spatial range of another entity in a different procedural instance, and associates each procedural instance with a time frame defined by the spatial range of the procedural instance.

[0161] FIG. 5 is a flowchart of an exemplary process 500 for selecting control settings for a set of current instances. For convenience, process 500 is described as being executed by one or more computer systems located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 programmed appropriately, can execute process 500.

[0162] The system determines a current procedural instance (step 502), as described above with reference to FIG. 4A.

[0163] Next, the system executes steps 504-514 for each controllable element to select the settings of the controllable elements for all current procedural instances.

[0164] Optionally, the system clusters the current procedural instances based on environmental characteristics, i.e., generates multiple clusters for the controllable elements (step 504). Since clustering is performed for each controllable element, the system can cluster the current procedural instances differently for different controllable elements. Clustering of procedural instances is described below with reference to FIG. 7.

[0165] That is, when the system is executing the current clustering phase, the system first determines the current cluster assignment for the current procedural instance. After the system determines the current cluster assignment, the system executes the iteration of steps 506-514 independently for each cluster.

[0166] When the system is not executing the current clustering phase, the system does not cluster the current procedural instances and executes a single iteration of steps 506-514 for all current procedural instances.

[0167] The system determines the current hybrid-to-baseline ratio (step 506). In particular, if the set of ratio parameters of the controllable elements contains only a single value, the system selects the current value of the ratio parameter as the current hybrid-to-baseline ratio. If the system of ratio parameters of the controllable elements defines a range of possible values, the system samples the value of the hybrid-to-baseline ratio from the current range of possible values defined by the ratio parameter based on the causal model of the set of ratio parameters.

[0168] The system identifies each instance as either a hybrid instance of the controllable element or a baseline instance of the controllable element based on the current hybrid-to-baseline ratio (step 508). For example, the system can assign each instance to a hybrid instance with a probability based on the ratio, or randomly divide the total number of instances to make the ratios as equal as possible. Alternatively, if the system is probabilistically changing at least one of the internal parameters based on the difference in performance between the hybrid instance and the baseline instance, the system applies an assignment scheme that takes into account the block scheme used when calculating the causal model that assigns instances based on the current ratio and measures the difference in performance (i.e., as described above).

[0169] The system selects the control settings for the controllable elements of the baseline instance according to the assignment scheme based on the baseline values of the internal parameters (step 512). In other words, the system selects the control settings for the baseline instance based on the baseline probability distribution for the possible values of the controllable elements determined at the start of the initialization phase.

[0170] The system selects the control settings for the hybrid instance according to the assignment scheme based on the current causal model (step 514).

[0171] In particular, the system maps the current causal model to a probability distribution over the possible settings of the controllable elements. For example, the system can apply probability matching to map the effect measures and confidence intervals of the controllable elements in the causal model to probabilities.

[0172] Generally, the system assigns control settings based on these probabilities such that a sufficient number of blocked groups are identified by the system when calculating the d-score. As a specific example, next, the system divides the hybrid instances into blocked groups (based on the same blocking scheme that will later be used in calculating the d-score), and then selects the control settings within each blocked group according to the probability distribution of the possible settings. That is, for each instance within a blocked group, any given possible setting is assigned the probability specified by the probability distribution.

[0173] FIG. 6 is a flowchart of an exemplary process 600 for updating a causal model of a given controllable element and a given type of environmental response. For convenience, process 600 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 programmed appropriately, can execute process 600.

[0174] The system can execute process 600 for each controllable element and for each type of environmental response for which the system maintains a causal model. For example, if the system maintains a causal model that models causal effects for only a single performance metric, the system executes process 600 only for the performance metric. Alternatively, if the system maintains a causal model that models causal effects for multiple different types of environmental responses, the system executes process 600 for each type of environmental response, such as for each different type of sensor reading or measurement.

[0175] When the system is currently clustering procedural instances into clusters, the system can execute process 600 independently for each cluster. That is, the system can independently maintain and update the causal model of each cluster.

[0176] The system determines the current data inclusion window of the controllable element (step 602), that is, based on the current data inclusion window parameter of the controllable element. In particular, if the set of data inclusion window parameters of the controllable element contains only a single value, the system selects the current value of the data inclusion window parameter as the current data inclusion window. If the set of data inclusion window parameters of the controllable element defines a range of possible values, the system samples the value of the data inclusion window from the range of values currently defined by the set of data inclusion window parameters. If the data inclusion window parameter is not changed by the system, the system sets the value to a fixed initial data inclusion window or samples the value from a fixed range of possible values.

[0177] For each possible value of the controllable element, the system obtains the environmental response of a given type recorded for the instance for which the possible value of the controllable element was selected (step 604). In particular, the system obtains only the environmental responses of the instances that occurred within the current data inclusion window.

[0178] The system updates the influence measurement value of the causal model based on the environmental responses for the possible settings of the controllable element (step 606).

[0179] That is, the system determines a set of blocked groups based on a block scheme, for example, one of the above block schemes.

[0180] Next, for each blocked group, the system determines the respective d-score for each possible setting selected in any of the instances within the blocked group. Generally, the system calculates the influence measurement value, i.e., the d-score, of a given controllable element based on the blocking scheme. That is, it calculates the d-score between the environmental responses of the instances assigned to the same blocked group.

[0181] As a specific example, the influence measurement value di of the possible setting i of a controllable element in a blocking scheme that assigns a blocked group and includes at least one instance with each possible setting can satisfy the following.

Equation

[0182] As another specific example, the influence measurement value di of the possible setting i of a controllable element in a blocking scheme that assigns a pair of instances to a blocked group can satisfy the following. d_i = x_i - x_(i + 1). Wherein, xi is an environmental response of a given type for an instance within the blocked group where setting i is selected, x_(i + 1) is an environmental response of a given type for an instance within the blocked group where setting i + 1 is selected, and in the formula, setting i + 1 is the setting immediately above the controllable element. In the case of the highest setting of the controllable element, setting i + 1 can be set to the lowest setting of the controllable element.

[0183] As yet another specific example, the influence measurement value di of the possible setting i of a controllable element in a blocking scheme that assigns a pair of instances to a blocked group can satisfy the following. d_i = x_i - x_1, where x_1 is the environmental response of a given type of instance for which one of the possible settings for a selected controllable element has been previously determined.

[0184] The system then calculates an overall impact measure updated for a given setting i as the average of the d-scores calculated for setting i.

[0185] In some cases, the calculation of the d-score can be by proportion rather than addition, i.e., any subtraction operation in the above definitions can be replaced by a division operation.

[0186] The system determines a confidence interval for the updated impact measure for each of the possible values of the controllable element (step 608). For example, the system can perform a t-test or other statistical hypothesis test to construct a p% confidence interval (i.e., near the average of the d-scores, where p is a fixed value, e.g., 95% or 97.5% or 99%) around the updated impact measure.

[0187] In some implementations, the system applies different p-values to different controllable elements, e.g., when external data specifies that different controllable elements have different levels of cost or risk associated with deviations from the baseline probability distribution of the different controllable elements.

[0188] In some implementations, if a particular setting of a controllable element is associated with various costs of implementation or a higher risk, the system applies a correction to the confidence interval, for example, the Bonferroni correction. In particular, in the Bonferroni correction, N confidence intervals are calculated for N possible settings of a controllable element, and the correction is applied such that the overall desired confidence level of that element is 95% (i.e., alpha = 0.05), in which case the alpha value used for each individual test to calculate the confidence interval is alpha / N. If a particular setting is associated with a higher risk or implementation cost, a "corrected" alpha value associated with a higher level of reliability may be specified for those settings. Thereby, the system is made to accumulate more data before utilizing these settings.

[0189] FIG. 7 is a flowchart of an exemplary process 700 for clustering a set of procedural instances of a given controllable element. For convenience, process 700 is described as being executed by one or more computer systems located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 appropriately programmed, can execute process 700.

[0190] The system selects the current hyperparameters of the clustering technique used by the system from the clustering parameters of the controllable element (step 702). In particular, each hyperparameter that can be changed by the system is defined by an individual set of internal parameters. That is, the clustering parameters include a separate set of internal parameters for each hyperparameter that is under the control of the system during operation.

[0191] The system can perform clustering using any of various clustering techniques. However, the hyperparameters that are changed by the system generally include the hyperparameters of the sizes of the clusters generated by the clustering technique, and in some cases, the environmental characteristics of the instances considered by the clustering technique when generating the clusters.

[0192] As an example, the system can use statistical analysis, such as analysis of variance (ANOVA), to generate clustering assignments. In particular, factor ANOVA is used to find the factors that explain the maximum amount of variance between clusters, i.e., the environmental characteristics. That is, since a D-score is calculated for each possible control setting, factor ANOVA can monitor the interaction terms between these treatment effects and external factors. As data accumulates and interactions begin to appear, factor ANOVA creates different clusters of instances over space and time, where each cluster represents the state or attributes of an individual external factor.

[0193] As another example, the system can generate clustering assignments using machine learning techniques. As a specific example of machine learning techniques, the system can use decision trees. Decision trees are classical machine learning algorithms used for classification and regression problems. Decision trees use a recursive partitioning scheme by sequentially identifying the best variable, i.e., the best environmental characteristic, to use for splitting using an information theoretic function such as the Gini coefficient. As another specific example of machine learning techniques, the system can use conditional inference trees. Similar to decision trees, conditional inference trees are a recursive binary partitioning scheme. The algorithm proceeds by selecting a sequence of variables to split based on a significance testing procedure that splits based on the most powerful environmental characteristic factors. As another specific example, the system can use a machine learning model, e.g., a deep neural network, to process data characterizing each of the procedural instances and their associated environmental characteristics, generate embeddings, and then cluster the procedural instances into the specified clusters based on the similarity between the embeddings (e.g., k-means clustering or another clustering technique). As a specific example, the embedding may be the output of an intermediate layer of a neural network trained to receive data characterizing a procedural instance and predict the value of a performance metric of the procedural instance.

[0194] In some cases, as the operation of the system proceeds, i.e., as more data becomes available, the system can switch clustering techniques. For example, the system can switch from using statistical techniques or decision trees to using a deep neural network when a threshold number of procedural instances becomes available.

[0195] The system clusters instances within the current data inclusion window using a clustering technique according to selected hyperparameters (step 704).

[0196] The system calculates the causal model for each cluster (step 706) (i.e., calculates as described above with reference to FIG. 6), but uses only the instances assigned to the cluster.

[0197] Next, the system independently assigns control settings for controllable elements within each cluster based on the calculated causal model of the cluster (i.e., as described above with reference to FIG. 5) (step 708). In particular, the system uses a clustering technique to cluster each current instance, assigns the control setting for a given current instance based on the cluster to which the current instance is assigned, and uses the causal model calculated for the cluster if the given current instance is not designated as a baseline instance.

[0198] Next, the system can determine whether the clustering parameters need to be adjusted (step 710), i.e., determine whether the current value of the clustering parameters is not optimal, and if not, update the clustering parameters for the controllable elements. In particular, during operation, the system updates the clustering parameters to balance the following two competing goals: (1) pooling instances into clusters such that there is maximum similarity within the cluster in the influence of the controllable elements on the performance metric and maximum difference between clusters in the influence of the controllable elements on the performance metric, and (2) maximizing the size of the clusters to maximize the sample size within the possible clusters and increase the accuracy of the causal model. The system can achieve this by adjusting the values using heuristics, using probabilistic sampling, or using both heuristics and probabilistic sampling.

[0199] The system can determine whether to change the number of clusters, i.e., whether to change the value of the clustering parameter for the controllable elements, in various ways, i.e., based on any of various heuristics.

[0200] More generally, as described above, for any given set of internal parameters that are changed by the system, the system can adjust the set of internal parameters in one of the following three ways: (i) adjusting a single value using a heuristic-based approach, (ii) using stochastic variation to adjust the likelihoods assigned to different values within a range of values, or (iii) adjusting the range of values using a heuristic-based approach and adjusting the likelihoods within the current range using stochastic variation.

[0201] The heuristic-based approach can include heuristics based on the properties of the current causal model, heuristics based on prior statistical analysis, or both.

[0202] In the stochastic variation approach, the system maintains a causal model that measures the causal effect between different values within the current range and a merit index of the set of internal parameters. The system then maps the causal model to the probabilities of the various values and, if necessary, selects the values of the internal parameters based on the probabilities. As will be described in more detail below, the merit index of any given set of internal parameters generally differs from the performance metric that is measured by a causal model that models the causal relationship between the control settings and the performance metric.

[0203] FIG. 8 is a flowchart of an exemplary process 800 for updating a set of internal parameters using stochastic variation. For convenience, process 800 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 appropriately programmed, can execute process 800.

[0204] Process 800 can be executed for any set of internal parameters that are updated using stochastic variations. Examples of such internal parameters can include any or all of a set of data inclusion window parameters, a set of clustering parameters, a set of ratio parameters, a set of spatial extent parameters, a set of temporal extent parameters, and the like.

[0205] As described above, during the clustering phase and for any set of internal parameters other than the clustering parameters, the system can execute Process 800 independently for each cluster or each controllable element and each cluster.

[0206] Furthermore, if the clustering parameters are changed using stochastic variations, the system can also execute Process 800 independently for each controllable element.

[0207] The system maintains (step 802) a causal model of the set of internal parameters that measures the causal relationship between different possible values of the internal parameters and a merit index of the set of internal parameters.

[0208] For example, the merit index of the set of internal parameters may be the difference between the performance of a hybrid instance and the performance of a baseline instance. In this example, the merit index measures the relative performance of the hybrid instance with respect to the baseline instance, and the system calculates an impact measurement value, i.e., a d-score, for this merit index for various values within the range defined by the internal parameters.

[0209] Thus, when calculating the causal model of a set of internal parameters, the system proceeds as described above with reference to FIG. 6, provided that (i) the possible settings are the possible values of the internal parameters, and (ii) each xi in the d - score calculation is the difference between (1) the performance metric of the hybrid instance to which the control setting is assigned with a possible value of the selected internal parameter and (2) the performance metric of the corresponding baseline instance.

[0210] As another example, the merit index of a set of internal parameters may be a measure of the accuracy of the causal model of the controllable elements, e.g., a measure of the width of the confidence intervals for various settings of the controllable elements.

[0211] This maintained causal model can be determined based on the data inclusion window of the set of internal parameters. If the set of internal parameters is actually the data inclusion window parameters, the data inclusion window will be different for different possible values within the current range. If the set of internal parameters is a different set of internal parameters, the data inclusion window may be fixed or may be changed based on a heuristic described below or also based on the probabilistic variations described in this figure, or may be a separate set of internal parameters.

[0212] The system maps the causal model to a probability distribution over the possible values within the range of values, for example, using probability matching (step 804). That is, the system uses probability matching or another suitable technique to map the impact measurements and confidence intervals to the probabilities of each possible value within the range of values.

[0213] If it is necessary to sample values from a range, the system samples values from the range of possible values according to a probability distribution (step 806). That is, if values from a range defined by internal parameters are required so that the system operates, for example, to assign a time range to a procedural instance, to assign a data inclusion window to a given controllable element, to determine hyperparameters of a clustering technique, or to assign a ratio of the current hybrid to the baseline of the current set of instances, the system samples values from the range of possible values according to a probability distribution. By sampling values in this way, the system is more likely to sample the values that are most likely to optimize the merit index of the set of internal parameters while still ensuring that the space of possible values is explored in order to maximize the delta between, for example, hybrid instances and baseline instances.

[0214] The system calculates a causal model update (step 808). That is, when a new environmental response for a new procedural instance is received, the system recalculates the causal model by calculating an overall impact measure, i.e., the average of the d-scores, and a confidence interval for the overall impact measure. The system can perform this calculation in the same way as the above-described causal model update with reference to FIG. 6, i.e., by selecting blocked groups and calculating the d-scores within those blocked groups (based on the merit index of the set of parameters described above), and then generating a causal model from those d-scores.

[0215] By repeatedly executing process 800, the system can repeatedly adjust the probabilities assigned to values within the range to prioritize the values that result in a more optimal merit index.

[0216] For example, when the set of internal parameters is a data inclusion window parameter, by maintaining a causal model that models the impact of various data inclusion window values on the hybrid-to-baseline performance, the system can select a data inclusion window, and a more accurate and robust causal model is calculated for the controllable elements.

[0217] As another example, when the set of internal parameters is a spatial or temporal range parameter, by maintaining a causal model that models the impact of different spatial or temporal range values on the hybrid-to-baseline performance, the system can select a spatial or temporal range that results in an orthogonal procedural instance that maximizes the performance of the hybrid instance compared to the performance of the baseline instance.

[0218] As another example, when the set of internal parameters defines clustering hyperparameters, by maintaining a causal model that models the impact of various hyperparameter values on the hybrid-to-baseline performance, the system can select a clustering assignment that maximizes the performance of the system, i.e., more effectively identify a clustering assignment that meets the above objectives with reference to FIG. 7.

[0219] In some implementations, the system determines whether to adjust the current range of possible values of the internal parameters (step 810). As described above, the range of possible values for any given internal parameter can be fixed or adjusted using heuristics to ensure that the space of possible values being explored throughout the operation of the system remains reasonable.

[0220] One example of a heuristic that can be used to adjust the current range of possible values is a heuristic that depends on the shape of the current causal model. In particular, the system can increase the upper bound of the range (or increase both the upper and lower bounds of the range) if the magnitude of the causal model's influence measure increases as the current upper bound of the range is approached, and can decrease the lower bound (or decrease both the upper and lower bounds) if the influence measure increases as the current lower bound of the range is approached.

[0221] Another example of a heuristic that can be used to adjust the current value of possible values is a heuristic that depends on power analysis.

[0222] For example, if the set of internal parameters is a set of clustering parameters that define the cluster size used in a clustering technique, the system can calculate a power curve that represents the effect of a changing sample size, i.e., the cluster size has a width of a confidence interval that the current causal model reflects in controllable elements. Considering the nature of the power curve, the confidence interval becomes more accurate more quickly at the smaller end of the sample size, but as the sample size increases, the improvement in the accuracy of the confidence interval becomes disproportionately small for each increase in the sample size (i.e., the decrease in the width of the confidence interval becomes disproportionately small). Therefore, exploring larger cluster sizes results in little improvement in power and a high risk of not accurately representing the current decision space. To account for this, the system can constrain the range of possible cluster sizes to a range between a lower threshold and an upper threshold of the power curve. By constraining the cluster size in this way, the system does not explore clusters that are so small that there is too little power to calculate a significant confidence interval. The system does not experiment with cluster sizes that are unnecessarily large, i.e., that result in only a small increase in power in exchange for the risk of not being able to capture all potential variation between instances.

[0223] As another example, when the set of internal parameters is a set of ratio parameters, the system can perform a power analysis to calculate the minimum number of baseline instances required to determine whether a hybrid instance is better than a baseline instance with a threshold power, assuming the current causal model of the ratio parameters. The system can then adjust the lower bound of the range of possible ratio values so that the number of baseline instances with a ratio below this minimum is not reached.

[0224] As another example of adjusting a range based on a heuristic, when the range of a temporal range parameter of an entity in an environment is updated based on a heuristic, the system can maintain a causal model for each entity that measures the causal relationship between (i) the control setting selected in a given control iteration and (ii) the environmental response obtained from the entity in a subsequent control iteration, i.e., the control iteration immediately following the given control iteration. Since the system is attempting to select a temporal range for the entity that ensures that the procedural instances are orthogonal, if the temporal range is appropriately selected, this causal model will indicate that there is a high likelihood that the causal effect between the current control setting and the environmental response to subsequent control settings is zero. Thus, if the causal model indicates that the confidence interval of the effect measurement for any control setting exceeds a threshold that overlaps with zero, the system can determine to increase the lower bound of the possible temporal range.

[0225] As another example of adjusting a range based on a heuristic, when the range of the spatial range parameter of an entity in the environment is updated based on a heuristic, the system can maintain, for each given entity, a causal model that measures the causal relationship between (i) a control setting selected in a given control iteration of a procedural instance that includes the given entity, and (ii) an environmental response obtained from an entity adjacent to the given entity in the current control iteration. The adjacent entity can be the entity closest to the given entity from among the entities included in the current set of instances of the current control iteration. Since the system is trying to select a spatial range for the entities that ensures orthogonality of the procedural instances, if the spatial range is appropriately selected, this causal model will likely indicate that the causal effect between the current control setting for a given entity and the environmental response for an adjacent entity is zero. Thus, if the causal model indicates that the confidence interval of the effect measurement for any control setting exceeds a threshold that overlaps with zero, the system can determine to increase the lower bound of the range of possible spatial ranges.

[0226] An additional example of a heuristic that can be used to adjust the range of possible values for the data inclusion window and ratio parameter is described in more detail below with reference to FIG. 12.

[0227] FIG. 9 is a flowchart of an exemplary process 900 for updating the value of the data inclusion value of a given controllable element based on a heuristic. For convenience, process 900 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 appropriately programmed, can execute process 900.

[0228] Generally, the system executes process 900 on the data inclusion window when the data inclusion window is a parameter that is changed heuristically without using stochastic variation.

[0229] When the system maintains multiple clusters for a given controllable element, the system can execute process 900 independently for each cluster. That is, the data inclusion window of a given controllable element within a cluster can be updated in a different way from the set of internal parameters of a given controllable element within another cluster.

[0230] The system accesses the current causal model of a given controllable element (step 902).

[0231] The system analyzes one or more properties of the current causal model (step 904). For example, the system can perform a normality test to determine whether the d-scores of the various possible control settings for a given controllable element are normally distributed (step 904). In particular, the system can perform a normality test (e.g., Shapiro-Wilk test) on the d-score distribution of a given controllable element within the current causal model. Generally, the system scales and pools the d-score distributions across different possible settings to generate a single distribution and then performs a normality test on the single distribution. The system performs this test on various data inclusion windows, e.g., the current causal model calculated using the current data inclusion window and one or more alternative causal models calculated using one or more alternative data inclusion windows, and can find the longest data inclusion window that satisfies the normality test at some defined p-value.

[0232] As another specific example, the system can measure the overlap of confidence intervals between different impact measurements in a given controllable element of the current causal model. The system can perform this test for various data inclusion windows, e.g., for the current causal model calculated using the current data inclusion window and one or more alternative causal models calculated using one or more alternative data inclusion windows, and find the data inclusion window closest to the desired degree of overlap.

[0233] As another specific example, the system can calculate a power analysis to identify a sample size that results in the current causal model having a desired power. The system can then adjust the data inclusion window such that the number of instances included in the adjusted window equals the identified sample size.

[0234] The system determines whether to adjust the data inclusion window parameter based on the results of the analysis (step 906). For example, the system can adjust the data inclusion window parameter to specify the longest data inclusion window that satisfies the above normality test, or the data inclusion window closest to the desired degree of overlap, or the data inclusion window that contains the number of instances equal to the identified sample size.

[0235] The example of FIG. 9 is an example of adjusting the data inclusion window based on heuristics. However, generally, any of the internal parameters can be adjusted based on heuristics (instead of being fixed and held or adjusted using stochastic variation). Some examples of setting internal parameters based on heuristics are shown below.

[0236] As an example, the system can set the value of the ratio parameter using power analysis. In particular, the system can perform a power analysis to calculate the minimum number of baseline instances needed to determine that a hybrid instance is better than a baseline instance at a threshold power. The system can then adjust the value of the ratio parameter to be equal to this minimum number.

[0237] As another example, to set the value of the hyperparameter for the cluster size, the system can perform a prior power analysis to determine the sufficient amount of environmental response needed for the causal model to have a desired power, i.e., instead of the above range, and set the value of the cluster size to this range.

[0238] The above description explains how the system can change internal parameters during operation of the system. By adjusting these internal parameters, the system can effectively account for changes in the properties of the environment. That is, it can account for the case where the mapping from control settings to environmental responses is not static and can change at various points during the operation of the system. Without being properly accounted for, changes in the properties of the environment that do not equally affect all possible control settings of all controllable elements can lead to an inaccurate causal model based on stale data that has become irrelevant, and thus potentially reduce the effectiveness of the system in controlling the environment.

[0239] FIG. 10 is a flowchart of an exemplary process 1000 for responding to changes in one or more properties of an environment. For convenience, process 1000 is described as being executed by a system of one or more computers located in one or more locations. For example, a control system, such as control system 100 of FIG. 1 properly programmed, can execute process 1000.

[0240] The system monitors the environmental response to the control settings selected by the system (step 1002). That is, as described above, the system repeatedly selects control settings and monitors the response to those selected control settings.

[0241] The system determines an indication that one or more properties of the environment have changed (step 1004). In particular, a change in the property of the environment is one that changes the relative impact that various settings for at least one of the controllable elements have on the environmental response being monitored by the system. That is, by determining an indication that one or more properties have changed, the system determines that it is likely that the relative causal effect of the various settings on the environmental response has changed. That is, it is in contrast to a global change that affects all possible control settings. The system does not have access to direct information specifying that a change has occurred, but the system can determine an indication that a change may have occurred based on the monitored environmental response.

[0242] For example, the system can determine an indication that a change has occurred when the difference between the current system performance and the baseline system performance is decreasing. In particular, as will be described in more detail below, the system can determine this based on a performance metric that increases as the possible value of the data inclusion window gets smaller, that is, as reflected in the causal model of the data inclusion window described above.

[0243] As another example, the system can determine an indication that a change has occurred when, as described above, the normality test determines that the d-scores of the possible settings of the controllable elements are no longer normally distributed.

[0244] In response to determining an indication that one or more properties of the environment have changed, the system adjusts the internal parameters of the system (step 1006).

[0245] Generally, the system adjusts the values of internal parameters to indicate an increasing level of uncertainty as to whether the causal model maintained by the system accurately captures the causal relationship between control settings and environmental responses.

[0246] For example, the system can adjust the data inclusion window parameter to narrow the data inclusion window. That is, when determining the causal model, only the more recent historical environmental responses are included. That is, the system can adjust the data inclusion window parameter to prioritize shorter data inclusion windows over the possible range of data inclusion windows.

[0247] As another example, the system can adjust the ratio parameter to decrease the hybrid pair search ratio, that is, the number of hybrid instances is reduced compared to the number of search instances. By decreasing the ratio, the system reduces its dependence on the current causal model when selecting control settings and instead explores the space of possible control settings more frequently. That is, the system can adjust the ratio parameter to prioritize smaller ratios over the possible range of ratios.

[0248] As another example, the system can adjust the clustering parameter to reduce the number of clusters into which instances are clustered. By reducing the number of clusters, the system prevents the causal model from clustering on potentially irrelevant characteristics when explaining differences in system performance between clusters.

[0249] FIG. 11 shows a representation 1100 of a data inclusion window for a given controllable element of the environment when a set of internal parameters defining data inclusion is changed probabilistically. As seen in the example of FIG. 11, the data inclusion window can range from zero (i.e., no data is included) to infinity (i.e., all procedural instances are included), but the current probabilistic variation range 110 within which the data inclusion window for a given controllable element is sampled is between a lower limit A 1102 and an upper limit B 1104. In some cases, the lower limit A 1102 and the upper limit B 1104 are fixed, and the system adjusts the probabilities assigned to different values between the lower limit A 1102 and the upper limit B 1104 by updating the causal model as described above. In other cases, the system can change the lower limit A 1102 and the upper limit B 1104 while also updating the causal model. In particular, the system can adjust the range 1110 based on the likelihood that the relative causal effects of different possible values of the controllable element are changing.

[0250] In particular, as shown in FIG. 11, the system maintains the range of possible values of the data inclusion window. That is, the data inclusion window parameters include the lower limit of the range, the upper limit of the range, and the possible values that the data inclusion window can take within the range. The data inclusion window parameters also include the probabilities of the possible values used when probabilistically sampling the values. As described above with reference to FIG. 8, these probabilities are adjusted by the system.

[0251] In some cases, the range of possible values is fixed. In other cases, however, the system changes the lower and upper limits of the range based on one or more heuristics, adjusts the possible data inclusion windows explored by the system, and prevents the system from exploring data inclusion windows that are too short or too long.

[0252] For example, the system can calculate a power curve that represents the impact of a change in sample size (due to a change in the data inclusion window) on the width of the confidence interval used by the current causal model for controllable elements. Considering the nature of the power curve, the confidence interval becomes more accurate more quickly at the smaller end of the sample size, but as the sample size increases, for each increase in the sample size, the improvement in the accuracy of the confidence interval becomes disproportionately small (i.e., the decrease in the width of the confidence interval becomes disproportionately small). Therefore, even if a longer data inclusion window is explored, there is little improvement in power, and the risk of not accurately representing the current decision space increases. To account for this, the system can constrain the range of the data inclusion window to obtain a number of samples that lies between a lower threshold and an upper threshold of the power curve. By constraining the data inclusion window in this way, the system does not explore data inclusion windows that are so short that they provide too little power to calculate a meaningful confidence interval, i.e., data inclusion windows that result in insufficient data to calculate a statistically significant confidence interval. The system also does not explore data inclusion windows that are unnecessarily long, i.e., that provide only a small increase in power in exchange for the risk of not being able to account for recent changes in the properties of the environment.

[0253] As another example, the system can calculate a measure of the stability of the interaction between time and a measured value of the relative impact of possible control settings of controllable elements, e.g., a factor analysis. That is, the system can determine the stability of the causal relationship over time. If the measure of stability indicates that the causal relationship is stable, the system can increase either the upper limit or both the upper and lower limits of the data inclusion window range, and if the measure of stability indicates that the causal relationship is unstable, i.e., dynamically changing, it can lower the upper limit or both the upper and lower limits. Thereby, the system can explore a smaller data inclusion window when the probability that the properties of the environment are changing is higher, ignore older data, and can explore a larger data inclusion window when the probability that the properties of the environment are stable is higher.

[0254] As yet another example, the system can adjust the range based on the shape of the causal model as described above. In particular, the system can explore a longer data inclusion window range when the magnitude of the impact measurement increases as the data inclusion window gets larger, and can explore a smaller data inclusion window range when the magnitude of the impact measurement increases as the data inclusion window gets shorter. In other words, the system can move the range down when the difference decreases and move the range up when the difference increases. Thereby, the system can explore a smaller data inclusion window and ignore older data when the probability that the properties of the environment are changing is high.

[0255] In some cases, the system can apply some combinations of these heuristics by increasing the upper limit, for example, based on either or both of the latter two examples, as long as the upper limit does not exceed a size corresponding to the upper threshold of the power curve, and by decreasing the lower limit, as long as the lower limit does not fall below a size corresponding to the lower threshold of the power curve, also based on either or both of the latter two examples.

[0256] Although these examples are described with respect to the data inclusion window, similar heuristics can also be used to adjust the ratio of the hybrid instance to the baseline instance, i.e., increase the number of baseline instances when the probability that the properties of the environment are changing or have recently changed is higher, and decrease the number of baseline instances when the probability that the properties of the environment are stable is higher.

[0257] Figure 12 shows the performance of the system described (shown as "DCL" in FIGS. 12-18) for controlling the same environment, compared to the performance of a system using an existing control scheme. In particular, FIG. 12 shows the performance of the described system compared to three different types of existing control schemes: (i) a "none" scheme where the system does not select a setting and receives only the baseline environmental response, (ii) a "random" scheme where the system randomly assigns control settings without replacement, and (iii) various state-of-the-art reinforcement learning algorithms.

[0258] In the example of FIG. 12, the environment being controlled has three controllable elements, each with five possible control settings, and the value of the performance metric at each iteration is obtained from a Gaussian distribution that is fixed overall. Applying a particular control setting changes the parameters of the Gaussian distribution from which the value of the performance metric is obtained. These characteristics are similar to those found in simple or highly controlled real-world environments, such as a particular manufacturing line, but lack the additional complexity that may be encountered in more complex real-world environments.

[0259] The upper set of plots in FIG. 12 shows the performance of each system in terms of the mean cumulative FOM ("MeanCumFOM"). The mean cumulative FOM at a given iteration is the performance metric received from the first iteration up to the given iteration, i.e., the average value of the FOM, i.e., the cumulative average performance metric value over time.

[0260] The lower set of plots in FIG. 12 shows the performance of each system by the mean FOM per instance ("MeanFOM"). The mean FOM per instance at any given iteration is the average of the performance metrics received for the instance at the given iteration, i.e., without considering previous iterations.

[0261] Generally, the first column ("DCL") shows the results of the described system, and the remaining columns show the results of the existing control schemes.

[0262] As shown above, the environment in which the results are shown in Figure 12 is not more complex than many real-world environments. For example, since the causal effects are fixed, there are no external uncontrollable characteristics that affect the performance measurement, and there is no uncertainty about the spatial or temporal extent. However, even in this relatively simple environment, the performance of the described system is equal to or better than the performance of state-of-the-art systems, regardless of whether advanced features are enabled.

[0263] The description of the state-of-the-art system used as a benchmark for the performance of the system is as follows.

[0264] ● BGE - Boltzmann - Gumbel Exploration [Cesa - Bianchi et al. Boltzmann Exploration Done Right, Conference on Neural Information Processing Systems (NeurIPS), 2017] is a multi - armed bandit algorithm that uses an exponentially weighted approach for the selection of control setting assignments. This maintains the variance of the entire FOM for each control setting assignment. At each step, samples are generated from each of these distributions, and the control setting assignment corresponding to the largest sample is selected by the algorithm. Then, using the received feedback, the internal parameters of the distribution are updated.

[0265] ● Ep Greedy - Epsilon Greedy is a general multi - armed bandit algorithm that selects a random control setting assignment with probability epsilon and selects the control setting assignment that gave the highest average FOM in the past with probability 1 - epsilon. In effect, it explores for epsilon percent of the time and exploits for 1 - epsilon percent of the time.

[0266] ● Upper Confidence Bound (UCB) of UCB [Auer et al. Finite-time Analysis of the Multiarmed Bandit Problem, Machine Learning, 2002] The multi-armed bandit algorithm is one of two basic approaches to solving the multi-armed bandit problem. It functions by calculating the average FOM and confidence intervals from historical data. It selects the assignment of control settings by calculating the assignment of the control setting with the highest average FOM + confidence interval. In this way, it operates optimistically about the potential FOM of the assignment of control settings and learns over time which assignment of control settings has the highest FOM.

[0267] ● Lin UCB - LinUCB [Li et al. A Contextual-Bandit Approach to Personalized News Article Recommendation, International World Wide Web Conference (WWW), 2010] is constructed based on UCB by maintaining the average FOM and confidence intervals, making the important assumption that the expected FOM is a linear function of the characteristics of the procedural instance in the experiment and the assignment of control settings. Then, the algorithm can select the best assignment of control settings for any individual procedural instance. Lin UCB is expected to exhibit the best performance in different situations for different groups of procedural instances for the ideal assignment of control settings.

[0268] ● Monitored UCB - Monitored UCB [Cao et al. Nearly Optimal Adaptive Procedure with Change Detection for Piecewise - Stationary Bandit, International Conference on Artificial Intelligence and Statistics (AISTATS), 2019] is constructed based on UCB by calculating the average FOM and confidence intervals, and is designed for environments where there may be a sudden change in FOM. Therefore, it incorporates a change - point detection algorithm that identifies the point in time when FOM changes and resets the internal parameters (effectively resetting the average FOM and confidence intervals) to start learning the new FOM. Monitored UCB is expected to perform better (than UCB and its variants) in environments where there are sudden changes in FOM.

[0269] ● ODAAF. Optimism for Delayed Aggregated Anonymous Feedback [Pike - Burke et al. Bandits with Delayed, Aggregated Anonymous Feedback, International Conference on Machine Learning (ICML), 2018] is a multi - armed bandit algorithm designed to function in settings where feedback is plagued by random bounded delays. The feedback is additionally aggregated and anonymized before being sent to the algorithm, making this setting very challenging. The algorithm proceeds in phases and maintains a set of candidates for the assignment of the optimal control settings. In each phase, it repeatedly executes a round - robin strategy among these candidates and updates the estimates of these performance metric values when it receives feedback. At the end of each phase, the algorithm eliminates candidates for which the estimated performance metric values are not significantly optimal.

[0270] ● Thompson Sampling. Thompson Sampling [Agrawal and Goyal. Analysis of Thompson Sampling for the Multi-armed Bandit Problem, Conference on Learning Theory (COLT), 2012] is a probability matching algorithm and another basic approach to solving the multi-armed bandit problem (the other is an optimism-based approach such as UCB). It works by maintaining the distribution of the estimated FOMs of all the allocation options for each control setting, sampling from each distribution, and then selecting the allocation option for the control setting with the highest sampled (estimated) FOM. When the true FOM is observed, the (posterior) distribution is updated using a Bayesian approach. The algorithm selects the allocation for each control setting in proportion to the probability that it is the best control setting allocation.

[0271] Figure 13 shows the performance of the described system compared to the performance of multiple other systems when controlling multiple different environments.

[0272] In particular, each of the other systems uses one of the corresponding existing control schemes described above to control multiple different environments.

[0273] The controlled environments each have three controllable elements with five possible settings, and the value of the performance metric optimized in each iteration is obtained from a Gaussian distribution.

[0274] The environments have various complexities due to the addition of various factors that cause variations among different procedural instances.

[0275] In particular, the basic environment shown in the top set of the graph changes the mean and variance of the Gaussian distribution according to the procedural instance, that is, different procedural instances can receive different performance metric values even when the same control setting is selected.

[0276] In other environments as well, time-based changes in the effects when various possible settings are applied to controllable elements, the basic sine-wave operation of performance metrics, and the various setting effects of various instance groups (i.e., representing the interaction between environmental characteristics and controllable elements) are introduced.

[0277] As can be seen from FIG. 13, many of the existing control schemes generally function well in the case of a simple baseline, and a given control scheme may work well with one additional factor of complexity, but none of the existing control schemes work well in all cases. On the other hand, the described system exhibits performance equal to or better than that of the best existing control schemes in all environments. Thus, the example of FIG. 13 shows the ability of the described system to automatically adapt to various complex environments for each different environment, without the need to manually select a model, i.e., by continuously changing the internal parameters of the system and considering the various properties of the various environments even when prior knowledge of the properties of the environment is not available, so that it can be run with possible settings similar to or better than those of other control schemes for each different environment.

[0278] A detailed description of each controlled environment is as follows.

[0279] ● 00_base. Three controllable elements with 100 procedural instances, each having five possible settings with performance metric values obtained from a Gaussian distribution. The selection of possible settings for different IVs changes the mean and / or standard deviation of the distribution. This environment is relatively simple but has many combinations of possible control settings, as is often seen in real-world environments.

[0280] ● 01_add_subject_var. Starting from 00_base, the procedural instance is divided into three groups with different base rate means and standard deviations of the distribution of these performance metric values. This introduces further variance into the data without changing the impact of the control setting assignment. This type of variance in procedural instances / EUs is very typical in the real world. For example, this particular configuration reproduces the sales behavior of a product assortment where a small group of products accounts for the majority of total sales (80 / 20 rule), the sales of more product groups are moderate, and the sales of most products are low.

[0281] ● 02_add_dynamic. Starting from 00_base, the effects of the possible settings of the IV change multiple times at a predetermined time (unknown to the algorithm) such that the impact of the possible settings of the IV is reversed. This changing behavior is very typical in the real world. For example, the effects of various advertising campaigns and technologies change periodically over space and time (what used to work may no longer work now). Similarly, the selection of the optimal control setting assignment on a manufacturing line changes due to factors such as temperature, humidity, and subtle differences in specific equipment (e.g., wear and tear).

[0282] ● 03_add_subject_var_dynamic. A combination of 01_add_subject_var and 02_add_dynamic. The combination of these two operations (described above) makes this environment more similar to many dynamic real-world environments.

[0283] ● 04_add_sine. Starting from 00_base, add the overall sine wave pattern to the performance metric values. This simulates the regular trends (e.g., seasonal, weekly) of the FOM that are independent of the influence of possible settings of the IV. In some algorithms, it is difficult to handle further data dispersion. This type of periodic behavior is very typical in the real world. For example, retail sales, supply chains, etc. often follow weekly, monthly, and seasonal cycles that bring significant dispersion to the performance metric. As another example, manufacturing and other processes affected by seasonal changes in weather may also have a similar effect. An important issue in these situations (which the described system addresses) is to be able to distinguish the influence of (e.g.) marketing activities from these basic behaviors.

[0284] ● 05_add_subject_var_sine. A combination of 01_add_subject_var and 04_add_sine. The combination of these two operations (described above) makes this environment more similar to a complex and dynamic real-world environment.

[0285] ● 06_add_ev_effects. The optimal combination of possible settings of the IV varies in some procedural instances. This variation in the assignment of control settings is very typical in real-world situations. For example, depending on the related products, recipients of the content, space, time, etc., different advertising or promotion approaches may work better than others.

[0286] ● 10_complex - A combination of 01_add_subject_var, 02_add_dynamic, 04_add_sine, and 06_add_ev_effects. This environment is most useful for capturing real-world behaviors in that it incorporates all of the above real-world operations and combines them into one environment.

[0287] FIG. 14 shows the performance of the described system compared to the performance of a plurality of other systems when controlling a plurality of different environments where the temporal effects vary.

[0288] In particular, each of the other systems uses a corresponding existing control scheme to control a plurality of different environments.

[0289] The environments being controlled each have four controllable elements with two possible settings each, and the value of the performance metric for each iteration is obtained from a Gaussian distribution. The environments have various time delays and durations, which affect the point in time at which the performance metric value is generated in relation to the initial application of the control settings for a given instance. For example, in the top environment, the environmental response to all effects is delayed by two iterations and persists for three iterations. In the next environment, all four controllable elements have different time delays and durations. The third and fourth environments add further complexity and variability.

[0290] As can be seen from the example of FIG. 14, the described system can execute with possible settings that are the same as or better than those of other control schemes for each of the various environments. This shows the ability of the described system to adapt dynamically to the temporal behavior of the effects of applying the control settings, i.e., by changing the temporal range parameters during operation.

[0291] Furthermore, the two environments include basic periodic operations that are not related to the assignment effect of the IV control settings. This operation is a typical situation encountered in the real world (e.g., advertisements, pharmaceuticals, etc.) where the executed action is delayed and does not have an immediate effect. At the same time, in such scenarios, there are often residual effects that continue after the assignment of the control settings has been aborted. Furthermore, it is rare to find these temporal operations alone. Rather, they mostly occur simultaneously with a basic operation similar to the sine wave pattern shown. As can be seen from Figure 14, the system described is superior to conventional systems. That is, by adjusting the temporal range parameter and other internal parameters to adapt to changes in basic behavior, different temporal operations can be better explained.

[0292] The details of the environment shown in Figure 14 are as follows.

[0293] ● 00_temporal. Four controllable elements with 500 procedural instances, each having two possible settings with performance metric values obtained from a Gaussian distribution. The selection of possible settings for different IVs changes the mean and / or standard deviation of the distribution. The performance metric values of all effects are delayed by two repetitions and persist for three repetitions.

[0294] ● 01_temporal_multi. The same as 00_temporal, except that the time delays and periods of the four controllable elements are different.

[0295] ● 02_temporal_sine. Starting from 00_base with sine wave operation added.

[0296] ● 03_temporal_delay_only. The same as 00_temporal, but the continuous operation is removed.

[0297] ● 04_temporal_multi_delay_only. The same as 01_temporal_multi, but the continuous operation is removed.

[0298] ● 05_temporal_sine_delay_only. It is the same as 02_temporal_sine, but continuous operation is removed.

[0299] Figure 15 shows the performance of the described system with and without clustering. The controlled environment has three controllable elements, each with five possible settings, and the value of the performance metric for each iteration is obtained from a Gaussian distribution fixed through the experiment. Depending on the characteristics of the procedure type instance / EU described by the environmental characteristics, there are different optimal control setting assignments (controllable elements) for the controlled environment. One set of control setting assignments produces generally good results but is actually negative for a sub-population. When a particular ideal control setting assignment is given to the sub-population, the overall usefulness is improved. This is typical of real-world situations where the optimal control setting assignment can vary significantly based on external characteristics. The left figure shows the performance of the described system that includes a clustering component. In this case, the described system assigns a particular control setting assignment to the procedure type instance / EU, which results in a higher overall figure of merit (FOM). The right figure shows the performance of the described system when the clustering component is not used, i.e., when it never enters the clustering phase. In this case, the algorithm uses a single overall control setting assignment approach for all procedure type instances. This results in the use of a non-optimal control setting assignment for a particular sub-population. As can be seen from Figure 15, the described system functions better when clustering is used.

[0300] FIG. 16 shows the performance of the described system with the ability to change data inclusion compared to the performance of the described system that controls the same environment while keeping the data inclusion window parameters fixed. In the example of FIG. 16, the controlled environment shows two gradual changes in the relative effect of the control settings on the performance measurement values. This is typical of the real world in two ways: 1) it is rare for the effects of actions (e.g., advertisements, manufacturing parameters, etc.) to always be static, and 2) when such changes occur, they are often essentially gradual and not sudden. The left figure shows the performance of the described system that includes the DIW component. In this case, the described system can quickly detect that the effect has changed, for example, through a hybrid baseline comparison, and the described system can immediately relearn the assignment of the best control settings by shrinking the data inclusion window. The right figure shows the performance of the described system when not using the DIW component. In this case, the algorithm adapts very gradually to changes in the treatment effect. By the time it does, the effect has already changed again.

[0301] Figure 17 shows the performance of the described system with and without time analysis, i.e., with and without the ability to change the time range. The controlled environment has four controllable elements, each with two possible settings, and the value of the performance metric for each iteration is obtained from a Gaussian distribution that is fixed through the experiment. The environment has various time delays and carry-over operations that affect the point in time at which the performance metric values are generated in relation to the initial application of the possible settings of the IV. Furthermore, the two environments include basic periodic operations that are not related to the effect. This operation is a typical situation encountered in the real world (e.g., advertising, pharmaceuticals), and often the actions taken have no immediate effect and often the effect remains even after the assignment of the control settings has been aborted. Furthermore, this temporal variability often exists in the context of other basic operations. This figure shows the value of the temporal optimization within the described system. The left column shows the performance of the described system when using the temporal component. The right column shows the performance of the described system when not using the temporal component. As can be seen from the example of Figure 17, the described system performs significantly better when using temporal analysis when these temporal characteristics are present in the environment.

[0302] Figure 18 shows the performance of the described system when controlling an environment, compared to the performance of a system that controls the same environment using an existing control scheme (“Lin UCB”). In the example of Figure 18, the environment being controlled has a periodic underlying behavior independent of the possible setting effects of the IV, along with changes in these effects, such that the assignment of optimal control settings changes over time. These characteristics are similar to those seen in many real-world environments, with periodic underlying dynamics (e.g., weekly, monthly, or seasonal patterns) along with changes over time in the impact of control setting assignments / actions. Figure 18 shows a subset of the time during which the impact of the possible settings of the IV changes in the underlying environment (between iterations 200 and 250). As can be seen from Figure 18, the performance of the existing control scheme remains in the exploitation phase based on the effects of previous control setting assignments and cannot adapt quickly to changes. On the other hand, the performance of the described system quickly adapts to the changing effects and finds stepwise improvements under the changed environmental effects (upper plot). This results in an increase in the stepwise benefit from the described system (lower plot). Note that the cumulative benefit from using the described system continues to increase over time.

[0303] In the foregoing description, specific terms have been used to refer to the functions of the described system or the actions performed by the described system, but it should be understood that these are not the only terms that can be used to describe the operation of the system. Some examples of alternative terms are given below. As an example, a controllable element may alternatively be referred to as an independent variable (IV). As another example, an environmental characteristic may alternatively be referred to as an external variable (EV). As another example, an environmental response may alternatively be referred to as a dependent variable (DV). As another example, a procedural instance may alternatively be referred to as an experimental unit or a self-organizing experimental unit (SOEU). As another example, a possible setting of a controllable element may alternatively be referred to as a level of the element (or IV). As yet another example, a control setting may alternatively be referred to as a process decision, and assigning a control setting to a procedural instance may alternatively be referred to as a process assignment.

[0304] The term "repeatedly" is used herein to generally mean that an action occurs multiple times regardless of a particular sequence, in the context of repeatedly performing an action. As an example, a process can follow a set of steps in a specified order, either continuously or repeatedly, or can follow the steps randomly or discontinuously. Further, not all steps are necessarily performed with the same frequency. For example, a process assignment may be performed more frequently than updating causal learning, and the latter frequency may change over time, e.g., as the exploitation phase becomes dominant and / or as computing capacity / speed requirements change over time.

[0305] In this specification, the term "configured" is used in relation to systems and computer program components. In the case of a system of one or more computers configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof is installed in the system and that the system performs the operation or action during operation. For one or more computer programs configured to perform a particular operation or action, it means that the one or more programs include instructions that cause the device to perform the operation or action when executed by a data processing device.

[0306] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of their combinations. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs. One or more computer programs can include one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded in an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus.

[0307] The term "data processing apparatus" refers to data processing hardware and encompasses any kind of apparatus, device, and machine for processing data, including, for example, one programmable processor, one computer, or a plurality of processors or computers. The apparatus may also be, or further include, dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Optionally, in addition to the hardware, the apparatus may include code for generating an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0308] A computer program may also be referred to as a program, software, software application, app, module, software module, script, or code, and may be described in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages, and may be deployed in any form, as a stand-alone program or including modules, components, subroutines, or other units suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program may be stored as part of a file that holds other programs or data, for example, in a markup language document, in a single file dedicated to the relevant program, or in one or more scripts stored in a plurality of coordinated files (for example, files that hold one or more modules, subprograms, or portions of code). A computer program may be executed on one computer or provided at one site, or may be deployed so as to be executed on a plurality of computers distributed across a plurality of sites and interconnected via a data communication network.

[0309] As used herein, the term "database" is widely used to refer to any collection of data, where the data need not be structured in any particular way, or structured at all, and can be stored in a memory device in one or more locations. Thus, for example, an index database can include multiple data collections, each of which can be organized and accessed in a different way.

[0310] Similarly, as used herein, the term "engine" is widely used to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine. In other cases, multiple engines can be installed and run on the same one or more computers.

[0311] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to function based on input data and generate output. The processes and logical flows can also be performed by, for example, a dedicated logic circuit such as an FPGA or ASIC, or by a combination of a dedicated logic circuit and one or more programmed computers.

[0312] A computer suitable for the execution of a computer program can be based on a general-purpose or special-purpose microprocessor, or both, or other types of central processing units. Generally, the central processing unit receives instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a central processing unit for executing or running instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be complemented by, or incorporated in, dedicated logic circuitry. Generally, a computer also includes, or is operably coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from, or transmit data to, or both, these devices. However, such devices are not necessarily required for a computer. Further, a computer can be embedded in another device, such as, by way of example only, a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.

[0313] Examples of computer-readable media suitable for storing computer program instructions and data include, but are not limited to, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0314] To provide interaction with a user, the described embodiments herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, a keyboard and a pointing device, such as a mouse or trackball, that enable the user to provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form including acoustic input, voice input, or tactile input. Additionally, the computer can interact with the user by sending documents to the devices used by the user and receiving documents from those devices, such as by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user in reply.

[0315] Embodiments of the subject matter described herein can be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a graphical user interface, a web browser, or an app that enables a user to interact with an implementation of the subject matter described herein, or a client computer having any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs) (e.g., the Internet).

[0316] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact over a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, such as an HTML page, to a user device for the purpose of receiving user input from a user, e.g., to display data and interact with a device that functions as a client. Data generated on the user device, e.g., as a result of user interaction, can be received on the device from the server.

[0317] This specification includes many specific implementation details, but these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may be described and even initially claimed as functioning in a particular combination, but one or more features from the claimed combination may, in some cases, be excised from the combination, and the claimed combination may be the subject of a sub-combination or variation of a sub-combination.

[0318] Similarly, operations are shown in the drawings and recited in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or sequence shown in order to achieve desirable results, or that all of the operations shown be performed. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0319] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. selecting a configuration of input settings for manufacturing a batch of a biological pharmaceutical based on a causal model that measures a current causal relationship between the input settings and measured values of the quality of the batch of the biological pharmaceutical; determining the measured values of the quality of the batch of the biological pharmaceutical produced using the configuration of the input settings; adjusting the causal model based on the measured values of the quality of the batch of the biological pharmaceutical; and repeatedly performing the above, a method.

2. selecting the configuration of the input settings includes selecting the configuration of the input settings based on a set of internal control parameters, and the method further includes adjusting the internal control parameters based on the measured values of the quality of the batch, the method according to claim 1.

3. the measured values of the quality of the batch include the yield of the batch, a measured value of the activity of the batch, a measured value of the effectiveness of the batch, a measured value of the impurities of the batch, a measured value of the viability of the transfected cells, a measured value of the transfection efficiency, a measured value of the purity of the biological pharmaceutical, a measured value of monodispersity, or a measured value of cell density, including one or more of the above, the method according to claim 1 or 2.

4. the purity of the biological pharmaceutical includes a measured value of host cell protein impurities, a measured value of DNA impurities, a measured value of RNA impurities, or a measured value of host cell metabolite impurities, including one or more of the above, the method according to claim 3.

5. the input settings include one or more settings related to processing steps, one or more settings related to the chemical composition of substances, one or more settings related to physical control during manufacturing, or one or more settings related to the purification of the batch of the biological pharmaceutical, including one or more of the above, the method according to any one of claims 1 to 4.

6. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ Temperature during manufacturing, Temperature during purification, Humidity during manufacturing, Humidity during purification, Flow rate used during manufacturing, Flow rate used during purification, G-force used during manufacturing, Magnetic measurement values used during manufacturing, Electrical measurement values used during manufacturing, or, Agitation measurement values used during manufacturing, The method according to any one of claims 5 to 7, comprising one or more of the above.

9. The settings related to the purification of the batch of the biological pharmaceutical, Electrostatic interaction, Hydrophobic interaction, Hydrophilic interaction, Affinity chromatography, or, Size exclusion chromatography, The method according to any one of claims 5 to 8, comprising generating the batch using one or more of the above.

10. Selecting the configuration includes selecting the configuration based on a predetermined set of the causal model and external variables, The method further includes adjusting control parameters that parameterize the influence of the predetermined set of external variables, according to any one of claims 1 to 9. [[ID= ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Fermentation pharmacy product quality soft measurement method based on GP learning model building

    CN101419216A

  • Probability model estimation device, probability model estimation method, and program

    JP2017204155A

  • Patient-specific immunotherapy to treat heterogeneous tumors

    JP2018500275A

  • Manufacturing biological medicines using causal models

    JP2022526260A

  • Systems and methods for disease control and management

    US20080306770A1