Safe Adaptive Machine Learning-Based Control for Cryogenic Atmospheric Pressure Plasma Biomedical Devices
The MOBO strategy addresses the challenges of data efficiency and safety in biomedical devices by adapting DNN-based control policies for CAP jets, enabling personalized and efficient plasma therapy with improved treatment times and safety.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- UNIVERSTIY OF CALIFORNIA
- Filing Date
- 2024-04-11
- Publication Date
- 2026-04-14
AI Technical Summary
Biomedical devices with multivariable, nonlinear, and fluctuating dynamics face challenges in achieving efficient and safe control policies due to limited data availability and the need for kHz sampling rates, especially in cold atmospheric plasma therapy, where interaction with humans requires stringent safety constraints.
A multi-objective Bayesian optimization (MOBO) strategy is employed to adapt DNN-based control policies using probabilistic surrogate models and global sensitivity analysis, enabling data-efficient and globally optimal adaptation of control policies for CAP jets, incorporating safety constraints and individual variability.
The MOBO approach allows for personalized and safe plasma therapy by efficiently adapting control policies to individual subjects, achieving faster treatment times with reduced constraint violations, thus enhancing therapeutic effects while ensuring patient safety and comfort.
Smart Images

Figure 2026512146000001_ABST
Abstract
Description
[Technical Field]
[0001] Statement on federally funded research and development This invention was made with government support under grants 1912772 and 2130734 from the U.S. National Science Foundation. The government has certain rights to this invention.
[0002] This invention relates to plasma medicine, and more specifically to controlling processes using model predictive controllers, deep learning controllers, and Bayesian optimization methods. [Background technology]
[0003] Biomedical devices often involve safety-critical constraints and can exhibit multivariable, nonlinear, and inherently fluctuating dynamics. Consequently, there is growing interest in advanced control of biomedical devices using model predictive control (MPC) and learning-based control strategies.
[0004] MPC generally includes control techniques that aim to minimize the cost function of a dynamic system with operational constraints over a finite receding horizon, where the calculated control actions are performed. Receding horizon generally refers to the change in the horizon at each time step or other iteration.
[0005] The concepts of learning, adapting, and autotuning control policies using closed-loop performance data have attracted attention. Reinforcement Learning (RL) is trained based on rewarding desirable behaviors and penalizing undesirable behaviors. In some cases, policy gradient methods are used as a common RL approach to guide policy search within continuous control-input spaces, and they have achieved remarkable results particularly in MPC policies.
[0006] Some people consider policy gradient RL to be a scalable alternative to the Bayesian Optimization (BO) strategy, which is becoming popular for controller autotuning, in order to utilize gradient information. However, it may come with the cost of reduced data efficiency, especially when initialization is insufficient.
[0007] Another consideration includes the safety of biomedical systems. The progress of data-driven control and decision-making capabilities has created important opportunities regarding the autonomy of biomedical devices. The fundamental feature of autonomy is to autonomously optimize what decisions to make next. Due to the interaction with humans, safety is a basic requirement in autonomous systems.
Brief Description of the Drawings
[0008] [Figure 1]FIG. 1 shows a schematic diagram of an embodiment of a system for a Cold Atmospheric Plasma (CAP) jet.
[0009] [Figure 2] FIG. 2 shows a graph of hypervolume improvement and observed Pareto frontier in the context of thermal dose delivery using a CAP jet.
[0010] [Figure 3] FIG. 3 shows a graph of the state of a closed-loop experiment and input profiles in iterations of an embodiment.
[0011] [Figure 4] FIG. 4 shows a graph of performance measures observed from exploration of an embodiment.
[0012] [Figure 5] FIG. 5 shows an exemplary feasible set and queries created using safe Bayesian optimization.
[0013] [Figure 6] FIG. 6 illustrates an observed closed-loop profile obtained from safe Bayesian optimization.
BRIEF DESCRIPTION OF THE DRAWINGS
[0014] This embodiment includes a strategy for approximate Model Predictive Control (MPC) based on adaptive deep learning for personalized and point-of-care biomedical plasma applications. Approximate MPC relies on approximating MPC laws via offline computation of an optimal control problem, enabling control of biomedical devices at kHz sampling rates. Approximating MPC laws based on deep neural networks (DNNs) is particularly attractive because it has a small memory footprint and allows for versatile embedded implementations on resource-constrained dedicated hardware such as field-programmable gate arrays (FPGAs).
[0015] This embodiment involves approximating an MPC using a DNN system in an environment with limited trial counts and datasets. There are several applications for controlling a “local” process using a DNN approximation of an MPC. As shown in Figure 1, the system includes one or more processors that receive data from a biomedical device resulting from each “run” or use of the device. The system provides run-to-run data obtained from multiple uses to one or more processors that execute code to perform Bayesian optimization. The processors then adapt a subset of the DNN system's parameters. Subsequently, the DNN system provides settings and also controls the plasma biomedical device in real time for the next run. Due to the nature of plasma biomedical applications, where the number of runs to generate data to provide to the DNN is limited, the nature of the interaction with the DNN differs significantly from that of typical machine learning applications, which have the advantage of training with large datasets.
[0016] While this type of architecture originated from specific solutions for the application of CAP jets to surfaces for wound care and other types of biomedical applications, this methodology and architecture have broader applications in cold atmospheric plasma devices.
[0017] Cold atmospheric plasmas (CAPs) have recently found promising applications in plasma medicine. CAPs are low-temperature, (partially) ionized gases that can be generated by applying an electric field to noble gases such as argon or helium, and the resulting discharge can be directed towards a target surface. The synergistic effects of CAPs, including the generation of reactive chemical species and ions, ultraviolet radiation, low-level electric fields, and thermal effects, can induce therapeutic outcomes. For this reason, portable CAP devices are considered promising for various point-of-care biomedical applications. Therefore, as described in "Mathematical modeling and control for cancer treatment with cold atmospheric plasma jet" by Y. Lyu, L. Lin, E. Gjika, T. Lee, and M. Keidar (Journal of Physics D: Applied Physics, vol. 52, no. 18, p. 185202, 2019), there is growing interest in the advanced control of biomedical CAP devices using model predictive control (MPC). Furthermore, as described in "Machine learning controlled self-adaptive plasma medicine" by L. Lin and M. Keidar (Proceedings of the IEEE International Conference on Plasma Science, pp. 561-561, 2020), there is growing interest in learning-based control strategies.
[0018] Two major challenges in CAPs' MPC stem from the following needs: (i) As described in "Fast approximate learning-based multistage nonlinear model predictive control using Gaussian processes and deep neural networks" by AD Bonzanini, JA Paulson, G. Makrygiorgos, and A. Mesbah (Computers & Chemical Engineering, vol. 145, p. 107174, 2021), it is necessary to address the fast dynamics of CAPs, and consequently, to support kilohertz (kHz) sampling rates. (ii) As stated in "Perspectives on machine learning-assisted plasma medicine: Toward automated plasma treatment" by AD Bonzanini, K. Shao, A. Stancampiano, DB Graves, and A. Mesbah (IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 6, no. 1, pp. 16-32, 2021), it is necessary to adapt the MPC policy parameters to take into account variations in the properties of the plasma and target surface.
[0019] In the field of autotuning control policies, policy gradient descent is used as a common reinforcement learning (RL) approach to guide policy search in a continuous control input space, and has shown particular success with MPC policies (for example, "Safe reinforcement learning using robust MPC" by M. Zanon and S. Gros (IEEE Transactions on Automatic Control, vol. 66, no. 8, pp. 3638-3652, 2020)). Because it uses gradient information, policy gradient RL is touted as a scalable alternative to Bayesian optimization (BO) strategies, which are gaining popularity for controller auto-tuning (e.g., "A data-driven automatic tuning method for MPC under uncertainty using constrained Bayesian optimization" by F. Sorourifar, G. Makrygiorgos, A. Mesbah, and JA Paulson, IFAC-PapersOnLine, vol. 54, no. 3, pp. 243-250, 2021). However, it comes at the cost of reduced data efficiency, especially when initialization is insufficient.
[0020] Instead, especially when interaction with performance data and / or the real environment is limited, BO can be a viable alternative to data-efficient policy exploration. BO is a derivative-free, probabilistically principled method for "global" optimization, capable of handling a mixture of continuous, discrete, and categorical decision variables. For example, in "Virtual vs. real: Trading off simulations and physical experiments in reinforcement learning with Bayesian optimization" (Proceedings of the IEEE International Conference on Robotics and Automation, pp. 1557-1563, 2017) by A. Marco, F. Berkenkamp, P. Hennig, AP Schoellig, A. Krause, S. Schaal, and S. Trimpe, an entropy-search BO approach was presented for performing sample-efficient policy learning on actual physical systems using prior information from an "inexpensive" simulation environment. Furthermore, when it is necessary to discover an optimal set of policies for multiple conflicting objectives, the multi-objective nature of policy search can be directly incorporated into BO.
[0021] In plasma therapy of complex interfaces, it is essential to adapt control policies to account for the time variability of plasma and surface properties, as well as the variability between different target surfaces. Furthermore, the adaptability of treatment policies is crucial in personalized plasma medicine, and CAP therapy must be tailored to each individual subject to enhance therapeutic effects without compromising patient safety and comfort. However, a major challenge stems from the limited number of treatments / trials that can be conducted in the biomedical context, making data efficiency a prerequisite for policy adaptation. CAPJ behavior is applicable beyond the biomedical context; the biomedical context is merely one example.
[0022] To this end, this embodiment presents a multi-objective BO (MOBO) strategy for data-efficient and globally optimal adaptation of DNN-based control policies in a run-to-run manner. MOBO uses probabilistic surrogate models of multiple closed-loop performance measures, such as plasma therapy outcomes, to systematically trade off between exploration and exploitation of subsets of DNN parameters. The selection of this parameter subset is guided by a global sensitivity analysis that quantifies the impact of each network parameter on the performance measure. Thus, MOBO provides a data-efficient scheme for performance-oriented adaptation of DNN-based control policies. In this embodiment, as an example of a biomedical device having a local control process managed by a DNN system, we experimentally demonstrate a proposed strategy for adaptive DNN-based approximate MPC for a CAP jet (CAPJ) with a prototype application in the processing of heat-sensitive biomaterials.
[0023] In the context of personalized plasma therapy, control problems for prototype CAPJs can be considered. As described in "Effective dose delivery in atmospheric pressure plasma jets for plasma medicine: A model predictive control approach" by D. Gidon, DB Graves, and A. Mesbah (Plasma Sources Science and Technology, vol. 26, no. 8, p. 085005, 2017), one embodiment uses a kHz-excited CAPJ consisting of a copper ring electrode wrapped around a quartz tube and heated in helium (He). Figure 1 shows a schematic diagram of the CAPJ system 10. A computer 30 controls the gas flow to a mass flow controller 40. In one embodiment, the gas is helium, and as the gas flows through the tube, plasma ignition is achieved by applying a high-frequency alternating current (AC) voltage 18 to the copper electrode 16 under the control of an amplifier 20. Plasma is directed from the tube 14 onto the target substrate. In one embodiment, the target substrate is a grounded, glass-covered metal plate 24 located 3 mm below the tip of the tube. The applied power P and gas flow rate q are manipulated inputs from a computer 30 having a processor 32 and memory 34. The manipulated inputs correspond to operating parameters, which are the evaluation results of the DNN. The maximum surface temperature T and total plasma light intensity I at the plasma-surface incident point are measured outputs from an IR camera 28 and a spectrometer 26, respectively. The IR camera and spectrometer represent different types of sensors that measure one or more properties of the plasma and surface.
[0024] Using data collected from CAPJ, the system dynamics can be modeled as follows using a linear time-invariant (LTI) state-space model.
number
number
[0025] Plasma therapy on complex surfaces relies on quantifying the plasma effect delivered to the surface. Here, as described in "Thermal dose determination in cancer therapy" by S.A. Sapareto and WC Dewey (International Journal of Radiation Oncology Biology Physics, vol. 10, no. 6, pp. 787-800, 1984), this process explains the accumulation of thermal effect on the target using a metric called cumulative equivalent minutes (CEM), as given by equation (2) below.
number
number
[0026] The goal of plasma treatment is to deliver the desired amount of plasma effect as quickly as possible without violating comfort and safety constraints. This embodiment aims to systematically consider the inherent uncertainties of CAPJ using the formulation of robust MPC, and for this purpose, scenario-based MPC (sMPC) is used. sMPC assumes that the system uncertainty is represented by a tree of discrete scenarios, and each branch originating from a node represents a specific scenario of uncertainty realization. Furthermore, to limit the number of scenarios, a robust horizon N r is often defined, bounding the propagation of uncertainty to a predetermined point. In this embodiment, the "worst-case" formulation of the scenario tree is selected, and the scenarios are generated based on the bounds of the worst-case range of uncertainty. To represent the trajectories generated by S scenarios, the notation (x j (i), u j (i)) is adopted in the following discussion. Here, the addition of the superscript j indicates a specific scenario j ∈ {1,..., S}. Therefore, the optimal control problem at time step k is formulated as follows.
Equation
number
[0027] Finally, in order to adapt policies for individualized CAP treatment, this process focuses on two closed-loop performance metrics: (i) delivery of the desired thermal dose and (ii) adherence to comfort / safety constraints. In this discussion, (i) is defined as the CEM tracking cost over treatment time N, as shown in (7) below. (ii) is defined as the sum of the degree of constraint violations in surface temperature over treatment time N, as shown in (8) below.
number
[0028] The requirements for embedded control for point-of-use CAPJ applications on low-cost and resource-constrained hardware present a major challenge in the online deployment of the sMPC law (5). This challenge stems from the high computational cost of scenario tree optimization in (4). For this purpose, this process approximates (5) using a DNN. We consider a dataset of the form shown in equation (9) below.
number
number
number
[0029] Generally, DNN parameters are fitted by minimizing the mean-squared error loss function. Meanwhile, DNN hyperparameters and training / fitting options (e.g., learning rate, optimizer solver) related to the architecture must be tuned. Examples of DNN hyperparameters related to the DNN architecture are shown below.
number
[0030] The adaptation of the control policy involves M closed-loop performance metrics φ as shown below. m It can be cast as a multi-objective (MO) problem characterized by the following:
number
number
number
number
number
[0031] If the performance metric is expensive and a black box, problem (11) cannot be solved directly. Therefore, a common approach in BO is to learn a stochastic surrogate model (typically a Gaussian process (GP) model) for the performance metric and select a set of points that jointly optimize the expected value of the current surrogate. This is done by solving a proxy problem that proposes points that the acquisition function queries to improve the surrogate representing the performance metric. The querying strategy is based on the search / exploit trade-off. That is, in this embodiment, the goal is to query the metric at points located in a neighborhood that may contain the optimal condition while reducing the prediction uncertainty of the surrogate model. Given newly observed data D shown below, each performance metric is updated using Bayesian inference. For example, GP regression is used for the GP surrogate.
number
[0032] Furthermore, in MO settings, there is no single best optimizer because performance metrics can be conflicting. Therefore, the goal is to discover a Pareto frontier consisting of a set of optimal points, i.e., Pareto optimal points. A Pareto frontier is a boundary in the performance metric space that improves one performance metric at the cost of degrading other metrics. The set of Pareto optimal solutions is denoted as P. Solutions included in P are known to be non-dominated by other solutions within the feasible region. For the application of control policies, Pareto dominance is defined as follows:
[0033] Definition 1: Given a set of parameters and corresponding performance metrics {θ, ▲h▼(θ)} as shown below, h i (θ A )≦h i (θ B ) holds for all i∈{1,...,M} and h in ∃i∈{1,...,M} i (θ A ) <h i (θ B When this is the case, the solution is ▲h▼(θ A ) is an alternative solution ▲h▼(θ B ) governs. Pareto dominance is shown as follows:
number
number
number
[0034] Establishing a Pareto frontier allows for the selection of an optimal control policy, each of which yields optimal performance with varying levels of trade-offs between performance metrics. In MOBO, the search for the Pareto frontier is typically facilitated by an expected hypervolume improvement (HVI) acquisition function. The expected HVI depends on the definition of an indicator known as hypervolume (HV), which quantifies the Pareto optimality of the estimated Pareto frontier ("Parallel Bayesian optimization of multiple noisy objectives with expected hypervolume improvement" by S. Daulton, M. Balandat, and E. Bakshy, "Advances in Neural Information Processing Systems, vol. 34, pp. 2187-2200, 2021").
[0035] Definition 2: HV is defined in the performance indicator space with respect to the reference point r shown below.
number
number
[0036] HV acts as a metric that quantifies the quality of the Pareto frontier and is influenced by the selection of reference points. Therefore, "convergence" to a single HV value means that MOBO performs sufficient sampling (based on some pre-specified budget) and constructs the best possible Pareto frontier, but this Pareto frontier is not necessarily a true Pareto frontier. Next, HVI represents the incremental improvement of P's HV when new points are added. The HVI of a newly observed set of indicators Ψ′ is given by equation (17) shown below.
number
number
[0037] Firstly, this process solves the sMPC problem (4) in a closed loop and collects training data for approximating the initial control law (5). In (4), this process predicts the horizon N p =5, Robust Horizon N r =2, and discrete uncertainty scenarios are [0.01w min ,0,0.01w max The control input is constrained by P∈[1.5,5]W and q∈[1.5,5]SLM, and the state is constrained by T∈[25,45]℃ and I∈[20,80] arbitrary units. sMPC was formulated using CasADi and solved using IPOPT. The inventors addressed the discrepancy between the plant and the control model and the normally distributed index noise N(0,(0.1) 2 The inventors simulated a true system with ). Sample n of the mapping from state to optimal input. sWe collected 5,000 data points and trained a fully-connected feedforward DNN architecture with L=4, H=7, and the ReLU activation function. The DNN was trained for 5,000 epochs using PyTorch with default optimizer settings. The resulting DNN-based policy achieved performance nearly equivalent to the implicit sMPC law. Furthermore, the computation time of the DNN was approximately three orders of magnitude faster on a standard CPU (2.4GHz quad-core Intel i5 processor) compared to solving (4), although this depended on the DNN architecture (~10). -5 s vs ~10 -2 s).
[0038] This process aims to cater to the individual needs of subjects by considering the treatment of subjects with characteristics different from the mean population values, and by adapting the initial policy designed for the population mean. The closed-loop performance index is based on subject-specific characteristics, namely K in the CEM setpoint tracking cost (7) and T in the comfort constraint cost (8). tol Note that it is parameterized by , where the parameter of the population mean is K pop =0.5 and T tol,pop =45℃, while the characteristics of individual subjects are K indiv =0.55 and T tol,indiv It is possible to verify the case where the temperature is 44.5℃.
[0039] First, it is necessary to examine the sensitivity of DNN-based policies to perturbations in subsets with different parameters. Given that the desire for individualized treatment is to minimize the number of trial-and-error treatments, a subset of DNN parameters was applied because the DNN parameters are high-dimensional (212), as shown below.
number
[0040] To evaluate this approach, the sensitivity of closed-loop performance metrics to the parameters of the first and last layers of a DNN-based policy was verified. In this process, sensitivity analysis tools from UQLab
[33] were used to perform the GSA. In this embodiment, the global sensitivity of selected DNN parameters to closed-loop performance metrics was analyzed using a moment-independent method (i.e., Borgonovo indices). The inventors generated 10,000 samples of 44 parameters encapsulated by the first and last layers of a 4-layer, 7-node DNN. The samples were selected from geometrically-bound values from the initial policy parameters. In this process, for each sample, three triplicate closed-loop simulations using the DNN-based policy were performed, and the mean plus and minus standard error values of the observed closed-loop metrics were evaluated. Table I summarizes the GSA results. [Table 1] Generally, the dose delivery index (7) is less sensitive to parameter changes compared to the temperature constraint index (8). Overall, both indices have equivalent sensitivity to all parameters selected for the GSA. Nevertheless, the parameters of the last layer have a slightly higher influence with fewer parameters. Therefore, this embodiment selects the last layer of the DNN-based policy (i.e., θ=W) as a subset of parameters to modify in the policy discovery procedure. L+1 nθ =16).
[0041] As a global optimization method, MOBO provides a means (means) for systematically exploring and detecting trade-offs between competing performance metrics. Figure 2 shows the results of 5 replications of 50 iterations of MOBO for a simulated CAPJ. Here, one iteration of MOBO is N d It consists of replications of =3. To implement MOBO, this process used Ax. Ax works in conjunction with BoTorch to perform BO, and BoTorch, in turn, works with GPyTorch for surrogate modeling using GPs. These tools were mainly used with default settings, using the Matern5 / 2 kernel for GPs, and an EHVI acquisition function with noise.
[0042] The HV profile in Figure 2(a) indicates the "convergence" of MOBO. The decrease in HV improvement after more than 30 iterations indicates that MOBO has achieved a somewhat optimal representation of the Pareto frontier illustrated in Figure 2(b). Although it takes more than 20 iterations to reach this Pareto frontier, the initial policy can be significantly improved in the first few iterations of MOBO. The rapid increase in HV suggests that the parameterization of the initial policy is suboptimal, and further suggests that even starting from a suboptimal solution, a new Pareto optimal point can be found in the first few iterations.
[0043] In the experimental demonstration of the proposed approach to CAPJ illustrated in Figure 1, sMPC is N p =N r= 2, and the input bounds were adjusted to P ∈ [1.5, 3.5]W and q ∈ [3.5, 7.5]SLM. Subsequently, 11 closed-loop experiments were performed to collect training data for DNN approximation, and as a result, n s A total of 1,378 samples were obtained. The DNN was trained with the same structure and procedure as described in the simulation study and achieved closed-loop performance similar to that of the implicit sMPC. Due to a limited budget of 45 treatments, a total of 15 iterations were performed in MOBO.
[0044] Figure 3 shows the state and input profiles for three specific iterations of the MOBO. The inventors have taken N during each iteration to account for system-specific variability. d We conducted 3 replicated real-time experiments. The CEM profile is plotted with the min-max bound of the three replicates represented by a shaded area. The solid line represents the median of the three runs. The temperature profile represents the mean of the three experiments (solid line) and the mean plus and minus two standard errors (shaded area). Both input profiles are plotted using the mean values obtained from the three experiments. The profiles shown in Figure 3 were determined to be some of the "best" treatment options encountered through the MOBO process.
[0045] In this case, "best" can be explained in one of the following two ways: (i) Where there was insufficient data to establish a clear Pareto frontier, the best treatment was considered to be the initial policy. (ii) Once the estimated Pareto optimal point was found, the best treatment was the policy that minimized the constraint violation. This series of treatment protocols follows intuition regarding “safe” treatment. As in (i), the initial treatment is considered safe for the general population and is considered “best” for the time being. In (ii), once the Pareto frontier is established, the treatment can then be switched to a more optimal treatment at the cost of minor temperature violations. Note that individualization of plasma treatment is possible by establishing trade-offs between different performance metrics via the estimated Pareto frontier.
[0046] From Figure 3, the initial "best" profile is the initial profile (solid blue line). After iteration 3, a new "best" profile (dashed orange line) is found. The dashed orange profile represents a new parameterization of the policy that surpasses the initial solid blue policy, achieving CEM faster (reducing median treatment time by approximately 8 seconds or 13%) with a slight constraint violation. After further MOBO iterations, a new policy (dotted green line) is discovered in iteration 12. In this case, CEM delivery is on average comparable to the orange (dashed) policy (reducing median treatment time by 10 seconds or 16%) while maintaining lower constraint costs.
[0047] Furthermore, in Figure 3, the He flow rate tends to saturate during treatment. Generally, higher He flow rates result in lower temperatures. Therefore, due to the temperature tolerance specification, the CAPJ requires a higher flow rate to remain within the desired operating temperature range. In Figure 4, the locations of these "best" points in the performance index space are indicated by red boxes. While it may not be possible to visually establish the Pareto frontier with a small number of iterations, some points can still be identified as Pareto optimal due to their proximity to the minimization of both performance indices. Figure 4 is consistent with the profile in Figure 3 in that Pareto optimal points are found in iterations 3 and 12 (indicated by red boxes). Thus, a strategy has been established that allows for trade-offs between multiple performance indices to tailor treatment to individual subjects.
[0048] One system constraint that may not exist in other situations but is present in biomedical devices is one related to patient safety. Here, we consider the problem of Bayesian optimization (BO) for systems with safety-critical constraints. Recent research has shown that a theoretically consistent way to consider constraints in BO without losing the guarantee of theoretical global optimality is to relax the constraint functions so that the feasible region has a high probability of containing the global solution. However, constructively, these approaches cannot guarantee safe and feasible behavior for all queries, which is unacceptable in safety-critical applications.
[0049] On the other hand, the safe-boost method forces query points to remain within a partially revealed safety zone, which can result in unacceptable (and unquantifiable) performance losses. The performance of safe-boost has been demonstrated for the safe control of cryogenic atmospheric pressure plasma jets for personalized plasma medicine (https: / / ieeexplore.ieee.org / abstract / document / 10384190).
[0050] As described above, while the CAPJ example serves as an example and situation to illustrate the principles of the MOBO embodiment, this discussion also applies to any low-temperature atmospheric-pressure plasma biomedical environment or other environment in which human safety must be taken into consideration. The safe BO restricts the search space to a known estimated feasible region. From a safety-critical standpoint, the constraints in (11) above must always be met. The BO provides a framework that this embodiment uses to modify the control of the plasma medical device. The BO has the flexibility to be tuned based on application and / or user preferences. This makes the BO applicable to plasma medical devices. Figure 5 illustrates how the BO restricts the search space to a known estimated feasible region.
[0051] Figure 6 graphically displays the observed closed-loop profile for safe Bayesian optimization. This was generated using a model predictive controller and does not include a neural network controller, but the principle of changing the controller parameters is the same as described above. The upper figure shows the change in CEM (evolution) over a 120-second treatment period. The lower figure shows the change in temperature over the same treatment period. The profile color / gradient indicates the change in the profile over 30 iterations of BO. The first two profiles, shown by the pink dotted lines, represent the initial data provided to BO. Figure 6 demonstrates that safe Bayesian optimization reliably proposes safe queries with a reasonable probability.
[0052] This embodiment demonstrates safeBays optimization for an exemplary application of personalized plasma medicine, which is an emerging research area including the use of low-temperature atmospheric-pressure plasma (CAPs) for various medical treatments, as described in "Low temperature plasma for biology, hygiene, and medicine: Perspective and roadmap" (IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 6, no. 2, pp. 127-157, 2022) by M. Laroussi, S. Bekeschus, M. Keidar, A. Bogaerts, A. Fridman, X. Lu, K. Ostrikov, M. Hori, K. Stapelmann, and V. Miller. As described in "The path to personalized medicine" by MA Hamburg and FS Collins (New England Journal of Medicine, vol. 363, no. 4, pp. 301-304, 2010), tailoring the plasma effect applied to specific surfaces / subjects is key to ensuring the effectiveness of plasma therapy.However, as described in "Multimodal action of atmospheric pressure plasma jets for biological applications" by J.-M. Pouvesle and E. Robert (ISPB 2017, 2017) and "Perspectives on machine learning-assisted plasma medicine: Toward automated plasma treatment" by AD Bonzanini, K. Shao, A. Stancampiano, DB Graves and A. Mesbah (IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 6, no. 1, pp. 16-32, 2021), the fundamental mechanisms of plasma surface interaction can only be quantified in relation to the population. Therefore, iterative improvements in automated treatment using BO (Body Oscillation) can enable individualization of CAP (Captive Pulsed Airway) treatment, with patient safety being of paramount importance.
[0053] It will be understood that the features and functions disclosed above, as well as variations or substitutes of other features and functions, can be combined into many other different systems or applications. Various currently unforeseen or unexpected substitutions, modifications, variations, or improvements will be made later by those skilled in the art, and these will also be included in this embodiment.
Claims
1. A system for controlling a low-temperature atmospheric pressure plasma device, Low-temperature atmospheric pressure biomedical device, One or more sensors for measuring one or more properties of the plasma and the surface to which the low-temperature atmospheric pressure plasma apparatus is directed, Includes one or more processors configured to execute code, The code is provided to one or more processors, A deep neural network (DNN) is used to provide operating parameters for the operation of the low-temperature atmospheric pressure plasma device. A system for controlling a low-temperature atmospheric pressure plasma device, characterized by receiving information from one or more sensors as feedback regarding the aforementioned operation.
2. The system according to claim 1, wherein the one or more processors are further configured to collect the results of the operation of the low-temperature atmospheric pressure plasma device in a run-to-run manner.
3. The system according to claim 2, wherein the one or more processors are further configured to execute code that causes the one or more processors to apply Bayesian optimization to run-to-run data collected from the operation of the low-temperature atmospheric pressure plasma device and to generate control parameters for the DNN for implementing a control policy.
4. The system according to claim 3, wherein the Bayesian optimization includes one of safe Bayesian optimization, multi-objective Bayesian optimization, or safe multi-objective Bayesian optimization.
5. The system according to claim 1, wherein the one or more processors are selected as resource-limited hardware for embedded control.
6. The system according to claim 1, wherein the one or more processors are further configured to ensure that the output of the deep neural network falls within a constrained range of values that is greater than an initial range of values that is considered safe.
7. The system according to claim 1, wherein the DNN receives the temperature and light intensity of a low-temperature atmospheric pressure plasma jet as input.
8. The system according to claim 1, wherein the DNN has one or more outputs of applied power, applied voltage, frequency, and gas flow rate.
9. A method for controlling a low-temperature atmospheric pressure plasma device, A step of identifying the operating parameters for the operation of the low-temperature atmospheric pressure plasma device that have the greatest influence on the good results of the operation of the low-temperature atmospheric pressure plasma device, The process involves providing a deep neural network (DNN) with data from each operation of the low-temperature atmospheric pressure plasma apparatus so that the DNN can provide an optimal value for the identified operating parameters, A step of operating the low-temperature atmospheric pressure plasma device using the aforementioned operating parameters, A method for controlling a low-temperature atmospheric pressure plasma apparatus, comprising the step of providing the data from the aforementioned operation back to the deep neural network.
10. The method according to claim 9, wherein the data includes light intensity and temperature.
11. The method according to claim 9, wherein the operating parameters include one or more of applied power, applied voltage, frequency, and gas flow rate.
12. The method according to claim 9, wherein the system for controlling the low-temperature atmospheric-pressure plasma device applies Bayesian optimization to the data from the operation of the low-temperature atmospheric-pressure plasma device and adjusts the control parameters for the DNN to implement the control policy.
13. The method according to claim 12, wherein the Bayesian optimization includes either safe Bayes optimization or multi-objective Bayes optimization.