Design method of nasal catheter based on reinforcement learning
By using a reinforcement learning-based nasal cannula design method, the problems of long design cycles and insufficient accuracy in existing technologies have been solved, enabling efficient and stable manufacturing of personalized nasal cannulas and improving product performance and production line stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU HUANXI ELECTRONIC MATERIALS CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-08
AI Technical Summary
Existing nasal cannula designs rely on empirical parameters, resulting in long design cycles, insufficient precision, and difficulty in adapting to individualized working conditions such as nasal cavity structure, oxygen flow rate, and pressure for different groups of people. Furthermore, the lack of simulation prediction and optimization mechanisms leads to unstable product performance.
A reinforcement learning-based nasal cannula design method is adopted. By collecting data on oxygen supply conditions, nasal cavity geometry, and manufacturing process boundaries, a parameter set is established and a simulation environment coupling fluid, heat, and contact pressure is constructed. Reinforcement learning algorithms are used to optimize geometric, material, and manufacturing process parameters, forming a closed-loop design and manufacturing system, thereby achieving automatic generation and calibration of the optimal design strategy.
It achieves high-precision and personalized design of nasal cannulas, shortens the design cycle, ensures stable performance under different oxygen supply conditions, improves manufacturing consistency and automation level, and reduces rework and scrap rates.
Smart Images

Figure CN121997486A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent design technology, and in particular to a design method for nasal cannulas based on reinforcement learning. Background Technology
[0002] As a commonly used consumable for clinical oxygen supply and respiratory assistance, the geometry, material properties, and manufacturing process of nasal cannulas directly affect oxygen supply efficiency, wearing comfort, and usage stability. Existing nasal cannula designs mostly rely on empirical parameters or simple trial-and-error methods, resulting in long design cycles, insufficient precision, and difficulty in adapting to individualized working conditions such as nasal cavity structure, oxygen flow rate, and pressure for different populations.
[0003] In traditional manufacturing processes, the geometric parameters of nasal cannulas are typically obtained through manual surveying or empirical setting, and material properties and manufacturing processes are also largely determined by manual adjustments. This approach is not only inefficient but also makes it difficult to achieve a high-precision mapping between parameters and performance. Furthermore, due to the lack of simulation prediction and optimization mechanisms, the stability of product performance is difficult to guarantee in actual production. Therefore, we propose a reinforcement learning-based nasal cannula design method.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a nasal cannula design method based on reinforcement learning, thereby solving the technical problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A reinforcement learning-based method for designing nasal cannulas includes the following steps:
[0008] S1. Collect oxygen supply conditions, nasal cavity geometry, material properties and manufacturing process boundaries, establish geometric parameter set, material parameter set, manufacturing process parameter set and quality inspection feature set, and unify parameterization to form a manufacturing feasible domain;
[0009] S2. Based on the parameter space, construct a simulation environment that couples fluid, heat and contact pressure, take the parameter set as input and the mass characteristics as output, establish a simulation prediction model and calibrate it experimentally.
[0010] S3. Determine the target performance indicators, construct a reward function with leakage rate, peak contact pressure and dimensional deviation as the core, and combine it with the manufacturing feasible domain to form a composite reward system;
[0011] S4. Using the simulation prediction model as the environment, a reinforcement learning algorithm is used to train the strategy on the parameter set. The search is guided by the reward function and constraints, and the optimal design strategy is output.
[0012] S5. Output the optimal geometry, material, and manufacturing process parameters according to the optimal design strategy, and map them into CAD mold models, material formulas, and equipment control instructions respectively to realize manufacturing implementation;
[0013] S6. Perform online monitoring of the manufacturing process. When the deviation between the measured and predicted values exceeds the threshold, calibrate the process parameters and feed the data back into the simulation model and reinforcement learning agent to achieve a closed-loop update of design-manufacturing-calibration.
[0014] S1 specifically includes:
[0015] Collect target oxygen supply parameters, including oxygen flow rate, oxygen pressure, nasal cavity geometric constraints, wearing time, and ambient humidity and temperature range;
[0016] Obtain the set of geometric parameters required for nasal cannula design, including nasal fork angle, end curvature, side opening ratio, internal cross-sectional dimensions, and external profile;
[0017] Obtain the set of material parameters used in the nasal cannula, including material hardness gradient, elastic modulus gradient, surface friction coefficient and deformation recovery characteristics;
[0018] Obtain the set of manufacturing process parameters, including extrusion temperature, traction speed, cooling curve, die compensation coefficient, and secondary shaping curve;
[0019] Obtain a quality inspection feature set, including leakage rate, peak contact pressure, dimensional deviation and batch consistency index; perform unified parameterization processing on the above geometric parameter set, material parameter set, manufacturing process parameter set and quality inspection feature set to form a unified parameter space and manufacturing feasible domain constraints.
[0020] S2 specifically includes:
[0021] Based on a unified parameter space, a multiphysics coupling model is established, covering gas flow, heat conduction, and nose contact pressure distribution processes.
[0022] The set of geometric parameters, the set of material parameters, and the set of manufacturing process parameters are used as input variables, and the set of quality inspection features is used as the output target.
[0023] Establish a simulation prediction model to quickly predict the quality inspection feature set under varying parameter combinations;
[0024] The simulation prediction model is calibrated using experimental data to ensure that the deviation between the prediction results and the actual manufacturing conditions is controlled within the set threshold.
[0025] The calibrated simulation prediction model is used as the environment model of the reinforcement learning agent to achieve a closed-loop mapping between parameter changes and quality feature prediction.
[0026] S3 specifically includes:
[0027] Based on the target oxygen supply condition parameters, a set of target performance indicators is defined; based on the quality detection feature set, a reward function is established, taking the reduction of leakage rate, the reduction of peak contact pressure, and the improvement of dimensional stability as positive incentives;
[0028] Using manufacturing feasible domain constraints as constraints, negative penalties are imposed on actions that violate processing or assembly limits; thus forming a unified multi-objective reward and constraint strategy system.
[0029] By synchronously binding the reward and constraint strategy with the simulation prediction model, a clear optimization direction is provided for reinforcement learning training.
[0030] S4 specifically includes:
[0031] Construct a reinforcement learning agent, defining the state as a combination of historical parameters and the corresponding predicted quality detection features;
[0032] The action is defined as a joint adjustment of the geometric parameter set, the material parameter set, and the manufacturing process parameter set;
[0033] A reinforcement learning algorithm based on policy gradient or proximal policy optimization is used to perform continuous action space search;
[0034] The agent iteratively optimizes based on the reward function and constraint policy, forming a convergence process that gradually approaches the optimal design.
[0035] The design strategy that converges under the target oxygen supply condition is obtained and used to infer the optimal parameter set.
[0036] S5 specifically includes:
[0037] Using the trained design strategy, inference is performed on the target oxygen supply condition parameters, and the optimal set of geometric parameters, optimal set of material parameters, and optimal set of manufacturing process parameters are output.
[0038] The optimal set of geometric parameters is mapped to the nasal cannula mold and the mold opening compensation dimensions, and specific machining drawings are generated.
[0039] The optimal material parameter set is mapped to manufacturing instructions for material formulation, hardness gradient, and elastic modulus gradient; the optimal manufacturing process parameter set is mapped to a complete process path and numerical control instructions for extrusion, cooling, traction, and secondary shaping.
[0040] The mapped parameter results are synchronously output to the production control system to form an integrated manufacturing solution.
[0041] S6 specifically includes:
[0042] Based on the above manufacturing scheme, conduct trial production or mass production, and collect measured data that are consistent with the quality inspection feature set;
[0043] The measured data and the output of the simulation prediction model are compared item by item to identify the sources of deviation in geometric, material and process parameters;
[0044] Based on the manufacturing feasible region constraint, online fine-tuning and calibration are performed on the continuous parameters of the optimal manufacturing process parameter set;
[0045] The calibrated results and measured data are fed back into the simulation prediction model and the reinforcement learning agent to form incremental training samples.
[0046] Update the agent strategy to enable the design strategy to adapt to fluctuations in the actual manufacturing environment and maintain output stability;
[0047] The final set of geometric parameters, the final set of material parameters, and the final set of manufacturing process parameters are solidified to form a closed-loop system for design, manufacturing, and calibration.
[0048] The beneficial effects of this invention are as follows:
[0049] This invention utilizes reinforcement learning algorithms to train and converge strategies within a unified space of geometric, material, and manufacturing process parameters. This automatically generates optimal design schemes that meet requirements for sealing, comfort, and dimensional stability, reducing manual trial-and-error and parameter tuning, and shortening the design cycle. Through multiphysics simulation coupling fluid, heat, and contact pressure, leakage rate, peak pressure, and dimensional deviations are predictable and controllable, ensuring stable sealing performance of the nasal cannula under different oxygen supply conditions and effectively reducing the risk of nasal pressure pain. The optimal design results are directly converted into CAD / CAM mold models, material formulations, and CNC equipment control instructions, streamlining the "algorithm, design, and manufacturing" chain. This avoids deviations caused by manual translation and secondary processing, improving manufacturing consistency and automation. By online collection and comparison of key indicators such as leakage rate, peak contact pressure, and dimensional deviations, manufacturing process parameters are automatically calibrated when deviations exceed thresholds. The calibration results are then fed back into the model and strategy for continuous adaptive optimization, ensuring long-term product performance stability. The closed-loop control mechanism reduces fluctuations in the production process, decreases rework and scrap rates, improves production line stability, and achieves highly consistent batch manufacturing. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of a nasal cannula design method based on reinforcement learning according to the present invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Example 1: As Figure 1 As shown, this embodiment provides a method for designing a nasal cannula based on reinforcement learning, including the following steps:
[0053] S1. Collect oxygen supply conditions, nasal cavity geometry, material properties and manufacturing process boundaries, establish geometric parameter set, material parameter set, manufacturing process parameter set and quality inspection feature set, and unify parameterization to form a manufacturing feasible domain;
[0054] S2. Based on the parameter space, construct a simulation environment that couples fluid, heat and contact pressure, take the parameter set as input and the mass characteristics as output, establish a simulation prediction model and calibrate it experimentally.
[0055] S3. Determine the target performance indicators, construct a reward function with leakage rate, peak contact pressure and dimensional deviation as the core, and combine it with the manufacturing feasible domain to form a composite reward system;
[0056] S4. Using the simulation prediction model as the environment, a reinforcement learning algorithm is used to train the strategy on the parameter set. The search is guided by the reward function and constraints, and the optimal design strategy is output.
[0057] S5. Output the optimal geometry, material and manufacturing process parameters according to the optimal design strategy, and map them into CAD mold models, material formulas and equipment control instructions respectively to realize manufacturing implementation;
[0058] S6. Perform online monitoring of the manufacturing process. When the deviation between the measured and predicted values exceeds the threshold, calibrate the process parameters and feed the data back into the simulation model and reinforcement learning agent to achieve a closed-loop update of design, manufacturing, and calibration.
[0059] S1 specifically includes the following sub-steps:
[0060] S110: Collect target oxygen supply parameters: Collect target oxygen supply parameters, including oxygen flow rate, oxygen pressure, nasal cavity geometric constraints, wearing time and ambient temperature and humidity range.
[0061] Oxygen supply flow rate and pressure are collected in real time by flow meter and pressure sensor; ambient temperature and humidity are collected by temperature and humidity sensor; nasal cavity geometric constraints are extracted by CT imaging or 3D scanning modeling to extract the shape of nasal cavity entrance, nasal alar distance and depth constraints; wearing time is set by clinical or usage scenario parameters.
[0062] All collected data are recorded in standard numerical form for subsequent determination of optimization targets for geometric, material, and process parameters.
[0063] S120: Obtain the set of geometric parameters: Based on the acquired nasal cavity geometric constraints, determine the set of geometric parameters required for the design of the nasal cannula, including at least: nasal fork angle; end radius of curvature; side hole opening ratio and arrangement angle; and internal cavity cross-sectional dimensions and shape.
[0064] External contour morphology. Geometric parameters are parametrically described using CAD modeling tools and encoded as continuous variables or finite sets, serving as one of the input dimensions for subsequent simulation and reinforcement learning training.
[0065] S130: Obtain the material parameter set: For the polymeric elastic material used in the nasal cannula, determine the material parameter set, which should include at least: hardness gradient (tested using a Shore hardness tester); elastic modulus gradient (tested using Dynamic Mechanical Analysis (DMA)); surface friction coefficient (tested using a standard friction coefficient test bench); and deformation recovery characteristics (obtained through stress-strain cyclic loading experiments).
[0066] Material parameters are input into a unified parameter space in numerical form and associated with a set of geometric parameters for subsequent predictive modeling of sealing performance and comfort.
[0067] S140: Obtain the set of manufacturing process parameters: Based on the capability boundary of the nasal cannula manufacturing equipment, determine the set of manufacturing process parameters, which shall include at least: extrusion temperature; traction speed; cooling curve (temperature-time-flow rate relationship); die compensation coefficient; and secondary shaping curve.
[0068] Manufacturing process parameters are collected in real time through the process monitoring system or preset by process engineers, and serve as an important component of the subsequent manufacturing feasibility domain constraints.
[0069] S150: Obtain a quality inspection feature set: Establish a quality inspection feature set for the performance of the nasal cannula, including at least: leakage rate (using pressure decay test); peak contact pressure (using pressure sensor pad or micro strain gauge); dimensional deviation (obtained using a high-precision coordinate measuring machine, CMM); batch-to-batch consistency index (calculating the standard deviation of dimensional and leakage data from multiple batches using statistical methods); the above inspection indicators correspond to the sealing performance, comfort, and stability evaluation of the final product, and will serve as the main evaluation objects of the reinforcement learning reward function.
[0070] S160: Parametricization and Manufacturing Feasibility Domain Constraints: The geometric parameter set, material parameter set, manufacturing process parameter set, and quality inspection feature set are uniformly parameterized to form a unified parameter space. Based on the equipment capability boundary and material physical limits, manufacturing feasibility domain constraints are defined, including but not limited to:
[0071] Extrusion temperature range (e.g., 150℃~210℃);
[0072] Traction speed range (e.g., 20 mm / s to 100 mm / s);
[0073] Die opening compensation coefficient range (e.g., ±5%);
[0074] Material elastic modulus range (e.g., 2MPa to 10MPa);
[0075] The range of nose fork angle variation (e.g., 20°-45°). Manufacturing feasibility domain constraints filter parameter combinations during reinforcement learning training and inference, ensuring that the output can be directly incorporated into the actual manufacturing process.
[0076] S2 specifically includes the following sub-steps:
[0077] S210: Establishing a multiphysics simulation environment based on a unified parameter space: based on the geometric parameter set defined in S160. Material parameter set and manufacturing process parameter set A multiphysics coupling simulation environment is built on the computer.
[0078] The simulation environment includes at least the following: a fluid field sub-model: used to simulate the flow rate distribution, pressure drop distribution, and turbulence characteristics of oxygen flowing through the nasal cannula and nasal cavity; a heat conduction sub-model: used to calculate the temperature and humidity changes of the airflow; and a contact pressure sub-model: used to calculate the contact pressure distribution generated at the interface between the outer contour of the nasal fork and the nasal wing.
[0079] Fluid field simulation can be achieved using CFD tools (such as Fluent, OpenFOAM, or a custom solver), contact pressure simulation can be achieved using FEM tools (such as Abaqus or ANSYS Mechanical), and thermal field analysis can be achieved by combining energy conservation equations.
[0080] The three sub-models mentioned above are coupled through a unified parameter interface to ensure that changes in geometric parameters can act synchronously on the fluid, thermal, and structural fields, forming a multiphysics simulation environment for nasal cannula performance.
[0081] S220: Simulation Input and Output Mapping: In the simulation environment described above, the geometric parameter set... Material parameter set With manufacturing process parameter set As an input variable.
[0082] Simulation output corresponds to the quality inspection feature set At least including: leakage rate L (%); peak contact pressure (Pa); Dimensional deviation D (mm); Temperature and humidity distribution (°C, %RH).
[0083] The input-output mapping relationship is as follows:
[0084]
[0085] The above represents a functional mapping relationship from multidimensional input parameters to multidimensional output performance indicators. Its core function is to describe the mathematical relationship between nasal cannula design parameters (input) and actual performance indicators (output), and it is the basic expression of the simulation prediction model. This represents the simulation-prediction mapping function, which is obtained through multiphysics coupling calculation and is a mathematical model of the relationship between nasal cannula design parameters and performance response. This represents the set of input parameters, including geometric parameters, material parameters, and manufacturing process parameters. This represents a set of output performance indicators, including leakage rate, peak contact pressure, dimensional deviation, and temperature and humidity distribution. This functional relationship is automatically solved by the coupled simulation process and used to train the simulation prediction surrogate model.
[0086] S230: Establish a simulation prediction model: In order to improve the efficiency of reinforcement learning training, a simulation prediction model is established on the basis of a multi-physics simulation environment, which serves as the "environment model" for the reinforcement learning agent.
[0087] The simulation prediction model can be implemented using any of the following methods or combinations thereof: Response Surface Method (RSM); Multilayer Perceptron Neural Network (MLP); Gaussian Process Regression (GPR); XGBoost Regression;
[0088] The training samples are derived from coupled simulation results, and the output corresponds to the quality detection features. This model allows reinforcement learning agents to quickly obtain performance responses to parameter changes without invoking costly simulations in each policy iteration.
[0089] S240: Experimental Calibration and Error Control: Regression calibration is performed on the test data of the nasal cannula physical sample and the output of the simulation prediction model to ensure that the prediction accuracy meets the design requirements. Calibration methods include: collecting measured values of leakage rate, peak contact pressure, and dimensional deviation; correcting the surrogate model weights using least squares regression or Bayesian calibration; and controlling the deviation between predicted and measured values within ±5%. After calibration, the simulation prediction model can stably reflect the performance characteristics of the nasal cannula during actual manufacturing.
[0090] S250: Using the simulation prediction model as the reinforcement learning environment: The calibrated simulation prediction model serves as the interaction environment for the reinforcement learning agent: The agent selects actions (i.e., the joint adjustment of geometric parameters, material parameters, and manufacturing process parameters); The simulation prediction model returns the corresponding quality characteristic response based on the actions. ;
[0091] The reward function calculates the reward value based on the target performance and constraints; the agent uses this feedback to update the policy. In this way, reinforcement learning training does not rely on large-scale physical trials, significantly improving the efficiency of design iteration and maintaining consistency with actual manufacturing performance.
[0092] S3 specifically includes the following sub-steps:
[0093] S310: Determination of the Target Performance Index Set: Based on the target oxygen supply condition parameters collected in S110–S160 and the simulation prediction models established in S210–S250, determine the target performance index set for the nasal cannula design, which should include at least the following three categories: Sealing index: Leakage rate L not exceeding 2%; Comfort index: Peak contact pressure. Not exceeding the set threshold Dimensional stability index: Dimensional deviation D does not exceed the allowable deviation. .
[0094] The above three types of target performance indicators obtain expected values through simulation prediction models and are compared with actual detection values for reward calculation and policy evaluation in reinforcement learning.
[0095] S320: Construction and Definition of Reward Function: For the above target performance indicators, a multi-objective reward function R is established to transform the changes in leakage rate, pressure peak, and size deviation into numerical reward signals.
[0096]
[0097] The reward function R is used to evaluate the performance of the design scheme for the reinforcement learning agent, mapping multiple key engineering indicators into a single scalar feedback signal, thereby achieving performance-driven policy optimization. The technical meaning and source of each variable are as follows:
[0098] , , is the weighting coefficient, used to control the proportion of different performance indicators in the total reward, ensuring that the sum of the weights is 1; L is the leakage rate, obtained through fluid simulation or pressure decay measurement, used to reflect the nasal cannula's sealing performance. It is the peak contact pressure between the nose fork and the nose wing, obtained through finite element contact simulation and actual measurement by pressure sensors, and is used to reflect wearing comfort. 1 is the pressure pain threshold, set according to medical safety standards or clinical experience. Exceeding this value will lead to discomfort during use. 2 is the dimensional deviation, obtained through a CMM coordinate measuring machine or vision inspection system, used to measure the consistency between the product and the design model. It is the maximum permissible dimensional deviation, set based on manufacturing process capabilities and quality standards.
[0099] In this embodiment, the three terms of the reward function correspond to three core performance objectives of the nasal cannula design: sealing performance (first term): the smaller the leakage rate L, the closer (1-L) is to 1, and the higher the reward value; wearing comfort performance (second term): when the peak contact pressure... Approaching or below the threshold The greater the contribution to comfort, the better; Manufacturing dimensional accuracy performance (third item): The smaller the dimensional deviation D, the closer this item is to 1, and when the deviation exceeds... When this value approaches 0 or becomes negative, it indicates an unacceptable design.
[0100] In practical engineering implementation, the reward function calculation process of this embodiment includes the following steps:
[0101] Simulation and testing data acquisition: Acquire the L and L corresponding to the design respectively. D; Reward calculation: Substitute the collected indicators into the above formula to calculate the comprehensive reward value R;
[0102] Policy Update: The reinforcement learning agent adjusts its design parameters based on the R-value. ;
[0103] Iterative convergence: Through multiple rounds of continuous iteration, it automatically converges to the optimal design solution that combines high sealing performance, comfort, and manufacturing precision.
[0104] This reward function achieves convergence from multi-objective engineering performance, single numerical signal, and policy optimization, avoiding the inefficiencies of traditional manual weighting and trial-and-error design. Furthermore, all parameters have clear physical definitions, detection methods, and engineering boundaries, ensuring the feasibility and verifiability of the algorithm and manufacturing process.
[0105] For example, in one specific implementation, the reward weight parameter can be set as follows: The tenderness threshold was set to 2.0 kPa, and the maximum allowable size deviation was set to 0.5 mm. Under these settings, the reward function can significantly guide the reinforcement learning algorithm to quickly converge to the optimal solution with low leakage rate, controllable tenderness peak, and high size accuracy.
[0106] S330: Introduction of Manufacturing Feasible Region Constraints: Based on the reward function calculation, manufacturing feasible region constraints are introduced to impose additional penalties on parameter combinations that do not conform to the actual manufacturing process boundaries.
[0107] The manufacturing feasibility domain constraint is derived from the threshold definition in S160, including but not limited to: extrusion temperature. ; traction speed Die compensation coefficient ; nose fork angle Material elastic modulus ;
[0108] When the parameters exceed the above feasible range, a negative reward is immediately triggered:
[0109]
[0110] in R is the final reward value; R is the base reward value, calculated as described in the aforementioned reward function. The penalty weighting coefficient is used to adjust the strength of the penalty term's influence in the total reward function. When the size is large, the system tends to satisfy process and safety constraints; when When the size is small, the system tends to explore performance-driven scenarios; This refers to the penalty value when the constraint is violated; It is mainly used to characterize the degree of deviation of a design scheme from its manufacturing feasibility, safety thresholds, and clinical applicability. In this embodiment, the penalty terms include at least the following three types of constraints:
[0111] Process feasibility constraints: When the selected manufacturing process parameters exceed the equipment capability boundaries (such as extrusion temperature, traction speed, die compensation coefficient, etc.), a high penalty value is applied;
[0112] Safety threshold constraint: When the peak contact pressure exceeds the safety threshold If the leakage rate L exceeds the allowable range, the penalty will be increased;
[0113] Structural constraints: When design geometry leads to localized stress concentration, material instability, or mold inability to be processed, a structural feasibility penalty is imposed.
[0114] The value of the above penalty term increases linearly or non-linearly depending on the degree of deviation, for example:
[0115]
[0116] in It is the peak value of the contact pressure between the nasal fork and the nasal ala; is the pressure pain threshold; L is the leakage rate, obtained through fluid simulation and pressure decay testing. This is the upper limit of the allowable leakage rate, determined based on the oxygen supply efficiency requirements. It is a geometric failure indicator function, which takes the value of 1 when the geometric structure does not meet the processing or wearing requirements, and 0 otherwise. , , where represents the weighting coefficients. This constraint strategy ensures that the reinforcement learning search space is limited to a practically manufacturable range.
[0117] Stress Punishment When the peak contact pressure Less than or equal to the threshold When the peak contact pressure exceeds the threshold, this item is 0; when the peak contact pressure exceeds the threshold, this item increases linearly by the amount of excess; this penalty ensures that there is no tenderness or discomfort when wearing the nasal cannula.
[0118] Leakage penalties When the leakage rate L is less than or equal to the allowable upper limit When the leakage rate exceeds the threshold, the penalty value increases with the extent of the exceedance; ensure that the nasal cannula has sufficient sealing when supplying oxygen.
[0119] Geometric failure penalty When the geometric parameters in the design cause the mold to be unmanufacturable, the nasal fork to be severely mismatched with the nasal cavity, or the structure to be unstable, the indicator function... A value of 1 applies a fixed penalty; this penalty is used to filter out structural designs that are unmanufacturable or do not meet clinical use requirements.
[0120] S340: Integration of Reward and Constraint Strategies: Integrating the reward function with the feasible region constraint to form a composite reward system.
[0121]
[0122] in This represents the final reward value received by the agent; when the parameters are within the process-feasible region, The system encourages high-quality design solutions; when exceeding process boundaries, The reward drops significantly, forcing the strategy to return to searching within the feasible region; this approach combines exploration efficiency and engineering feasibility within the RL framework.
[0123] S350: Binding the reward system to the RL agent: Embedding the above-mentioned composite reward system into the training process of the reinforcement learning agent:
[0124] The intelligent agent executes actions (i.e., adjusts geometric parameters, material parameters, and manufacturing process parameters); the simulation and prediction models output corresponding... Reward function calculation The agent updates its policy parameters based on the reward; the reward function and feasible region constraints operate synchronously throughout the training cycle. Through this binding method, the optimization direction of reinforcement learning is entirely constrained by technical goals and process constraints, fundamentally avoiding the search process for ineffective designs.
[0125] S4 specifically includes the following sub-steps:
[0126] S410: Building a reinforcement learning agent:
[0127] Based on the simulation prediction model of S210–S250 and the reward / constraint system of S310–S350, a reinforcement learning agent is constructed. The elements of the agent are defined as follows:
[0128] State space S: consisting of the set of historical geometric parameters Material parameter set Manufacturing process parameter set and the corresponding quality inspection feature set constitute;
[0129] Action space A: A continuous adjustment vector for geometric parameters, material parameters, and manufacturing process parameters. , , ;
[0130] Reward function R: A composite reward function defined by S320–S340 Environment Model E: Composed of the simulation prediction model of S250, responsible for outputting a response based on the action.
[0131] This definition method ensures that the input, actions, and environmental responses of the intelligent agent are all based on clear physical meaning, avoiding the emergence of abstract and unimplementable "black box" situations.
[0132] S420: Training Algorithm and Parameter Sampling
[0133] To ensure the agent has convergence stability and high search efficiency in a continuous multidimensional design space, the training algorithm preferably adopts a continuous action space reinforcement learning algorithm, including but not limited to: Proximal Policy Optimization (PPO); Deep Deterministic Policy Gradient (DDPG); and Soft Actor Critics Algorithm (SAC).
[0134] The initial strategy exploration employs Latin hypercube sampling or Bayesian optimized sampling to ensure that the initial training samples are uniformly distributed in the parameter space, reducing the risk of local convergence.
[0135] During training, the policy network and value network are optimized separately, and the expected loss function is minimized through gradient updates:
[0136]
[0137] in It is the loss function value; (s,a): represents state s and action a respectively, corresponding to the current design state and design parameter action of the nasal cannula; Indicates by parameters Control policy network; This represents the final reward value corresponding to the action a performed in state s. The specific calculation method is described in the aforementioned reward and penalty formula.
[0138] The state space s represents the multidimensional features of the nasal cannula design, including but not limited to the nasal fork angle, die compensation coefficient, elastic modulus gradient, extrusion temperature, and traction speed.
[0139] The action space 'a' represents the design decisions that a reinforcement learning agent can adjust in the current state; for example, adjusting the angle, changing the material distribution, and correcting process parameters; strategies. From neural network parameters The probability distribution of the representation; it determines the probability of choosing an action in different states and is the core object of training; the role of the loss function: to minimize the loss function. This is equivalent to maximizing the reward value. As the design scheme approaches high sealing performance, high comfort, high precision, and high manufacturability, The smaller the value, the greater the penalty when the design violates constraints. Increase, thereby guiding the agent away from unacceptable design regions.
[0140] S430: Training Process and Simulation Interaction: During the training iteration process: the agent selects actions in the state space S. (Adjustments to nasal cannula parameters); Environmental model E calculates the corresponding physical response, including leakage rate L and peak contact pressure. Dimensional deviation D and temperature and humidity distribution ;
[0141] The reward function module calculates the corresponding Value; the agent updates its policy based on the reward value. The system generates new actions in the next iteration; if an action exceeds the feasible region constraint, the reward function automatically deducts a penalty, forcing the agent to converge towards the feasible region. In this way, the reinforcement learning training process, simulation environment, reward function, and constraint policy form a stable closed-loop structure, ensuring that the training process has a clear technical implementation path.
[0142] S440: Training Convergence and Stability Control: To ensure the reliability of the agent's policy, a convergence criterion and stability control mechanism are set during the training process.
[0143] When the cumulative reward curve fluctuates by less than 1% over N=50 consecutive iterations, the training is considered converged; if the agent's actions continuously trigger penalty terms, the exploration rate is automatically adjusted. Alternatively, the step size can be reduced to avoid training oscillations; if a local optimum trap occurs, the local optimum can be escaped by restarting sampling or using gradient perturbation mechanisms. This mechanism ensures that the training convergence process is controllable and verifiable, rather than a "black box convergence".
[0144] S450: Output the optimal policy: Output the final policy after the training process meets the convergence condition. This strategy corresponds to the optimal set of geometric parameters, the optimal set of material parameters, and the optimal set of manufacturing process parameters for the nasal cannula.
[0145]
[0146] in This represents the optimal strategy after reinforcement learning converges; It is a comprehensive reward function (including performance and penalty terms). It is a set of geometric parameters, including nose fork angle, end curvature, cross-sectional dimensions, etc. It is a set of material parameters, including elastic modulus gradient, hardness, friction coefficient, etc. It is a set of manufacturing process parameters, including extrusion temperature, traction speed, die compensation, etc.
[0147] The essence of the objective function: to maximize The reinforcement learning agent can adaptively find combinations of geometric, material and process parameters to obtain a nasal cannula design with high sealing performance, low pressure pain, high dimensional accuracy and high manufacturability.
[0148] Search space coverage: The strategy avoids the limitations of traditional manual design by sampling actions in a continuous parameter space and traversing a large number of candidate solutions in the design space.
[0149] The engineering significance of the optimization objective: Maximizing the reward is equivalent to achieving the overall optimality of the following engineering objectives: minimizing the leakage rate L; controlling the peak contact pressure. Not exceeding the safety threshold Control the dimensional deviation D within Within; satisfying process and geometric constraints, penalty term Minimum. This optimal strategy does not rely on manual parameter tuning and can directly enter the S510 inference output and manufacturing mapping stage, completing the closed-loop design path from algorithm to production.
[0150] S5 specifically includes the following sub-steps:
[0151] S510: Inference Execution and Optimal Parameter Set Output: The Final Policy Obtained Based on Training and Convergence of S410–S450 The inference process is executed under given target oxygen supply parameters. The input for the inference is:
[0152] Target oxygen supply parameters (flow rate, pressure, nasal cavity geometric constraints, etc.); equipment process capacity constraints (extrusion temperature range, traction speed range, etc.);
[0153] The output of the inference is: the optimal set of geometric parameters. Optimal material parameter set Optimal manufacturing process parameter set .
[0154]
[0155] in "Optimal strategy after training convergence" indicates the optimal strategy; "Operating conditions" indicates the actual operating conditions of the oxygen supply scenario, including at least oxygen flow rate, oxygen pressure, ambient temperature and humidity, patient nasal cavity geometry, and wearing time. It is the optimal set of geometric parameters; It is the optimal set of material parameters; : Optimal manufacturing process parameter set.
[0156] Operating parameters: Oxygen supply related parameters: including flow rate, pressure, humidity, temperature, etc., which directly affect leakage rate and comfort; Human adaptation parameters: including nasal cavity channel geometry, wearing time, and usage scenario (such as bedridden / active); Environmental constraint parameters: including external temperature and humidity, oxygen supply stability, etc., which affect material performance and sealing.
[0157] Inference process: Input conditions: Oxygen supply pressure, flow rate, nasal cavity geometry, etc. are used as reinforcement learning strategies. Input; Strategy decision: Based on the optimal policy distribution learned during training, respond to the input conditions and output the optimal design parameters;
[0158] Parameter output: Output This includes: geometric structure (such as nose fork angle and aperture distribution); material parameters (such as elastic modulus gradient and coefficient of friction); and process parameters (such as extrusion temperature, traction speed, and cooling time). These steps transform the reinforcement learning training results into executable manufacturing design parameters.
[0159] S520: Mapping of geometric parameter sets to CAD / CAM files: Mapping the optimal geometric parameter set It is mapped to CAD models and machining geometry information.
[0160] Specifically, this includes: nose fork angle, end curvature, and internal cavity cross-sectional dimensions → establishing a three-dimensional solid model; side hole opening ratio and arrangement → automatically generating distribution characteristics; and die orifice compensation amount → directly affecting the geometric compensation design of the extrusion die.
[0161] Automatically generate 3D geometric models and 2D machining drawings using CAD parametric modeling software (such as Siemens NX, SolidWorks, CATIA, or open-source FreeCAD), and export intermediate formats (such as STEP, IGES, STL) or CAM process files.
[0162] S530: Mapping of material parameter sets to material processing instructions: Mapping the optimal material parameter set... Mapped to material formulation and process execution instructions:
[0163] Material hardness gradient and elastic modulus gradient → determine the material co-extrusion ratio and proportion table; surface friction coefficient and texture information → determine the in-mold texture treatment or spraying parameters; deformation recovery characteristics → corresponding to the length and tension adjustment of the material extrusion cooling section.
[0164] Material process instructions can be converted into material batching process cards, co-extrusion runner temperature control instruction files, or material database call parameters to achieve automatic docking with manufacturing equipment.
[0165] S540: Mapping of manufacturing process parameter sets to equipment control commands: Mapping the optimal manufacturing process parameter set... Converted into CNC instructions and process curves that the equipment can recognize:
[0166] Extrusion temperature → Temperature control module setting command; Traction speed → Motor controller setting command; Cooling curve → Cooling circuit and flow control; Secondary shaping → Thermal shaping and tension control curve.
[0167] The manufacturing control system accepts CNC / Gcode / Mcode, PLC control logic instructions, or custom JSON / XML process instruction formats. These instructions, along with the geometry / material mapping results, form a complete production process package that can be directly deployed to the nasal cannula production line.
[0168] S550: Linkage between inference results and manufacturing process: The geometric model, material formula, and manufacturing process instructions generated by inference are synchronously distributed to the extrusion equipment, cooling line, and secondary shaping unit through the Manufacturing Execution System (MES) or production control interface to achieve automated manufacturing. Specifically, this includes: the CAD model is automatically converted into extrusion die processing files; the material formula retrieves raw materials through the batching system; the process curve is set by the control system for the extrusion, traction, and cooling processes; the equipment executes production tasks and provides real-time feedback of process signals such as temperature, speed, and pressure, providing data sources for S610 online calibration.
[0169] In this way, the results of reinforcement learning inference are directly used to guide manufacturing, realizing a complete implementation chain from algorithm → design → process → manufacturing.
[0170] S6 specifically includes the following sub-steps:
[0171] S610: Measured Data Acquisition: After completing the S550 inference mapping, the actual production process of the nasal cannula is monitored online, and measured data consistent with the quality inspection feature set is collected, including: real-time leakage rate. The contact pressure peak value is obtained through a pressure attenuation sensor or flow monitoring device. Dimensional deviations are measured using flexible pressure sensing pads or micro-strain gauges. Data is collected in real time through an online vision measurement system or a coordinate measuring machine (CMM);
[0172] Temperature and humidity distribution Data is collected via thermocouples and humidity sensors.
[0173] These measured data are automatically aggregated to the data acquisition terminal through the Manufacturing Execution System (MES) and compared synchronously with the prediction results of the simulation prediction model.
[0174] S620: Deviation Calculation and Anomaly Identification: Compare the measured data with the predicted values of the simulation prediction model item by item, and calculate the residuals:
[0175]
[0176] in This is the leakage rate prediction error; This is the peak contact pressure prediction error; It is the error in predicting dimensional deviations; , , Measured values obtained from actual nasal cannula samples; , , The predicted values are obtained through simulation-prediction model calculation.
[0177] Error calculation: For each set of design parameters, the measured performance value and the simulation prediction value are obtained respectively; the three types of errors are calculated by the above formula to quantitatively reflect the accuracy of the simulation model.
[0178] Error source analysis: It mainly reflects the difference between fluid simulation and actual leakage path; This mainly reflects the difference between simulated nasal ala contact pressure and actual tenderness; It mainly reflects the dimensional deviation of the finished product caused by process deformation.
[0179] Calibration Mechanism: When the error exceeds a preset threshold, the model's internal parameters are adjusted using the least squares method or Bayesian parameter update method; iterative calibration continues until the error converges to within ±5%, ensuring that the model's prediction accuracy meets the requirements of the reinforcement learning environment. The calibration process is automatically triggered by setting a deviation threshold. When the measured value is within the threshold range, the system maintains the existing production parameters unchanged; when any deviation exceeds the threshold, online adjustment is initiated in S630.
[0180] It should be noted that the deviation threshold setting refers to ISO 13485, specifically the leakage rate deviation threshold. Pick scope( The contact pressure peak and dimensional deviation thresholds are controlled within ±5% and ±0.1 mm, respectively, based on historical production standard deviation. A calibration process is automatically triggered when the deviation exceeds the limit; otherwise, the existing process is maintained. This threshold strategy effectively reduced invalid alarms by 70% during a 6-month process validation, ensuring the stability and efficiency of closed-loop adjustment.
[0181] S630: Online calibration of manufacturing process parameters: Once the deviation exceeds the threshold, an automatic process fine-tuning strategy is activated. If the deviation originates from an abnormal leakage rate, the die compensation amount and traction speed are adjusted first; if the deviation originates from an abnormal contact pressure peak, the nose fork angle and material elastic gradient are adjusted; if the deviation originates from an excessive dimensional deviation, the cooling curve and forming tension are adjusted; if multiple deviations occur simultaneously, the parameter correction vector is solved using the least squares method to achieve global minimum deviation adjustment.
[0182] Adjusted parameters It will write back to the production control system in real time to ensure that the manufacturing process stably converges to the target state; among which: This represents the candidate values of geometric parameters in the current iteration, such as nose fork angle, end curvature, orifice opening, etc. This represents the candidate values of material parameters in the current iteration round, such as the elastic modulus gradient, friction coefficient, and hardness distribution. This represents the candidate values for manufacturing process parameters in the current iteration round, such as extrusion temperature, traction speed, and cooling profile.
[0183] Candidate parameter set technical logic: Policy sampling phase: The agent selects parameters based on the current policy. A set of actions is sampled in the state space to form a candidate design parameter set. Performance evaluation phase: Input candidate parameters into the simulation-prediction model to calculate leakage rate L and peak contact pressure. Size deviation D; based on the reward function Calculate the comprehensive score corresponding to the candidate solution; Policy update phase: Sort, filter and update the gradient of the candidate solutions according to the reward value; Candidate solutions with higher reward values drive the convergence direction of the policy parameters, while solutions with lower reward values are eliminated or penalized.
[0184] S640: Incremental Model Learning and Feedback Update: The calibrated measured data and corrected parameters are synchronously fed back into the simulation prediction model and reinforcement learning agent to achieve incremental updates.
[0185] The simulation prediction model uses online regression or neural network weight fine-tuning to ensure that the predicted values are consistent with the measured values; the reinforcement learning agent uses new samples to incrementally train the policy network to avoid policy failure due to environmental drift.
[0186] The incremental update frequency can be dynamically adjusted based on production batches, equipment stability, and error trends.
[0187] This step ensures the long-term self-adaptability of the entire system, preventing it from failing due to equipment aging or changes in operating conditions.
[0188] S650: Stability Control and Dynamic Threshold Management: To avoid frequent and ineffective adjustments, the system introduces a stability control mechanism: If the deviation of 5 consecutive batches is less than the set threshold, the system locks the current parameter as the "stable condition"; if the deviation shows a fluctuating trend, the deviation threshold is dynamically adjusted to prevent excessive parameter tuning from causing oscillations; for sudden large deviations, an abnormal alarm and manual review mechanism are triggered.
[0189] By employing a control strategy that balances stability and flexibility, the system can achieve a balance between automatic adjustment and stable production.
[0190] S660: Closed-loop curing and continuous optimization: After calibration and updates, the final set of geometric parameters, material parameters, and manufacturing process parameters will be determined.
[0191]
[0192] This configuration is then solidified as the "standard manufacturing configuration" for the current batch. Simultaneously, this configuration and corresponding measured performance indicators are stored in the production database to accumulate data for future strategy optimization and production verification.
[0193] In this way, the system achieves: a closed loop from training inference, production and manufacturing, online detection, calibration feedback, and policy evolution; continuous collaborative optimization of reinforcement learning strategies and manufacturing equipment; and robust control that adapts to equipment aging, material fluctuations, and changes in the external environment.
[0194] In detail: Each batch collects more than 50 sets of measured data and feeds them back into the simulation-prediction model and reinforcement learning agent. When the cumulative number of samples reaches 300 sets, the policy network is automatically triggered for incremental updates. The database records include geometric, material, and process parameters and corresponding performance indicators, deviation correction vectors, and policy version information. As the data scale expands, this policy library can be used for rapid transfer training of new nasal cannulas, significantly shortening the development cycle of new processes and realizing a continuously evolving adaptive design and manufacturing system.
[0195] It should be noted that "reinforcement learning" in this embodiment is equivalent to "reinforcement learning", "machine learning" and "deep learning". Its essence is a nasal cannula design method based on reinforcement learning / reinforcement learning / machine learning / deep learning.
[0196] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0197] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0198] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0199] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0200] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0201] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0202] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0203] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0204] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for designing a nasal cannula based on reinforcement learning, characterized in that, Includes the following steps: S1. Collect oxygen supply conditions, nasal cavity geometry, material properties and manufacturing process boundaries, establish geometric parameter set, material parameter set, manufacturing process parameter set and quality inspection feature set, and unify parameterization to form a manufacturing feasible domain; S2. Based on the parameter space, construct a simulation environment that couples fluid, heat and contact pressure, take the parameter set as input and the mass characteristics as output, establish a simulation prediction model and calibrate it experimentally. S3. Determine the target performance indicators, construct a reward function with leakage rate, peak contact pressure and dimensional deviation as the core, and combine it with the manufacturing feasible domain to form a composite reward system; S4. Using the simulation prediction model as the environment, a reinforcement learning algorithm is used to train the strategy on the parameter set. The search is guided by the reward function and constraints, and the optimal design strategy is output. S5. Output the optimal geometry, material and manufacturing process parameters according to the optimal design strategy, and map them into CAD mold models, material formulas and equipment control instructions respectively to realize manufacturing implementation.
2. The nasal cannula design method based on reinforcement learning according to claim 1, characterized in that, It also includes S6, which performs online monitoring of the manufacturing process, calibrates process parameters when the deviation between actual measurement and prediction exceeds a threshold, and feeds the data back into the simulation model and reinforcement learning agent to achieve a closed-loop update of design-manufacturing-calibration.
3. The nasal cannula design method based on reinforcement learning according to claim 1, characterized in that, S1 specifically includes: Collect target oxygen supply parameters, including oxygen flow rate, oxygen pressure, nasal cavity geometric constraints, wearing time, and ambient humidity and temperature range; Obtain the set of geometric parameters required for nasal cannula design, including nasal fork angle, end curvature, side opening ratio, internal cross-sectional dimensions, and external profile; Obtain the set of material parameters used in the nasal cannula, including material hardness gradient, elastic modulus gradient, surface friction coefficient and deformation recovery characteristics; Obtain the set of manufacturing process parameters, including extrusion temperature, traction speed, cooling curve, die compensation coefficient, and secondary shaping curve.
4. The nasal cannula design method based on reinforcement learning according to claim 3, characterized in that, S1 also includes: acquiring a quality inspection feature set, including leakage rate, peak contact pressure, dimensional deviation and batch consistency index; and performing unified parameterization processing on the above geometric parameter set, material parameter set, manufacturing process parameter set and quality inspection feature set to form a unified parameter space and manufacturing feasible domain constraints.
5. The nasal cannula design method based on reinforcement learning according to claim 4, characterized in that, S2 specifically includes: Based on a unified parameter space, a multiphysics coupling model is established, covering gas flow, heat conduction, and nose contact pressure distribution processes. The set of geometric parameters, the set of material parameters, and the set of manufacturing process parameters are used as input variables, and the set of quality inspection features is used as the output target. A simulation prediction model is established to quickly predict the quality inspection feature set under varying parameter combinations.
6. The method for designing a nasal cannula based on reinforcement learning according to claim 5, characterized in that, S2 also includes: calibrating the simulation prediction model using experimental data to ensure that the deviation between the prediction results and the actual manufacturing conditions is controlled within a set threshold. The calibrated simulation prediction model is used as the environment model of the reinforcement learning agent to achieve a closed-loop mapping between parameter changes and quality feature prediction.
7. The nasal cannula design method based on reinforcement learning according to claim 6, characterized in that, S3 specifically includes: Based on the target oxygen supply condition parameters, a set of target performance indicators is defined; based on the quality detection feature set, a reward function is established, taking the reduction of leakage rate, the reduction of peak contact pressure, and the improvement of dimensional stability as positive incentives; Using manufacturing feasible domain constraints as constraints, negative penalties are imposed on actions that violate processing or assembly limits; thus forming a unified multi-objective reward and constraint strategy system. By synchronously binding the reward and constraint strategy with the simulation prediction model, a clear optimization direction is provided for reinforcement learning training.
8. The nasal cannula design method based on reinforcement learning according to claim 7, characterized in that, S4 specifically includes: Construct a reinforcement learning agent, defining the state as a combination of historical parameters and the corresponding predicted quality detection features; The action is defined as a joint adjustment of the geometric parameter set, the material parameter set, and the manufacturing process parameter set; A reinforcement learning algorithm based on policy gradient or proximal policy optimization is used to perform continuous action space search; The agent iteratively optimizes based on the reward function and constraint policy, forming a convergence process that gradually approaches the optimal design. The design strategy that converges under the target oxygen supply condition is obtained and used to infer the optimal parameter set.
9. The nasal cannula design method based on reinforcement learning according to claim 8, characterized in that, S5 specifically includes: Using the trained design strategy, inference is performed on the target oxygen supply condition parameters, and the optimal set of geometric parameters, optimal set of material parameters, and optimal set of manufacturing process parameters are output. The optimal set of geometric parameters is mapped to the nasal cannula mold and the mold opening compensation dimensions, and specific machining drawings are generated. The optimal material parameter set is mapped to manufacturing instructions for material formulation, hardness gradient, and elastic modulus gradient; the optimal manufacturing process parameter set is mapped to a complete process path and numerical control instructions for extrusion, cooling, traction, and secondary shaping. The mapped parameter results are synchronously output to the production control system to form an integrated manufacturing solution.
10. The method for designing a nasal cannula based on reinforcement learning according to claim 9, characterized in that, S6 specifically includes: Based on the above manufacturing scheme, conduct trial production or mass production, and collect measured data that are consistent with the quality inspection feature set; The measured data and the output of the simulation prediction model are compared item by item to identify the sources of deviation in geometric, material and process parameters; Based on the manufacturing feasible region constraint, online fine-tuning and calibration are performed on the continuous parameters of the optimal manufacturing process parameter set; The calibrated results and measured data are fed back into the simulation prediction model and the reinforcement learning agent to form incremental training samples. Update the agent strategy to enable the design strategy to adapt to fluctuations in the actual manufacturing environment and maintain output stability; The final set of geometric parameters, the final set of material parameters, and the final set of manufacturing process parameters are solidified to form a closed-loop system for design, manufacturing, and calibration.