Pressure relief valve working condition air tightness detection method based on reinforcement learning
By using a reinforcement learning-based method to drive the valve core's working state, collecting pressure data, performing leakage parameter inversion, and optimizing adaptive test formulas, the problem of inconsistency between test results and actual operating conditions and misjudgment in the airtightness testing of pressure reducing valves is solved, achieving higher detection accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN TRINOVA AUTOMOTIVE TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-21
AI Technical Summary
Existing methods for testing the air tightness of pressure reducing valves mainly rely on static sealing tests. Since the valve core is in a non-working state, it is difficult to cover the actual leakage channels and sealing contact conditions. Fixed test formulas are difficult to adapt to pressure reducing valves with different structures and batches, resulting in inconsistent test results and misjudgments compared to actual operating conditions.
A reinforcement learning-based approach is adopted, which uses a valve core drive signal to put the valve core into working state, collects pressure data sequences, uses a physical constraint neural network of a stage-gated structure to invert leakage parameters, and generates adaptive test formula parameters through a constraint reinforcement learning strategy network. Combined with a safety filtering module for correction, the test results are adaptively optimized and safety constraints are achieved.
It covers the actual leakage path and sealing contact state after the valve core moves, improving the accuracy and stability of detection, shortening the detection time, and reducing the risk of misjudgment.
Smart Images

Figure CN122430002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of airtight detection and intelligent control, and particularly to an airtight detection method for the working condition of a pressure reducing valve based on reinforcement learning. Background Art
[0002] Pressure reducing valves are widely used in scenarios such as vehicle braking systems and industrial pneumatic control systems. Their airtight performance directly affects the system pressure maintaining ability and safety. The airtight detection of existing pressure reducing valves is usually completed on the production line or test bench. After the detection gas path is connected to and sealed with the pressure reducing valve to be tested, the gas path is pressurized by a pressure regulating element and enters the pressure holding process. The pressure sensor is used to collect the data of pressure change over time, and the leakage amount is calculated by the pressure drop method, differential pressure method or flow method, and it is judged whether it is qualified according to a preset threshold. With the development of automated detection, some detection devices further integrate actuators such as solenoid valves and proportional valves, which can realize the process control of pressurization, pressure holding and exhaust, and filter, fit and judge the sampled data through a controller to improve the detection rhythm and consistency.
[0003] However, the above existing technologies still have the following deficiencies:
[0004] 1. Most detection schemes are mainly for static seal testing. When detecting, the valve core is often in a non-working state or not effectively driven, and the real leakage channels and seal contact states formed after the valve core is opened or actuated are difficult to cover, resulting in the detection results may not be consistent with the actual working conditions.
[0005] [[ID=ID=17]]2. Parameters such as pressurization rate, target pressure, pressure holding duration, and exhaust throttling are coupled with the valve core state. Existing technologies mostly adopt fixed test recipes and fixed thresholds, which are difficult to take into account pressure reducing valves with different batches and different structural differences, and are prone to misjudgment or insufficient stability.
[0006] 3. Existing judgments are mostly based on empirical thresholds or simple fitting of pressure curves, lacking the ability to invert leakage mechanism parameters and predict subsequent pressure evolution, and it is difficult to perform adaptive optimization and safety constraint control during the detection process.
[0007] Therefore, an airtight detection method for the working condition of a pressure reducing valve that can solve the above deficiencies of the existing technologies is needed. Summary of the Invention
[0008] One objective of this invention is to propose a reinforcement learning-based method for detecting the airtightness of pressure reducing valves under operating conditions. Addressing the problems of existing technologies that often employ static sealing tests, resulting in the valve core being in a non-operating state, difficulty in covering the actual leakage path and sealing contact state after valve core operation, and the inability of fixed test formulas to adapt to different structures and batches of pressure reducing valves, leading to inconsistencies between test results and actual operating conditions, and misjudgments, the following technical solution is proposed: The pressure reducing valve under test is connected to the test gas path, and an operating condition driving signal is applied to put the valve core into an operating state, collecting pressure and its corresponding time to form a sampling data sequence; the current sampling data subsequence is input into a stage judgment model to obtain a detection stage identifier; the sampling data subsequence and stage identifier are input into a physical constraint neural network with a stage-gated structure, and leakage parameters are inverted based on the stage to generate predicted values for the sampling data of the next detection period; a detection state vector is constructed based on the sampling data subsequence, stage identifier, and estimated leakage parameters, and input into a constraint reinforcement learning strategy network to output candidate test formula parameters for the next detection period. A safety filtering module corrects the candidate test formula parameters based on hard constraints to obtain an executable formula, iterating repeatedly until the judgment conditions are met and the detection result is output. This invention has the technical advantages of covering leakage paths under real working conditions, achieving adaptive optimization of test formulas while taking into account safety constraints, improving detection accuracy and stability, and shortening detection time.
[0009] This invention provides a reinforcement learning-based method for detecting the airtightness of a pressure reducing valve under operating conditions, comprising:
[0010] S1. Connect the pressure reducing valve under test to the detection gas circuit, apply gas pressure to the detection gas circuit through the pressure regulating component, and exhaust gas from the detection gas circuit through the exhaust component. Apply a working condition drive signal to the pressure reducing valve under test through the valve core drive component to make the valve core work. The sensor collects and forms a sampling data sequence, including pressure data and time information corresponding to the pressure data, and obtains the current sampling data subsequence.
[0011] S2. Input the current sampled data subsequence into the stage determination model to obtain the stage identifier of the current detection process;
[0012] S3. Input the current sampled data subsequence and stage identifier into the physical constraint neural network, including the stage gating structure, to select the physical constraint terms corresponding to the current detection stage based on the stage identifier. The physical constraint neural network performs leakage parameter inversion on the current sampled data subsequence to obtain the leakage parameter estimate and generate the sampled data prediction value for the next detection period.
[0013] S4. Construct a detection state vector based on the current sampled data subsequence, stage identifier, and leakage parameter estimate, input it into the constraint reinforcement learning policy network, obtain the candidate test formula parameters for the next detection period, and use it to generate the parameter set of the pressure setting curve for the next detection period.
[0014] S5. Input the candidate test formula parameters and the predicted values of the sampled data into the safety filtering module. The safety filtering module performs constraint correction on the candidate test formula parameters according to the preset hard constraint conditions to obtain the execution test formula parameters.
[0015] S6. Compare the estimated leakage parameters with the preset judgment conditions. If the preset judgment conditions are met, output the airtightness test result. If the preset judgment conditions are not met, control the pressure regulating component and the exhaust component to perform at least one of the pressure control and exhaust control in the next detection period according to the test formula parameters, so as to update the sampling data sequence and iterate.
[0016] Optionally, S1 includes:
[0017] Connect the pressure reducing valve to be tested to the test gas circuit and seal the test gas circuit;
[0018] The exhaust component is controlled to exhaust gas from the detection gas path, so that the detection gas path reaches the preset initial pressure;
[0019] Then, the pressure regulating component is controlled to apply gas pressure to the detection gas path, so that the pressure of the detection gas path changes according to the preset pressurization rate and reaches the preset target pressure.
[0020] During the process of exhausting and applying gas pressure, a sensor collects pressure data at a preset sampling period and records the time information corresponding to the pressure data. The pressure data is filtered to obtain processed pressure data, and the processed pressure data and the time information are combined in chronological order to form a sampling data sequence.
[0021] Optionally, S2 includes:
[0022] Input the current sampled data subsequence into the stage determination model to obtain the stage scores corresponding to the pressurization stage identifier, the pressure holding stage identifier, and the exhaust stage identifier, respectively;
[0023] The stage identifier of the current detection process is determined based on the stage score according to a preset selection rule, wherein the preset selection rule includes selecting the stage identifier with the highest stage score.
[0024] Output the identified stage identifiers.
[0025] Optionally, S3 includes:
[0026] Input the current sampled data subsequence and the stage identifier into the physical constraint neural network;
[0027] The stage gating structure in the physical constraint neural network selects physical constraint terms corresponding to the current stage based on the stage identifier. When the stage identifier is a pressurization stage identifier, physical constraint terms containing constraints based on mass conservation and throttling flow rate are selected. When the stage identifier is a pressure holding stage identifier, physical constraint terms containing constraints based on the gas equation of state and mass conservation are selected. When the stage identifier is an exhaust stage identifier, physical constraint terms containing constraints based on mass conservation and throttling flow rate are selected.
[0028] The stage gating structure includes a gating network or a gating vector. The stage gating structure outputs a selection signal or weight vector for the residuals of each physical constraint item based on the stage identifier, so as to enable or disable the residuals of different physical constraint items under different stage identifiers, and assign different weights to the residuals of different physical constraint items.
[0029] Based on the current sampled data subsequence and the selected physical constraint terms, the physical constraint neural network outputs an estimated value of the leakage parameter.
[0030] The estimated leakage parameters are leakage parameters that characterize the leakage channel of the pressure reducing valve under test. The leakage parameters include at least one of the following: equivalent leakage area, equivalent leakage orifice diameter, leakage flow coefficient, and leakage mass flow coefficient. The throttling flow relationship includes a formula for calculating the leakage flow based on the leakage parameters.
[0031] Based on the estimated leakage parameters and the current sampled data subsequence, the predicted sampled data value for the next time period is generated according to the relationship with the selected physical constraint term, and the estimated leakage parameters and the predicted sampled data value for the next time period are output.
[0032] The predicted value of the sampling data for the next time period includes a pressure prediction sequence corresponding to multiple predicted sampling times within the next detection time period, and the maximum predicted pressure and the predicted pressure rise rate are determined based on the pressure prediction sequence.
[0033] Furthermore, the leakage parameter estimate output by the physical constraint neural network also includes an uncertainty parameter corresponding to the leakage parameter estimate. The uncertainty parameter includes at least one of variance, upper bound of confidence interval, and lower bound of confidence interval. The predicted value of the sampled data in the next time period generates a predicted pressure upper bound sequence and a predicted pressure lower bound sequence based on the uncertainty parameter.
[0034] Optionally, S4 includes:
[0035] Extract the current pressure value and pressure change rate from the current sampled data subsequence according to the preset feature extraction rules;
[0036] The stage identifier is encoded to obtain the stage identifier code;
[0037] The current pressure value, the pressure change rate, the estimated leakage parameter value, and the stage identifier code are concatenated to construct a detection state vector;
[0038] The detection state vector is input into the constrained reinforcement learning policy network, which outputs candidate test recipe parameters for the next detection period.
[0039] The candidate test recipe parameters are a set of time-level parameters used to generate the pressure setting curve for the next testing period, and the pressure setting curve is a piecewise linear curve or a parameterized function curve.
[0040] The constrained reinforcement learning policy network outputs the time-level parameter set, but does not output the instantaneous control quantity per sampling period;
[0041] The candidate test recipe parameters include at least one of the following: target pressure value, pressure rise rate, pressure holding duration, and exhaust throttling parameters; and the candidate test recipe parameters are output.
[0042] The exhaust throttling parameter is a parameter characterizing the throttling capability of the exhaust component. The exhaust component includes at least one of an exhaust valve, an exhaust solenoid valve, a proportional exhaust valve, or a throttle valve. The exhaust throttling parameter includes at least one of the following: the valve opening degree of the exhaust component, the pulse width modulation duty cycle of the exhaust component, the equivalent throttling area, and the equivalent throttling orifice diameter.
[0043] Optionally, S5 includes:
[0044] Input the candidate test formula parameters and the predicted values of the sampling data from the next time period into the safety filtering module;
[0045] The safety filtration module determines the target pressure value and pressure rise rate for the next detection period based on the candidate test formula parameters, and determines the predicted maximum pressure value and predicted pressure rise rate for the next detection period based on the predicted value of the sampled data for the next period.
[0046] The predicted maximum pressure is compared with the maximum allowable pressure constraint, and the predicted pressure rise rate is compared with the maximum allowable pressure rise rate constraint.
[0047] If the comparison results show that at least one hard constraint condition is not met, the pressure target value and pressure rise rate in the candidate test recipe parameters are constrained and corrected so that the constrained and corrected pressure target value does not exceed the maximum allowable pressure constraint and the constrained and corrected pressure rise rate does not exceed the maximum allowable pressure rise rate constraint, thus obtaining the execution test recipe parameters.
[0048] If the comparison results show that all hard constraints are met, the candidate test recipe parameters are determined as the execution test recipe parameters;
[0049] Furthermore, constraining the candidate test formulation parameters includes solving an optimization problem with the candidate test formulation parameters as the objective parameters. The objective function of the optimization problem is used to minimize the deviation between the constrained test formulation parameters and the candidate test formulation parameters, and the constraints of the optimization problem include the hard constraints.
[0050] Optionally, S6 includes:
[0051] The estimated value of the leakage parameter is compared with a preset judgment condition, wherein the preset judgment condition includes a comparison condition between the estimated value of the leakage parameter and a preset leakage threshold, and a stability condition in which the change of the estimated value of the leakage parameter within a preset time window does not exceed a preset change threshold.
[0052] When the comparison results show that the preset judgment condition is met, an airtightness detection result is generated based on the estimated value of the leakage parameter and the airtightness detection result is output.
[0053] If the comparison results show that the preset judgment condition is not met, the pressure regulating component and the exhaust component are controlled to perform the pressurization operation, pressure holding operation or exhaust operation in the next detection period according to the test formula parameters, and the sensor collects the sampling data of the next detection period to obtain the updated value of the sampling data sequence.
[0054] The updated values of the sampled data sequence are merged and updated with the sampled data sequence in chronological order to obtain the updated sampled data sequence. Based on the updated sampled data sequence, the updated current sampled data subsequence is obtained by updating it according to a sliding window.
[0055] Optionally, the stage determination model, the physical constraint neural network, and the constraint reinforcement learning policy network are obtained through the following training steps:
[0056] A training sample data sequence is acquired, and stage identifiers are labeled onto the training sample data sequence. The stage identifiers include a pressurization stage identifier, a pressure holding stage identifier, and an exhaust stage identifier. The stage determination model is trained based on the training sample data sequence and the stage identifiers to obtain the stage determination model. The training sample data sequence and the pressure boundary conditions corresponding to the training sample data sequence are acquired, and physical constraint terms corresponding to the pressurization stage identifier, the pressure holding stage identifier, and the exhaust stage identifier are constructed respectively. A loss function is constructed based on the data fitting error and the residuals of the physical constraint terms, and the physical constraint neural network is made more sensitive to the stage through the stage gating structure in the loss function. The network employs different physical constraints under different stage identifiers. Based on the loss function, the physical constraint neural network is trained to obtain the physical constraint neural network. A training environment is constructed, which outputs a sampled data sequence corresponding to the test formula parameters when given test formula parameters. The detection state vector is used as the state input, and the test formula parameters are used as the action output. A reward function is established with the airtightness detection accuracy and detection time as optimization objectives, and a constraint function corresponding to the hard constraint conditions is established. The reward function and the constraint function are jointly optimized using Lagrange multipliers, and the Lagrange multipliers are updated to train the constraint reinforcement learning policy network.
[0057] The beneficial effects of this invention are:
[0058] 1. By applying a working condition drive signal to the pressure reducing valve under test through the valve core drive component, the valve core is put into working state and pressure time series is collected during the pressurization, pressure holding and venting process. This can cover the actual leakage channel and sealing contact state formed after the valve core is activated, thereby reducing the risk of missed detection and misjudgment caused by the inconsistency between static sealing test and actual use conditions.
[0059] 2. A physical constraint neural network with a stage-gated structure is adopted. Based on the detection stage, the corresponding physical constraint terms are selected to perform leakage parameter inversion and output the pressure prediction for the next period. This changes the leakage judgment from empirical pressure drop thresholds to parameterized representation based on physical constraints, improving the detection stability and accuracy under noise, batch differences and structural differences.
[0060] 3. The constrained reinforcement learning strategy network outputs time-period test recipe parameters, and combines them with a safety filtering module to correct the pressure target value and pressure rise rate according to hard constraints, so as to realize the adaptive optimization and safe and controllable execution of the test recipe, shorten the detection time and improve the detection efficiency while ensuring the maximum allowable pressure and pressure rise rate constraints. Attached Figure Description
[0061] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0062] Figure 1 This is a flowchart of the pressure reducing valve airtightness detection method based on reinforcement learning proposed in this invention;
[0063] Figure 2 This is a flowchart illustrating the execution of the physical constraint neural network in step S3 of the present invention. Detailed Implementation
[0064] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0065] refer to Figure 1-2 A reinforcement learning-based method for detecting the airtightness of pressure reducing valves under operating conditions includes:
[0066] S1. Connect the pressure reducing valve under test to the detection gas circuit, apply gas pressure to the detection gas circuit through the pressure regulating component, and exhaust gas from the detection gas circuit through the exhaust component. Apply a working condition drive signal to the pressure reducing valve under test through the valve core drive component to make the valve core work. The sensor collects and forms a sampling data sequence, including pressure data and time information corresponding to the pressure data, and obtains the current sampling data subsequence.
[0067] S2. Input the current sampled data subsequence into the stage determination model to obtain the stage identifier of the current detection process;
[0068] S3. Input the current sampled data subsequence and stage identifier into the physical constraint neural network, including the stage gating structure, to select the physical constraint terms corresponding to the current detection stage based on the stage identifier. The physical constraint neural network performs leakage parameter inversion on the current sampled data subsequence to obtain the leakage parameter estimate and generate the sampled data prediction value for the next detection period.
[0069] S4. Construct a detection state vector based on the current sampled data subsequence, stage identifier, and leakage parameter estimate, input it into the constraint reinforcement learning policy network, obtain the candidate test formula parameters for the next detection period, and use it to generate the parameter set of the pressure setting curve for the next detection period.
[0070] S5. Input the candidate test formula parameters and the predicted values of the sampled data into the safety filtering module. The safety filtering module performs constraint correction on the candidate test formula parameters according to the preset hard constraint conditions to obtain the execution test formula parameters.
[0071] S6. Compare the estimated leakage parameters with the preset judgment conditions. If the preset judgment conditions are met, output the airtightness test result. If the preset judgment conditions are not met, control the pressure regulating component and the exhaust component to perform at least one of the pressure control and exhaust control in the next detection period according to the test formula parameters, so as to update the sampling data sequence and iterate.
[0072] In this specific embodiment, S1 includes:
[0073] The test is performed in a closed test gas circuit, which includes a gas source interface, a pressure regulating component, a pressure reducing valve interface under test, an exhaust component, and a pressure sensor pressure tap in series. The controller connects the inlet and outlet of the pressure reducing valve under test to the test gas circuit through an airtight connector and completes the sealing. The sealing method is to set a sealing ring that mates with the interface on each connection end face and tighten the interface to a specified torque to eliminate axial clearance. At the same time, the controller puts the pressure regulating component and the exhaust component in the closed state to confirm that the test gas circuit is isolated from the outside world only through the pressure reducing valve under test.
[0074] The controller then drives the exhaust components to open and maintains this position until the pressure in the detection air circuit reaches the preset initial pressure. The preset initial pressure An absolute pressure of 101.3 kPa was used to ensure that the detection gas path was consistent with the ambient pressure. The pressure sensor was a range... And accuracy The absolute pressure sensor acquires its current signal through a 16-bit analog-to-digital converter to obtain pressure data;
[0075] During and after the exhaust process, the controller applies a working condition drive signal to the pressure reducing valve under test through the valve core drive component to put the valve core into the working state. The valve core drive component is an electromagnetic actuator driver that outputs a two-stage current drive sequence. In each drive cycle, the drive sequence first outputs a pull-in current. continued Then output holding current continued This allows the valve core to complete the suction and maintain the target open state during the holding phase, and to remain in the working state during continuous drive cycles;
[0076] After exhaust is complete, the controller shuts off the exhaust components and controls the pressure regulating components to apply gas pressure to the detection gas path, ensuring that the detection gas path pressure is maintained at a preset pressurization rate. From the preset initial pressure linearly rise to the preset target pressure The pressure regulating component is an electronically controlled proportional pressure regulating valve, and the controller adjusts its control voltage in a closed-loop manner based on feedback from the pressure sensor to track the preset pressurization rate. With preset target pressure And when the preset target pressure is reached Then, the proportional pressure regulating valve is kept at the holding opening to proceed to the subsequent pressure holding process;
[0077] Throughout the entire process of exhausting and applying gas pressure, the controller operates at a sampling period. Read discrete pressure data from the pressure sensor And record the time information corresponding to the pressure data. ,in Indicates the sampling sequence number and from Start incrementing, the time information Generated by the controller with the sampling trigger time aligned with the internal clock and satisfying The sampling period is the interval between adjacent sampling times. ;
[0078] To suppress high-frequency noise introduced by valve spool movement, electromagnetic interference, and gas path pulsation, the controller processes the pressure data. Perform a first-order discrete low-pass filter to obtain the processed pressure data. The filtering relationship is as follows:
[0079] ;
[0080] in Indicates the first Processed pressure data at each sampling time point. Indicates the first Processed pressure data at each sampling time point. Indicates the first The original pressure data at each sampling time point, Represents the filter coefficients and takes And set the initial value to ;
[0081] The controller will assign each sampling number Corresponding processed pressure data With time information The sampled data is combined and stored in ascending order of sampling sequence number to form a sampled data sequence. The sampled data sequence contains multiple consecutive sampling points, and each sampling point consists of time information and processed pressure data under the same sampling sequence number. After forming the sampled data sequence, the latest consecutive sampling points are selected from the sequence. Each sampling point captures a subsequence of the currently sampled data.
[0082] In this specific embodiment, S2 includes:
[0083] Execute based on the current sampled data subsequence, which is composed of the most recent It consists of 10 sampling points, and each sampling point contains time information. Compared with processed pressure data ,in The sampling sequence number is the sampling period, and the time interval between adjacent sampling points is the sampling period. ;
[0084] The controller constructs a phase-based input for the current sampled data subsequence and calculates the pressure change rate sequence according to the sampling sequence number. And order And let ,in Indicates the first The rate of pressure change corresponding to each sampling point Indicates the first Processed pressure data corresponding to each sampling point Indicates the first Processed pressure data corresponding to each sampling point Indicates the sampling period;
[0085] The controller will process the pressure data. With the pressure change rate A two-dimensional feature sequence is formed by concatenating data at each sampling point and then fed into a stage determination model. This stage determination model is a time-series classification neural network with fixed parameters, deployed on the controller for inference and execution. The model input is a dimensionless vector. The feature tensors are arranged in chronological order. The model structure includes a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a temporal pooling layer, a gated recurrent unit layer, and a fully connected output layer. The first one-dimensional convolutional layer has 16 convolutional kernels with a kernel length of 5 and uses effective convolution with a stride of 1 and ReLU activation. The second one-dimensional convolutional layer has 32 convolutional kernels with a kernel length of 5 and uses effective convolution with a stride of 1 and ReLU activation. The temporal pooling layer is a max pooling layer of length 2 with a stride of 2 for downsampling the time dimension. The gated recurrent unit layer is a single-layer GRU with 32 hidden units to extract stage features across time. The fully connected output layer outputs three stage scores and obtains the stage score vector through Softmax normalization. ,in This indicates the stage score corresponding to the pressurization stage identifier. This indicates the stage score corresponding to the pressure holding stage identifier. This indicates the stage score corresponding to the exhaust stage identifier, and the sum of the three scores is 1;
[0086] The controller determines and outputs the stage identifier according to a preset selection rule, which is to select the stage identifier with the highest stage score and is expressed by the formula:
[0087] ;
[0088] in The stage identifier indicates the current testing phase of the testing process. The values {press, hold, exh} represent the stage category indexes and correspond to the pressurization stage identifier, the holding stage identifier, and the exhaust stage identifier, respectively. The stage category is indicated as Stage score;
[0089] When multiple stage scores are tied for the highest, the controller selects the stage identifier that is the same as the stage identifier output in the previous sampling period as the highest score. When the previous sampling period has not yet output a stage identifier, Set as the pressurization stage identifier.
[0090] In this specific embodiment, S3 includes:
[0091] In the stage marker Based on the current sampled data subsequence, the controller will process the pressure data from the current sampled data subsequence. Convert kPa to Pa and compare with pressure change rate series The dimensions are composed in order of sampling sequence number. The input feature sequence, where and The sampling sequence number;
[0092] The controller will identify the stage. The encoding is a one-hot vector of length 3 and used as the selection signal for stage gating, enabling the physical constraint neural network to... When pressing, enable "Mass Conservation Relationship Constraint Residual" and "Throttling Flow Relationship Constraint Residual" and disable "Gas Equation of State Constraint Residual". Enable "Gas state equation constraint residual" and "mass conservation relation constraint residual" while disabling "throttling flow relation constraint residual". Enable "Mass Conservation Relationship Constraint Residual" and "Throttling Flow Relationship Constraint Residual" and disable "Gas Equation of State Constraint Residual";
[0093] The physical constraint neural network is an edge-side inference model and consists of a shared temporal feature extractor, a leakage parameter output head, and a physical predictor. The shared temporal feature extractor includes a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a temporal pooling layer, and a gated recurrent unit layer. The first one-dimensional convolutional layer has 32 channels and a kernel length of 7, and uses ReLU activation. The second one-dimensional convolutional layer has 64 channels and a kernel length of 5, and uses ReLU activation. The temporal pooling layer is a max pooling layer with a length of 2 and a stride of 2. The gated recurrent unit layer is a single-layer GRU with 64 hidden units and outputs the final hidden state as a temporal feature vector.
[0094] The leakage parameter output header consists of two fully connected layers. The first layer has an output dimension of 32 and uses ReLU activation, while the second layer has an output dimension of 2, representing the estimated leakage parameter and the uncertainty parameter, respectively. The estimated leakage parameter is selected as the leakage mass flow rate coefficient. And it is guaranteed through Softplus activation. The uncertainty parameter is selected as being similar to... corresponding variance And it is guaranteed through exponential mapping ;
[0095] Physics predictor in obtaining Then, a pressure prediction sequence for the next detection period is generated, the length of which is fixed. There are 1 sampling point and the sampling period is 1 And use the processed pressure data corresponding to the end of the current sampled data subsequence as the initial condition. ,in Indicates the last sample number of the current sampled data subsequence;
[0096] The pressure prediction sequence is obtained using the isothermal ideal gas assumption and discrete integral with mass conservation, and is determined according to the stage identifier at each prediction step. Set the boundary quality flow rate, where when Pressing will increase the inlet quality flow rate. Based on the linear throttling relationship, the gas supply pressure Current forecast pressure Calculate and take the inlet throttling coefficient. and make the export quality flow rate ,when Hold the season and ,when season And will Based on the current predicted pressure according to the linear throttling relationship Environmental pressure Calculate and take the outlet throttling coefficient. and the leaked mass flow rate According to the linear throttling relationship and Calculate and in season ;
[0097] The discrete integral adopts the following pressure update relation:
[0098] ;
[0099] in Indicates the first The unit of the predicted pressure value corresponding to each predicted sampling time is: Indicates the prediction step number and takes Denotes the gas specificity gas constant and takes Indicates the temperature of the gas being detected and taken This represents the equivalent volume formed by the test gas path and the pressure reducing valve under test, and takes... Indicates the first The inlet mass flow rate of each prediction step, in units of Indicates the first The output mass flow rate of each prediction step, in units of , Indicates the first Leakage mass flow rate for each prediction step, in units of , Indicates the sampling period, with the unit being seconds (s).
[0100] After obtaining the pressure prediction sequence, the controller calculates the maximum predicted pressure as the largest element of the pressure prediction sequence and calculates the rate of increase of the predicted pressure as the pressure difference between adjacent prediction sampling times divided by the sampling period. The maximum positive value;
[0101] To output the uncertainty corresponding to the leakage parameter estimate and to form the upper and lower bound sequences of the predicted pressure, the controller uses... As the mean, with Construct a truncated normal distribution for variance and generate A set of leakage mass flow coefficient samples were obtained, and each sample was substituted into the above physical predictor to obtain... Group pressure prediction sequence, and at each prediction sampling time... The 2.5 percentile of each predicted pressure value is used as the lower bound for the predicted pressure, and the 97.5 percentile is used as the upper bound for the predicted pressure. The final output is the estimated leakage parameter value. Uncertainty parameters The pressure prediction sequence for the next detection period, as well as the upper and lower bound sequences for predicted pressure.
[0102] In this specific embodiment, S4 includes:
[0103] In the current sampled data subsequence, stage identifier and estimated leakage parameters With uncertainty parameter Based on this, the controller extracts the current pressure value from the end of the currently sampled data subsequence. With pressure change rate ,in Take the processed pressure data corresponding to the end of the current sampled data subsequence and convert it from kPa to Take the pressure change rate corresponding to the end sampling point in S2 and use it as... Converted to ;
[0104] The controller will identify the stage. Encoding as stage identifier encoding , among which when press ,when hold ,when hour ;
[0105] The controller will The detection state vector is formed by concatenating the stage identifier encoding with the detection state vector and inputting it into the constrained reinforcement learning policy network. The detection state vector is defined as follows:
[0106] ;
[0107] in Represents the detection state vector. This indicates the current pressure value, with the unit being Pa. Represents the rate of change of pressure and the unit is This represents an estimated value of the leakage mass flow coefficient, with units of . Indicates and The corresponding variance and its unit is This indicates the coded component representing the pressurization phase. This indicates the coded component representing the pressure holding stage. This indicates the exhaust stage identification code component. Indicates transpose;
[0108] The constrained reinforcement learning policy network is a deterministic policy network deployed on the controller for inference. The network input dimension is 7, and a fixed-value normalization process is performed before input. The normalization rule is to... Divide by To obtain the pressure characteristic with dimensionless accuracy, Divide by To obtain the characteristic of the rate of change with dimensionless 1, Divide by To obtain leakage characteristics with a dimension of 1, Divide by To obtain the uncertainty feature with dimension 1 while keeping the stage identifier code unchanged;
[0109] The policy network structure sequentially includes a first fully connected layer, a second fully connected layer, and an output layer. The first fully connected layer has 64 neurons and uses ReLU activation; the second fully connected layer also has 64 neurons and uses ReLU activation; and the output layer has 4 neurons and uses ReLU activation. Activate to limit each output component to ;
[0110] The controller linearly maps the four output components obtained from the output layer to candidate test recipe parameters for the next detection period, forming a time-level parameter set. The candidate test recipe parameters include the pressure target value. Rate of pressure rise Pressure holding duration With exhaust throttling parameters ,in The mapping range is taken The mapping range is taken The mapping range is taken The mapping range is taken and Defined as the pulse width modulation duty cycle of the exhaust component;
[0111] The candidate test recipe parameters are used to generate the pressure setting curve parameter set for the next testing period, and the pressure setting curve is a piecewise linear curve. The controller fixes the duration of the next testing period to [value missing]. Within this duration, only time-level parameters are issued, without issuing instantaneous control quantities for each sampling period, where the stage identifier... When pressing The pressure setpoint was adjusted to change the slope from... Linear change to And in reaching Then keep until End, when stage marker Hold the pressure setting value. And at least continue And continue to maintain this for the remaining time, when the stage marker... Keep the pressure regulating component closed and operate at duty cycle. Drive exhaust components continuously To achieve time-level exhaust throttling control.
[0112] In this specific embodiment, S5 includes:
[0113] The process is executed on the controller by the safety filtering module, whose input is the candidate test recipe parameters. and the pressure prediction sequence for the next detection period With the upper bound sequence of predicted pressure ,in Indicates the candidate pressure target value and the unit is Indicates the candidate pressure rise rate and the unit is Indicates the candidate holding pressure duration and the unit is This represents the candidate exhaust throttling parameter and is a pulse width modulation duty cycle. Indicates the predicted sampling point number and takes and Indicates the first The predicted pressure values at each predicted sampling point are in units of... Indicates the first The upper bound of the predicted pressure at each predicted sampling point, with the unit being Pa;
[0114] The security filtering module first relies on Calculate the predicted maximum pressure And calculate the predicted rate of pressure rise. ,in Defined as exist to The maximum value within the range, Defined as the upper bound difference in pressure between adjacent predicted sampling points exist to The maximum value within the range and The sampling period;
[0115] The safety filtering module reads the preset hard constraint conditions and sets the maximum allowable pressure constraint to... The maximum permissible rate of pressure rise constraint is set to And set the lower limit of the pressure target value to And the lower limit of the rate of pressure rise is set to To ensure executability;
[0116] The security filtering module will and Compare and and The comparison shows that at least one hard constraint is not met. The safety filtering module then converts the predicted information into a tightened execution upper limit and sets:
[0117] And let ;
[0118] in Indicates the upper limit of the effective pressure used for correction, and the unit is . Indicates the upper limit of the effective rise rate used for correction, and the unit is . ;
[0119] Based on this, the safety filtering module performs constraint correction on the pressure target value and pressure rise rate in the candidate test formula parameters. The constraint correction is achieved by solving a quadratic convex optimization problem and is defined as follows:
[0120] st ;
[0121] in This represents the target value of the execution pressure after constraint correction, and the unit is... Indicates the rate of increase of the execution pressure after constraint correction, and the unit is . This represents the target pressure value in the optimization variables, and the unit is... Represents the rate of increase of pressure in the optimization variables, with units of . Indicates the weight of the deviation from the pressure target value and takes Indicates the weight of the deviation in the rate of pressure rise and takes and These represent the candidate pressure target value and the candidate pressure rise rate, respectively.
[0122] Since this optimization problem is a quadratic convex optimization with box constraints, the safety filtering module uses the analytical projection method to solve it and implements it at the controller end using fixed-point arithmetic. Projected onto interval get And will Projected onto interval get ;
[0123] When the comparison results show that all hard constraints are met, the safety filtering module directly sets... And let and maintain and ,in Indicates the duration of pressure holding, and the unit is... This indicates that exhaust throttling parameters are being applied and that the duty cycle is pulse width modulation.
[0124] The safety filtration module will execute the test formula parameters. The output to S6 is used to drive the pressure regulating component and the exhaust component to perform pressure control and exhaust control for the next detection period.
[0125] In this specific embodiment, S6 includes:
[0126] The process is executed iteratively on the controller and linked in a closed loop with the data acquisition process. After each S3 operation, the controller reads the estimated leakage parameters. And compare it with the preset judgment conditions, while simultaneously setting the current The leak parameter history buffer, along with the controller timestamp of its generation time, is written to the leak parameter history buffer, which is of length [length missing]. A circular queue that only stores the most recent The estimated leakage parameters corresponding to the next iteration are used for stability determination, and the preset leakage threshold is set as follows: The preset change threshold is set to And the stability time window is defined as the most recent The time range covered by each iteration;
[0127] The controller determines whether the preset judgment conditions are met based on the following rules:
[0128] ;
[0129] in This is a Boolean value indicating whether a preset condition is met, and its value is either true or false. This represents the estimated leakage mass flow rate coefficient for this iteration, with units of 1. Indicates the preset leakage threshold and the unit is , Indicates the number of leaked parameters in the history buffer. The estimated value of the leakage mass flow coefficient is in units of Indicates the length of the history buffer for leaked parameters and takes... , Indicates the preset change threshold and the unit is , and These represent operators that retrieve the maximum and minimum values within a given set of indices, respectively. This indicates a logical AND operation.
[0130] when The real-time controller generates and outputs an airtightness detection result, which includes estimated leakage parameters. Stage identifiers The judgment conclusion is defined as qualified and simultaneously archives and stores the sampling data sequence, leakage parameter historical buffer, and test formula parameters of this test and ends the iteration;
[0131] when When the result is false, the controller executes the test recipe parameters. Execute the control actions for the next detection period and update the sampling data sequence, where the stage identifier... When pressed, the controller switches the pressure regulation component to closed-loop pressure control mode and uses a piecewise linear pressure set curve as a reference input, so that the detected gas pressure is at the same rate as the execution pressure rise. From current pressure to the target pressure level And remain there until the end of the next inspection period, with exhaust components kept closed, when the stage indicator is displayed. The time controller maintains the pressure regulating component at the specified opening and closes the exhaust component, ensuring that the pressure in the detection air path is at least... Continue sampling throughout the duration until the end of the next testing period, when the phase marker... The time controller shuts off the pressure regulating component and drives the exhaust component in pulse width modulation mode with the duty cycle set to [value missing]. This allows for time-level throttling and exhaust, which continues until the end of the next detection period. Meanwhile, the valve core drive component maintains a two-stage current drive sequence to ensure that the valve core remains in a working state.
[0132] During the next detection period, the sensor continues at the sampling period. Pressure data is collected and processed using the same first-order discrete low-pass filtering as S1 to obtain processed pressure data. Time information corresponding to each sampling point is recorded to form an updated sampling data sequence. The controller appends the updated sampling data sequence values to the existing sampling data sequence in chronological order to obtain an updated sampling data sequence. Based on the updated sampling data sequence, the controller starts from the latest continuous... Each sampling point is used to capture and update the current sampling data subsequence to enter the next round of iterative processing from S2 to S6.
[0133] In this specific embodiment, the stage determination model, the physical constraint neural network, and the constraint reinforcement learning strategy network are obtained through offline training on the same set of working condition testing benches and then used for online inference after the parameters are fixed at the controller. The training data is collected by a device consisting of a detection air path, a pressure regulating component, an exhaust component, a valve core driving component, and a pressure sensor, and the collection process is consistent with S1 and the sampling period is maintained. Sliding window length Predicted time domain length Each training sample data sequence in the training dataset contains time information arranged by sample number. Compared with processed pressure data The controller synchronously records the pressure regulating component setting, the opening and closing status of the exhaust component, and the valve core drive signal for annotation and boundary condition construction.
[0134] The training of the stage determination model is completed using supervised learning. The controller labels each sampling point of the training sampled data sequence with stage identifiers based on the control command time periods of the pressure regulating component and the exhaust component. The condition where the pressure regulating component is in pressurization mode and the exhaust component is closed is marked as follows: "press" is indicated when the pressure regulating component is in the pressure holding command and the venting component is closed. hold, marked as when the exhaust component is in the open command position. and based on Calculate the rate of change of pressure As a second input channel, thus forming a dimension of The input feature sequence, the stage determination model structure is consistent with S2 and the parameters are fixed as two layers of one-dimensional convolution plus temporal pooling plus a single layer of GRU plus a fully connected output layer, and outputs three stage scores. During training, cross-entropy loss was used, and the Adam optimizer was employed. The batch size was set to 128, and the learning rate was set to [value missing]. The number of training rounds is set to 50, and after each round of training, the optimal parameters are selected based on the accuracy of the validation set and solidified into a stage judgment model.
[0135] The training of the physical constraint neural network employs a joint loss of "data fitting error + physical constraint term residuals" and uses a stage gating structure to enable different physical constraint term residuals under different stage identifiers. The training input is constructed using the same method as the stage determination model. Feature sequence with accompanying stage identifier One-hot encoding is used as the gating input, and the training labels are consecutive samples of the same data sequence after the end of the window. The actual post-processed pressure data corresponding to each sampling point is used to constrain the pressure prediction sequence output by the network. The boundary conditions are determined by the known and recorded parameters of the test bench and are fixed by the supply gas pressure. Environmental pressure Inlet throttling coefficient Export throttling coefficient Equivalent volume Gas ratio gas constant Gas temperature Among them, the stage gating structure is in press and When the residuals constrained by the mass conservation relation and the throttling flow rate relation are enabled, and the residuals constrained by the gas equation of state are disabled, in When holding, enable the residual constraint of the gas state equation and the residual constraint of the mass conservation relationship, and disable the residual constraint of the throttling flow relationship. The residuals of the enabled physical constraint terms are all weighted at 1, and the residuals of the disabled physical constraint terms are weighted at 0.
[0136] The physical constraint neural network structure is consistent with S3, and the leakage parameter output is an estimated value of the leakage mass flow coefficient. Its variance During training, the Adam optimizer was used, with a batch size of 64 and a learning rate of [value missing]. The number of training rounds was set to 80, and the minimum weighted sum of "mean square error of pressure prediction + mean value of physical constraint residual" on the validation set was used as the model selection criterion and solidified into a physical constraint neural network.
[0137] The training of the constrained reinforcement learning policy network employs a model-based closed-loop training environment, wherein the training environment operates under given test formula parameters. The system invokes the established stage determination model and physical constraint neural network to generate a sampling data sequence corresponding to the test formula parameters and updates the detection state vector. in Its structure is consistent with S4 and includes the current pressure value, pressure change rate, With phase identifier encoding, the action space is defined as a time-period level test recipe parameter set rather than a sample-period control quantity, and uses the same mapping range as S4 to ensure that the action is executable;
[0138] The reward function of the training environment uses the airtightness detection accuracy and detection time as joint optimization objectives. The reward for each detection period is defined as follows: a positive reward of 10 when a stable judgment is achieved and a leak is correctly determined; a time penalty of -0.1 for each detection period; and a penalty of -10 for an incorrect leak judgment. This drives the strategy to shorten the detection time and avoid false positives. The constraint function is consistent with the hard constraint conditions and is defined as the excess amount calculated from the upper limit sequence of predicted pressure. The excess amount also includes the maximum allowable pressure constraint. With maximum permissible rate of pressure rise constraint Exceeding either or both limits incurs a positive constraint cost.
[0139] The policy network parameters are denoted as Furthermore, it employs two layers of 64-dimensional fully connected ReLU hidden layers and The output layer is trained using Lagrange multipliers. The reward function and constraint function are jointly optimized, and the policy is updated according to the degree of constraint violation after each policy update. The joint optimization objective is defined as:
[0140] ;
[0141] in Indicates the joint optimization objective. This represents the network parameters that constrain the reinforcement learning policy. Describes a Lagrange multiplier with a range of values of 1. This represents the expectation operator for the sampled trajectory of the training environment. Indicates the detection time period number. Indicates the maximum number of detection periods per round and takes... Represents the discount factor and takes Indicates the first The return for each testing period, Indicates the first The constraint cost of each detection period;
[0142] Training employs a deterministic policy gradient algorithm with a double-Q network and sets the empirical replay capacity to [value missing]. The batch size is 256, and the learning rates for both the policy network and the Q network are 1. The target network soft update coefficient is 0.005, and the Lagrange multiplier update step size is set to... Furthermore, the objective constraint cost is set to 0 to drive the policy to tend to satisfy the hard constraints during training. The total number of training rounds is set to 5000, and the model with the "shortest average detection time and constraint cost of 0" in the verification environment is used as the final constraint reinforcement learning policy network and is solidified and deployed on the controller for S4 inference output of candidate test recipe parameters.
[0143] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0144] This invention establishes a closed-loop coupling between pressure-time sampling data driven by operating conditions and stage identification, leakage parameter inversion, and adaptive test formula optimization: The pressure-reducing valve under test is brought into actual working condition by valve core actuation; the stage determination model segmentally identifies the pressurization, pressure holding, and venting processes; a physical constraint neural network inverts leakage parameters under physical mechanism constraints and provides pressure evolution predictions for the next detection period; a constraint reinforcement learning strategy network outputs test formula parameters for the next period based on leakage parameters and stage information, and executes the formula after constraint correction by a safety filtering module. Therefore, detection no longer relies on static sealing and fixed threshold formulas, but can perform online evaluation and dynamic decision-making based on the actual leakage channel and sealing contact state after valve core movement, thereby improving the consistency between detection results and actual operating conditions, while shortening detection time and improving detection stability while ensuring safety boundaries.
[0145] To address the aforementioned technical issues, this invention improves the algorithm structure for condition-based detection in several ways: First, it employs a physical constraint neural network with condition-segmented gating, enabling matching physical constraint residuals at different detection stages to reduce inversion bias caused by inconsistencies in mechanisms across stages, making leakage parameter estimation more stable and reliable. Second, the reinforcement learning action space uses time-level outputs parameterized with test recipes, replacing the instantaneous control quantities of each sampling cycle with parameters from the entire pressure setting curve, reducing control jitter and better adapting to the production line detection cycle. Third, it constructs a two-layer constraint structure, achieving real-time correction on the execution side through prediction-based hard constraint safety filtering, and achieving long-term constraint satisfaction and performance optimization of the strategy on the training side through Lagrange soft constraints, thereby better balancing detection accuracy, efficiency, and safety under different batch and structural differences.
Claims
1. A method for detecting the airtightness of a pressure reducing valve under operating conditions based on reinforcement learning, characterized in that, include: S1. Connect the pressure reducing valve under test to the detection gas circuit, apply gas pressure to the detection gas circuit through the pressure regulating component, and exhaust gas from the detection gas circuit through the exhaust component. Apply a working condition drive signal to the pressure reducing valve under test through the valve core drive component to make the valve core work. The sensor collects and forms a sampling data sequence, including pressure data and time information corresponding to the pressure data, and obtains the current sampling data subsequence. S2. Input the current sampled data subsequence into the stage determination model to obtain the stage identifier of the current detection process; S3. Input the current sampled data subsequence and stage identifier into the physical constraint neural network, including the stage gating structure, to select the physical constraint terms corresponding to the current detection stage based on the stage identifier. The physical constraint neural network performs leakage parameter inversion on the current sampled data subsequence to obtain the leakage parameter estimate and generate the sampled data prediction value for the next detection period. S4. Construct a detection state vector based on the current sampled data subsequence, stage identifier, and leakage parameter estimate, input it into the constraint reinforcement learning policy network, obtain the candidate test formula parameters for the next detection period, and use it to generate the parameter set of the pressure setting curve for the next detection period. S5. Input the candidate test formula parameters and the predicted values of the sampled data into the safety filtering module. The safety filtering module performs constraint correction on the candidate test formula parameters according to the preset hard constraint conditions to obtain the execution test formula parameters. S6. Compare the estimated leakage parameters with the preset judgment conditions. If the preset judgment conditions are met, output the airtightness test result. If the preset judgment conditions are not met, control the pressure regulating component and the exhaust component to perform at least one of the pressure control and exhaust control in the next detection period according to the test formula parameters, so as to update the sampling data sequence and iterate.
2. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 1, characterized in that, S1 includes: Connect the pressure reducing valve to be tested to the test gas circuit and seal the test gas circuit; The exhaust component is controlled to exhaust gas from the detection gas path, so that the detection gas path reaches the preset initial pressure; Then, the pressure regulating component is controlled to apply gas pressure to the detection gas path, so that the pressure of the detection gas path changes according to the preset pressurization rate and reaches the preset target pressure. During the process of exhausting and applying gas pressure, a sensor collects pressure data at a preset sampling period and records the time information corresponding to the pressure data. The pressure data is filtered to obtain processed pressure data, and the processed pressure data and the time information are combined in chronological order to form a sampling data sequence.
3. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 1, characterized in that, S2 include: Input the current sampled data subsequence into the stage determination model to obtain the stage scores corresponding to the pressurization stage identifier, the pressure holding stage identifier, and the exhaust stage identifier, respectively; The stage identifier of the current detection process is determined based on the stage score according to a preset selection rule, wherein the preset selection rule includes selecting the stage identifier with the highest stage score. Output the identified stage identifiers.
4. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 1, characterized in that, S3 includes: Input the current sampled data subsequence and the stage identifier into the physical constraint neural network; The stage gating structure in the physical constraint neural network selects physical constraint terms corresponding to the current stage based on the stage identifier. When the stage identifier is a pressurization stage identifier, physical constraint terms containing constraints based on mass conservation and throttling flow are selected. When the stage identifier is a pressure holding stage identifier, physical constraint terms containing constraints based on the gas equation of state and mass conservation are selected. When the stage identifier is an exhaust stage identifier, physical constraint terms containing constraints based on mass conservation and throttling flow are selected. The stage gating structure includes a gating network or a gating vector. The stage gating structure outputs a selection signal or weight vector for the residuals of each physical constraint item based on the stage identifier, so as to enable or disable the residuals of different physical constraint items under different stage identifiers, and assign different weights to the residuals of different physical constraint items. Based on the current sampled data subsequence and the selected physical constraint terms, the physical constraint neural network outputs an estimated value of the leakage parameter. The estimated leakage parameters are leakage parameters that characterize the leakage channel of the pressure reducing valve under test. The leakage parameters include at least one of the following: equivalent leakage area, equivalent leakage orifice diameter, leakage flow coefficient, and leakage mass flow coefficient. The throttling flow relationship includes a formula for calculating the leakage flow based on the leakage parameters. Based on the estimated leakage parameters and the current sampled data subsequence, the predicted sampled data value for the next time period is generated according to the relationship with the selected physical constraint term, and the estimated leakage parameters and the predicted sampled data value for the next time period are output. The predicted value of the sampling data for the next time period includes a pressure prediction sequence corresponding to multiple predicted sampling times within the next detection time period, and the maximum predicted pressure and the predicted pressure rise rate are determined based on the pressure prediction sequence.
5. The method for detecting the airtightness of a pressure reducing valve under operating conditions based on reinforcement learning according to claim 1, characterized in that, S4 includes: Extract the current pressure value and pressure change rate from the current sampled data subsequence according to the preset feature extraction rules; The stage identifier is encoded to obtain the stage identifier code; The current pressure value, the pressure change rate, the estimated leakage parameter value, and the stage identifier code are concatenated to construct a detection state vector; The detection state vector is input into the constrained reinforcement learning policy network, which outputs candidate test recipe parameters for the next detection period. The candidate test recipe parameters are a set of time-level parameters used to generate the pressure setting curve for the next testing period, and the pressure setting curve is a piecewise linear curve or a parameterized function curve. The constrained reinforcement learning policy network outputs the time-level parameter set, but does not output the instantaneous control quantity per sampling period; The candidate test recipe parameters include at least one of the following: target pressure value, pressure rise rate, pressure holding duration, and exhaust throttling parameters; and the candidate test recipe parameters are output. The exhaust throttling parameter is a parameter characterizing the throttling capability of the exhaust component. The exhaust component includes at least one of an exhaust valve, an exhaust solenoid valve, a proportional exhaust valve, or a throttle valve. The exhaust throttling parameter includes at least one of the following: the valve opening degree of the exhaust component, the pulse width modulation duty cycle of the exhaust component, the equivalent throttling area, and the equivalent throttling orifice diameter.
6. The method for detecting the airtightness of a pressure reducing valve under operating conditions based on reinforcement learning according to claim 1, characterized in that, S5 include: Input the candidate test formula parameters and the predicted values of the sampling data from the next time period into the safety filtering module; The safety filtration module determines the target pressure value and pressure rise rate for the next detection period based on the candidate test formula parameters, and determines the predicted maximum pressure value and predicted pressure rise rate for the next detection period based on the predicted value of the sampled data for the next period. The predicted maximum pressure is compared with the maximum allowable pressure constraint, and the predicted pressure rise rate is compared with the maximum allowable pressure rise rate constraint. If the comparison results show that at least one hard constraint condition is not met, the pressure target value and pressure rise rate in the candidate test recipe parameters are constrained and corrected so that the constrained and corrected pressure target value does not exceed the maximum allowable pressure constraint and the constrained and corrected pressure rise rate does not exceed the maximum allowable pressure rise rate constraint, thus obtaining the execution test recipe parameters. If the comparison results show that all hard constraints are met, the candidate test recipe parameters are determined as the execution test recipe parameters.
7. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 1, characterized in that, S6 include: The estimated value of the leakage parameter is compared with a preset judgment condition, wherein the preset judgment condition includes a comparison condition between the estimated value of the leakage parameter and a preset leakage threshold, and a stability condition in which the change of the estimated value of the leakage parameter within a preset time window does not exceed a preset change threshold. When the comparison results show that the preset judgment condition is met, an airtightness detection result is generated based on the estimated value of the leakage parameter and the airtightness detection result is output. If the comparison results show that the preset judgment condition is not met, the pressure regulating component and the exhaust component are controlled to perform the pressurization operation, pressure holding operation or exhaust operation in the next detection period according to the test formula parameters, and the sensor collects the sampling data of the next detection period to obtain the updated value of the sampling data sequence. The updated values of the sampled data sequence are merged and updated with the sampled data sequence in chronological order to obtain the updated sampled data sequence. Based on the updated sampled data sequence, the updated current sampled data subsequence is obtained by updating it according to a sliding window.
8. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 1, characterized in that, The stage determination model, the physical constraint neural network, and the constraint reinforcement learning policy network are obtained through the following training steps: A training sampling data sequence is obtained, and stage identifiers are labeled on the training sampling data sequence. The stage identifiers include a pressurization stage identifier, a pressure holding stage identifier, and an exhaust stage identifier. The stage determination model is trained based on the training sampling data sequence and the stage identifiers to obtain the stage determination model. The training sampling data sequence and the pressure boundary conditions corresponding to the training sampling data sequence are obtained, and physical constraint terms corresponding to the pressurization stage identifier, the pressure holding stage identifier, and the exhaust stage identifier are constructed respectively. A loss function is constructed based on the data fitting error and the residual of the physical constraint terms. In the loss function, the physical constraint neural network adopts different physical constraint terms under different stage identifiers through the stage gating structure. The physical constraint neural network is trained based on the loss function to obtain the physical constraint neural network. A training environment is constructed to output a sampled data sequence corresponding to the test recipe parameters when given test recipe parameters. The detection state vector is used as the state input, and the test recipe parameters are used as the action output. A reward function is established with the optimization objectives of airtightness detection accuracy and detection time, and a constraint function corresponding to the hard constraint conditions is established. The reward function and the constraint function are jointly optimized using Lagrange multipliers, and the Lagrange multipliers are updated to train the constraint reinforcement learning policy network.
9. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 4, characterized in that, The leakage parameter estimate output by the physical constraint neural network also includes an uncertainty parameter corresponding to the leakage parameter estimate. The uncertainty parameter includes at least one of variance, upper bound of confidence interval, and lower bound of confidence interval. The predicted value of the sampled data in the next time period generates a predicted pressure upper bound sequence and a predicted pressure lower bound sequence based on the uncertainty parameter.
10. The method for detecting the airtightness of a pressure reducing valve based on reinforcement learning according to claim 6, characterized in that, Constraint correction of the candidate test formulation parameters includes solving an optimization problem with the candidate test formulation parameters as the objective parameters. The objective function of the optimization problem is used to minimize the deviation between the constrained test formulation parameters and the candidate test formulation parameters, and the constraints of the optimization problem include the hard constraints.