Double q-table reinforcement learning combined with dynamic attention for semiconductor gate analysis system
The semiconductor valve analysis system, which combines dual-Q table reinforcement learning with dynamic attention, solves the problems of traditional testing, such as the inability to dynamically adjust and reliance on human experience. It achieves full-process automation and intelligence in semiconductor valve testing, improves testing efficiency and decision accuracy, and provides efficient parameter optimization and visualization tools.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI JUKE FLUID CONTROL CO LTD
- Filing Date
- 2026-02-14
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional semiconductor valve testing relies on fixed test procedures and manual analysis. It cannot dynamically adjust the test excitation according to the real-time response of the valve under test, makes it difficult to identify the critical failure mode of the valve, and lacks continuous evaluation of the valve's dynamic operating domain. Parameter optimization depends on the engineer's experience, which is inefficient and makes it difficult to obtain the global optimal solution.
A semiconductor valve analysis system combining dual-Q table reinforcement learning and dynamic attention is adopted. Through a main gas path pressure stabilization module, a parallel multi-test branch module, a central control and data acquisition unit, a data preprocessing module, an adaptive state representation module, and a reinforcement learning decision and optimization module, adaptive state representation and optimized control are achieved. Combined with the dynamic attention mechanism and dual-Q table reinforcement learning algorithm, key features are automatically extracted and test strategies are optimized.
It achieves full automation and intelligence in semiconductor valve testing, solving the bottleneck of traditional testing that relies on fixed procedures and human experience, improving testing efficiency and decision accuracy, ensuring system robustness and safety, and providing efficient parameter optimization and visualization tools.
Smart Images

Figure CN121706040B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semiconductor valve performance analysis technology, and more specifically, to a semiconductor valve analysis system that combines dual Q-table reinforcement learning with dynamic attention. Background Technology
[0002] The performance of industrial fluid valves, especially their response characteristics under dynamic loads, is crucial for ensuring the stability of precision processes. Traditional testing methods rely on fixed test procedures (such as step pressure tests) and manual data analysis.
[0003] Shortcomings of existing technology:
[0004] The test sequence is preset and cannot dynamically adjust the test excitation (such as pressure change rate and flow scan speed) according to the real-time response of the valve under test, resulting in low test efficiency and difficulty in actively stimulating and identifying the critical failure mode of the valve.
[0005] The pass / fail assessment is usually based on static data from a few discrete test points (such as pressure regulation accuracy), lacking a continuous and comprehensive evaluation model for the entire dynamic working domain of the valve.
[0006] Matching valves with optimal upstream pressure, load curves, and other operating parameters relies heavily on engineers' experience and is a time-consuming "black box" process that makes it difficult to obtain a globally optimal solution.
[0007] To address the above problems, this invention proposes a solution. Summary of the Invention
[0008] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a semiconductor valve analysis system that combines dual-Q table reinforcement learning with dynamic attention. This semiconductor valve analysis system, which combines dual-Q table reinforcement learning with dynamic attention, solves the problems mentioned in the background art.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] A semiconductor valve analysis system combining dual-Q table reinforcement learning and dynamic attention includes a main gas path pressure stabilization module, a parallel multi-test branch module, a central control and data acquisition unit, a data preprocessing module, an adaptive state representation module, a reinforcement learning decision-making and optimization module, and a human-computer interaction and visualization module. These modules are interconnected.
[0011] The main air circuit pressure regulator module is used to adjust the input air source pressure to the first-level set pressure;
[0012] The parallel multi-test branch module is used to connect each test branch in series with a secondary voltage regulation unit, a precision flow control unit, and a high-precision sensing unit;
[0013] The central control and data acquisition unit is used to receive control commands to drive each adjustment unit and to synchronously acquire multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit.
[0014] The data preprocessing module is used to perform outlier detection, filtering, and normalization on the collected multidimensional time-series observation data;
[0015] The adaptive state representation module receives standardized multidimensional time-series observation data output from the data preprocessing module and automatically calculates the attention weights of each observation parameter and different timestamp data using a neural network model based on a dynamic attention mechanism. The process is as follows: The module has a built-in neural network model based on a dynamic attention mechanism. Through the built-in dynamic attention mechanism of the model, spatial attention weights and temporal attention weights are calculated on the input standardized multidimensional time-series observation data. The spatial and temporal attention weights are then fused to obtain joint attention weights.
[0016] Dynamically focus on the core features related to the current decision-making task, extract and output the state feature vector representing the current operating state of the gas path test system;
[0017] The reinforcement learning decision and optimization module is based on a double Q-table reinforcement learning model and integrates an adaptive priority experience replay mechanism to output the optimal control action according to the state feature vector.
[0018] The human-computer interaction and visualization module is used to display the system's operating status, attention weight distribution, learning process, and optimization results.
[0019] In a preferred embodiment, the main gas path pressure stabilizing module includes a main diaphragm valve and a main supply pressure reducing valve. The main diaphragm valve acts as a master switch to control the gas supply on / off of the entire test system. The main supply pressure reducing valve stabilizes the upstream high-pressure, fluctuating gas source pressure at a first-level set pressure according to a preset value.
[0020] In a preferred embodiment, the process of adjusting the input air source pressure to the first-level set pressure is as follows:
[0021] High-pressure gas flows through the main supply pressure reducing valve, which starts working according to the preset value. Its diaphragm senses the outlet pressure and compares it with the preset value.
[0022] If the outlet pressure is lower than the preset value, the valve core opening of the main supply pressure reducing valve will increase, allowing more high-pressure gas to pass through in order to increase the downstream pressure;
[0023] If the outlet pressure is higher than the preset value, the valve core opening of the main supply pressure reducing valve will be reduced or closed to prevent gas from flowing in, and slight venting will be carried out through the internal vent to reduce the downstream pressure.
[0024] Until the outlet pressure of the main supply pressure reducing valve is maintained within the preset error range.
[0025] In a preferred embodiment, the secondary voltage regulating unit is implemented as follows:
[0026] The secondary pressure regulation value is extracted according to the test requirement command sent by the central control unit to the programmable precision pressure reducing valve R1. The secondary pressure regulation value is compared with the actual outlet pressure value sensed by its internal or external feedback loop through the built-in controller of R1, and the deviation is calculated. A control signal is output according to the PID algorithm.
[0027] The control signal drives the actuator of R1 to move the valve core, changing the throttling area at the valve seat;
[0028] If the measured pressure is lower than the set value, the valve opening increases and the downstream pressure begins to rise.
[0029] The pressure sensor monitors the inlet pressure changes in real time and feeds the data back to the controller and central acquisition system of R1;
[0030] The controller continuously adjusts dynamically until the inlet pressure stabilizes within the allowable error range, thus completing the construction of the static pressure environment and forming an inlet pressure rise curve.
[0031] In a preferred embodiment, the process of receiving control commands to drive each adjustment unit is as follows:
[0032] It receives digital commands from the host computer and controls the precise actions of each valve actuator through digital-to-analog conversion and drive circuit.
[0033] At a fixed and uniform time interval, all branch sensors are simultaneously triggered to sample and stamped with a uniform timestamp;
[0034] All sensor readings at each timestamp are combined into a multidimensional data packet and transmitted to the intelligent analysis software platform in real time.
[0035] In a preferred embodiment, the process of acquiring multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit is as follows:
[0036] Allocate a data acquisition card channel to each analog sensor and configure a real-time industrial Ethernet node for each intelligent digital device;
[0037] The clocks of the real-time network master station and the host computer are synchronized to the microsecond level, defining the absolute time base of the system;
[0038] In the same hardware cycle that the central controller sends action commands to the actuators, it sends a global hardware trigger pulse through the digital I / O lines;
[0039] After receiving the trigger signal, the data acquisition card immediately starts synchronous sampling and analog-to-digital conversion of all analog input channels at a preset fixed sampling rate. The real-time network master station synchronously sends a broadcast command to all slave stations to freeze and upload time-stamped data.
[0040] The data acquisition card packages all analog values obtained at each sampling moment and adds a precise timestamp. The flow meter and valve controller upload their internal sampled values and internal timestamps via a real-time network and correct all received internal timestamps to a unified time axis.
[0041] All aligned time-series data are written to a circular memory buffer in real time. The system uses the action trigger time as a reference point and determines the observation window according to preset dynamic calculations. When the time range of the data in the buffer covers the window, the system extracts all data rows within this time period to form a multidimensional observation dataset, thus obtaining multidimensional time-series observation data.
[0042] In a preferred embodiment, the process of obtaining the state feature vector is as follows:
[0043] Based on the calculated joint attention weights, the standardized multidimensional time-series observation data are weighted to automatically enhance the feature representation of high-weight data. The weighting calculation formula is as follows:
[0044]
[0045] in This represents the core time-series data matrix after focusing. For element-wise multiplication, Represents the joint attention weight matrix. This is the original input data;
[0046] The core features after focusing are further extracted, integrated, and optimized. First, the focused data is mapped to a high-dimensional embedding space through a linear projection layer to obtain the embedding features, as shown in the formula:
[0047]
[0048] in Represents a high-dimensional embedding feature matrix. The weight matrix of the linear projection layer. For bias terms, It is the transpose of the core time-series data matrix;
[0049] Next, global pooling is performed on the embedded features to extract the final state feature vector, as shown in the formula:
[0050]
[0051] In the formula, This represents a state feature vector that can accurately characterize the current operating state of the gas path test system. Indicates the number of timestamps. express The high-dimensional embedding feature of the t-th timestamp.
[0052] The technical effects and advantages of the semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention in this invention are as follows:
[0053] 1. This invention achieves full automation and intelligence from data acquisition to autonomous parameter optimization by constructing an intelligent closed loop of "physical testing platform - dynamic attention perception - double Q-table learning optimization," solving the fundamental bottleneck of traditional valve testing that relies on fixed procedures and human experience. Traditional semiconductor valve testing relies on preset, static test sequences (such as fixed pressure steps and flow scans), resulting in a rigid testing process that cannot be dynamically adjusted according to individual valve differences or complex operating conditions. Data analysis remains at the level of threshold judgment and simple trend observation, making it difficult to uncover deep performance correlations. Parameter optimization relies entirely on engineer trial and error, which is inefficient and makes it difficult to obtain the globally optimal solution. This invention transforms the modular multi-station testing platform into an interactive environment for reinforcement learning agents. It uses a dynamic attention mechanism to adaptively extract key features from high-dimensional time-series data, and then uses a double Q-table reinforcement learning algorithm to autonomously learn the optimal testing strategy or operating parameters in real-time interaction with the physical environment. This system can proactively design tests to explore performance boundaries and accurately locate the optimal operating point, upgrading testing from "data collection" to "cognitive discovery and optimization," achieving a qualitative leap.
[0054] 2. This invention effectively addresses two core challenges—the difficulty of feature extraction from industrial time-series data and the instability and low sample efficiency of reinforcement learning training in physical systems—by introducing a core algorithm architecture that combines dynamic attention mechanisms with double Q-table reinforcement learning. This ensures the robustness, efficiency, and safety of the system in practical applications. Industrial valve test data is high-dimensional, long-term, noisy, and contains complex physical couplings. Traditional feature engineering methods (such as manually calculating overshoot and rise time) are laborious, incomplete, and difficult to generalize. Furthermore, directly training reinforcement learning agents on expensive physical systems faces risks such as high sample acquisition costs, easy divergence during training due to Q-value overestimation, and potential equipment damage during exploration. The innovative solution of this invention is to employ a dynamic attention model based on Transformer or multi-head attention at the perception layer. This model can automatically learn and focus on the time-series segments and sensor channels most relevant to the current decision-making task without prior knowledge. For example, when evaluating dynamic response, it focuses on the transient pressure process after the action command; when evaluating steady-state accuracy, it focuses on the long-term stable segment. This dynamic focusing capability generates highly condensed, task-relevant state representations, greatly improving the information quality and dimensionality of the input decision network, forming the foundation for the system's intelligent perception. At the decision optimization layer, a composite reinforcement learning algorithm combining Double Q-learning (DQN) with a dueling network architecture and prioritized experience replay is employed. Double Q-learning effectively alleviates the problem of overestimating Q-values by decoupling action selection and value assessment, making the training process more stable and the final policy more reliable. The dueling network structure enables the agent to better assess the value of the state itself, thereby making finer distinctions among a large number of similar actions and improving learning efficiency. The prioritized experience replay mechanism intelligently selects and prioritizes learning those "unexpected" or "information-rich" historical experiences, significantly improving sample utilization efficiency and making it possible to train high-performance policies within a limited number of physical interactions. The combination of these two mechanisms ensures that the agent can steadily converge to a superior policy through safe and efficient exploration.
[0055] 3. This invention, through a parallel multi-station testing architecture and modular design, cleverly balances the data requirements of reinforcement learning with the cost and risks of physical testing in engineering, and endows the system with excellent scalability, interpretability, and practicality. Multiple structurally consistent test branches allow for simultaneous execution of multiple testing or training tasks, improving data acquisition efficiency several times over and directly meeting the needs of reinforcement learning for massive interactive data. More importantly, physical isolation can be achieved between branches; for example, high-risk exploration tasks and robust verification tasks can be assigned to different branches, achieving a safe separation of "exploration" and "exploitation," effectively protecting valuable equipment under test, and making proactive exploratory learning feasible in industrial scenarios. The system not only outputs optimization results but also provides in-depth visualization tools such as attention weight heatmaps, learning curves, and state space projections through a human-computer interaction module. Engineers can intuitively understand "why the AI focuses on a certain data segment" and "how the strategy evolves and optimizes," breaking the "black box" of the AI model, establishing human-machine trust, and making advanced analysis results easy to understand and adopt. The system's intelligent software platform and physical testing platform adopt a loosely coupled design. Its core algorithm framework (dynamic attention + double Q learning) can be transferred to other types of industrial equipment testing and optimization scenarios, requiring only a redefinition of the state, action, and reward functions. The modular hardware design also facilitates customization and expansion to suit different testing pressures and traffic volumes. Attached Figure Description
[0056] Figure 1 This is a schematic diagram of the semiconductor valve analysis system that combines dual Q-table reinforcement learning and dynamic attention according to the present invention.
[0057] Figure 2 This is a schematic diagram of the semiconductor valve analysis method combining dual Q-table reinforcement learning and dynamic attention according to the present invention.
[0058] Figure 3 This is a schematic diagram of an electronic device structure according to the present invention. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0060] Example 1, Figure 1 This invention presents a semiconductor valve analysis system that combines dual Q-table reinforcement learning with dynamic attention.
[0061] A semiconductor valve analysis system combining dual-Q table reinforcement learning and dynamic attention includes:
[0062] A modular physics testing platform, serving as an interactive environment for reinforcement learning agents, executes control actions and provides real-time feedback of high-dimensional sensor data. The platform includes a main gas path pressure stabilization module, a parallel multi-test branch module, and a central control and data acquisition unit. This platform provides a highly controllable, high-precision, and repeatable interactive "physical simulator" for subsequent machine learning algorithms. It replaces the difficult-to-build digital simulation models required for traditional reinforcement learning, directly using the real physical world as the training and evaluation environment, ensuring the ultimate practicality and robustness of the learning strategy.
[0063] The main air circuit pressure regulating module includes a main diaphragm valve and a main supply pressure reducing valve, which are used to regulate and stabilize the input air source pressure to the first-level set pressure.
[0064] The main diaphragm valve (V0) acts as the master switch, controlling the gas supply to and from the entire testing system. Adjusting and stabilizing the input gas pressure to the primary set pressure means that the main supply pressure reducing valve (R0), based on a preset value, stabilizes the upstream high-pressure, potentially fluctuating gas pressure (e.g., 6000 psig nitrogen) onto a uniform and precise primary pressure platform. The specific process is as follows:
[0065] High-pressure gas flows through the main supply pressure reducing valve, which starts working according to the preset value. Its diaphragm senses the outlet pressure and compares it with the preset value.
[0066] If the outlet pressure is lower than the preset value, the valve core opening of R0 increases, allowing more high-pressure gas to pass through in order to increase the downstream pressure;
[0067] If the outlet pressure is higher than the preset value (which may be caused by upstream pressure fluctuations or sudden changes in downstream gas consumption), the valve core opening of R0 will be reduced or closed to prevent gas from flowing in, or even slightly vented through the internal vent to reduce downstream pressure.
[0068] This process is a dynamic and continuous negative feedback adjustment process until the R0 outlet pressure is stably maintained within the preset error range.
[0069] It should be noted that the main gas path pressure stabilization module provides a "clean" input baseline for subsequent tests, ensuring that the initial conditions of all test branches are consistent and avoiding external gas source pressure fluctuations from contaminating the test data; it is the first line of defense for system safety and the basis for achieving standardized test procedures.
[0070] The parallel multi-test branch module contains at least two independent test branches with identical structures. Each test branch is connected in series with a secondary pressure regulating unit, a precision flow control unit, and a high-precision sensing unit for monitoring the inlet and outlet pressures of the valve under test.
[0071] The parallel multi-test branch module consists of at least two independent test branches that are completely identical in terms of pneumatic circuit structure and electrical interface connected in parallel. Each branch is a standardized valve test unit with complete pressure regulation, load simulation and high-precision sensing capabilities. Its core function is to create an independent, controllable and accurately measurable "working environment" for the "valve under test" and to serve as a parallel channel for standardized interaction between the reinforcement learning agent and the physical world.
[0072] The secondary pressure regulating unit performs a second, precise adjustment based on testing requirements, transitioning the outlet pressure to the new set value. The process is as follows:
[0073] The secondary pressure regulation value is extracted according to the test requirement command sent by the central control unit to the programmable precision pressure reducing valve R1. The secondary pressure regulation value is compared with the actual outlet pressure value sensed by its internal or external feedback loop through the controller built into R1 (usually a PID controller), and the deviation is calculated. A control signal is output according to the PID algorithm.
[0074] The control signal drives the actuator of R1 (such as a piezoelectric driver or stepper motor) to precisely move the valve core and change the throttling area at the valve seat;
[0075] If the measured pressure is lower than the set value, the valve core opening increases, allowing more high-pressure gas to pass through, and the downstream (i.e., the inlet pipe of the valve being tested) pressure begins to rise;
[0076] The pressure sensor monitors the inlet pressure changes in real time and feeds the data back to the controller and central acquisition system of R1;
[0077] The controller continuously adjusts dynamically until the inlet pressure stabilizes within the allowable error range, completing the construction of the static pressure environment. This process may last from hundreds of milliseconds to several seconds, forming an inlet pressure rise curve.
[0078] The outlet of the valve under test is connected to a precision flow control unit, in which a regulating valve (such as a needle valve V16) changes its opening according to instructions to simulate the dynamic changes in gas flow demand of downstream equipment.
[0079] High-frequency pressure sensors installed at the valve inlet and outlet, along with a mass flow meter integrated into the flow path, synchronously collect three core physical quantities with millisecond-level accuracy: the curves of inlet pressure change over time, outlet pressure change over time, and volumetric flow rate through the valve change over time.
[0080] It's important to note that the design of at least two test branches is essentially aimed at transforming the high-cost physical testing system into a highly efficient "reinforcement learning training ground." It meets the fundamental requirement of reinforcement learning for massive interaction samples through parallel data acquisition, achieves the safe separation and synchronization of "high-risk exploration" and "robust optimization" through physical isolation between branches, and utilizes multiple environment replicas to provide cross-validation and system fault tolerance. Thus, under realistic time and cost constraints, it enables the agent to learn and optimize reliably, efficiently, and safely in the real physical world. This is no longer a simple parallel connection of traditional testing equipment "to increase output," but a systematic architecture tailored to serve an "intelligent agent with autonomous learning and optimization capabilities." It transforms the physical testing system from a passive "data generator" into an active "experience factory" and "policy training ground," which is one of the most crucial design aspects that propelled this invention from concept to practical application and engineering.
[0081] The parallel multi-test branch module is far more than a simple pneumatic path replication; it's an engineered solution specifically designed to achieve the core goal of "data-driven, efficient reinforcement learning." It addresses the physical constraints of sample efficiency through parallelization, ensures environmental consistency through standardization, and enables precise interaction with the agent algorithm through meticulous control. It serves as the core bridge and key enabler connecting AI intelligence with the physical performance of valves. Training reinforcement learning agents on real physical systems requires tens of thousands or even millions of "action-state" interactions. Performing such numerous dynamic tests on a single valve on a single-station system is extremely time-consuming and causes severe wear and tear on the tested valve and equipment. The parallel multi-branch design effectively increases sample acquisition speed by N times (where N is the number of branches). The agent can simultaneously explore multiple valves (or different samples of the same valve), improving training efficiency over calendar time by several orders of magnitude, making reinforcement learning training on physical systems feasible.
[0082] The central control and data acquisition unit is used to receive control commands to drive the main gas path pressure stabilization module and each adjustment unit in the parallel multi-test branch module, and to simultaneously acquire multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit.
[0083] It receives digital commands from the host computer and controls the precise actions of each valve actuator through digital-to-analog conversion and drive circuit.
[0084] At a fixed and uniform time interval, all branch sensors are simultaneously triggered to sample and stamped with a uniform timestamp;
[0085] All sensor readings at each timestamp are combined into a multidimensional data packet and transmitted to the intelligent analysis software platform in real time.
[0086] The process of acquiring multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit is as follows:
[0087] Allocate a dedicated data acquisition card channel with synchronous sample-and-hold function to each analog sensor (pressure, temperature), and configure a deterministic real-time industrial Ethernet node for each intelligent digital device;
[0088] The IEEE 1588 precision time protocol is adopted to synchronize the clocks of the data acquisition card, real-time network master station and host computer to the microsecond level deviation, defining a unique absolute time reference for the system.
[0089] In the same hardware cycle that the central controller sends action commands to the actuators, it sends a global hardware trigger pulse through the digital I / O line. After receiving the trigger signal, the data acquisition card immediately starts synchronous sampling and analog-to-digital conversion of all analog input channels at a preset fixed sampling rate. The real-time network master station synchronously sends a broadcast command to all slave stations (flow meters, valve controllers) to "freeze and upload time-stamped data".
[0090] The acquisition card packages all analog values obtained at each sampling moment and adds a precise timestamp for that moment. The flow meter and valve controller upload their internal sampled values and their internal timestamps at the moment they receive the "freeze" command via a real-time network. The central acquisition software corrects all received internal timestamps to a unified time axis based on the global time reference.
[0091] All aligned time-series data are written in real time to a first-in-first-out circular memory buffer. The system uses the action trigger time as a reference point and determines the observation window according to preset dynamic calculations. When the time range of the data in the buffer covers the window, the system accurately extracts all data rows within this time period to form a complete multidimensional observation dataset.
[0092] The segmented dataset, along with metadata (test ID, branch number, action description, sampling rate, etc.), is packaged into a standard format data packet.
[0093] The intelligent analysis software platform is the "brain" of the system. Its role is to transform raw, high-dimensional, and noisy sensor data into an "understanding" of the valve's state, and make intelligent decisions based on this understanding, ultimately achieving automated performance evaluation and autonomous parameter optimization. It solves the core pain points of traditional testing: "rich data but scarce information" and "automated testing but still requiring manual decision-making."
[0094] Multidimensional time-series observation data includes core physical quantity time-series data (three-dimensional vector sequence), control command and state feedback time-series data, system environment and auxiliary monitoring data, event markers and metadata; a complete "multidimensional time-series observation data package" is a matrix strictly aligned on the time axis, where each row is a snapshot of the overall system at the moment of synchronous sampling, and each column is the evolution of a physical or logical variable over time.
[0095] The data preprocessing module is used to perform outlier detection, filtering, and normalization on the collected multidimensional time-series observation data;
[0096] For the time-series data of each sensor channel, using a sliding time window (e.g., 100ms), calculate the mean μ and standard deviation σ of the data within the window. This will exceed... Data points within a certain range are marked as transient outliers;
[0097] Based on the physical limits of the valve and system (such as the pressure not being negative and the maximum rate limit for flow change), mark the data points that violate the physical laws.
[0098] For the marked outliers, linear interpolation of the valid data points before and after them is used to replace them to ensure the continuity of the time series.
[0099] The curves of inlet pressure change over time, outlet pressure change over time, and volumetric flow rate through the valve change over time are used to filter out high-frequency electrical noise and mechanical vibration noise from the core dynamic signals "outlet pressure change over time" and "volumetric flow rate through the valve change over time" while retaining the true physical dynamics. The cutoff frequency of the low-pass filter should be higher than the effective bandwidth of the valve under test, but much lower than half of the sampling frequency.
[0100] The transfer function of the filter is realized in the digital domain through a bilinear transformation, and a smooth sequence is obtained by recursively calculating the original sequence.
[0101] Verify that the timestamps of data from different hardware channels are within the allowable synchronization tolerance (e.g., <1ms). If individual channels deviate slightly due to communication delays, use a timestamp-based linear interpolation method to resample the data from all channels onto a completely uniform, equally spaced new time axis.
[0102] For each variable, normalization is performed according to its physical meaning and test range, with the pressure / flow signal using minimum-maximum normalization;
[0103] The output of the preprocessing module is a clean tensor of shape [sequence length L, feature dimension D], where L corresponds to the total number of sampling timestamps, D is the total number of all sensors and state variables, and represents the dimension of the observation parameters.
[0104] An adaptive state representation module is used to dynamically extract state feature vectors related to the current decision-making task from preprocessed multidimensional time-series observation data;
[0105] A neural network model based on dynamic attention mechanism is adopted to receive standardized multi-dimensional time-series observation data output by the data preprocessing module. The dynamic attention mechanism automatically calculates the attention weight of each observation parameter and data with different timestamps, dynamically focuses on the core features related to the current decision task, suppresses the interference of redundant information, and extracts and outputs the state feature vector that can accurately characterize the current operating state of the gas path test system.
[0106] The module incorporates and employs a neural network model based on a dynamic attention mechanism. This model is designed to adapt to multi-dimensional time-series observation data in gas path testing. Its core includes a linear projection layer, a dynamic attention calculation layer, and a feature extraction layer, which are used to achieve accurate extraction and filtering of subsequent data features.
[0107] The system receives standardized multidimensional time-series observation data (i.e., time-series data of gas path pressure, flow rate, temperature, etc., after outlier detection, filtering and denoising, and normalization to eliminate dimensional differences and invalid interference) transmitted from the data preprocessing module, as the raw input data for feature extraction. Where T represents the number of timestamps, N is the dimension of the observation parameters, and R represents the set of real numbers;
[0108] By using the built-in dynamic attention mechanism of the model, dual weight calculation is performed on the input standardized multidimensional time series observation data. On the one hand, the spatial attention weight of each observation parameter (such as pressure, flow rate, and temperature) is calculated, and on the other hand, the temporal attention weight of the same parameter at different timestamps is calculated, so as to achieve quantitative differentiation of the importance of the data.
[0109] The spatial attention weight calculation process is as follows:
[0110] First, global pooling is performed on the standardized time series data to extract global features for each parameter, using the following formula:
[0111]
[0112] in This represents the global feature vector of each observed parameter. Indicates the first All observation parameter data for each timestamp Indicates the number of timestamps;
[0113] The spatial attention weights are then calculated using a fully connected layer and a sigmoid activation function, as shown in the formula:
[0114]
[0115] in Represents the spatial attention weight vector. These are the learnable parameters of the fully connected layer. ReLU represents the Sigmoid activation function. It is the transpose of the global eigenvectors of each observation parameter;
[0116] The calculation process for temporal attention weights is as follows:
[0117] After performing dimensionality transformation on the standardized time-series data, temporal features are extracted through a fully connected layer, and then the temporal attention weights are calculated using the following formula:
[0118]
[0119] in Represents the time series feature matrix. Represents the temporal attention weight vector. These are the learnable parameters of the fully connected layer. This is the transpose of the original input data;
[0120] By fusing spatial and temporal attention weights, we obtain the joint attention weight, as shown in the formula:
[0121]
[0122] in Represents the joint attention weight matrix. This is the transpose of the temporal attention weight vector. For outer product operations, a dual weighting of the time and parameter dimensions is implemented.
[0123] Based on the joint attention weights calculated above, the standardized multidimensional time-series observation data are weighted to automatically enhance the feature representation of high-weight data (core data relevant to the current decision task) while suppressing low-weight redundant data (redundant interference data unrelated to the current decision task), thus achieving dynamic focusing of core features. The weighting calculation formula is as follows:
[0124]
[0125] in This represents the core time-series data matrix after focusing. For element-wise multiplication, Represents the joint attention weight matrix. This is the original input data;
[0126] The core features after focusing are further extracted, integrated, and optimized. First, the focused data is mapped to a high-dimensional embedding space through a linear projection layer to obtain the embedding features, as shown in the formula:
[0127]
[0128] in Represents a high-dimensional embedding feature matrix. The weight matrix of the linear projection layer. For bias terms, It is the transpose of the core time-series data matrix;
[0129] Next, global pooling is performed on the embedded features to extract the final state feature vector, as shown in the formula:
[0130]
[0131] In the formula, This represents a state feature vector that can accurately characterize the current operating state of the gas path test system. Indicates the number of timestamps. express The high-dimensional embedding feature of the t-th timestamp.
[0132] Traditional methods require engineers to define and calculate dozens of features (such as rise time, overshoot, and integral error). This module, through a multi-head self-attention mechanism, automatically learns hundreds or even thousands of implicit features. These features may correspond to complex physical patterns (such as resonance at specific frequencies or the phase relationship between pressure and flow changes), which are difficult to design manually. When evaluating "opening characteristics," the model allocates more than 80% of its attention weight to the data segment within 200ms after the action command; while when evaluating "closing sealing," it focuses on the pressure decay curve after the flow rate drops to zero. This dynamic and target-correlated information compression results in extremely high "information entropy" in the state vector transmitted to the decision-making module, directly leading to improved decision accuracy and more efficient search of the exploration space. The generated state vector is an abstract and general "valve dynamic performance descriptor." It can be used as a pre-trained feature for different tasks such as downstream fault classification and life prediction, achieving "one-time perception, multiple uses," reducing the data requirements for new tasks.
[0133] The reinforcement learning decision and optimization module adopts a reinforcement learning model based on a double Q table and integrates an adaptive priority experience replay mechanism to output the optimal control action for the physical test platform based on the state feature vector.
[0134] Electrically connected to the adaptive state characterization module, it employs a dual-Q-table reinforcement learning model combined with an adaptive priority experience playback mechanism to receive the state feature vector output by the adaptive state characterization module. With the goal of optimizing the operational stability and testing accuracy of the gas path test system, the dual-Q-table reinforcement learning model avoids the overestimation bias of traditional reinforcement learning, and the adaptive priority experience playback mechanism improves learning efficiency and decision accuracy. It outputs the optimal control action commands for each adjustment unit in the main gas path stabilization module and the parallel multi-test branch module, realizing the adaptive dynamic optimization control of the gas path test system.
[0135] An electrical connection is established between the reinforcement learning decision and optimization module and the adaptive state representation module to ensure that the state feature vector output by the adaptive state representation module can be stably transmitted to this module as the basis for decision input. This module has a built-in core model and adopts the "double Q table reinforcement learning model" as the basic decision model. At the same time, it integrates the "adaptive priority experience playback mechanism" to improve decision accuracy and learning efficiency and avoid the defects of traditional reinforcement learning.
[0136] Through electrical connection with the adaptive state characterization module, the system receives the state feature vector transmitted from it. This vector is an accurate representation of the current operating state of the gas path test system and serves as the core input data for module decision-making.
[0137] The module sets a core decision objective, namely, to ensure the stable operation of the gas path testing system and achieve optimal testing accuracy, and all decision actions are carried out around this objective.
[0138] By calling the built-in double-Q table reinforcement learning model and using the core logic of alternating action selection and cross-evaluation of the double-Q table, the overestimation bias of action value in traditional single-Q table reinforcement learning is avoided, thereby improving decision accuracy.
[0139] By integrating an adaptive priority experience replay mechanism, the experience data during the model learning process is prioritized and key experiences (such as abnormal working conditions and experiences with large decision-making biases) are learned first, further improving the model learning efficiency and decision-making accuracy.
[0140] Based on the input state feature vector and preset target, combined with the model calculation results, the optimal control action command is generated. The controlled objects are clearly defined as: the main diaphragm valve and the main supply pressure reducing valve in the main gas circuit pressure stabilization module, and the secondary pressure regulating unit and the precision flow control unit in the parallel multi-test branch module.
[0141] The generated optimal control action command is sent to the corresponding adjustment unit. After each unit executes the command, it adjusts its operating state, thus realizing the adaptive dynamic optimization control of the gas path test system and ensuring that the system is always in a stable and high-precision operating state.
[0142] It should be noted that employing a reinforcement learning model based on a double Q-table and incorporating an adaptive priority experience replay mechanism has the following advantages:
[0143] The policy stability brought by Double Q learning: In traditional DQN, the Q-value is easily overestimated due to maximizing operations, causing the policy to oscillate wildly between "aggressive" and "conservative" approaches. Double Q learning, by decoupling action selection and evaluation, reduces the volatility of the policy in the later stages of training by about 70%, resulting in a more reliable policy that can be deployed directly.
[0144] The exceptional sample efficiency brought by prioritized experience replay: the intelligent system repeatedly "reviews" experiences with large TD errors (i.e., "unexpected" or "highly rewarding" moments). This is equivalent to focusing 80% of the learning effort on 20% of the key events. Real-world testing shows that, while achieving the same performance level, compared to uniform replay, it can reduce the number of environmental interactions by 40%-60%, which translates to significant time and cost savings for physical testing.
[0145] Resolving multi-objective conflict optimization: By carefully designing the reward function, the agent can autonomously find the optimal trade-off point for these three conflicting indicators and provide clear and quantitative optimization suggestions, which is almost impossible to accomplish systematically by manual parameter tuning.
[0146] The human-computer interaction and visualization module is used to display the system's operating status, attention weight distribution, learning process, and optimization results.
[0147] It is electrically connected to the central control and data acquisition unit, the adaptive state representation module, and the reinforcement learning decision and optimization module, respectively, to display the overall system operation status, working parameters of each module, dynamic attention weight distribution, reinforcement learning model learning process, and optimization results of optimal control actions in real time. It also supports users to input control commands, query test data, and adjust system parameters, realizing two-way human-machine interaction.
[0148] The human-computer interaction and visualization module is electrically connected to the central control and data acquisition unit, the adaptive state representation module, and the reinforcement learning decision and optimization module, respectively, to ensure that various types of data, parameters and operating information transmitted by the three modules can be received synchronously, providing data support for subsequent display and interaction.
[0149] The module's built-in visual interface displays five types of core information in real time: ① the overall operating status of the gas path testing system; ② the working parameters of each functional module of the system; ③ the dynamic attention weight distribution of the adaptive state characterization module; ④ the model learning process of the reinforcement learning decision and optimization module; and ⑤ the optimization results corresponding to the optimal control action output by reinforcement learning, ensuring that users can intuitively grasp the overall picture of system operation.
[0150] While providing display functions, it supports users to perform three core operations: ① inputting external control commands (used to regulate the overall operation of the system); ② querying various historical and real-time test data during the gas path test process; ③ adjusting relevant system parameters according to needs, so as to realize the user's active control of the system.
[0151] Through bidirectional transmission of data between the module and the user, and user inputting operation commands to the module, two-way human-machine interaction is achieved. This ensures real-time monitoring of the system by the user and allows the user to make precise adjustments to the system based on the displayed information, thus realizing human-machine collaboration.
[0152] The human-computer interaction and visualization module provides a "transparent window" and a "bridge of trust" for the system. It presents complex internal states (such as where attention is focused, learning progress, and decision-making criteria) to engineers in intuitive charts (heatmaps, learning curves, and decision traceability). This not only enables effective monitoring and debugging of the system's operational status, but more importantly, it greatly enhances the interpretability of AI decisions, dispels doubts about the "black box," and makes human-computer collaboration possible. It is a key interface for the understanding and adoption of technological achievements.
[0153] The modules are not isolated; their collaborative work produces a system-level benefit of "1+1>2":
[0154] From automation to intelligence: The system has completed a paradigm shift from "executing tests according to scripts" to "autonomously designing experiments to seek the optimal solution".
[0155] From data to knowledge: The system transforms massive amounts of raw test data into actionable performance profiles, optimization suggestions, and predictive insights through intelligent analysis.
[0156] From cost center to value engine: It not only improves testing efficiency, but also creates potential value for improving the stability and yield of core processes such as downstream semiconductor manufacturing by optimizing valve performance.
[0157] In summary, each module enhances a specific weakness in traditional testing methods, and through precise integration, they collectively build a next-generation intelligent testing and optimization platform that can adapt to complex needs, learn autonomously, and collaborate with humans and machines.
[0158] Example 2, Figure 2 This invention presents a semiconductor valve analysis method combining dual Q-table reinforcement learning and dynamic attention, comprising the following steps:
[0159] Adjust and stabilize the input air pressure to the first-level set pressure;
[0160] Each test branch is connected in series with a secondary voltage regulation unit, a precision flow control unit, and a high-precision sensing unit;
[0161] It receives control commands to drive the main gas path pressure stabilization module and each adjustment unit in the parallel multi-test branch module, and simultaneously collects multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit.
[0162] Outlier detection, filtering, and normalization are performed on the collected multidimensional time-series observation data;
[0163] The process of receiving standardized multidimensional time-series observation data and automatically calculating the attention weights of each observation parameter and different timestamp data using a neural network model based on dynamic attention mechanism is as follows: The module has a built-in neural network model based on dynamic attention mechanism. Through the built-in dynamic attention mechanism of the model, spatial attention weights and temporal attention weights are calculated on the input standardized multidimensional time-series observation data. The spatial and temporal attention weights are then fused to obtain joint attention weights.
[0164] Dynamically focus on the core features related to the current decision-making task, extract and output the state feature vector representing the current operating state of the gas path test system;
[0165] A reinforcement learning model based on a double Q-table and an adaptive priority experience replay mechanism are adopted to output the optimal control action for the physical test platform based on the state feature vector.
[0166] It displays the system's operating status, attention weight distribution, learning process, and optimization results.
[0167] Figure 3 An electronic device, comprising:
[0168] At least one processor; and,
[0169] A memory communicatively connected to the at least one processor; wherein,
[0170] The memory stores computer programs that can be executed by the at least one processor.
[0171] The computer program is executed by the at least one processor.
[0172] This enables the at least one processor to execute the method of the semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention as described in the present invention.
[0173] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0174] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0175] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0176] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0177] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0178] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A semiconductor valve analysis system combining dual-Q table reinforcement learning and dynamic attention, characterized in that, It includes a main gas path pressure stabilization module, a parallel multi-test branch module, a central control and data acquisition unit, a data preprocessing module, an adaptive state representation module, a reinforcement learning decision-making and optimization module, and a human-computer interaction and visualization module. These modules are interconnected. The main air circuit pressure regulator module is used to adjust the input air source pressure to the first-level set pressure; The parallel multi-test branch module is used to connect each test branch in series with a secondary voltage regulation unit, a precision flow control unit, and a high-precision sensing unit; The central control and data acquisition unit is used to receive control commands to drive each adjustment unit and to synchronously acquire multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit. The data preprocessing module is used to perform outlier detection, filtering, and normalization on the collected multidimensional time-series observation data; The adaptive state representation module receives standardized multidimensional time-series observation data output from the data preprocessing module and automatically calculates the attention weights of each observation parameter and different timestamp data using a neural network model based on a dynamic attention mechanism. The process is as follows: The module has a built-in neural network model based on a dynamic attention mechanism. Through the built-in dynamic attention mechanism of the model, spatial attention weights and temporal attention weights are calculated on the input standardized multidimensional time-series observation data. The spatial and temporal attention weights are then fused to obtain joint attention weights. Dynamically focus on the core features related to the current decision-making task, extract and output the state feature vector representing the current operating state of the gas path test system; The reinforcement learning decision and optimization module is based on a double Q-table reinforcement learning model and integrates an adaptive priority experience replay mechanism to output the optimal control action according to the state feature vector. The human-computer interaction and visualization module is used to display the system's operating status, attention weight distribution, learning process, and optimization results.
2. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 1, characterized in that, The main gas path pressure stabilization module includes a main diaphragm valve and a main supply pressure reducing valve. The main diaphragm valve acts as a master switch to control the gas supply on and off of the entire test system. The main supply pressure reducing valve stabilizes the upstream high pressure and fluctuating gas source pressure at the first-level set pressure according to a preset value.
3. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 2, characterized in that, The process of adjusting the input air source pressure to the first-level set pressure is as follows: High-pressure gas flows through the main supply pressure reducing valve, which starts working according to the preset value. Its diaphragm senses the outlet pressure and compares it with the preset value. If the outlet pressure is lower than the preset value, the valve core opening of the main supply pressure reducing valve will increase, allowing more high-pressure gas to pass through in order to increase the downstream pressure; If the outlet pressure is higher than the preset value, the valve core opening of the main supply pressure reducing valve will be reduced or closed to prevent gas from flowing in, and slight venting will be carried out through the internal vent to reduce the downstream pressure. Until the outlet pressure of the main supply pressure reducing valve is maintained within the preset error range.
4. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 3, characterized in that, The secondary voltage regulation unit is implemented as follows: The secondary pressure regulation value is extracted according to the test requirement command sent by the central control unit to the programmable precision pressure reducing valve R1. The secondary pressure regulation value is compared with the actual outlet pressure value sensed by its internal or external feedback loop through the built-in controller of R1, and the deviation is calculated. A control signal is output according to the PID algorithm. The control signal drives the actuator of R1 to move the valve core, changing the throttling area at the valve seat; If the measured pressure is lower than the set value, the valve opening increases and the downstream pressure begins to rise. The pressure sensor monitors the inlet pressure changes in real time and feeds the data back to the controller and central acquisition system of R1; The controller continuously adjusts dynamically until the inlet pressure stabilizes within the allowable error range, thus completing the construction of the static pressure environment and forming an inlet pressure rise curve.
5. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 4, characterized in that, The process of receiving control commands to drive each adjustment unit is as follows: It receives digital commands from the host computer and controls the precise actions of each valve actuator through digital-to-analog conversion and drive circuit. At a fixed and uniform time interval, all branch sensors are simultaneously triggered to sample and stamped with a uniform timestamp; All sensor readings at each timestamp are combined into a multidimensional data packet and transmitted to the intelligent analysis software platform in real time.
6. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 5, characterized in that, The process of acquiring multi-dimensional time-series observation data from the high-precision sensing unit and the precision flow control unit is as follows: Allocate a data acquisition card channel to each analog sensor and configure a real-time industrial Ethernet node for each intelligent digital device; The clocks of the real-time network master station and the host computer are synchronized to the microsecond level, defining the absolute time base of the system; In the same hardware cycle that the central controller sends action commands to the actuators, it sends a global hardware trigger pulse through the digital I / O lines; After receiving the trigger signal, the data acquisition card immediately starts synchronous sampling and analog-to-digital conversion of all analog input channels at a preset fixed sampling rate. The real-time network master station synchronously sends a broadcast command to all slave stations to freeze and upload time-stamped data. The data acquisition card packages all analog values obtained at each sampling moment and adds a precise timestamp. The flow meter and valve controller upload their internal sampled values and internal timestamps via a real-time network and correct all received internal timestamps to a unified time axis. All aligned time-series data are written to a circular memory buffer in real time. The system uses the action trigger time as a reference point and determines the observation window according to preset dynamic calculations. When the time range of the data in the buffer covers the window, the system extracts all data rows within this time period to form a multidimensional observation dataset, thus obtaining multidimensional time-series observation data.
7. The semiconductor valve analysis system combining dual Q-table reinforcement learning and dynamic attention according to claim 6, characterized in that, The process of obtaining the state feature vector is as follows: Based on the calculated joint attention weights, the standardized multidimensional time-series observation data are weighted to automatically enhance the feature representation of high-weight data. The weighting calculation formula is as follows: in This represents the core time-series data matrix after focusing. For element-wise multiplication, Represents the joint attention weight matrix. This is the original input data; The core features after focusing are further extracted, integrated, and optimized. First, the focused data is mapped to a high-dimensional embedding space through a linear projection layer to obtain the embedding features, as shown in the formula: in Represents a high-dimensional embedding feature matrix. The weight matrix of the linear projection layer. For bias terms, It is the transpose of the core time-series data matrix; Next, global pooling is performed on the embedded features to extract the final state feature vector, as shown in the formula: In the formula, This represents a state feature vector that can accurately characterize the current operating state of the gas path test system. Indicates the number of timestamps. express The high-dimensional embedding feature of the t-th timestamp.
Citation Information
Patent Citations
Valve actuator self-adaptive control method and system based on artificial intelligence
CN120686589A
LLM-based intelligent analysis system and method for air leakage of compressed air pipeline
CN121117474A