Method and apparatus for quality control of monitoring instrument manufacture

By constructing a dynamic diagram of the monitoring instrument manufacturing process and extracting spatiotemporal anomaly features and modeling causal transmission, combined with small-sample quality risk transfer reasoning and hybrid strategy optimization of process parameters, the problem of poor adaptability of process parameter optimization in the monitoring instrument manufacturing process was solved, the accuracy and robustness of quality control were improved, and the defect rate was reduced.

CN122432977APending Publication Date: 2026-07-21SHENZHEN ZHAOXING BOTUO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ZHAOXING BOTUO TECH CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Poor adaptability of process parameters during the manufacturing of monitoring instruments affects the accuracy and robustness of quality control, resulting in a high defect rate. This is especially true in the early stages of mass production of new models, where there are large blind spots in quality control, making it impossible to guarantee product quality stability.

Method used

A dynamic diagram of the manufacturing process of monitoring instruments is constructed. Causal transmission modeling is performed by extracting spatiotemporal anomaly features and embedding physical information. Small sample quality risk transfer reasoning is carried out. A hybrid strategy of model predictive control and flexible reinforcement learning is combined to optimize process parameters and achieve optimal parameter adjustment.

Benefits of technology

It significantly improved the adaptability of process parameter optimization, shortened the quality convergence time during the introduction of new products, improved the control accuracy and robustness of the manufacturing quality of monitoring instruments, and reduced the defect rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432977A_ABST
    Figure CN122432977A_ABST
Patent Text Reader

Abstract

The application discloses a kind of monitoring instrument manufacturing quality control method and device, comprising: based on the operation parameters and quality detection data of each working procedure equipment in monitoring instrument manufacturing process, manufacturing process dynamic diagram is constructed;Based on dynamic diagram structure data, space-time abnormal feature extraction is carried out;Based on the causal conduction modeling of space-time abnormal state feature tensor, causal structure parameter and residual term are obtained;Based on causal structure parameter and residual term, small sample quality risk migration inference is carried out, and quality risk prediction value and uncertainty measure value are obtained;Based on quality risk prediction value, uncertainty measure value and space-time abnormal state feature tensor, process parameter optimization is carried out, and optimal process parameter adjustment amount is obtained and issued to corresponding equipment station execution, can effectively improve process parameter optimization adaptability, to improve the control precision and robustness of monitoring instrument manufacturing quality, guarantee the product quality stability of monitoring instrument in mass production process, reduce the rate of defective products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring instrument manufacturing technology, and in particular to a method and apparatus for quality control in the manufacturing of monitoring instruments. Background Technology

[0002] With the rapid development of medical electronics technology, portable monitoring instruments, such as electrocardiogram (ECG) monitors and pulse oximeters, are playing an increasingly important role in clinical applications. The manufacturing quality of these monitoring instruments directly affects patient safety, thus requiring extremely high technical precision in controlling their manufacturing process. Current monitoring instrument manufacturing involves multiple complex processes, including surface mount technology (SMT), die-in-the-loop (DIP), and assembly. Each process involves numerous interconnected and interdependent equipment operating parameters, forming a highly complex manufacturing system. However, monitoring instrument product models iterate rapidly, and new models often lack sufficient historical quality data. Traditional supervised learning methods struggle to achieve effective knowledge transfer from the source model (rich in data) to the target new model (scarce in data), resulting in significant quality control blind spots in the early stages of mass production. Furthermore, the manufacturing process involves frequent switching between steady-state and disturbed dynamic conditions. In steady-state conditions, the system state changes gradually, requiring fine-tuning; in disturbed dynamic conditions, such as equipment failure or material changes, the system state changes drastically, requiring global re-optimization. A single optimization strategy is difficult to meet the needs of both operating conditions, which can easily lead to control oscillations or slow convergence. The process parameter optimization has poor adaptability, which affects the control accuracy and robustness of the manufacturing quality of the monitoring instrument. It cannot guarantee the product quality stability of the monitoring instrument during mass production, resulting in a high defect rate caused by process fluctuations and equipment abnormalities. Summary of the Invention

[0003] The main purpose of this application is to provide a quality control method for the manufacturing of monitoring instruments, which aims to solve the technical problem that the poor adaptability of process parameters in the manufacturing process affects the control accuracy and robustness of the manufacturing quality of monitoring instruments, and makes it impossible to guarantee the product quality stability of monitoring instruments during mass production, resulting in a high defect rate.

[0004] To achieve the above objectives, this application proposes a quality control method for the manufacturing of monitoring instruments, which includes: Based on the operating parameters and quality inspection data of equipment in each process of the monitoring instrument manufacturing process, a dynamic diagram of the manufacturing process is constructed to obtain the dynamic diagram structure data; Based on the dynamic graph structure data, spatiotemporal anomaly features are extracted to obtain a spatiotemporal anomaly state feature tensor. Based on the spatiotemporal anomaly state feature tensor, causal transmission modeling of physical information embedding is performed to obtain the causal structure parameters and residual terms of physical information embedding; Based on the causal structure parameters and residual terms embedded in the physical information, small-sample quality risk transfer reasoning is performed to obtain the quality risk prediction value and uncertainty measure value. Based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal anomaly state feature tensor, the process parameters are optimized to obtain the optimal process parameter adjustment amount, and the optimal process parameter adjustment amount is sent to the corresponding equipment station for execution to achieve quality control of the monitoring instrument manufacturing.

[0005] Optionally, the construction of a dynamic diagram of the manufacturing process based on the operating parameters and quality inspection data of equipment at each stage of the monitoring instrument manufacturing process includes: The operating parameters and quality inspection data of each equipment in the surface mount process, insertion process and assembly process are collected, and the operating parameters and quality inspection data are time aligned and missing values ​​are imputed to obtain a standardized time series dataset. Based on the standardized time-series dataset, each process equipment is mapped as a heterogeneous node, the material flow relationship between processes is mapped as a time-varying edge, and an initial adjacency matrix is ​​constructed based on the time-varying edge to obtain the initial graph topology. The node attributes in the initial graph topology are feature-encoded, and the equipment operating parameters and quality inspection data are converted into node feature vectors to generate a node attribute matrix. Based on the time-series records of material flow, the weight values ​​of the time-varying edges under each time slice are calculated, a time-varying edge weight sequence is generated, and a time-varying adjacency tensor is constructed based on the time-varying edge weight sequence. By fusing the node attribute matrix with the time-varying adjacency tensor, a dynamic graph of the manufacturing process is obtained, and the time continuity of the dynamic graph is verified, and the dynamic graph structure data is output.

[0006] Optionally, the step of extracting spatiotemporal anomaly features based on the dynamic graph structure data to obtain a spatiotemporal anomaly state feature tensor includes: Extract the node attribute matrix and time-varying adjacency tensor from the dynamic graph structure data; The time-varying adjacency tensor is spectrally decomposed to obtain the graph Laplacian matrix, and an adaptive wavelet kernel function is constructed based on the eigenvalue distribution of the graph Laplacian matrix. The node attribute matrix is ​​convolved with the adaptive wavelet kernel function to extract spatial structural features at different scales and generate a multi-scale spatial feature map. The multi-scale spatial feature map is divided into overlapping time window sequences according to the time dimension, and the time window sequences are input into a gated recurrent unit network for temporal dependency feature extraction to obtain the hidden state sequence. The hidden state sequence is input into a multi-head attention mechanism to determine the attention weights at each time step. The hidden state sequence is then weighted and fused based on the attention weights to obtain temporal fusion features. By concatenating the spatial feature slices corresponding to the latest sampling time in the multi-scale spatial feature map with the temporal fusion features, a spatiotemporal anomaly state feature tensor is obtained.

[0007] Optionally, the causal transmission modeling based on the spatiotemporal anomaly state feature tensor for physical information embedding, to obtain the causal structure parameters and residual terms of the physical information embedding, includes: Based on a preset abnormal deviation threshold, the feature tensor of the spatiotemporal abnormal state is used to locate abnormal nodes, extract abnormal nodes and their corresponding neighborhood abnormal correlation subgraphs, parse the node attributes of the abnormal correlation subgraphs, and obtain a set of candidate process parameters. The set of candidate process parameters includes solder joint cold solder rate, board separation stress peak, locking torque deviation, ion concentration and temperature deviation of temperature zone. A set of key performance indicators for monitoring instruments is defined, and an initial fully connected causal model is constructed between the set of candidate process parameters and the set of performance indicators. The set of performance indicators includes common-mode rejection ratio of electrocardiogram signal, blood oxygen saturation measurement accuracy, baseline drift, and power consumption leakage current. The causal directions in the initial fully connected causal model are identified using the Hilbert-Schmidt independence criterion, resulting in the causal adjacency matrix. The physical constraint equations of the monitoring instrument are constructed based on the gain-bandwidth product equation of the electrocardiogram amplifier circuit and the optical path coupling efficiency equation of the blood oxygen probe. An initial physical information neural network is constructed, using the causal adjacency matrix as the initial mask matrix for network connections. The physical constraint equation is transformed into a physical residual regularization term and embedded in the total loss function of the initial physical information neural network to obtain the target physical information neural network. The total loss function is the weighted sum of the data fitting mean square error and the physical residual regularization term. Using the spatiotemporal anomaly state feature tensor as input features, the target physical information neural network is iteratively trained using the gradient descent algorithm to obtain the physical information neural network after training convergence. The causal structure parameters and residual terms of the physical information embedding are determined based on the converged physical information neural network.

[0008] Optionally, the step of performing small-sample quality risk transfer inference based on the causal structure parameters and residual terms embedded in the physical information to obtain the quality risk prediction value and uncertainty measure value includes: The directed edge adjacency matrix in the causal structure parameters is flattened into a one-dimensional vector, and then added element by element to the residual term to generate a fused causal feature vector. The fused causal feature vector is input into a fully connected mapping layer for dimensionality reduction, generating a feature representation vector that represents the quality pattern of the current manufacturing batch. Obtain historical manufacturing batch data of the source model monitoring instrument, extract feature representation vector sets under different quality modes based on the historical manufacturing batch data, and construct a source domain prototype library based on the feature representation vector sets under different quality modes. Obtain labeled samples of the target new model of monitoring instrument, extract the corresponding target domain feature representation vector and quality label, and construct the target support set; Based on the optimal transport theory, the transport cost matrix is ​​constructed using the cosine distance between each prototype vector in the source domain prototype library and each feature representation vector in the target support set as the transport cost. Based on the transport cost matrix, the minimum transport cost vector from each target sample to each source domain prototype is determined. The minimum transportation cost vector is normalized to generate a prototype deviation vector, which represents the similarity distribution between the target sample and the normal / abnormal quality pattern of the source domain. The prototype deviation vector is input into a Bayesian neural network for variational inference to obtain the quality risk prediction value and uncertainty measure value.

[0009] Optionally, the step of optimizing process parameters based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal anomaly state feature tensor to obtain the optimal process parameter adjustment includes: Set an uncertainty threshold range, and determine the working condition judgment label based on the uncertainty metric value and the threshold range; When the working condition determination label is a steady-state working condition, the process parameters are optimized by a hybrid strategy of model predictive control and flexible actor-critic reinforcement learning based on the quality risk prediction value, uncertainty measurement value and spatiotemporal abnormal state feature tensor to obtain the first process parameter adjustment amount; When the working condition determination label is a disturbance working condition, the process parameters are optimized by the opposition learning-gray wolf optimization algorithm based on the quality risk prediction value, uncertainty measurement value and the spatiotemporal abnormal state feature tensor to obtain the second process parameter adjustment amount. Obtain the historical optimal process parameters from the previous moment's optimization output, calculate the smoothness constraint factor based on the historical optimal process parameters and the adjustment amount of the first process parameter / the adjustment amount of the second process parameter, and generate transition constraint conditions. Based on the transition constraints, the adjustment amounts of the first and second process parameters are smoothly corrected to generate the optimal process parameter adjustment amounts.

[0010] Optionally, the step of optimizing process parameters based on the predicted quality risk value, uncertainty metric value, and spatiotemporal anomaly state feature tensor using a hybrid strategy of model predictive control and flexible actor-critic reinforcement learning to obtain the first process parameter adjustment amount includes: The uncertainty metric is broadcast and multiplied with all channels of the spatiotemporal anomaly state feature tensor to generate a weighted spatiotemporal feature; The weighted spatiotemporal features, the predicted quality risk values, and the current process parameter set are concatenated to obtain a state space vector; A model predictive control rolling optimization mechanism is constructed. The state space vector is used as the initial state. Multiple sets of candidate adjustment actions are simulated within a preset prediction time domain to generate a performance index prediction deviation matrix. Construct a policy network and a value network, input the state space vector into the policy network, output action probability distribution parameters, and sample candidate process parameter adjustment actions from the action probability distribution parameters; A reward function is constructed that includes a cumulative deviation penalty term, an uncertainty penalty term, and a policy entropy incentive term. Based on the reward function and the performance index, the deviation matrix is ​​predicted. The policy network and the value network are updated by maximizing the expected return. An entropy regularization mechanism with automatic temperature coefficient adjustment is used to balance exploration and utilization, thereby generating convergent policy network parameters. The optimal action is output using the converged strategy network, and the first process parameter adjustment amount is generated based on the optimal action and the controllable parameter range in the current process parameter set.

[0011] Optionally, the step of optimizing process parameters using the oppositional learning-gray wolf optimization algorithm based on the predicted quality risk value, uncertainty metric value, and spatiotemporal anomaly state feature tensor to obtain the second process parameter adjustment amount includes: Based on the directed edge adjacency matrix in the causal structure parameters, the causal effect coefficients between key process parameters and performance indicators are extracted, and a causal-guided initial population generation space is constructed. The initial position of the gray wolf population is generated in the initial population generation space using chaotic Tent mapping, and the position of each gray wolf in the initial gray wolf population represents a set of candidate process parameter adjustment schemes. Based on the deviation between the predicted quality risk value and the target performance index value, the uncertainty metric value, and the smoothness of the process parameter adjustment, a fitness function is constructed. The fitness function is then used to calculate the fitness value of each individual gray wolf in the initial gray wolf population, generating a fitness value sequence. The fitness value sequence is arranged in ascending order of fitness value, and gray wolf individuals with fitness values ​​higher than a preset fitness threshold are extracted as individuals to be optimized; Perform an opposition learning operation on the individual to be optimized to generate an opposition individual, and calculate the fitness value of the opposition individual; The population is updated based on the fitness values ​​of the individuals to be optimized and the opposing individuals, resulting in the gray wolf population after the first evolution. Based on the fitness value sequence, the positions of the alpha wolf, the co-alpha wolf, and the third-level wolf are determined. The positions of the alpha wolf, co-alpha wolf, and the third-level wolf guide the gray wolf population after the first evolution to surround and attack prey, resulting in the gray wolf population after the second evolution. This process continues until the preset maximum number of iterations or convergence conditions are reached. The position of the alpha wolf in the gray wolf population after the second evolution is then decoded into the second process parameter adjustment amount.

[0012] Furthermore, to achieve the above objectives, this application also proposes a quality control device for the manufacture of monitoring instruments, which includes: The module is used to construct a dynamic diagram of the manufacturing process based on the operating parameters and quality inspection data of equipment in each process of the monitoring instrument manufacturing process, and to obtain the dynamic diagram structure data. The extraction module is used to extract spatiotemporal anomaly features based on the dynamic graph structure data to obtain a spatiotemporal anomaly state feature tensor. The modeling module is used to perform causal transmission modeling of physical information embedding based on the spatiotemporal anomaly state feature tensor, and to obtain the causal structure parameters and residual terms of physical information embedding. The reasoning module is used to perform small-sample quality risk transfer reasoning based on the causal structure parameters and residual terms embedded in the physical information, and to obtain the quality risk prediction value and uncertainty measure value. The control module is used to optimize process parameters based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal abnormal state feature tensor to obtain the optimal process parameter adjustment amount, and to send the optimal process parameter adjustment amount to the corresponding equipment station for execution, so as to realize the quality control of the monitoring instrument manufacturing.

[0013] In addition, to achieve the above objectives, this application also proposes a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the monitoring instrument manufacturing quality control method as described above.

[0014] This application proposes one or more technical solutions that, by constructing a dynamic diagram of the manufacturing process and combining adaptive graph wavelet transform and gated attention mechanism for spatiotemporal anomaly feature extraction, can accurately capture cross-process spatiotemporal anomaly propagation patterns, significantly improving the sensitivity to potential quality risks. By introducing the Hilbert-Schmidt independence criterion for causal discovery and embedding the core physical equations of the monitoring instrument into the neural network loss function, spurious correlations in the data-driven model are effectively eliminated, ensuring physical consistency of the causal transmission path. Furthermore, prototype networks and optimal transport theory are used to achieve small-sample domain adaptation, significantly shortening the quality convergence time during new product introduction. Finally, a dual-strategy hybrid optimization mechanism based on uncertainty measurement is introduced, using reinforcement learning for fine-tuning under steady-state conditions and global optimization algorithms for rapid relocation under disturbed conditions, balancing control stability and robustness. This effectively improves the adaptability of process parameter optimization, thereby enhancing the control accuracy and robustness of the monitoring instrument manufacturing quality, ensuring product quality stability during mass production, and reducing the defect rate. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating the first embodiment of the quality control method for the manufacturing of monitoring instruments in this application; Figure 2 This is a schematic diagram of the module structure of the quality control device for the manufacturing of monitoring instruments according to an embodiment of this application.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or a quality control device for monitoring instrument manufacturing capable of the above functions. The following description uses a quality control device for monitoring instrument manufacturing as an example to illustrate this embodiment and the subsequent embodiments.

[0022] Based on this, the embodiments of this application provide a quality control method for the manufacturing of monitoring instruments, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the monitoring instrument manufacturing quality control method of this application.

[0023] In this embodiment, the quality control method for the manufacturing of monitoring instruments includes steps S10 to S50: Step S10: Construct a dynamic diagram of the manufacturing process based on the operating parameters and quality inspection data of the equipment in each process of the monitoring instrument manufacturing process, and obtain the dynamic diagram structure data.

[0024] It should be noted that the manufacturing process dynamic graph is a mathematical model that transforms multi-source heterogeneous observation data in discrete manufacturing systems into a structured representation with topological evolution characteristics. Its essence is a graph theory abstraction with process equipment as information nodes, material flow relationships as information transmission channels, and operating parameters and quality inspection results as node state variables. The dynamic aspect is reflected in the time dependence of edge weights on production cycle time and material transfer intensity. The structured data is a unified encoding of graph topology, node attributes, and edge weight temporal sequence, forming a complete input for subsequent spectral analysis and spatiotemporal modeling.

[0025] In actual implementation, operating parameters of equipment at each process stage, such as temperature field, pressure value, movement speed, and vibration spectrum, as well as quality inspection results, such as geometric dimensional deviations, electrical characteristic values, and defect category markings, are collected through industrial data interfaces. The collected data undergoes timestamp alignment and interpolation to eliminate timing biases introduced by asynchronous multi-source sampling. Furthermore, using equipment workstations as node sets and material transfer paths between processes as directed edge sets, a time-series sequence of edge weights is calculated based on real-time material transfer frequencies. Node attribute vectors are constructed from equipment operating parameters and quality inspection results after feature encoding. Finally, the node attribute matrix is ​​fused with the adjacency tensor representing the connection strength at each time segment to output a complete dynamic graph structure data package.

[0026] In this embodiment, by abstracting the state evolution of the physical production line into a dynamic graph model, the fluctuations in equipment state and the effects of material flow during the manufacturing process are uniformly incorporated into the mathematical framework of graph topology. This abstraction enables the originally isolated point monitoring data to acquire dual structural constraints of spatial adjacency relationships and temporal evolution laws, providing the necessary mathematical foundation for subsequent multi-scale feature decomposition in non-Euclidean space, thereby overcoming the inherent defect of traditional statistical process control methods in their insufficient ability to capture inter-process correlation effects.

[0027] In one feasible implementation, step S10 may include: collecting operating parameters and quality inspection data of each device in the surface mount process, insertion process, and assembly process, and performing time alignment and missing value imputation on the operating parameters and quality inspection data to obtain a standardized time-series dataset; based on the standardized time-series dataset, mapping each process device to heterogeneous nodes, mapping the material flow relationship between processes to time-varying edges, and constructing an initial adjacency matrix based on the time-varying edges to obtain an initial graph topology; performing feature encoding on the node attributes in the initial graph topology, converting the device operating parameters and quality inspection data into node feature vectors to generate a node attribute matrix; calculating the weight values ​​of the time-varying edges under each time slice based on the time-series records of material flow, generating a time-varying edge weight sequence, and constructing a time-varying adjacency tensor based on the time-varying edge weight sequence; fusing the node attribute matrix and the time-varying adjacency tensor to obtain a manufacturing process dynamic graph, and performing time continuity verification on the manufacturing process dynamic graph to output dynamic graph structure data.

[0028] It should be noted that the standardized time-series dataset is a normalized data matrix formed by aligning and quality-restoring multi-source heterogeneous manufacturing data along the time dimension. Multi-source heterogeneity refers to differences in data sampling frequency, transmission protocols, and data formats among different devices; time alignment refers to unifying data with different sampling frequencies to the same time reference coordinate system; and missing value imputation refers to reasonably filling in data gaps caused by sensor failures, network jitter, etc., in a manner consistent with physical continuity. The standardized time-series dataset constitutes the original signal source for dynamic graph modeling, and its data quality directly affects the accuracy of subsequent graph structure representation.

[0029] In actual implementation, operating parameters and quality inspection data are collected from the PLCs and testing instruments of various equipment in the surface mount, insertion, and assembly processes via the industrial Ethernet protocol. Operating parameters include, but are not limited to, printer squeegee pressure and speed, pick-and-place machine placement pressure and nozzle vacuum, set and actual temperatures of each zone in the reflow oven, peak stress and feed rate of the depaneling machine, and mounting torque and ion concentration detection values ​​at the assembly station. Quality inspection data includes, but is not limited to, solder paste thickness and volume from the solder paste inspector, coordinates and categories of solder joint defects from the automatic optical inspection instrument, and electrical continuity signals and impedance values ​​from the online testing instrument. Due to the different sampling periods of each device, alignment is performed using a unified reference timestamp. For the offset between the original sampling time and the reference time, a linear interpolation method is used for resampling, calculated using the following formula: in, The target reference time, and These are the original sampling times before and after the target time, respectively. and This corresponds to the original collected value.

[0030] For consecutive missing data segments, a physical constraint-based interpolation method is used, such as exponentially smoothing filling of temperature parameters by considering the heat conduction time constant. After the above processing, a standardized time-series dataset with unified timestamps, complete data, and consistent dimensions is formed.

[0031] Understandably, the initial graph topology abstracts the physical production line's equipment layout and material flow paths into a skeleton network formed by nodes and edges in graph theory. Heterogeneous nodes refer to the different node attribute spaces of equipment in different processes due to their varying functions, parameter types, and quality impact mechanisms, requiring differentiation using different node types. Time-varying edges refer to the fact that the material flow relationships between processes are not constant or have a constant intensity, but rather change dynamically with production cycle time, batch switching, and equipment status. The initial adjacency matrix is ​​a binary representation of whether a material flow relationship exists at a given moment, forming the connectivity basis of the graph topology.

[0032] In actual execution, based on the device identifier field in the standardized time-series dataset, all unique device identifiers are extracted, and a unique integer index i∈{1,2,...,N} is assigned to each device, where N is the total number of devices on the production line, forming a node index list. According to the process flow definition in the manufacturing execution system, the material transfer relationships between processes are identified: if a material transfer path exists from device i to device j, a directed edge e is defined. ij The initial adjacency matrix A0∈{0,1} N×N The definition of is: Combine the node set, the directed edge set, and the initial adjacency matrix to form the initial graph topology G0=(V,E,A0), where V is the node set and E is the edge set. Understandably, the node attribute matrix is ​​a numerical matrix formed by vectorizing the state information of the equipment workstations represented by each node in the graph at a specific moment. Its core function is to compress multi-dimensional equipment operating parameters and quality inspection results into fixed-dimensional feature vectors, allowing the state of each node to be processed by subsequent graph neural networks or spectral analysis methods. Each row of the node attribute matrix corresponds to a device node, and each column corresponds to a state feature dimension. The matrix element values ​​reflect the quantization level of the equipment in that feature dimension.

[0033] In actual execution, for each device node v i The system extracts operational parameters and quality inspection data at corresponding times from a standardized time-series dataset to form the original attribute set. Due to significant differences in the physical dimensions and numerical ranges of different parameters, each parameter is first normalized using the Min-Max normalization method to linearly map its values ​​to the [0,1] interval. For discrete categorical variables in the quality inspection data, such as defect types, one-hot encoding is used to convert them into multidimensional binary vectors. The normalized parameter values ​​are then concatenated in a preset order to form node v. i eigenvector x i ∈R d , where d is the total dimension of the features. The feature vectors of all nodes are vertically stacked in node index order to generate a node attribute matrix.

[0034] Understandably, the time-varying adjacency tensor is a multidimensional numerical representation of the evolution of information transmission intensity on each edge of a graph topology over time. Unlike the initial adjacency matrix, which only characterizes the existence of edges, the time-varying adjacency tensor quantifies the real-time intensity of material flow by assigning a weight value to each edge that changes over time. The weight values ​​are calculated based on the actual physical quantities of material transfer, such as batch quantity and workpiece quantity, and are normalized to eliminate the impact of capacity differences between different processes. The three-dimensional structure of the time-varying adjacency tensor—time, source node, and target node—completely preserves the temporal information of the dynamic evolution of the manufacturing system.

[0035] In actual implementation, the timing records of material flow between processes are extracted from the manufacturing execution system, and the number of material batches transferred from equipment i to equipment j within each fixed time window Δt is counted. The size of the time window is determined based on the production line cycle time and is taken as an integer multiple of the production cycle time of a single unit. For each time slice t, the time-varying edge weight value is calculated: in, This represents the time-varying edge weight from device i to device j at time t. Let represent the set of all outgoing neighbor nodes of node i, that is, the set of the next process equipment from which material flows from equipment i. This represents the actual amount of material transferred from device i to device j within the time window t. It is a very small positive number, used to prevent the denominator from being zero.

[0036] This formula normalizes the edge weights, ensuring that the sum of the weights of all outgoing edges originating from the same node is always 1, thus guaranteeing the numerical stability of information propagation during graph convolution operations.

[0037] Stack the edge weight matrices of all time slices t∈{1,2,...,T} along the time dimension to construct a time-varying adjacency tensor.

[0038] Understandably, dynamic graph structure data is a unified data structure formed by spatiotemporal registration and fusion of node attribute matrices and time-varying adjacency tensors. Fusion refers to aligning static node attribute representations with dynamic edge weight evolution on the time axis, ensuring a complete description of the graph state at any given time. Temporal continuity verification involves checking and repairing the integrity of the fused data in the time dimension, filling in temporal breaks caused by data acquisition anomalies, and ensuring the continuity of dynamic graph evolution. Dynamic graph structure data serves as an information carrier for a complete mathematical description of the manufacturing process. In actual execution, the node attribute matrix and the time-varying adjacency tensor are aligned along the time dimension. Since the sampling frequency of node attributes may not be entirely consistent with the edge weight statistical window, the node attributes need to be resampled or interpolated based on the time slices of the time-varying adjacency tensor, ensuring that each time slice t corresponds to a set of node attributes. Subsequently, a time continuity check is performed, i.e., traversing each time slice of the time-varying adjacency tensor to check for anomalies such as rows or columns of all zeros in the adjacency matrix, and for abrupt values ​​in node attributes that exceed the physical feasible region. For detected abnormal time slices, linear interpolation is used to repair them using data from adjacent slices. For missing multiple consecutive time slices, physical constraint-based estimation is used to fill in the missing slices. After verification and repair are completed, the node index list, the node attribute matrix of each time slice, and the adjacency matrix of each time slice are packaged to output the complete dynamic graph structure data.

[0039] Step S20: Extract spatiotemporal anomaly features based on the dynamic graph structure data to obtain the spatiotemporal anomaly state feature tensor.

[0040] It should be noted that the spatiotemporal anomaly state feature tensor is a high-order representation of the anomaly patterns implicit in the dynamic diagram of the manufacturing process, which is jointly decoupled and encoded on both spatial and temporal scales. Spatial feature extraction refers to using graph wavelet transform to perform multi-resolution analysis of node attribute signals in the graph domain to separate local structural variations at different process aggregation levels. Temporal feature extraction refers to using gated recurrent units and attention mechanisms to perform dependency modeling of the evolution sequence of spatial features on the time axis to capture the long-term memory effect of quality drift and the ability to trace key events. The tensor formed by the fusion of the two achieves a unified representation of the ability to locate spatial anomaly sources and track temporal propagation paths.

[0041] In practice, node attribute matrices and time-varying adjacency tensors are parsed from dynamic graph structure data. For each time slice, the normalized graph Laplacian matrix is ​​first calculated and its eigenvalues ​​are decomposed to obtain the basis functions and frequency distribution in the graph frequency domain. Based on the eigenvalue distribution, a wavelet kernel function with adaptive scale parameters is constructed. This kernel function is used to perform convolution operations on the graph signal at multiple scales to obtain detail coefficients and approximation coefficients at different resolutions. The coefficients at each scale are aggregated along the channel dimension to generate a multi-scale spatial feature map sequence. Subsequently, the feature map sequence is segmented into overlapping time windows and input into a gated recurrent unit network to extract the hidden state sequence. The hidden state adaptively retains key historical information through a gating mechanism. Then, the hidden state sequence is input into a multi-head attention layer to calculate the contribution weight of each time step to the current state and perform weighted fusion to obtain a temporal fusion feature vector. Finally, the spatial feature slices at the current time and the temporal fusion features are extracted and concatenated to form a spatiotemporal anomaly state feature tensor.

[0042] In this embodiment, the introduction of adaptive graph wavelet transform enables the feature extraction process to automatically adjust the analysis scale according to the spectral characteristics of the graph itself, thereby simultaneously capturing abnormal patterns at three granularities: process level, workstation level, and production line level, avoiding the omission of local details or global trends by fixed-scale methods. The combination of gated recurrent units and attention mechanisms endows the model with the ability to remember and backtrack on the quality evolution trajectory over a long period of time, enabling the feature tensor at the current moment to reflect the cumulative effect and propagation path of abnormal states over a past period of time, providing rich input with spatiotemporal contextual information for subsequent causal inference.

[0043] In one feasible implementation, step S20 may include: extracting a node attribute matrix and a time-varying adjacency tensor from the dynamic graph structure data; performing spectral decomposition on the time-varying adjacency tensor to obtain a graph Laplacian matrix, and constructing an adaptive wavelet kernel function based on the eigenvalue distribution of the graph Laplacian matrix; using the adaptive wavelet kernel function to perform multi-scale graph convolution on the node attribute matrix to extract spatial structure features at different scales and generate a multi-scale spatial feature map; dividing the multi-scale spatial feature map into overlapping time window sequences according to the time dimension, and inputting the time window sequences into a gated recurrent unit network for temporal dependency feature extraction to obtain a hidden state sequence; inputting the hidden state sequence into a multi-head attention mechanism to determine the attention weights at each time step, and performing weighted fusion on the hidden state sequence based on the attention weights to obtain a temporal fusion feature; and concatenating the spatial feature slices corresponding to the latest sampling time in the multi-scale spatial feature map with the temporal fusion feature to obtain a spatiotemporal anomaly state feature tensor.

[0044] It should be noted that the graph Laplacian matrix is ​​a mathematical tool for analyzing the graph topology in the spectral domain. Its eigenvalues ​​reflect the distribution characteristics of the graph signal at different frequency components. Spectral decomposition refers to decomposing the graph Laplacian matrix into a product of eigenvalues ​​and eigenvectors, providing orthogonal basis functions for the graph Fourier transform. The adaptive wavelet kernel function is a filter function with tight support or fast attenuation characteristics in the graph frequency domain. Its scaling parameter is automatically adjusted according to the eigenvalue distribution to achieve multi-resolution bandpass filtering of graph signals in different frequency ranges.

[0045] In actual implementation, for each time slice t, row normalization is performed on the adjacency matrix A(t) to obtain the normalized adjacency matrix. Then, the normalized graph Laplacian matrix is ​​calculated, as follows: in, Let be the normalized graphical Laplacian matrix at time t. It is the identity matrix. Let be the normalized adjacency matrix at time t.

[0046] Eigenvalue decomposition is performed on the normalized graph Laplacian matrix to obtain the eigenvalue diagonal matrix and the corresponding orthogonal eigenvector matrix. The eigenvalue distribution is then obtained from the eigenvalue diagonal matrix, and an adaptive wavelet kernel function is constructed. This embodiment employs a learnable parameterization form based on the Mexican hat wavelet basis. in, For wavelet kernel function with scale parameter s, the eigenvalues The value at that location, The Laplacian eigenvalues characterize the frequency levels in the graph frequency domain. Low eigenvalues correspond to low-frequency smooth components, and high eigenvalues correspond to high-frequency detail components. s is the scale parameter that controls the frequency domain passband width of the wavelet kernel. The smaller the scale, the more sensitive to high-frequency components, and the larger the scale, the more sensitive to low-frequency components. is the normalization constant, such that, ensuring the energy comparability of wavelet coefficients at each scale.

[0047] The adaptive selection strategy for the scale parameter s is as follows: First, determine the kernel density estimate of the eigenvalue distribution, identify the clustering intervals of eigenvalues, set a central frequency for each clustering interval, and take the scale parameter as the reciprocal of the central frequency.

[0048] It can be understood that multi-scale graph convolution means using graph wavelet kernels of different scales to perform graph frequency domain filtering operations on the node attribute matrix to extract local structure features at different spatial granularities. Small-scale wavelet kernels extract subtle abnormal changes at the internal equipment level of the process, and large-scale wavelet kernels extract macroscopic quality drift patterns across processes. The multi-scale spatial feature map is an aggregated representation of the filtering results at all scales along the channel dimension.

[0049] In actual implementation, for time slice t, the filtering of the node attribute matrix X(t) in the graph frequency domain is achieved through graph Fourier transform and inverse transform. First, project the node attribute matrix into the graph frequency domain, construct frequency domain filters for each scale, perform frequency domain filtering, and then convert the filtered signal back to the spatial domain through graph Fourier inverse transform to obtain the detail coefficient matrices at all scales. At the same time, use the low-pass scale function to determine the approximation coefficients, retain the overall smooth trend of the graph, obtain the approximation coefficient matrix, and splice the detail coefficient matrices and the approximation coefficient matrix at all scales along the channel dimension to generate the multi-scale spatial feature map.

[0050] It can be understood that time window segmentation is to divide the continuous multi-scale spatial feature map sequence into subsequences with overlapping segments to capture the dynamic patterns of quality evolution within local time periods. The gated recurrent unit is a recurrent neural network with a gating mechanism that controls the retention ratio of historical information through the update gate and the forgetting ratio of historical information through the reset gate, thus effectively modeling the dependencies in long time series and alleviating the vanishing gradient problem.

[0051] In actual implementation, set the time window size to W, indicating the number of consecutive time steps included in each window, and the window sliding step size to Stride, and Stride < W to ensure overlap between windows. The multi-scale spatial feature map sequence is segmented into a window sequence {W1, W2,..., W M}, where each window W m = [M(t m ), M(t ... m+1),...,M(t m+W 1)], where tm is the starting time index of the m-th window. The multi-scale spatial feature map within each window is flattened, reshaping M(t) into a one-dimensional vector v(t). The flattened vectors of each time step within the window are sequentially input into a gated recurrent unit (GRU) network. After processing by the GRU, the hidden state of the last time step is taken as the temporal coding feature of that window. The coding features of all windows constitute the hidden state sequence.

[0052] As can be understood, multi-head attention is an information aggregation method that assigns differentiated importance weights to elements at different positions in a sequence. By mapping input features to multiple different representation subspaces and independently calculating attention weights, the multi-head mechanism can capture various types of dependency patterns in the sequence in parallel. Temporal fusion features are compact vector representations obtained by adaptively weighting and summing the hidden states of each window according to their correlation strength with the current state.

[0053] In actual implementation, the hidden state sequence is input into a multi-head attention layer. The number of attention heads is set to H, and the dimension of each attention head is d. k =d h / H. For the h-th attention head, the hidden state sequence is mapped to the query matrix Q through a linear transformation. h Key matrix K h Sum matrix V h Then, the scaling dot product attention weights are determined and weighted summation is performed to obtain the output of the attention head. The outputs of all attention heads are concatenated along the channel dimension and linearly transformed through the output projection matrix to obtain the temporal fusion feature vector.

[0054] Understandably, the spatiotemporal anomaly feature tensor is a multidimensional data representation that jointly encodes the spatial structural features at the current moment and the historical temporal evolution features. The spatial feature slice at the current moment reflects the spatial anomaly distribution pattern of the manufacturing system at the latest sampling moment; the temporal fusion feature, on the other hand, encapsulates the regularity information of the evolution of anomaly states over a period of time. By concatenating the two in the channel dimension, the ability to locate spatial anomalies and track temporal propagation is unified.

[0055] In practice, spatial feature slices corresponding to the latest sampling time are extracted from the multi-scale spatial feature map sequence. These slices are then flattened into vectors along the spatial dimension. The flattened spatial feature vector is then concatenated with the temporal fusion feature vector along the channel dimension to obtain the fused feature vector. Depending on the specific needs, the fused feature vector can be reshaped into a tensor form and output as a spatiotemporal anomaly state feature tensor.

[0056] Step S30: Based on the spatiotemporal anomaly state feature tensor, perform causal transmission modeling of physical information embedding to obtain the causal structure parameters and residual terms of physical information embedding.

[0057] It should be noted that the causal structure parameters of the physical information embedding are a set of causal effect coefficients that possess both statistical significance and mechanistic consistency, obtained by introducing the inherent physical laws of the manufactured object as constraints, based on data-driven causal discovery. Causal transmission modeling refers to identifying the directed causal relationship between process parameters and performance indicators from observed variables and expressing it in the form of a graphical model. Physical information embedding refers to incorporating the physical equations followed by the core functional modules of the monitoring instrument, such as circuit gain-bandwidth product constraints and optical path attenuation laws, into the learning objective function of the causal model as regularization terms, forcing the estimated values ​​of the causal coefficients to satisfy the rigid constraints of physical laws while fitting the observed data. The residual term is an explicit modeling output of complex nonlinear perturbation components that the physical equations cannot fully explain.

[0058] In actual implementation, abnormal nodes are located based on spatiotemporal anomaly state feature tensors. Abnormal equipment nodes are identified using feature deviation measurements. Anomaly association subgraphs are extracted by expanding the neighborhood around the abnormal nodes, and the node attributes of these subgraphs are analyzed to obtain a set of candidate process parameters. Simultaneously, a set of key performance indicators to be controlled for the monitoring instruments is defined. An initial fully connected causal model between process parameters and performance indicators is constructed. For each pair of variables, under the assumption of an additive noise model, the kernel method is used to map the variables to a regenerative kernel Hilbert space. The Hilbert-Schmidt independence criterion statistic is calculated to test the independence between noise and causal variables, thereby identifying and retaining statistically significant causal directions, forming a causal adjacency matrix. Subsequently, the gain-bandwidth product equation of the ECG amplifier circuit and the optical path coupling efficiency equation of the pulse oximeter probe are extracted as physical constraint equations. A neural network containing a causal main output branch and a residual auxiliary output branch is constructed, using the causal adjacency matrix as the connection mask for the main output branch. The physical constraint equations are transformed into physical residual regularization terms, which are weighted and combined with the data fitting loss to form the total loss function. The network is trained using the spatiotemporal anomaly state feature tensor as input, and the physical constraint weights are dynamically adjusted during training. After training converges, the connection weight matrix is ​​extracted from the causal main output branch as the causal structure parameters for physical information embedding, and the output vector is extracted from the residual auxiliary output branch as the residual term for physical information embedding.

[0059] In this embodiment, by embedding the physical equations into the causal discovery process, the spurious correlation problem commonly found in purely data-driven causal inference methods is effectively suppressed. Since the physical equations are directly derived from the fundamental electromagnetic and optical principles of monitoring instruments, their constraints on causal coefficients possess first-principles rigidity. This ensures that the final causal structure parameters not only reflect the true causal chain in the manufacturing process but also meet the consistency requirements of the underlying physical mechanisms. The independent output of the residual terms further explicitly separates manufacturing disturbances not covered by the physical model, providing supplementary information dimensions for subsequent quality prediction.

[0060] In one feasible implementation, step S30 may include: locating abnormal nodes in the spatiotemporal abnormal state feature tensor based on a preset abnormal deviation threshold; extracting abnormal nodes and their corresponding neighborhood abnormal correlation subgraphs; parsing the node attributes of the abnormal correlation subgraphs to obtain a candidate process parameter set, the candidate process parameter set including solder joint cold solder joint rate, board separation stress peak, locking torque deviation, ion concentration, and temperature deviation of the temperature zone; setting a set of key performance indicators for the monitoring instrument, and constructing an initial fully connected causal model between the candidate process parameter set and the performance indicator set, the performance indicator set including ECG signal common-mode rejection ratio, blood oxygen saturation measurement accuracy, baseline drift, and power consumption leakage current; and using the Hilbert-Schmidt independence criterion to identify the causal direction in the initial fully connected causal model. The physical constraint equations of the monitoring instrument are constructed based on the gain-bandwidth product equation of the ECG amplifier circuit and the optical path coupling efficiency equation of the pulse oximeter probe, obtained from the causal adjacency matrix. An initial physical information neural network is constructed, using the causal adjacency matrix as the initial mask matrix for network connections. The physical constraint equations are transformed into physical residual regularization terms and embedded into the total loss function of the initial physical information neural network to obtain the target physical information neural network. The total loss function is the weighted sum of the mean square error of data fitting and the physical residual regularization term. Using the spatiotemporal abnormal state feature tensor as input features, the target physical information neural network is iteratively trained using the gradient descent algorithm to obtain the converged physical information neural network. Based on the converged physical information neural network, the causal structure parameters and residual terms of the physical information embedding are determined.

[0061] It is understandable that abnormal node localization refers to identifying equipment workstations in an abnormal state at the current moment based on the degree of deviation between the feature vectors of each node in the spatiotemporal abnormal state feature tensor and the center of the normal operating mode. Here, the preset abnormal deviation threshold is the boundary value that defines normal fluctuations and abnormal states. The abnormal correlation subgraph is a local subgraph structure composed of several orders of neighbors centered on the abnormal node and extended along the material flow direction. It is used to trace the upstream and downstream process equipment related to the abnormality. The candidate process parameter set is a set of process variables that have potential causal relationships with quality abnormalities, extracted from the node attributes of the abnormal correlation subgraph.

[0062] In actual implementation, the node feature vectors of the spatiotemporal abnormal state feature tensor are first extracted from historical normal operating condition samples, and the cluster centers of the normal pattern are calculated. For node v i Its normal pattern center is denoted as The characteristic covariance matrix is ​​denoted as Σ i For the spatiotemporal anomaly state feature tensor at the current moment, extract the feature vector of each node. Calculate its Mahalanobis distance to the center of the corresponding normal pattern: in, The Mahalanobis distance between a node feature and the center of a normal pattern measures the degree to which the current feature deviates from the normal state. The larger the Mahalanobis distance, the higher the probability that the current node is abnormal. For node v i The spatiotemporal anomaly feature vector, This is the feature mean vector under normal operating conditions of this node. This is the inverse covariance matrix of the feature distribution under normal operating conditions of this node, used to eliminate the influence of correlation between feature dimensions.

[0063] The Mahalanobis distance is compared with a preset anomaly deviation threshold τ, which is determined based on the quantiles of the Mahalanobis distance distribution of the normal operating condition samples, such as the 95th quantile. >τ, then node v i Mark nodes as anomalous. Using each anomalous node as the center, perform a breadth-first search to extract all nodes within its L-order neighborhood and the edges connecting these nodes, forming an anomalous association subgraph. L is typically 2 or 3 to ensure coverage of upstream and downstream processes directly related to the anomalous node. Analyze the original attributes of each node in the anomalous association subgraph to extract the corresponding process parameter variables, forming a candidate process parameter set. This set includes, but is not limited to, solder joint failure rate, peak stress during board separation, mounting torque deviation, ion concentration, and temperature deviation within the temperature zone.

[0064] Understandably, the set of key performance indicators for monitoring instruments is a set of clinical functional parameters that have a decisive impact on the quality of the final product. For example, the common-mode rejection ratio (CMRR) of ECG signals directly affects the ability of ECG detection to suppress environmental noise; the accuracy of blood oxygen saturation measurement determines the accuracy of blood oxygen parameter detection; baseline drift is related to the morphological reproduction of the ECG waveform; and power consumption leakage current directly affects the instrument's battery life and standby safety. The initial fully connected causal model refers to an assumption, without any causal direction constraints, that there may be a directed causal relationship between each variable in the candidate process parameter set and each variable in the performance indicator set, thus forming a fully connected bipartite graph structure from the process parameter node to the performance indicator node.

[0065] In actual implementation, a set of key performance indicators Q for monitoring instruments is set. target ={q1,q2,q3,q4}, where q1 is the common-mode rejection ratio of the ECG signal, which measures the ECG amplifier circuit's ability to suppress power frequency common-mode interference; q2 is the blood oxygen saturation measurement accuracy, reflecting the deviation between the measured blood oxygen value and the true value; q3 is the baseline drift, characterizing the slow fluctuation amplitude of the ECG signal baseline over time; and q4 is the power consumption leakage current, which is related to the electrical safety of the equipment and the battery life.

[0066] Construct the variable set V=P candidate ∪Q target ={v1,v2,...,vm,vm+1,...,vm+4}. The initial fully connected causal model is defined as a directed graph. edge set Includes each process parameter node p i Pointing to each performance metric node q j All directed edges, i.e., the adjacency matrix The corresponding element in the middle is 1.

[0067] Understandably, the Hilbert-Schmidt independence criterion is a nonparametric independence measure based on the kernel method, used to test whether a statistical dependency exists between two random variables. Under the assumptions of an additive noise model, if variable X is a direct cause of variable Y, then Y can be expressed as a function of X plus a noise term independent of X. In this case, the noise term and the causal variable X should satisfy the independence condition. By testing whether the HSIC statistic of the noise and the candidate causal variable is significant, the direction of the causal relationship can be determined.

[0068] In actual implementation, for each pair of process parameter variables p i ∈P candidate and performance index variable q j ∈Q targetThe following causal direction identification process is executed: First, based on the additive noise model assumption... Where f is an unknown nonlinear function, This represents the noise term. The residual estimate is obtained by fitting the function f using Gaussian process regression or a neural network. . Variable p i and residual The samples are mapped to the reproducing kernel Hilbert space, respectively. A Gaussian radial basis function is used: in, The kernel function outputs the values ​​at input samples x and y. Let be the Euclidean distance between the sample points. The kernel bandwidth parameter is determined using the median heuristic method, i.e. The median of the distances between samples.

[0069] Calculate the kernel matrix The elements are respectively and For the kernel matrix Centralized processing: Among them, the centralized matrix , It is an n×n identity matrix. It is a column vector of all 1s, and n is the number of samples.

[0070] Estimate the Hilbert-Schmidt independence criterion statistic: in, Here, is the Hilbert-Schmidt independence criterion statistic, and tr represents the trace operation of the matrix.

[0071] The significance threshold is determined using a permutation test, which is to determine the residual. The order of the HSIC values ​​is randomly shuffled B times, usually B=1000. Each time, the HSIC values ​​after shuffling are calculated to obtain the null distribution. The empirical probability p of the original HSIC values ​​within the null distribution is determined. If p < α, where α is the significance level (e.g., 0.05), the independence hypothesis is rejected, and p is considered to be true. i and If not independent, this direction does not satisfy the additive noise model conditions; if p ≥ α, then the independence assumption is accepted, and the causal relationship direction is determined to be p. i →q j After traversing all variable pairs, the causal adjacency matrix A is obtained. causal , where element a ij =1 indicates that there is a causal directed edge from node i to node j, otherwise it is 0.

[0072] Understandably, the physical constraint equations of monitoring instruments are deterministic mathematical equations derived from the fundamental principles of electromagnetics, optics, and circuitry, describing the input-output relationships of key functional modules. These equations constitute hard constraints in the mass transfer process and are unaffected by random factors such as manufacturing batches or environmental noise. Introducing them into the causal modeling process can effectively constrain spurious causal relationships that violate physical laws, which may arise from purely data-driven methods.

[0073] In actual implementation, two core physical constraint equations are extracted. Equation one is the gain-bandwidth product equation of the ECG amplifier circuit: in, The voltage gain of the ECG amplifier circuit characterizes its ability to amplify weak ECG signals. The -3dB bandwidth of the amplifier circuit represents the range of signal frequencies that can be effectively amplified. The transconductance of the input stage field-effect transistor is determined by the device process parameters. The equivalent load capacitance at the output includes wiring parasitic capacitance and subsequent input capacitance.

[0074] This equation shows that, given device parameters, the product of gain and bandwidth is a constant, and there is a trade-off between the two.

[0075] Equation 2 is the optical path coupling efficiency equation for the pulse oximeter. in, Optical path coupling efficiency, dimensionless. For the output current of the photodetector, This is the driving current for the light-emitting diode. This is the attenuation coefficient of biological tissue to light of a specific wavelength, which is related to tissue type and wavelength. The equivalent optical path from the light source to the detector is affected by the probe bonding pressure and the structural assembly precision.

[0076] This equation describes the exponential decay of light signals during tissue penetration, and the coupling efficiency directly affects the signal-to-noise ratio and accuracy of blood oxygen saturation measurement.

[0077] Understandably, a Physical Information Neural Network (PIN) is a neural network model that incorporates physical laws into the training process as a regularization term in the loss function. Unlike traditional neural networks that only minimize data fitting errors, the total loss function of a PIN is a weighted sum of the mean squared error of data fitting and the physical residual regularization term. Here, the causal adjacency matrix, used as the initial mask matrix, applies sparsity constraints to network connections based on the causal structure identified by HSIC, allowing information transfer only between variables with causal edges. The physical residual regularization term uses the squared difference between the left and right sides of the physical equation as a penalty, forcing the causal effect coefficients output by the network to satisfy the physical constraints.

[0078] In actual implementation, an initial physical information neural network is constructed. Its network architecture includes an input layer, hidden layers, and an output bifurcation. The input layer receives the flattened vector of the spatiotemporal anomaly feature tensor, with a dimension of 1. The hidden layer contains three fully connected layers with 128, 128, and 64 neurons per layer, respectively, and the activation function is ReLU. The output bifurcation occurs after the last hidden layer, with the network branching into two parallel output branches: a causal main output branch and a residual auxiliary output branch. The causal main output branch is a fully connected layer with an output dimension of 4 (the number of performance indicators) and no activation function, used to predict the values ​​of each performance indicator. The residual auxiliary output branch is also a fully connected layer with an output dimension of 4 and the activation function is tanh, used to output the residual terms that the physical equations cannot explain.

[0079] Using the causal adjacency matrix A causal As the initial mask for the connection weights of the causal main output branch: Let the weight matrix of this branch be W. causal ∈R dhid×4 , where dhid is the dimension of the last hidden layer. For the path from the i-th process parameter to the j-th performance index, if A causal If [i,j]=0, then the corresponding weights will be fixed to 0 during training to block information transmission.

[0080] The physical constraint equations are transformed into physical residual regularization terms. Let the performance index of the network causal main output branch prediction be... The intermediate variables required for the physical equations can be extracted or deduced from the input features. The physical residual regularization term is defined as: in, The loss value is the physical residual regularization term. The first term is the squared residual of the ECG amplifier circuit gain-bandwidth product constraint, and the second term is the squared residual of the blood oxygen probe optical path coupling efficiency constraint. The two residuals together measure the degree to which the network prediction results violate known physical laws. The larger the residual, the less the prediction results conform to the underlying physical principles.

[0081] The mean square error of data fitting is defined as: in, The mean square error of the data fit. The number of samples in each training batch. Let be the performance metric vector for the network's prediction of the k-th sample. This represents the actual measured value of the performance index for the k-th sample.

[0082] The total loss function is obtained by a weighted combination of the mean squared error of data fitting and the physical residual regularization term. The weighting coefficient λ is the physical constraint weight coefficient, used to adjust the relative importance of the physical regularization term in the total loss. During training, λ adopts a dynamic scheduling strategy, initially set to a small value λ0 to allow the network to quickly fit the data distribution. As the training epochs increase, λ is gradually increased to strengthen the physical consistency constraint. A gradient descent algorithm, such as the Adam optimizer, is used to iteratively update the network parameters. The learning rate is set to η. lr Training continues until the loss function converges or the preset maximum number of training epochs is reached. The converged network is the target physical information neural network.

[0083] Understandably, the causal structure parameters of the physical information embedding refer to the connection weight matrix extracted from the causal main output branch of the physical information neural network after training convergence. The element values ​​characterize the strength of the causal effect of process parameters on performance indicators under the constraints of the physical equations. The residual term of the physical information embedding refers to the output vector extracted from the residual auxiliary output branch. Its value reflects the deviation components contributed by complex nonlinear perturbation factors in the manufacturing process that cannot be fully explained by the physical equations.

[0084] In actual implementation, causal structure parameters are extracted from the converged target physical information neural network, that is, the connection weight matrix of the causal main output branch is extracted. The element in the i-th row and j-th column of this matrix represents the connection weight matrix from the process parameter p. i To performance indicators qj The causal effect coefficient. Since the weight matrix is ​​subject to both data fitting loss and physical regularization term during training, its value reflects both the statistical correlation in the observed data and the consistency requirement of the monitoring instrument's inherent physical mechanism.

[0085] Extracting the residual term involves extracting the output vector of the residual auxiliary output branch from the input spatiotemporal anomaly feature tensor after training convergence. Each component of the residual term corresponds to a performance index, the value of which represents the nonlinear perturbation component that the physical constraint equations fail to explain.

[0086] Step S40: Perform small-sample quality risk transfer inference based on the causal structure parameters and residual terms embedded in the physical information to obtain the quality risk prediction value and uncertainty measure value.

[0087] It should be noted that the quality risk prediction value is a quantitative estimate of the probability that the performance of the final product of the current manufacturing batch will deviate from the target value, while the uncertainty measure value is a quantitative expression of the reliability of the estimate itself.

[0088] This implementation addresses the problem of insufficient model generalization ability caused by the scarcity of training samples in new models or small-batch production scenarios. The technical approach is to transfer the rich quality pattern knowledge accumulated in the source domain (historically produced models) to the current new model (target domain) using distribution alignment techniques in the metric space. This enables reliable risk inference even with very few labeled samples in the target domain. Optimal transfer theory is used to find the minimum-cost alignment scheme between the prototype distribution of the source domain and the feature distribution of the target domain, while the Bayesian inference framework is used to simultaneously provide the prediction mean and its cognitive uncertainty.

[0089] In practice, the causal structure parameters are flattened and element-wise added to the residual terms to form a fused causal feature vector. This vector is then reduced to the embedding space via a mapping layer to obtain the feature representation vector. For historical data of the source model, the mean of the embedding vectors is calculated according to the quality result category to serve as the source domain prototype, thus constructing a prototype library. For a small number of labeled samples of the target new model, embedding vectors are extracted to form the target support set. A transportation cost matrix is ​​constructed using the cosine distance between the source domain prototype vector and the target feature vector. The optimal transportation plan is obtained by solving the entropy-regularized optimal transportation problem, thereby obtaining the minimum transportation cost vector from each target sample to each source domain prototype. This cost vector is normalized to generate a prototype bias vector, representing the similarity between the target sample and the normal or abnormal quality pattern of the source domain. Finally, the prototype bias vector is input into a Bayesian neural network. The posterior distribution of the approximate network weights is approximated through variational inference. Multiple forward propagation sampling is performed, with the sample mean used as the quality risk prediction value and the sample variance used as the uncertainty measure.

[0090] In this embodiment, the domain alignment mechanism based on optimal transmission can explicitly solve for the minimum transport cost between the source and target domain distributions. Compared to traditional domain adaptation methods based on adversarial or moment matching, it has a stronger ability to preserve the distribution geometry and exhibits better stability under small sample conditions. The introduction of Bayesian variational inference enables the model to self-evaluate the reliability of its predictions. The output uncertainty metric accurately reflects the degree of distributional difference between the current sample and the source domain knowledge, providing a reliable trigger signal for the adaptive switching of subsequent control strategies. This small-sample transfer inference framework provides a novel technical path for quality control in the rapid introduction of new models in the manufacturing field.

[0091] In one feasible implementation, step S40 may include: flattening the directed edge adjacency matrix in the causal structure parameters into a one-dimensional vector, adding it element-wise with the residual term to generate a fused causal feature vector; inputting the fused causal feature vector into a fully connected mapping layer for dimensionality reduction to generate a feature representation vector representing the quality mode of the current manufacturing batch; acquiring historical manufacturing batch data of the source model monitoring instrument, extracting feature representation vector sets under different quality modes based on the historical manufacturing batch data, and constructing a source domain prototype library based on the feature representation vector sets under different quality modes; acquiring labeled samples of the target new model monitoring instrument, and extracting the corresponding... The target domain feature representation vector and quality label are used to construct a target support set. Based on optimal transport theory, the cosine distance between each prototype vector in the source domain prototype library and each feature representation vector in the target support set is used as the transport cost to construct a transport cost matrix. Based on the transport cost matrix, the minimum transport cost vector from each target sample to each source domain prototype is determined. The minimum transport cost vector is normalized to generate a prototype deviation vector, which represents the similarity distribution between the target sample and the normal / abnormal quality patterns of the source domain. The prototype deviation vector is input into a Bayesian neural network for variational inference to obtain the quality risk prediction value and uncertainty measure value.

[0092] It should be noted that the fused causal feature vector is a joint feature vector obtained by flattening the directed edge adjacency matrix in the causal structure parameters embedded with physical information and adding it element-wise with the residual terms embedded with physical information. Flattening refers to rearranging the elements in the multidimensional matrix into a one-dimensional vector in the order of rows or columns to eliminate the influence of the original data structure on subsequent processing. Element-wise addition is a feature fusion strategy. Its physical meaning is to numerically superimpose the linear causal effect components under the constraints of the physical equation with the nonlinear perturbation components that the physical equation cannot explain, forming a complete description of the causal transmission characteristics of the current manufacturing state.

[0093] In actual execution, a directed edge adjacency matrix is ​​extracted from the causal structure parameters embedded in physical information. This matrix is ​​then flattened into a one-dimensional vector in row-major order. Simultaneously, the residual terms embedded in physical information are obtained. To align the dimensions of the residual terms with the flattened causal vector, the residual terms are broadcast and expanded along the process parameter dimension. Subsequently, element-wise addition is performed to obtain the fused causal feature vector.

[0094] Understandably, the feature representation vector is a compact vector representation obtained in a low-dimensional embedding space after the fused causal feature vector is dimensionality-reduced through a fully connected mapping layer. The purpose of the dimensionality reduction mapping is to eliminate redundant information in the fused causal feature vector and project data from different manufacturing batches onto a unified low-dimensional manifold, so that samples with similar quality patterns are close to each other in the embedding space, creating conditions for subsequent prototype matching based on distance metrics. The parameters of the fully connected mapping layer are pre-trained on historical data in the source domain using the backpropagation algorithm.

[0095] Understandably, the source domain prototype library is a set of representative reference vectors constructed in the embedding space based on historical manufacturing batch data of source model monitoring instruments, according to different quality modes. Here, the source model refers to a historical model that has accumulated a large amount of production data and is similar to the target new model in terms of process and structure; the quality mode refers to the category divided according to the final quality inspection results, such as qualified and abnormal, or a more refined quality level; and the prototype vector refers to the mean of all sample feature representation vectors under a certain quality mode, serving as the representative center point of that class in the embedding space.

[0096] In actual implementation, historical manufacturing batch data of the source model monitoring instrument is retrieved, including process parameter records, quality inspection results, and corresponding measured performance indicators for each batch. Based on the final quality judgment results, historical batches are divided into different quality mode categories, for example, a category set C={c1,c2,...,cK}, where K is the number of categories. For each category ck, the fused causal feature vectors of all historical batches belonging to that category are extracted, and a corresponding feature representation vector set is generated through the aforementioned fully connected mapping layer. The prototype vector for that category is then calculated. in, Let be the prototype vector of the quality pattern of class ck. This refers to the number of historical batch samples included in this category. Let be the feature representation vector of the i-th sample in this category. By summarizing and storing the prototype vectors of all categories, we obtain the source domain prototype library.

[0097] Understandably, the target support set is a small dataset formed by extracting corresponding feature vectors from a limited number of manufacturing batch samples with existing quality labels for a new model of monitoring instrument. In the early stages of introducing a new model, full-scale quality testing has not yet been completed, and only a small number of trial production or first-piece inspection samples have reliable quality labels. The target support set is thus the reference data constructed using these scarce labeled samples to guide transfer reasoning.

[0098] In actual execution, the N of the target new model monitoring instrument is obtained. tgt There are labeled samples, typically N.tgt The sample size is much smaller than the source domain sample size, such as 5 to 20. For each sample, its fused causal feature vector is extracted according to the aforementioned process, and a feature representation vector is generated through the same fully connected mapping layer, with its corresponding quality label recorded. The sample feature vector and label are combined to form the target support set.

[0099] Understandably, optimal transport theory is a mathematical theory that studies how to transform one probability distribution into another with minimal cost. In this implementation, the alignment problem between the discrete distribution of the source domain prototype library and the feature distribution of the target support set is modeled as an optimal transport problem. The cosine distance between the source domain prototype vector and the target domain feature vector is used as the transport cost. Solving for the optimal transport plan yields the minimum transport cost vector from each target sample to each source domain prototype. This vector quantifies the distribution distance between the target sample and each source domain quality pattern, serving as the basis for subsequent similarity determination.

[0100] In practical implementation, the marginal weights of the source domain prototype distribution and the target domain feature distribution are first defined. Since each prototype may have a different sample size in the source domain, its marginal weight can be defined as the sample proportion of the corresponding category. Uniform weights are assigned to each sample in the target domain, and a transportation cost matrix is ​​constructed, where each element represents the cost of transporting a unit mass from the source domain prototype to the target sample. Cosine distance is used as the cost metric. The optimal transportation problem with entropy regularization is solved, and the Sinkhorn algorithm is used iteratively to calculate the approximately optimal transportation plan. For the j-th target sample, the minimum transportation cost vector to each source domain prototype is defined as the j-th column of the transportation plan.

[0101] Understandably, the prototype bias vector is a probability distribution vector obtained by normalizing the minimum transportation cost vector. Its components represent the similarity distribution between the target sample and each quality pattern in the source domain. Normalization transforms transportation cost, i.e., distance, into probabilistic similarity, making the bias vectors between different samples comparable. Essentially, the prototype bias vector is a soft-assignment representation of the target sample in the source domain quality pattern space.

[0102] Understandably, a Bayesian neural network is a neural network model that treats network weights as random variables rather than deterministic values. By imposing a prior distribution on the weights and using variational inference methods to approximate the true posterior distribution of the weights with a simple variational distribution, a Bayesian neural network can output a predicted distribution, not just point estimates. Variational inference is an optimization method for approximating the posterior distribution; by minimizing the KL divergence between the variational distribution and the true posterior, a posterior approximation is obtained at an acceptable computational cost. The final output quality risk prediction is the mean of the predicted distribution, and the uncertainty metric is the variance of the predicted distribution.

[0103] In practical implementation, a Bayesian neural network is constructed, consisting of an input layer, hidden layers, and an output layer. The input layer receives the prototype bias vector; the hidden layer comprises two fully connected hidden layers, each with 32 neurons and a ReLU activation function; the output layer has a dimension of 1, representing the quality risk prediction value, a continuous value, typically normalized to [0,1], and has no activation function. All network weights are assumed to be parameters, and a standard Gaussian prior distribution is applied to each weight. Then, a mean-field variational distribution is used to approximate the posterior, assuming that the weights are independent and follow a Gaussian distribution. The optimization objective of variational inference is to maximize the lower bound of evidence.

[0104] Gradient estimation is performed using a reparameterization technique: for each weight, samples are taken from a standard normal distribution. Calculate weights Therefore, variational parameters can be used. , Find the gradient.

[0105] During training, the Bayesian neural network is pre-trained using labeled samples from the source domain to obtain initial variational parameters. In the target domain inference phase, for a given prototype bias vector π, S Monte Carlo forward propagations are performed to obtain S predicted outputs.

[0106] Calculate the predicted value of quality risk, i.e., the predicted mean: in, This represents the average value of the quality risk forecast, i.e., the predicted quality risk value. Let S be the predicted output obtained from the s-th forward propagation, where S is the number of Monte Carlo samplings.

[0107] Calculate the uncertainty measure, i.e., the prediction variance: in, This is a measure of uncertainty.

[0108] Step S50: Based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal abnormal state feature tensor, optimize the process parameters to obtain the optimal process parameter adjustment amount, and send the optimal process parameter adjustment amount to the corresponding equipment station for execution, so as to realize the quality control of the monitoring instrument manufacturing.

[0109] It should be noted that process parameter optimization is a process of finding a set of process parameter corrections that minimize future quality risks through decision-making algorithms. Using uncertainty metrics as the criterion for strategy selection, a reinforcement learning strategy that balances safety constraints and long-term rewards is employed for fine-tuning when the system is in a steady state. When the system encounters significant disturbances, a metaheuristic search strategy with strong global exploration capabilities is switched for rapid repositioning. The resulting adjustment instructions are then issued to physical equipment for execution, and the feedback data is fed back to the dynamic graph construction module, forming a complete closed loop of data acquisition, state awareness, causal reasoning, risk assessment, decision optimization, and execution feedback.

[0110] In actual execution, a threshold range for uncertainty measurement is set, and the current uncertainty measurement value is compared with it to determine the current operating condition of the system. When the operating condition is determined to be steady state, the uncertainty measurement value is used as a weighting factor on the spatiotemporal abnormal state feature tensor, and concatenated with the quality risk prediction value and the current process parameters to form a state vector. Subsequently, a hybrid solution process of model predictive control and flexible actor-critic reinforcement learning is invoked: a prediction model is constructed using causal structure parameters, and the response of candidate actions is simulated in the prediction time domain to generate a deviation prediction matrix; at the same time, the state vector is input into the policy network to sample actions, and the network parameters are updated with a reward function that includes deviation penalty, risk penalty, smoothing penalty and policy entropy, outputting the first process parameter adjustment amount. When the operating condition is determined to be perturbation, the gray wolf optimization process of opposition learning reinforcement is invoked: the search space is constrained using causal effect coefficients, the population is initialized with chaotic mapping, a fitness function is constructed with quality risk deviation, uncertainty and adjustment smoothness, and the position update of opposition learning and hierarchy guidance is iteratively executed to decode the leader individual into the second process parameter adjustment amount. At the moment of strategy switching, the current output is smoothly weighted using historical adjustment values ​​to avoid parameter step shocks. The final adjustment value is then sent to the device controller for execution, and the feedback parameters after execution are collected to update the dynamic graph node attributes before entering the next loop.

[0111] In this embodiment, by using the uncertainty metric as the driving signal for strategy switching, autonomous perception and adaptive response to different manufacturing conditions are achieved. The hybrid strategy under steady-state conditions utilizes the physical constraints of model predictive control to ensure adjustment safety, while simultaneously employing the exploration mechanism of maximum entropy reinforcement learning to continuously search for a better process window within the safety boundary. Under perturbation conditions, the global search strategy can quickly escape local optima and efficiently locate new feasible solutions within a finite search space guided by causal priors. The strategy switching buffer mechanism further eliminates the risk of abrupt changes in control variables caused by discrete switching. The entire closed-loop architecture enables the quality control model to continuously optimize itself through continuous perception-decision-execution-feedback iterations, achieving a paradigm shift from passive detection to proactive regulation.

[0112] This embodiment provides a quality control method for the manufacturing of monitoring instruments. By constructing a dynamic graph of the manufacturing process and combining adaptive graph wavelet transform and gated attention mechanism for spatiotemporal anomaly feature extraction, it can accurately capture the spatiotemporal anomaly propagation pattern across processes, significantly improving the sensitivity to potential quality risks. By introducing the Hilbert-Schmidt independence criterion for causal discovery and embedding the core physical equations of the monitoring instrument into the neural network loss function, it effectively eliminates spurious correlations in the data-driven model, ensuring physical consistency of the causal transmission path. Furthermore, it utilizes prototype networks and optimal transport theory to achieve small-sample domain adaptation, significantly shortening the quality convergence time during the new product introduction period. Finally, it introduces a dual-strategy hybrid optimization mechanism based on uncertainty measurement. Under steady-state conditions, it uses reinforcement learning for fine-tuning, and under perturbed conditions, it uses a global optimization algorithm for rapid relocation, balancing control stability and robustness. This effectively improves the adaptability of process parameter optimization, thereby enhancing the control accuracy and robustness of the monitoring instrument manufacturing quality, ensuring the product quality stability of the monitoring instrument during mass production, and reducing the defect rate.

[0113] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, step S30 includes steps S501 to S505: Step S501: Set an uncertainty threshold range, and determine the working condition judgment label based on the uncertainty metric value and the threshold range.

[0114] Understandably, the uncertainty threshold range is a numerical boundary used to distinguish the operating condition type of the current operating state of a manufacturing system. A steady-state condition refers to a state where the manufacturing system operates smoothly, the parameters of each process fluctuate within a small range around the setpoint, and the quality risk prediction model has a high degree of confidence in its output. A disturbed state refers to a state where the manufacturing system is subjected to significant internal or external disturbances, such as changes in raw material batches, sudden equipment drift, or fluctuations in environmental temperature and humidity, leading to a significant increase in the prediction uncertainty of the quality risk model. The operating condition label is a discrete identifier generated based on the comparison between the uncertainty metric and the threshold range, used to trigger subsequent differentiated optimization strategies.

[0115] In actual implementation, an uncertainty threshold range [θlow, θhigh] is set, where θlow is the upper limit of the steady-state threshold, representing the maximum allowable uncertainty when the model's prediction confidence is high, and θhigh is the lower limit of the perturbation threshold, representing the minimum uncertainty that requires triggering a global re-optimization strategy. The thresholds are determined based on the statistical distribution of prediction uncertainties of historical normal operating batches: the uncertainty values ​​of the Bayesian neural network outputs of historical production batches of the source model are collected, their empirical cumulative distribution function is determined, and the 80th percentile is taken as θlow, and the 95th percentile is taken as θhigh.

[0116] Obtain the quality risk prediction value and uncertainty metric value for the current control cycle. Compare the uncertainty metric value with the threshold range to generate a condition label: when the uncertainty metric value ≤ θlow, the condition label is determined to be a steady-state condition; when the uncertainty metric value ≥ θhigh, the condition label is determined to be a disturbed condition; when θlow < uncertainty metric value < θhigh, the label of the previous cycle is maintained. For cases where the uncertainty is in the middle of the threshold range, a hysteresis control strategy is adopted to maintain the label of the previous cycle to avoid control oscillations caused by frequent switching near the threshold boundary.

[0117] Step S502: When the working condition determination label is a steady-state working condition, the process parameters are optimized by using a hybrid strategy of model predictive control and flexible actor-critic reinforcement learning based on the quality risk prediction value, uncertainty metric value and spatiotemporal abnormal state feature tensor to obtain the first process parameter adjustment amount.

[0118] It should be noted that the hybrid strategy of model predictive control and flexible actor-critic reinforcement learning is a composite control method that combines physical model-based rolling optimization with data-driven policy learning. Model predictive control uses causal structure parameters embedded with the physical information obtained in the preceding steps to construct a quality response prediction model. Within a finite prediction time domain, it simulates the future impact of different process parameter adjustment schemes on performance indicators, generating a deviation prediction matrix that provides a physically prior reference trajectory for reinforcement learning. Flexible actor-critic reinforcement learning is a maximum entropy reinforcement learning algorithm that directly outputs the probability distribution of process parameter adjustment actions through a policy network, evaluates the long-term rewards of state-action pairs through a value network, and introduces a policy entropy term into the reward function to encourage exploration, thereby continuously searching for a better process window within the safety constraints.

[0119] Optionally, the predicted quality risk value, uncertainty metric, current adjustable process parameter vector, and spatiotemporal anomalous state feature tensor are concatenated to obtain the state input vector for reinforcement learning. The uncertainty metric is used as a weighting factor to weight the spatiotemporal anomalous state feature tensor, amplifying the impact of high uncertainty features on the optimization decision. A pre-trained policy network is invoked to process the state vector, sampling to obtain a set of candidate process parameter adjustment actions. A quality response prediction model is constructed based on causal structure parameters. Within N prediction time domains, rolling simulations are performed on the future quality risk changes and process parameter fluctuations corresponding to each candidate action, generating a prediction matrix containing the deviations of each time domain step. A comprehensive reward function is constructed, including deviation penalty, risk penalty, smoothing penalty, and policy entropy: the deviation penalty term constrains the deviation of process parameters from the target setpoint; the risk penalty term penalizes the accumulated quality risk in the future time domain; the smoothing penalty term constrains the difference between the current adjustment and the previous cycle's adjustment to avoid excessive adjustment; and the policy entropy term encourages policy exploration and avoids premature convergence to a local optimum. The value of state-action pairs is evaluated using an evaluation network. The parameters of the policy network and the evaluation network are updated based on the reward function and the output of the value network. The optimal action output is the first process parameter adjustment amount.

[0120] In one feasible implementation, step S502 may include: broadcasting and multiplying the uncertainty metric value with all channels of the spatiotemporal abnormal state feature tensor to generate a weighted spatiotemporal feature; concatenating the weighted spatiotemporal feature, the quality risk prediction value, and the current process parameter set to obtain a state space vector; constructing a model predictive control rolling optimization mechanism, using the state space vector as the initial state, simulating multiple sets of candidate adjustment actions within a preset prediction time domain to generate a performance index prediction deviation matrix; constructing a policy network and a value network, inputting the state space vector into the policy network, outputting action probability distribution parameters, and sampling candidate process parameter adjustment actions from the action probability distribution parameters; constructing a reward function containing a cumulative deviation penalty term, an uncertainty penalty term, and a policy entropy incentive term, updating the policy network and the value network by maximizing the expected return based on the reward function and the performance index prediction deviation matrix, and using an entropy regularization mechanism with automatic temperature coefficient adjustment to balance exploration and utilization, generating converged policy network parameters; using the converged policy network to output the optimal action, and generating the first process parameter adjustment amount based on the optimal action and the controllable parameter range in the current process parameter set.

[0121] It should be noted that the uncertainty-weighted spatiotemporal feature is a feature representation obtained by multiplying the uncertainty metric as a scaling factor with all feature channels of the spatiotemporal anomaly state feature tensor element-wise. Here, broadcast multiplication refers to expanding the scalar uncertainty metric to the same dimension as each channel of the spatiotemporal anomaly state feature tensor before performing element-wise multiplication. Its physical meaning is: when the uncertainty of the model's prediction of the current state is high, the spatiotemporal anomaly feature is globally amplified, giving it greater weight in subsequent decisions, thereby enhancing the system's sensitivity to uncertain states and its response priority; when the uncertainty is low, the feature amplitude is suppressed accordingly to avoid overreaction.

[0122] Understandably, the state-space vector is a comprehensive vector representation that fully describes the current state and risk situation of the manufacturing system, formed by splicing and fusing the uncertainty-weighted spatiotemporal characteristics, the predicted quality risk value, and the current set of process parameters. The predicted quality risk value provides a quantitative estimate of the deviation of the final quality of the current batch from the target; the current set of process parameters provides the real-time set values ​​of each adjustable process variable. The combination of these three elements constitutes the complete state input for the reinforcement learning agent to perceive the environment.

[0123] It is understandable that model predictive control (MMC) rolling optimization mechanism is an advanced control method that determines the current optimal control action by solving an optimization problem within a finite prediction time domain, based on an explicit prediction model of the controlled object. In this embodiment, the prediction model is constructed from causal structure parameters embedded with physical information, used to simulate the future impact trajectory of different process parameter adjustment schemes on the performance indicators of the monitoring instrument. The performance indicator prediction deviation matrix is ​​the set of cumulative performance deviation values ​​corresponding to multiple candidate adjustment actions within the prediction time domain, providing a reference evaluation signal with physical priors for reinforcement learning.

[0124] In actual implementation, causal structure parameters embedded in physical information are used to determine a linearized response model of the performance index vector to the adjustment amount of the process parameters under the current process parameters. The prediction time domain is set to H steps, and the control time domain to Hc steps, where Hc ≤ H, and the adjustment amount is assumed to remain zero after exceeding the control time domain. Within the adjustable range of the process parameters, M candidate adjustment action sequences are generated through random sampling or Latin hypercube sampling. For the k-th candidate action sequence, the performance index trajectory is recursively calculated along the prediction time domain. Then, based on the performance index trajectory, the cumulative performance index prediction deviation of the k-th candidate action in the prediction time domain is calculated. The cumulative deviations of all candidate actions are arranged to generate a performance index prediction deviation matrix.

[0125] Understandably, the policy network is a neural network in reinforcement learning used to directly output action decisions. Its input is a state space vector, and its output is the parameters of the action probability distribution, such as the mean and log-standard deviation of a Gaussian distribution. The value network is a neural network used to estimate the long-term expected reward of the state-action pair, providing the gradient direction for updating the policy network. The flexible actor-critic algorithm uses a stochastic policy to sample specific actions from the distribution output by the policy network, thus preserving the ability to explore the environment.

[0126] In practical implementation, a reinforcement learning module based on the flexible actor-critic algorithm is constructed, comprising a policy network and a value network. The policy network takes a state space vector as input, passes through two hidden layers (256 neurons per layer, ReLU activation), and its output layer branches into two parallel fully connected layers, outputting the action mean and log-standard deviation, respectively. To ensure a positive standard deviation, the actual network output is the log-standard deviation. To mitigate the Q-value overestimation problem, a double-Q network structure is used in the value network. Each value network takes a concatenation of the state space vector and the action vector as input, passes through two hidden layers, and outputs a scalar Q-value. The action sampling process is as follows: noise is sampled from a standard normal distribution. The uncompressed action is identified, and then compressed to [ ] using the tanh function. 1,1] m Finally, the normalized action is mapped to the actual adjustable range of the process parameters to obtain the candidate process parameter adjustment actions.

[0127] Understandably, the reward function is a scalar signal used in reinforcement learning to evaluate the quality of an agent's actions. Its design directly determines the direction of policy optimization. In this implementation, the reward function is composed of three weighted core terms: the cumulative bias penalty term guides the policy to optimize in the direction of reducing quality risk; the uncertainty penalty term prompts the policy to tend to explore in regions with high model confidence, avoiding radical adjustments in cognitive blind spots; and the policy entropy encouragement term gives the policy an inherent exploratory drive, preventing premature convergence to local optima.

[0128] In actual execution, for the current state and the sampled candidate actions, the cumulative deviation of the model predictive control corresponding to the action is first determined. If the action is consistent with an action in the model predictive control candidate action set or can be corresponding through interpolation, its cumulative deviation E(a) is taken; otherwise, the action is input into the predictive model to calculate its cumulative deviation online.

[0129] Understandably, updating the policy network and value network involves iteratively optimizing the network parameters using gradient descent to maximize the expected return objective function. The core of the flexible actor-critic algorithm lies in simultaneously maximizing cumulative reward and policy entropy, thereby maintaining sufficient exploration capability while pursuing high performance. The automatic temperature coefficient adjustment mechanism is a technique that adaptively adjusts the weights of the entropy regularization term, automatically maintaining an appropriate level of randomness during training by pre-setting the target entropy.

[0130] In practice, an experience replay pool D is maintained to store transition samples generated by the agent's interactions with the environment. At each training step, a batch of data is randomly sampled from the replay pool for network updates. For value network updates, the loss function is the mean squared error of the soft Bellman residuals. The smaller of the two Q-values ​​in the double Q-network is used to mitigate overestimation. For policy network updates, the policy network is updated by maximizing the sum of the expected Q-value and the entropy regularization term. Since the sampling process involves reparameterization, the gradient of the policy parameters can be directly calculated. Automatic temperature coefficient adjustment involves adaptively updating the temperature coefficient by minimizing the loss function. The loss functions of the three are optimized alternately until the network parameters converge, yielding the converged policy network parameters.

[0131] Understandably, the first process parameter adjustment is the optimal process parameter correction value obtained by deterministic inference of the current state using the converged policy network under steady-state conditions. During the inference phase, to obtain stable control commands, random sampling is no longer performed from the policy distribution; instead, the mean action value output by the policy network is directly used, and after the same mapping transformation, it becomes the final first process parameter adjustment.

[0132] Input the current state space vector into the converged policy network and take the mean of its output actions. Compress this mean using the tanh function to [...]. 1,1] m Then, by mapping the values ​​to the actual adjustable range of the process parameters, the adjustment amount of the first process parameter can be obtained.

[0133] Step S503: When the working condition determination label is a disturbance dynamic working condition, the process parameters are optimized by the opposition learning-gray wolf optimization algorithm based on the quality risk prediction value, uncertainty measurement value and the spatiotemporal abnormal state feature tensor to obtain the second process parameter adjustment amount.

[0134] It should be noted that the Opposition Learning-Grey Wolf Optimization Algorithm is a swarm intelligence optimization method that integrates opposition learning strategies. The Grey Wolf Optimization Algorithm simulates the social hierarchy and cooperative hunting behavior of a grey wolf pack, guiding the population towards the optimal solution region through alpha wolves, beta wolves, and delta wolves—that is, the alpha wolf, deputy alpha wolf, and third-level wolves. The opposition learning strategy simultaneously evaluates candidate solutions and their opposites, increasing the probability of covering potentially high-quality regions in the search space, accelerating population convergence, and improving global search capabilities. Under perturbation conditions, the high uncertainty of the quality risk model indicates that the current combination of process parameters has deviated significantly from the original optimization trajectory, requiring a global search strategy with strong capabilities to escape local optima and relocate feasible solutions.

[0135] Optionally, a gray wolf population is randomly initialized within the feasible search range of the current process parameters to obtain an initial candidate solution set. Simultaneously, a corresponding opposing candidate solution is generated for each initial candidate solution. The original candidate solutions and opposing candidate solutions are merged, and their fitness is evaluated. Individuals with higher fitness are selected as the initial population. The fitness function is a weighted average of the predicted quality risk, parameter constraint penalty, and uncertainty penalty; lower quality risk and less uncertainty result in higher fitness. Subsequently, the population is ranked according to a social hierarchy. The top three individuals in terms of fitness are designated as α, β, and δ wolves, respectively, and the remaining individuals are designated as ω wolves. The α, β, and δ wolves jointly guide the search direction of the population. The search step size and direction of each gray wolf individual relative to the positions of α, β, and δ wolves are calculated, and the positions of all gray wolf individuals are updated, completing one iteration. After each iteration, opposing solutions are generated again for each individual in the current population, and their fitness is evaluated. Individuals with lower fitness in the population are replaced to maintain population diversity, and the ranking of α, β, and δ wolves is updated. Repeat the above iterative process until the maximum number of iterations is reached or the fitness meets the preset threshold. Finally, decode the position parameters corresponding to α wolf to obtain the process parameter adjustment amount, which is the second process parameter adjustment amount.

[0136] In one feasible implementation, step S503 may include: extracting the causal effect coefficients between key process parameters and performance indicators based on the directed edge adjacency matrix in the causal structure parameters, and constructing a causally guided initial population generation space; generating the initial position of the gray wolf population within the initial population generation space using a chaotic Tent mapping, thereby generating the initial gray wolf population, wherein the position of each gray wolf individual in the initial gray wolf population represents a set of candidate process parameter adjustment schemes; constructing a fitness function based on the deviation between the quality risk prediction value and the target performance indicator value, the uncertainty metric value, and the smoothness of the process parameter adjustment amount, and using the fitness function to calculate the fitness value of each gray wolf individual in the initial gray wolf population, thereby generating a fitness value sequence; and processing the fitness value sequence... Sort the gray wolves in ascending order of fitness value, and extract those with fitness values ​​higher than a preset fitness threshold as individuals to be optimized; perform an adversarial learning operation on the individuals to be optimized to generate adversarial individuals, and calculate the fitness value of the adversarial individuals; perform population update based on the fitness values ​​of the individuals to be optimized and the adversarial individuals to obtain the gray wolf population after the first evolution; determine the positions of the alpha wolf, the co-alpha wolf, and the third-level wolf based on the fitness value sequence, and guide the gray wolf population after the first evolution to surround and attack prey based on the positions of the alpha wolf, the co-alpha wolf, and the third-level wolf to obtain the gray wolf population after the second evolution, until the preset maximum number of iterations or convergence condition is reached, and decode the position of the alpha wolf in the gray wolf population after the second evolution into the second process parameter adjustment amount.

[0137] It should be noted that the initial population generation space guided by causality refers to the feasible solution space formed by extracting the causal effect coefficients between key process parameters and performance indicators using the directed edge adjacency matrix embedded in the causal structure parameters of physical information, and then applying differentiated constraints to the search boundary of the Grey Wolf optimization algorithm based on these coefficients. Here, the causal effect coefficient represents the intensity of the impact of a unit change in process parameter on the performance indicator. The larger the absolute value of the coefficient, the more significant the impact of adjusting the process parameter on the final quality, and a wider exploration range should be assigned to it in the search space. Causal guidance is reflected in the fact that the search range is no longer uniformly distributed, but rather weighted and scaled according to the causal strength, so that the algorithm can concentrate more search resources on the key process parameter dimensions that have a greater impact on quality, thereby improving search efficiency.

[0138] Understandably, the chaotic Tent map is a deterministic iterative sequence generation function with ergodicity and pseudo-randomness, used to generate initial population positions that are uniformly distributed and have good diversity within the search space. Compared to uniform random sampling, the chaotic map can more evenly cover all regions of the search space, reducing the clustering of the initial population in local areas and improving the convergence stability of the global search algorithm. Each gray wolf individual's position encodes a complete set of process parameter adjustment schemes.

[0139] In practice, the population size is set to Np, typically between 30 and 50. Initial values ​​for the chaotic variables of the first individual are randomly generated within the interval (0,1). For subsequent individuals in the population, the chaotic variables of each dimension are recursively generated using the Tent mapping. To prevent the chaotic sequence from falling into a periodic loop, small perturbations are applied to the minute values ​​near the boundary during iteration. The generated chaotic variables are mapped into the search space to obtain the position coordinates of the k-th gray wolf individual in the i-th dimension. Thus, the position vector of each gray wolf individual represents a set of candidate process parameter adjustment schemes. The position vectors of all individuals constitute the initial gray wolf population.

[0140] Understandably, the fitness function is a quantitative indicator used to evaluate the merits of candidate process parameter adjustment schemes represented by individual gray wolves. A higher fitness value indicates a better overall performance of the scheme in reducing quality risk, maintaining model confidence, and controlling the smoothness of adjustments. In this implementation, the fitness function integrates three dimensions: the deviation between the predicted quality risk value and the target value reflects the effectiveness of the scheme; the uncertainty metric reflects the reliability of the scheme; and the smoothness of the process parameter adjustment reflects the feasibility of the scheme, as shown in the following formula: in, Let be the fitness value of the candidate process parameter adjustment schemes corresponding to the kth gray wolf individual. Let w1, w2, and w3 be the process parameter adjustment vector corresponding to the kth gray wolf individual, where w1, w2, and w3 are preset positive weight coefficients. The deviation between the current predicted quality risk value and the preset target quality risk value. This is the uncertainty measure of the quality risk prediction corresponding to the current plan.

[0141] Taking a negative sign for the fitness function transforms the original minimization problem into a maximization problem, consistent with the convention in the gray wolf optimization algorithm that higher fitness equates to better performance. This fitness function is then used to calculate the fitness value of each individual in the initial gray wolf population, generating a fitness value vector.

[0142] Understandably, fitness-ranked selection of individuals to be optimized is a strategy that selectively applies stronger optimization operations to inferior individuals with lower fitness, based on their relative performance within the current population. Individuals with fitness values ​​above a preset threshold are considered to be optimized; these individuals are in a disadvantaged position within the current population and have significant potential for improvement. By concentrating computational resources on improving inferior individuals, rather than applying equal effort to all individuals, the convergence speed of the population can be accelerated.

[0143] In actual implementation, the elements in the fitness value vector are sorted in ascending order to obtain the sorted fitness value sequence and the corresponding individual index mapping. A preset fitness threshold is set, which can be the median of the current population fitness values, or the p-th percentile after sorting, such as the 50th percentile. The gray wolf individuals with fitness values ​​higher than the preset fitness threshold are extracted to form a set of individuals to be optimized. Individuals with fitness values ​​not higher than the threshold are considered as dominant individuals in the current population, and in this embodiment, adversarial learning operations are not applied to preserve their superior genes.

[0144] Understandably, the complement learning operation is an optimization strategy that accelerates search space coverage by simultaneously evaluating candidate solutions and their complements. For a given candidate solution X, its complement in the search space is defined as Xopp = lb + ub. X. The basic assumption of opposition learning is that if a solution performs poorly, its dual position has a high probability of being located in a better region. By simultaneously evaluating the original solution and its opposition and selecting the better one, two symmetrical positions can be explored in a single iteration with low computational cost, significantly improving search efficiency. For each gray wolf individual in the set of individuals to be optimized, its opposition individual position vector is generated, and the opposition individual is substituted into the aforementioned fitness function to calculate its fitness value.

[0145] Understandably, population update based on fitness comparison refers to a process where, for each individual to be optimized, its original fitness value is compared with the fitness value of its counterpart, and the superior individual is retained for the next generation—a process of natural selection. This operation determines whether to replace the original individual with the counterpart, thereby achieving targeted improvement of the population's genes. For each gray wolf individual in the set of individuals to be optimized, its original fitness value and the fitness of its counterpart are obtained. If the fitness of the counterpart is greater than the original fitness value, the counterpart replaces the original individual's position in the population; otherwise, the original individual remains unchanged. For dominant individuals not in the set of individuals to be optimized, their positions remain unchanged. After this update, the gray wolf population after its first evolution is obtained.

[0146] Understandably, the social hierarchy in the gray wolf optimization algorithm serves as an optimization mechanism simulating the strict hierarchical system of the gray wolf pack. The three individuals with the best fitness values ​​in the population are labeled as α wolf, β wolf, and δ wolf, representing the alpha wolf, deputy alpha wolf, and third-level wolf, respectively. These individuals are considered to be closest to prey, i.e., the global optimum. The position updates of the remaining individuals, i.e., ω wolf, are jointly guided by the positions of these three alpha wolves, achieving gradual population convergence by moving closer to α, β, and δ wolves. The convergence factor decays linearly with the number of iterations, giving the algorithm strong global exploration capabilities in the early stages of the search.

[0147] In practice, the fitness value of each individual is recalculated based on the gray wolf population after the first evolution, resulting in an updated fitness value vector. Individuals with the best fitness values ​​are sorted in descending order and labeled as α wolf, β wolf, and δ wolf. The convergence factor 'a' corresponding to the current iteration number 't' is determined, linearly decreasing from 2 to 0 with each iteration. For each individual X in the population, the update amount by which it moves towards α, β, or δ wolves is determined. The new position of the individual after moving towards the three wolves is the arithmetic mean of these three values. To ensure the new position remains within the search space, components exceeding the boundary are pruned, and the original individual position is replaced with the updated position, resulting in the gray wolf population after the second evolution.

[0148] Understandably, the iteration termination condition refers to the criterion for determining whether the algorithm has found a satisfactory solution or whether further iteration will not lead to significant improvement. Termination conditions include reaching a preset maximum number of iterations, or the change in the optimal fitness value of the population over several consecutive generations being less than a preset convergence threshold. When the termination condition is met, the algorithm stops iterating and decodes the position of the individual with the highest fitness value in the current population, i.e., the α wolf, into the final process parameter adjustment scheme.

[0149] Step S504: Obtain the historical optimal process parameters from the previous moment's optimization output, calculate the smoothness constraint factor based on the historical optimal process parameters and the adjustment amount of the first process parameter / the adjustment amount of the second process parameter, and generate transition constraint conditions.

[0150] It should be noted that the smoothness constraint factor is a buffer correction coefficient introduced to prevent abrupt changes in process parameters during operating condition switching. When switching between steady-state and disturbed states, due to the different solution mechanisms and output characteristics of the two optimization strategies, directly adopting the adjustment amount output by the new strategy may cause excessive changes in process parameters within adjacent control cycles, resulting in secondary shocks to the manufacturing process. The smoothness constraint factor generates a transition command that balances response speed and stability by integrating the historical optimal adjustment amount with the current strategy output.

[0151] In actual execution, the historical optimal process parameter adjustment amount ΔPprev, the final output of the previous control cycle, is obtained. For the first process parameter adjustment amount ΔP1 or the second process parameter adjustment amount ΔP2 obtained under the current operating condition, they are collectively denoted as ΔPcurr, and the difference metric between them is determined. A maximum permissible single-step change amount δmax is set, determined based on the physical constraints of the process parameters on the equipment actuators, such as the upper limit of temperature regulation rate and the upper limit of pressure change rate. A smoothness constraint factor is determined based on the difference metric, and then transition constraints are generated: if the difference metric between the current adjustment amount and the historical adjustment amount does not exceed the maximum permissible single-step change amount δmax, the smoothness constraint factor is set to 1, and the current adjustment amount ΔPcurr is directly output as the final adjustment amount; if the difference metric exceeds δmax, the smoothness constraint factor is set to δmax / difference metric, and the current adjustment amount is scaled proportionally to ΔPfinal = ΔPprev + α(ΔPcurr - ΔPprev), where α is the smoothness constraint factor, ensuring that the change amplitude of process parameters in adjacent cycles does not exceed the maximum allowable range of the equipment and process, avoiding the impact of parameter mutations on the stability of the manufacturing system. Finally, the adjusted process parameters after smoothing constraint correction are output to the manufacturing execution system to complete the quality closed-loop optimization of the current control cycle and enter the cyclic optimization process of the next control cycle.

[0152] Step S505: Based on the transition constraint, the adjustment amount of the first process parameter / the adjustment amount of the second process parameter is smoothly corrected to generate the optimal adjustment amount of the process parameter.

[0153] It should be noted that smoothing correction is a process of reconstructing the output of the original optimization strategy under transitional constraints. Its purpose is to adopt the optimization direction of the current strategy as much as possible while ensuring the stability of the manufacturing process. The adjustment amount after smoothing correction is the optimal process parameter adjustment amount finally issued to the equipment actuator.

[0154] In this embodiment, a dual-strategy hybrid optimization mechanism based on uncertainty measurement is introduced. Reinforcement learning is used for fine-tuning under steady-state conditions, while global optimization algorithm is used for rapid relocation under disturbed conditions. This balances the stability and robustness of control, effectively improving the control accuracy and robustness of the monitoring instrument manufacturing quality and ensuring the product quality stability of the monitoring instrument during mass production.

[0155] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the quality control method for manufacturing monitoring instruments of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0156] This application also provides a quality control device for the manufacture of monitoring instruments; please refer to [reference needed]. Figure 2 The quality control device for the manufacture of monitoring instruments includes: Module 10 is used to construct a dynamic diagram of the manufacturing process based on the operating parameters and quality inspection data of the equipment in each process of the monitoring instrument manufacturing process, and to obtain the dynamic diagram structure data.

[0157] The extraction module 20 is used to extract spatiotemporal anomaly features based on the dynamic graph structure data to obtain a spatiotemporal anomaly state feature tensor.

[0158] Modeling module 30 is used to perform causal transmission modeling of physical information embedding based on the spatiotemporal anomaly state feature tensor, and to obtain the causal structure parameters and residual terms of physical information embedding.

[0159] The reasoning module 40 is used to perform small-sample quality risk transfer reasoning based on the causal structure parameters and residual terms embedded in the physical information to obtain the quality risk prediction value and uncertainty measure value.

[0160] The control module 50 is used to optimize process parameters based on the predicted quality risk value, the uncertainty metric value and the spatiotemporal abnormal state feature tensor, obtain the optimal process parameter adjustment amount, and send the optimal process parameter adjustment amount to the corresponding equipment station for execution, so as to realize the quality control of the monitoring instrument manufacturing.

[0161] The monitoring instrument manufacturing quality control device provided in this application, employing the monitoring instrument manufacturing quality control method described in the above embodiments, can solve the technical problem of poor adaptability of process parameters in the manufacturing process, which affects the control accuracy and robustness of the monitoring instrument manufacturing quality, and fails to guarantee the product quality stability of the monitoring instrument during mass production, resulting in a high defect rate. Compared with the prior art, the beneficial effects of the monitoring instrument manufacturing quality control device provided in this application are the same as those of the monitoring instrument manufacturing quality control method provided in the above embodiments, and other technical features in the monitoring instrument manufacturing quality control device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0162] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the monitoring instrument manufacturing quality control method as described above.

[0163] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A quality control method for the manufacture of a monitoring instrument, characterized in that, The method includes: Based on the operating parameters and quality inspection data of equipment in each process of the monitoring instrument manufacturing process, a dynamic diagram of the manufacturing process is constructed to obtain the dynamic diagram structure data; Based on the dynamic graph structure data, spatiotemporal anomaly features are extracted to obtain a spatiotemporal anomaly state feature tensor. Based on the spatiotemporal anomaly state feature tensor, causal transmission modeling of physical information embedding is performed to obtain the causal structure parameters and residual terms of physical information embedding; Based on the causal structure parameters and residual terms embedded in the physical information, small-sample quality risk transfer reasoning is performed to obtain the quality risk prediction value and uncertainty measure value. Based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal anomaly state feature tensor, the process parameters are optimized to obtain the optimal process parameter adjustment amount, and the optimal process parameter adjustment amount is sent to the corresponding equipment station for execution to achieve quality control of the monitoring instrument manufacturing.

2. The method as described in claim 1, characterized in that, The manufacturing process dynamic diagram is constructed based on the operating parameters and quality inspection data of equipment at each stage of the monitoring instrument manufacturing process, including... The operating parameters and quality inspection data of each equipment in the surface mount process, insertion process and assembly process are collected, and the operating parameters and quality inspection data are time aligned and missing values ​​are imputed to obtain a standardized time series dataset. Based on the standardized time-series dataset, each process equipment is mapped as a heterogeneous node, the material flow relationship between processes is mapped as a time-varying edge, and an initial adjacency matrix is ​​constructed based on the time-varying edge to obtain the initial graph topology. The node attributes in the initial graph topology are feature-encoded, and the equipment operating parameters and quality inspection data are converted into node feature vectors to generate a node attribute matrix. Based on the time-series records of material flow, the weight values ​​of the time-varying edges under each time slice are calculated, a time-varying edge weight sequence is generated, and a time-varying adjacency tensor is constructed based on the time-varying edge weight sequence. By fusing the node attribute matrix with the time-varying adjacency tensor, a dynamic graph of the manufacturing process is obtained, and the time continuity of the dynamic graph is verified, and the dynamic graph structure data is output.

3. The method as described in claim 1, characterized in that, The process of extracting spatiotemporal anomaly features based on the dynamic graph structure data to obtain a spatiotemporal anomaly state feature tensor includes: Extract the node attribute matrix and time-varying adjacency tensor from the dynamic graph structure data; The time-varying adjacency tensor is spectrally decomposed to obtain the graph Laplacian matrix, and an adaptive wavelet kernel function is constructed based on the eigenvalue distribution of the graph Laplacian matrix. The node attribute matrix is ​​convolved with the adaptive wavelet kernel function to extract spatial structural features at different scales and generate a multi-scale spatial feature map. The multi-scale spatial feature map is divided into overlapping time window sequences according to the time dimension, and the time window sequences are input into a gated recurrent unit network for temporal dependency feature extraction to obtain the hidden state sequence. The hidden state sequence is input into a multi-head attention mechanism to determine the attention weights at each time step. The hidden state sequence is then weighted and fused based on the attention weights to obtain temporal fusion features. By concatenating the spatial feature slices corresponding to the latest sampling time in the multi-scale spatial feature map with the temporal fusion features, a spatiotemporal anomaly state feature tensor is obtained.

4. The method as described in claim 1, characterized in that, The causal transmission modeling based on the spatiotemporal anomaly state feature tensor for physical information embedding yields the causal structure parameters and residual terms of the physical information embedding, including: Based on a preset abnormal deviation threshold, the feature tensor of the spatiotemporal abnormal state is used to locate abnormal nodes, extract abnormal nodes and their corresponding neighborhood abnormal correlation subgraphs, parse the node attributes of the abnormal correlation subgraphs, and obtain a set of candidate process parameters. The set of candidate process parameters includes solder joint cold solder rate, board separation stress peak, locking torque deviation, ion concentration and temperature deviation of temperature zone. A set of key performance indicators for monitoring instruments is defined, and an initial fully connected causal model is constructed between the set of candidate process parameters and the set of performance indicators. The set of performance indicators includes common-mode rejection ratio of electrocardiogram signal, blood oxygen saturation measurement accuracy, baseline drift, and power consumption leakage current. The Hilbert-Schmidt independence criterion is used to identify the causal directions in the initial fully connected causal model, and the causal adjacency matrix is ​​obtained. The physical constraint equations of the monitoring instrument are constructed based on the gain-bandwidth product equation of the electrocardiogram amplifier circuit and the optical path coupling efficiency equation of the blood oxygen probe. An initial physical information neural network is constructed, using the causal adjacency matrix as the initial mask matrix for network connections. The physical constraint equation is transformed into a physical residual regularization term and embedded in the total loss function of the initial physical information neural network to obtain the target physical information neural network. The total loss function is the weighted sum of the data fitting mean square error and the physical residual regularization term. Using the spatiotemporal anomaly state feature tensor as input features, the target physical information neural network is iteratively trained using the gradient descent algorithm to obtain the physical information neural network after training convergence. The causal structure parameters and residual terms of the physical information embedded are determined based on the converged physical information neural network.

5. The method as described in claim 1, characterized in that, The method of performing small-sample quality risk transfer inference based on the causal structure parameters and residual terms embedded in the physical information to obtain quality risk prediction values ​​and uncertainty measures includes: The directed edge adjacency matrix in the causal structure parameters is flattened into a one-dimensional vector, and then added element by element to the residual term to generate a fused causal feature vector. The fused causal feature vector is input into a fully connected mapping layer for dimensionality reduction, generating a feature representation vector that represents the quality pattern of the current manufacturing batch. Obtain historical manufacturing batch data of the source model monitoring instrument, extract feature representation vector sets under different quality modes based on the historical manufacturing batch data, and construct a source domain prototype library based on the feature representation vector sets under different quality modes. Obtain labeled samples of the target new model of monitoring instrument, extract the corresponding target domain feature representation vector and quality label, and construct the target support set; Based on the optimal transport theory, the transport cost matrix is ​​constructed using the cosine distance between each prototype vector in the source domain prototype library and each feature representation vector in the target support set as the transport cost. Based on the transport cost matrix, the minimum transport cost vector from each target sample to each source domain prototype is determined. The minimum transportation cost vector is normalized to generate a prototype deviation vector, which represents the similarity distribution between the target sample and the normal / abnormal quality pattern of the source domain. The prototype deviation vector is input into a Bayesian neural network for variational inference to obtain the quality risk prediction value and uncertainty measure value.

6. The method as described in claim 1, characterized in that, The process parameter optimization based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal anomaly state feature tensor to obtain the optimal process parameter adjustment includes: Set an uncertainty threshold range, and determine the working condition judgment label based on the uncertainty metric value and the threshold range; When the working condition determination label is a steady-state working condition, the process parameters are optimized by a hybrid strategy of model predictive control and flexible actor-critic reinforcement learning based on the quality risk prediction value, uncertainty measurement value and spatiotemporal abnormal state feature tensor to obtain the first process parameter adjustment amount; When the working condition determination label is a disturbance working condition, the process parameters are optimized by the opposition learning-gray wolf optimization algorithm based on the quality risk prediction value, uncertainty measurement value and the spatiotemporal abnormal state feature tensor to obtain the second process parameter adjustment amount. Obtain the historical optimal process parameters from the previous moment's optimization output, calculate the smoothness constraint factor based on the historical optimal process parameters and the adjustment amount of the first process parameter / the adjustment amount of the second process parameter, and generate transition constraint conditions. Based on the transition constraints, the adjustment amounts of the first and second process parameters are smoothly corrected to generate the optimal process parameter adjustment amounts.

7. The method as described in claim 6, characterized in that, The first process parameter adjustment amount is obtained by optimizing the process parameters using a hybrid strategy of model predictive control and flexible actor-critic reinforcement learning based on the predicted quality risk value, uncertainty metric value, and spatiotemporal anomaly state feature tensor, including: The uncertainty metric is broadcast and multiplied with all channels of the spatiotemporal anomaly state feature tensor to generate a weighted spatiotemporal feature; The weighted spatiotemporal features, the predicted quality risk values, and the current process parameter set are concatenated to obtain a state space vector; A model predictive control rolling optimization mechanism is constructed. The state space vector is used as the initial state. Multiple sets of candidate adjustment actions are simulated within a preset prediction time domain to generate a performance index prediction deviation matrix. Construct a policy network and a value network, input the state space vector into the policy network, output action probability distribution parameters, and sample candidate process parameter adjustment actions from the action probability distribution parameters; A reward function is constructed that includes a cumulative deviation penalty term, an uncertainty penalty term, and a policy entropy incentive term. Based on the reward function and the performance index, the deviation matrix is ​​predicted. The policy network and the value network are updated by maximizing the expected return. An entropy regularization mechanism with automatic temperature coefficient adjustment is used to balance exploration and utilization, thereby generating convergent policy network parameters. The optimal action is output using the converged strategy network, and the first process parameter adjustment amount is generated based on the optimal action and the controllable parameter range in the current process parameter set.

8. The method as described in claim 6, characterized in that, The process parameter optimization based on the predicted quality risk value, uncertainty metric value, and spatiotemporal anomaly state feature tensor using the opposition learning-gray wolf optimization algorithm to obtain the second process parameter adjustment amount includes: Based on the directed edge adjacency matrix in the causal structure parameters, the causal effect coefficients between key process parameters and performance indicators are extracted, and a causal-guided initial population generation space is constructed. The initial position of the gray wolf population is generated in the initial population generation space using chaotic Tent mapping, and the position of each gray wolf in the initial gray wolf population represents a set of candidate process parameter adjustment schemes. Based on the deviation between the predicted quality risk value and the target performance index value, the uncertainty metric value, and the smoothness of the process parameter adjustment, a fitness function is constructed. The fitness function is then used to calculate the fitness value of each individual gray wolf in the initial gray wolf population, generating a fitness value sequence. The fitness value sequence is arranged in ascending order of fitness value, and gray wolf individuals with fitness values ​​higher than a preset fitness threshold are extracted as individuals to be optimized; Perform an opposition learning operation on the individual to be optimized to generate an opposition individual, and calculate the fitness value of the opposition individual; The population is updated based on the fitness values ​​of the individuals to be optimized and the opposing individuals, resulting in the gray wolf population after the first evolution. Based on the fitness value sequence, the positions of the alpha wolf, the co-alpha wolf, and the third-level wolf are determined. The positions of the alpha wolf, co-alpha wolf, and the third-level wolf guide the gray wolf population after the first evolution to surround and attack prey, resulting in the gray wolf population after the second evolution. This process continues until the preset maximum number of iterations or convergence conditions are reached. The position of the alpha wolf in the gray wolf population after the second evolution is then decoded into the second process parameter adjustment amount.

9. A quality control device for the manufacture of monitoring instruments, characterized in that, The device includes: The module is used to construct a dynamic diagram of the manufacturing process based on the operating parameters and quality inspection data of equipment in each process of the monitoring instrument manufacturing process, and to obtain the dynamic diagram structure data. The extraction module is used to extract spatiotemporal anomaly features based on the dynamic graph structure data to obtain a spatiotemporal anomaly state feature tensor. The modeling module is used to perform causal transmission modeling of physical information embedding based on the spatiotemporal anomaly state feature tensor, and to obtain the causal structure parameters and residual terms of physical information embedding. The reasoning module is used to perform small-sample quality risk transfer reasoning based on the causal structure parameters and residual terms embedded in the physical information, and to obtain the quality risk prediction value and uncertainty measure value. The control module is used to optimize process parameters based on the predicted quality risk value, the uncertainty metric value, and the spatiotemporal abnormal state feature tensor to obtain the optimal process parameter adjustment amount, and to send the optimal process parameter adjustment amount to the corresponding equipment station for execution, so as to realize the quality control of the monitoring instrument manufacturing.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the quality control method for manufacturing monitoring instruments as described in any one of claims 1 to 8.