A defect property guided process optimization method and system
By extracting real-time defect features from workpiece images and constructing a quality prediction model, combined with reinforcement learning, a precise mapping from defect features to process parameter adjustment amounts is achieved. This solves the problems of low optimization efficiency in high-dimensional parameter spaces and the inability of static optimization systems to autonomously adapt to changes, thereby improving the real-time control capability of intelligent manufacturing systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies lack methods that can use real-time defect information as a direct guiding signal and quickly and accurately lock key adjustment parameters in a high-dimensional parameter space for efficient optimization, resulting in intelligent manufacturing systems being unable to achieve real-time adaptive control.
By extracting and quantifying real-time defect features from workpiece images, a quality prediction model is constructed. Furthermore, through reinforcement learning, a mapping from defect features to process parameter adjustment amounts is established, enabling adaptive control based on real-time defect detection.
It enables precise control of process parameters, improves the targeting and efficiency of control, and the system has adaptive capabilities, enabling continuous optimization in dynamic environments, reducing computing resource consumption and long-term maintenance costs.
Smart Images

Figure CN121504941B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of industrial process parameter optimization and intelligent manufacturing, and particularly relates to a defect characteristic guided process optimization method and system. BACKGROUND
[0002] In the industrial manufacturing process, process parameter optimization is a key link to ensure the stability of product quality. Currently, there are mainly two types of optimization methods: a method based on a physical model and a data-driven method.
[0003] The method based on the physical model optimizes by establishing a mathematical relationship between the process parameters and the product quality. This method depends on a deep understanding of the process mechanism, but has obvious limitations in actual application. Taking typical manufacturing processes such as stamping and welding as examples, the process involves the coupling of multiple physical fields, and it is difficult to establish an accurate model. At the same time, with the influence of factors such as equipment wear and tear and raw material batch changes, the model needs to be continuously modified, and the maintenance cost is high. The adaptability and generalization ability of this method in actual production lines are insufficient.
[0004] The data-driven method optimizes by analyzing historical production data to establish a statistical relationship between process parameters and quality indicators. This method reduces the dependence on mechanism models, but in actual application, it also faces two outstanding problems: first, existing methods usually take historical pass rates and other macro indicators as optimization targets, and cannot perceive and respond to real-time specific defects of various forms that appear on the production line. When a new defect mode appears or the equipment drifts slightly, the optimal parameter combination obtained based on the historical static model often cannot effectively eliminate the current actual defects; second, when the number of process parameters is large, the optimization process needs to search in a high-dimensional space, resulting in low computational efficiency and difficulty in meeting the real-time optimization requirements.
[0005] Therefore, the existing technology lacks a method that can use real-time defect information as a direct guide signal and quickly and accurately lock the key adjustment parameters in a high-dimensional parameter space and perform efficient optimization, which restricts the ability of intelligent manufacturing systems to evolve towards real-time self-adaptive regulation and control. SUMMARY
[0006] To solve the problems existing in the prior art, the application provides a defect characteristic guided process optimization method and system, which realizes intelligent screening and accurate regulation and control of process parameters driven by real-time defect visual features.
[0007] To achieve the above-mentioned purpose, the application provides the following solutions:
[0008] A defect characteristic guided process optimization method, the method comprising:
[0009] extracting and quantifying real-time defect features of a workpiece image;
[0010] Construct a quality prediction model;
[0011] Process parameters are selected based on real-time defect characteristics and quality prediction models;
[0012] By establishing a mapping from defect features to process parameter adjustment amounts through reinforcement learning, adaptive control based on real-time defect detection can be achieved.
[0013] Preferred methods for extracting and quantifying real-time defect features from workpiece images include:
[0014] A deep learning-based defect detection model is used to process workpiece images to obtain the bounding box coordinates and defect category labels of defects;
[0015] Deep convolutional neural networks are used to extract depth, morphological, and texture features of defect regions.
[0016] The depth features, morphological features, and texture features are spliced together to obtain the fused features;
[0017] Principal component analysis was used to reduce the dimensionality of the fused features to obtain the defect feature vector.
[0018] Preferred methods for screening process parameters based on real-time defect characteristics and quality prediction models include:
[0019] Based on the learned Q-function, select the one with the highest value. One parameter:
[0020] ;
[0021] Output key parameter index set The corresponding subset of key parameters ;
[0022] in, Indicates state, Indicates an action, For the total number of actions, This indicates the original process parameter vector. Parameter index in Represents the first element in the original process parameter vector. The specific values of each parameter These are network parameters.
[0023] Preferably, methods for establishing a mapping from defect features to process parameter adjustment amounts through reinforcement learning to achieve adaptive control based on real-time defect detection include:
[0024] ;
[0025] wherein, is a key parameter adjustment amount, s represents a state, denotes a network parameter, represents a maximum adjustment amplitude vector, represents an optimal strategy.
[0026] The application also provides a defect characteristic guided process optimization system, which is used to implement the foregoing method, and comprises an extraction module, a construction module, a screening module and a regulation and control module.
[0027] The extraction module is used to extract real-time defect features of a workpiece image.
[0028] The construction module is used to construct a quality prediction model.
[0029] The screening module is used to screen process parameters based on the real-time defect features and the quality prediction model.
[0030] The regulation and control module is used to establish a mapping from defect features to process parameter adjustment amounts through reinforcement learning, so as to realize adaptive regulation and control based on real-time defect detection.
[0031] Preferably, the extraction module comprises a preprocessing unit, a feature extraction unit, a fusion unit and a dimension reduction unit.
[0032] The preprocessing unit is used to process a workpiece image by using a defect detection model based on deep learning, so as to obtain a bounding box coordinate of a defect and a defect category label.
[0033] The feature extraction unit is used to extract deep features, morphological features and texture features of a defect area by using a deep convolutional neural network.
[0034] The fusion unit is used to splice the deep features, the morphological features and the texture features to obtain fused features.
[0035] The dimension reduction unit is used to perform dimension reduction on the fused features by using principal component analysis, so as to obtain a defect feature vector.
[0036] Preferably, the process of screening process parameters based on real-time defect features and a quality prediction model comprises:
[0037] According to the learned Q function, the parameter with the highest value is selected.
[0038] ;
[0039] Output a key parameter index set , corresponding to a key parameter subset ;
[0040] wherein, represents a state, represents an action, is the total number of actions, represents the parameter index in the original process parameter vector , represents the specific value of the th parameter in the original process parameter vector, is the network parameter.
[0041] Preferably, the mapping from defect features to process parameter adjustment amounts is established by reinforcement learning, and the process of realizing adaptive regulation based on real-time defect detection comprises:
[0042] ;
[0043] wherein, is the key parameter adjustment amount, s represents a state, represents a network parameter, represents a maximum adjustment amplitude vector, represents an optimal strategy.
[0044] Compared with the prior art, the present application has the following beneficial effects:
[0045] 1. Significant improvement in regulation accuracy and pertinence
[0046] The traditional method adjusts parameters based on macro quality indicators, and the regulation action is disconnected from the specific morphological characteristics of defects. The present application constructs a precise mapping of "defect-strategy" by taking real-time defect visual features as direct input and guidance. The system can dynamically lock key process parameters and perform directional optimization according to the unique causes and characteristics of defects, so that each adjustment is "targeted". This avoids the blindness of traditional regulation in mechanism and improves the accuracy of defect elimination and the final quality of products from the source.
[0047] 2. Substantial growth in system optimization and decision-making efficiency
[0048] The traditional full-parameter optimization method has inherent defects of high computational resource consumption and long optimization period due to the large search space. The present application converts the high-dimensional optimization problem into efficient search on a small subset of key parameters through the two-stage architecture of "preliminary selection and then optimization", fundamentally avoiding "brute force search". This design greatly reduces the computational overhead and improves the convergence speed, making real-time online optimization and closed-loop control of complex manufacturing processes possible, and significantly improving the production rhythm and resource utilization efficiency.
[0049] 3. System self-adaptation and intelligent level of fundamental evolution
[0050] The traditional method based on fixed rules or static model cannot adapt to the dynamic changes of the production environment. The reinforcement learning core of the present application has the ability of online learning and self-optimization, and can continuously adjust and improve its regulation strategy through continuous interaction. This learning mechanism guided by defect characteristics enables the system to accumulate experience in dealing with different defect patterns, thereby having strong adaptability and long-term robustness to complex dynamic working conditions, and realizing the fundamental change from "depending on artificial intervention" to "autonomous intelligent decision-making". BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the present application, the following briefly introduces the drawings needed in the embodiments. Obviously, the drawings described in the following are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0052] Figure 1 A defect characteristic guided process optimization method flowchart of an embodiment of the present application. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0054] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0055] Technical problems solved:
[0056] 1. Solving the problem of low optimization efficiency in high-dimensional parameter space:
[0057] When existing optimization algorithms deal with dozens to hundreds of process parameters, they search in the whole parameter space, resulting in "combinatorial explosion", which leads to unaffordable calculation amount and extremely slow convergence speed. This makes the existing method unable to be used for real-time optimization of production line, and can only be used for time-consuming offline analysis. The core problem to be solved by the present application is to break the limitation of high-dimensional parameter space on optimization efficiency, so that online and real-time precise regulation becomes possible.
[0058] 2. Solving the problem that static optimization system cannot adapt to changes autonomously, leading to long-term performance decline:
[0059] Static models trained based on historical data, the core of which is the lack of ability to learn from real-time production results. It regards the production process as an unchanging static system. Therefore, when the production environment changes due to equipment wear, raw material fluctuations, and other inevitable factors, the static model cannot perceive these changes, nor can it autonomously adjust its strategy, and its recommended parameters will quickly become outdated, leading to quality fluctuations. This forces engineers to frequently and manually re-tune or re-train the model, resulting in high long-term maintenance costs for the system and the inability to achieve truly "unmanned" intelligent operation. The invention aims to solve the fundamental contradiction between this static system and the dynamic environment, aiming to create an intelligent system that can continuously optimize itself as experienced engineers do with changing production conditions.
[0060] The invention is applied in adaptive process control systems, intelligent manufacturing execution systems, high-end material and precision machining production lines. To detect whether the invention is used, check whether the system dynamically filters key process parameters based on the visual features of the defects after identifying the defects, and outputs accurate numerical adjustment instructions in real time. Having such a "defect feature guided closed-loop parameter adaptive regulation" capability is considered as using the invention scheme.
[0061] Embodiment one
[0062] As shown in Figure 1 The invention provides a defect characteristic guided process optimization method, which comprises:
[0063] extracting and quantifying real-time defect features of workpiece images;
[0064] constructing a quality prediction model;
[0065] filtering process parameters based on real-time defect features and quality prediction model;
[0066] establishing a mapping from defect features to process parameter adjustment amounts through reinforcement learning to achieve adaptive regulation based on real-time defect detection. The specific implementation process is as follows:
[0067] Step 1: Extraction and quantification of real-time defect features
[0068] This step is the basis for realizing defect guided optimization. This step requires not only to identify where the defects are in the image and what the defects are, but more importantly, to convert the visual appearance of the defects into a quantitative defect feature vector that a computer can understand and process. The specific implementation is as follows:
[0069] (1) Defect detection and positioning: input the workpiece image collected by the industrial camera on the production line in real time , wherein H and Wrespectively. The embodiment adopts an anchor-based single-stage target detection model paradigm as the defect detection model The input image is processed. The model extracts multi-scale feature maps of the image through a shared weight backbone network ResNet . Subsequently, through a region proposal network (RPN) or directly applying a detection head on the feature map, the boundary box and the category of the defect are predicted, and the boundary box coordinates of the defect are output , wherein is the center coordinates of the th boundary box, is its width and height; the defect category label is output , wherein represents the defect type, is the total number of preset defect categories, is the number of detected defects.
[0070] (2) Feature extraction and quantification: the visual features of the defect region are extracted by using a deep convolutional neural network, and the specific process is as follows: for each detected defect boundary box , first, the corresponding region of interest is cropped from the original image , the region of interest is scaled to a fixed size, such as pixels, and is input into a pre-trained deep convolutional neural network ResNet, VGG, etc. The convolutional neural network is stacked by multiple convolutional layers, pooling layers and nonlinear activation functions, and its forward propagation process can be abstractly represented as:
[0071] ,
[0072] wherein is the feature map output by the last layer of the network. denotes the convolutional neural network. In order to obtain a global feature representation of a fixed length, the feature map is usually subjected to a global average pooling (GAP) operation in the spatial dimension :
[0073] ,
[0074] Thus, the deep feature vector of the defect region is obtained. Meanwhile, the morphological feature of the defect is calculatedincluding area, perimeter, circularity, etc. and texture features including contrast, entropy, etc. .
[0075] (3) Feature fusion and dimension reduction: concatenate three types of features to get complete feature vector . Use principal component analysis to reduce the dimension of the fused features, retain the most important 512-dimensional features, and get defect feature vector .
[0076] (4) Feature standardization: standardize the reduced features:
[0077] ,
[0078] where and are the feature mean and standard deviation vectors calculated by a large number of samples. The final output defect feature vector uniquely represents the visual characteristics of the current defect.
[0079] Step 2: Construction and application of quality prediction model:
[0080] This step builds a key quality prediction model, i.e. scoring the severity of the current detected defect. This model will provide the core evaluation standard for subsequent parameter screening and optimization. The specific implementation is as follows:
[0081] (1) Model architecture design: build a quality prediction model , where is the model parameter. This model uses a multi-layer feedforward neural network (Multi-Layer Perceptron, MLP) architecture, which is characterized by early fusion of defect feature vector and process parameter vector , and learning the non-linear mapping from the fused features to the quality score. Specifically, first concatenate the defect feature and process parameter into a joint feature vector:
[0082] , ,
[0083] Then, the joint feature vector is input to a network containing fully connected layers. The forward propagation formula of the first layer is as follows:
[0084] ,
[0085] where, and are the weight matrix and bias vector of the layer, is a nonlinear activation function (such as ReLU, LeakyReLU, etc.). The final output layer does not use an activation function to perform a regression task:
[0086] ,
[0087] Thus, the output is the predicted quality score The present invention first proposes to fuse defect visual features, morphological texture features, and cross-process process parameters in this model to directly predict macro quality scores and provide high-quality correlation signals for subsequent causal discovery.
[0088] (2) Training data preparation: Collect a large number of sample data from the historical production database , is the number of samples, where is the corresponding true quality label, which can be a binary label, i.e. pass / fail (1 / 0); continuous quality score, i.e. score between 0-1 based on defect severity; multi-level quality rating, i.e. numerical values corresponding to levels such as excellent, good, medium, and poor.
[0089] (3) Model training: Train the quality prediction model using supervised learning, with the loss function as:
[0090] ,
[0091] Optimize model parameters through backpropagation algorithm so that the model can accurately predict the product quality under given defect features and process parameters.
[0092] (4) Model verification and application: Verify the prediction accuracy of the model on an independent test set to ensure its reliability in actual application. The trained quality prediction model will serve as the basis for reward calculation in the subsequent steps.
[0093] Step 3: Defect-guided key parameter screening
[0094] This step is to solve the problem of low optimization efficiency in high-dimensional parameter space. Innovatively, it maps defect features and process parameter sequences into a dynamic decision-making process, constructing an intelligent decision-making problem for selecting which parameter to optimize. The question to be answered is which parameter to adjust for the specific defect at hand to bring the best long-term quality return. The specific implementation is as follows:
[0095] (1) State construction: The defect feature vector obtained in the first step and all the process parameter values of the current production line Concatenating to form the state vector:
[0096] ,
[0097] where is the total number of process parameters.
[0098] (2) Deep Q-network design: A deep Q-network is constructed to evaluate the long-term value of each parameter. The network structure contains three fully connected layers, outputting the Q-value of each action , where is the network parameter. The forward propagation of the network is represented as:
[0099] ,
[0100] ,
[0101] ,
[0102] where is the output feature vector of the first hidden layer, is the output feature vector of the second hidden layer, , , is the weight matrix, , , is the bias vector. ReLU is the activation function.
[0103] (3) Bellman equation and long-term value evaluation: The training of the deep Q-network is based on the Bellman optimal equation:
[0104] ,
[0105] where represents the optimal action value function of performing action in state , represents the expectation about the next state , is the immediate reward, calculated by the quality prediction model:
[0106] ,
[0107] is the discount factor, which quantifies the weight of long-term impact. The larger the value, the more the system values future returns, i.e., the more attention is paid to the long-term impact of parameters. represents the next state The starting point is the best cumulative return that can be achieved in the future. Thus, the learned Q value is itself an evaluation metric that incorporates long-term impact weights. It not only considers the immediate quality improvement brought by adjusting parameters , but also considers how this adjustment action affects the future state trajectory and cumulative return of the entire system.
[0108] (4) Network training: The mean squared error loss function is used:
[0109] ,
[0110] where represents the parameters of the current Q network, which is the network being trained and updated in real time. represents the target network parameters, which are used to calculate the target Q value in the Bellman equation. represents the expectation with respect to the state transition sequence .
[0111] This training process forces the network to learn to accurately estimate the long-term value of each parameter adjustment action.
[0112] (5) Key parameter selection: According to the learned Q function, select the parameters with the highest value:
[0113] ,
[0114] Finally, output the key parameter index set , and the corresponding key parameter subset . represents the parameter index in the original process parameter vector , and represents the specific value of the th parameter in the original process parameter vector. These parameters are selected because the deep Q network evaluates that adjusting them can bring the greatest long-term quality return to the entire system, not just instantaneous defect improvement.
[0115] Step 4: Defect-guided parameter adaptive control
[0116] This step is to solve the problem that static models are difficult to adapt to dynamic environments. Through reinforcement learning, an accurate mapping from defect features to process parameter adjustment amounts is established, realizing adaptive control based on real-time defect detection. According to the severity and type of defects, accurate adjustment instructions need to be generated for each key parameter. Overall, the input is the defect feature vector , the current key parameters , and the output is the key parameter adjustment amount The implementation is as follows:
[0117] (1) State space construction: combine defect features and current parameter state to construct a regulatory state vector:
[0118] ,
[0119] Splice the defect feature vector from step 1 with the current K key process parameters to form a complete state description. The state vector comprehensively represents the "current product defect features" and "current production settings".
[0120] (2) Intelligent decision-making and action generation: the decision-making core is a deep neural network (Actor network) that maps continuous states to continuous action space:
[0121] ,
[0122] The Actor network receives the state , represents a deterministic policy function (Actor network) with parameters , outputting a K-dimensional action vector normalized in the range [-1, 1]. Each dimension corresponds to the adjustment direction and relative amplitude of a key parameter (-1 indicates adjustment to the lower limit, and 1 indicates adjustment to the upper limit). is an exploration noise artificially added in the training phase to encourage the agent to try different actions to find better policies. It can be removed in deployment applications, represents network parameters.
[0123] (3) Action execution and instruction conversion: the normalized action output by the network must be converted into parameter adjustment instructions with actual physical meaning:
[0124] ,
[0125] By multiplying each element by a preset maximum adjustment amplitude vector , the normalized action is interpreted as the actual parameter change .
[0126] (4) Virtual parameter generation and reward calculation: to evaluate the goodness of the action without interfering with actual production, the system uses a "virtual exploration" mechanism to calculate the reward:
[0127] ,
[0128] First, calculate the "virtual" key parameters assuming this adjustment is performed, i.e. the current parameters plus the recommended adjustment amount Then, through the mapping function Map, it is restored to the complete process parameter vector required by the quality prediction model This step is because the quality prediction model is trained on the complete parameter space, and the complete N-dimensional parameter vector must be provided for the model to give an accurate quality prediction. Only part of the parameters will lead to inaccurate prediction or model error. Then the reward function is calculated:
[0129] ,
[0130] The reward function formula is divided into two parts. First, the quality improvement reward part: , represents the complete process parameter vector at time , represents the predicted output of the quality prediction model at time , i.e. the predicted quality score based on the current defect features and process parameters , represents the adjustment amplitude penalty coefficient, which is a hyperparameter, used to balance the quality improvement and parameter stability; the expected improvement in quality after using the virtual parameters. The value is positive, indicating quality improvement, and negative, indicating quality decline. This is the core incentive for the agent to improve quality. The following is the adjustment amplitude penalty part: , the L2 norm of the adjustment amount is penalized. The coefficient is used to balance "quality improvement" and "control stability" to avoid parameter shock and encourage fine tuning.
[0131] (5) Strategy optimization and network update: the model constantly optimizes its decision-making strategy (Actor network) and evaluation ability (Critic network) through interaction with the environment, i.e. rewards. For Critic network, i.e. value function , the input is state s and action a , and the output is the long-term value estimate of the state-action pair. represent the network parameters. The Critic network (value evaluator) update formula is:
[0132]
[0133] ,
[0134] Critic network The learning objective is to make the predicted Q-value, i.e., the estimate of the long-term value of an action, as accurate as possible. This is achieved by minimizing the temporal difference error (i.e., the difference between the predicted and target values). The target value is achieved by (mean squared error). From instant rewards The target Q-value for the next state (calculated by the target network) is the target Q-value calculated for the i-th sample, which is the target learned by the Critic network. This value is a more stable long-term value estimate calculated by the target network for the i-th sample using the Bellman equation. This is a discount factor that assigns a certain weight to future returns. Where i is the index of the mini-batch sample, such as when calculating the target value... hour The target Q value is calculated based on this sample.
[0135] For Actor network (policy controller) updates:
[0136] ,
[0137] The optimization objective of the Actor network is to maximize the long-term value of its decisions. It utilizes guidance signals provided by the updated Critic network, specifically the gradient of the Q-value with respect to the action. This allows the Actor network to adjust its parameters. The above formula implies that the Actor network updates its strategy in directions that improve the Q-value (i.e., long-term return). This indicates the parameters of the Actor network. The gradient. Indicates the action The gradient. It is the objective function, representing the expected total reward of the policy, for example, in a method. Specifically, the Actor network parameters are The expected cumulative discount return that this strategy can achieve is given by N. N represents the mini-batch size, indicating the number of independent empirical samples used in a single parameter update.
[0138] Next, perform a soft update on the target network:
[0139] ,
[0140] By using tiny weights Slowly track the current network parameters and the target network parameters. Maintaining relative stability greatly improves the stability of the entire training process.
[0141] (6) Final output and application: After training converges, the model learns an optimal policy. When online regulation, the system directly uses the strategy:
[0142] ,
[0143] The system finally realizes the end-to-end accurate mapping from the defect feature to the parameter adjustment amount , and can adaptively adjust the production process according to the different characteristics of the defects, such as type, severity, etc., so as to realize the quality closed-loop intelligent control of the manufacturing process.
[0144] Embodiment two
[0145] The application also provides a defect characteristic guided process optimization system, which is used to realize the method of embodiment one, and the system comprises an extraction module, a construction module, a screening module and a regulation module.
[0146] The extraction module is used to extract real-time defect features of a workpiece image.
[0147] The construction module is used to construct a quality prediction model.
[0148] The screening module is used to screen process parameters based on real-time defect features and a quality prediction model.
[0149] The regulation module is used to establish a mapping from defect features to process parameter adjustment amounts through reinforcement learning, so as to realize adaptive regulation based on real-time defect detection.
[0150] In this embodiment, the extraction module comprises a preprocessing unit, a feature extraction unit, a fusion unit and a dimension reduction unit.
[0151] The preprocessing unit is used to process a workpiece image by using a defect detection model based on deep learning, so as to obtain the bounding box coordinates of defects and defect category labels.
[0152] The feature extraction unit is used to extract deep features, morphological features and texture features of a defect area by using a deep convolutional neural network.
[0153] The fusion unit is used to splice the deep features, morphological features and texture features to obtain fused features.
[0154] The dimension reduction unit is used to perform dimension reduction on the fused features by using principal component analysis, so as to obtain a defect feature vector.
[0155] In this embodiment, the process of screening process parameters based on real-time defect features and a quality prediction model comprises:
[0156] According to the learned Q function, the parameter with the highest value is selected:
[0157] ;
[0158] Output key parameter index set , corresponding key parameter subset ;
[0159] wherein, represents a state, represents an action, is the total number of actions, represents the parameter index in the original process parameter vector , represents the specific value of the th parameter in the original process parameter vector, is the network parameter.
[0160] In this embodiment, the mapping from defect features to process parameter adjustment amounts is established through reinforcement learning, and the process of adaptive regulation based on real-time defect detection comprises:
[0161] ;
[0162] wherein, is the key parameter adjustment amount, s represents a state, represents a network parameter, represents a maximum adjustment amplitude vector, represents an optimal strategy.
[0163] In summary, the beneficial effects of the present application are:
[0164] 1. A new path of "defect feature driven" reverse regulation is proposed: The present application breaks through the traditional "parameter-quality" forward optimization paradigm and creates a technical path of directly guiding "process parameter optimization" from "defect visual features". The quantified defect features are taken as the core input, and precise adjustment instructions are directly generated through an intelligent decision-making model, realizing an end-to-end closed loop from "defect recognition" to "defect elimination", greatly improving the pertinence and intuitiveness of control.
[0165] 2. An efficient decision-making architecture of "pre-filtering and post-optimization" is constructed: The present application establishes a two-stage architecture of "key parameter dynamic screening" and "deep reinforcement learning optimization". The architecture compresses the high-dimensional optimization problem to a very small key parameter subset through front-end screening, avoids global brute force search from the root, enables the back-end optimization to be efficient and accurate, and thus solves the efficiency bottleneck problem of real-time optimization in complex manufacturing systems.
[0166] 3. A long-term value dynamic screening mechanism based on real-time defect features is introduced: In the screening stage, the application innovatively uses real-time defect features as the core decision basis. By introducing the value function in reinforcement learning, the long-term impact weight of each parameter on suppressing the current specific defect is evaluated, thereby completing the screening. This ensures that the key variables screened by the system each time are the optimal solution for the current defect, achieving intelligentization and precision of the screening process.
[0167] Embodiment three
[0168] The parameter adaptive regulation scheme based on defect features and reinforcement learning provided by the application realizes the leap from "diagnosis" to "treatment" in the intelligent manufacturing scene. The following is illustrated by a typical user scenario:
[0169] A high-end electronic board card manufacturer's surface mount (SMT) production line, its optical detection system frequently detects "solder bridge" defects. In the traditional control mode, even if the system can accurately identify and classify the defect, it can only trigger an alarm or perform sorting. Process engineers must manually try to adjust a few parameters that may be related from hundreds of variables such as the temperature setting of the reflow soldering furnace, the solder paste printing parameter, and the mounting precision. This trial-and-error regulation not only responds slowly, but also often leads to "over-regulation" or "under-regulation" due to the difficulty in grasping the precise adjustment amount, and even causes new quality problems, resulting in continuous waste of materials and working hours.
[0170] After applying the application, when the online detection system captures the "solder bridge" defect image again, the system first extracts the quantitative visual features of the defect (such as bridge area, shape, and position distribution) in real time. Then, based on the defect features, the system dynamically analyzes and locks the most relevant key parameter subset to the current bridge mode - in this case, the "peak temperature of the fifth temperature zone of the reflow soldering furnace" and the "printing thickness of the solder paste".
[0171] Subsequently, these key parameters and their current set values, together with the defect features, form a state vector, which is input into the pre-trained reinforcement learning agent (Actor-Critic network). The agent, in a short time, combines the prior knowledge of the quality prediction model and the long-term optimization goal to calculate the optimal adjustment amount for the above two key parameters, and directly generates control instructions: "lower the peak temperature of the fifth temperature zone by 3.5℃, and reduce the solder paste printing thickness by 5μm".
[0172] The production line control system automatically executes the adjustment after receiving the instructions. Then the board produced immediately enters the next round of detection, and the system automatically updates its policy network by comparing the changes in defect features and the improvement in quality prediction scores before and after adjustment, so that it can make faster and more accurate decisions when encountering the same defect next time.
[0173] The above-described embodiments are merely intended to describe the preferred modes of the present application, and are not intended to limit the scope of the present application. Various modifications and improvements of the present application made by those skilled in the art based on the above-described embodiments should fall within the scope of the present application defined by the claims.
Claims
1. A defect-characteristic-guided process optimization method, characterized in that, The method includes: Extract and quantify real-time defect features from workpiece images; Build a quality prediction model; Process parameters are selected based on real-time defect characteristics and quality prediction models; By establishing a mapping from defect features to process parameter adjustment amounts through reinforcement learning, adaptive control based on real-time defect detection can be achieved. Methods for screening process parameters based on real-time defect characteristics and quality prediction models include: Based on the learned Q-function, select the one with the highest value. One parameter: ; Output key parameter index set The corresponding subset of key parameters ; in, Indicates state, Indicates an action, For the total number of actions, This indicates the original process parameter vector. Parameter index in Represents the first element in the original process parameter vector. The specific values of each parameter These are network parameters.
2. The method according to claim 1, characterized in that, Methods for extracting and quantifying real-time defect features from workpiece images include: A deep learning-based defect detection model is used to process workpiece images to obtain the bounding box coordinates and defect category labels of defects; Deep convolutional neural networks are used to extract depth, morphological, and texture features of defect regions. The depth features, morphological features, and texture features are spliced together to obtain the fused features; Principal component analysis was used to reduce the dimensionality of the fused features to obtain the defect feature vector.
3. The method according to claim 1, characterized in that, Methods for adaptive control based on real-time defect detection, which establish a mapping from defect features to process parameter adjustment amounts through reinforcement learning, include: ; in, For key parameter adjustment amounts, s Indicates state, Represents network parameters, Represents the maximum adjustment magnitude vector. This represents the optimal strategy.
4. A defect-characteristic-guided process optimization system, said system being used to implement the method according to any one of claims 1-3, characterized in that, The system includes: an extraction module, a construction module, a filtering module, and a control module; The extraction module is used to extract and quantify real-time defect features of the workpiece image; The building module is used to build a quality prediction model; The filtering module is used to filter process parameters based on real-time defect characteristics and quality prediction models; The control module is used to establish a mapping from defect features to process parameter adjustment amounts through reinforcement learning, thereby achieving adaptive control based on real-time defect detection.
5. The system according to claim 4, characterized in that, The extraction module includes: a preprocessing unit, a feature extraction unit, a fusion unit, and a dimensionality reduction unit; The preprocessing unit is used to process the workpiece image using a deep learning-based defect detection model to obtain the bounding box coordinates of the defect and the defect category label. The feature extraction unit is used to extract the depth features, morphological features and texture features of the defect region using a deep convolutional neural network. The fusion unit is used to stitch together depth features, morphological features and texture features to obtain the fused features; The dimensionality reduction unit is used to reduce the dimensionality of the fused features using principal component analysis to obtain the defect feature vector.
6. The system according to claim 4, characterized in that, The process of screening process parameters based on real-time defect characteristics and quality prediction models includes: Based on the learned Q-function, select the one with the highest value. One parameter: ; Output key parameter index set The corresponding subset of key parameters ; in, Indicates state, Indicates an action, For the total number of actions, This indicates the original process parameter vector. Parameter index in Represents the first element in the original process parameter vector. The specific values of each parameter These are network parameters.
7. The system according to claim 4, characterized in that, The process of establishing a mapping from defect features to process parameter adjustment amounts through reinforcement learning, and realizing adaptive control based on real-time defect detection, includes: ; in, For key parameter adjustment amounts, s Indicates state, Represents network parameters, Represents the maximum adjustment magnitude vector. This represents the optimal strategy.
Citation Information
Patent Citations
Aluminum alloy surface defect detection method and system using deep learning
CN120411061A