Copper plate surface impurity identification and decision control intelligent sampling method
The deep reinforcement learning model enhances copper plate surface defect detection by accurately identifying defects and optimizing sampling paths, improving efficiency and reducing equipment wear.
Patent Information
- Application Number
- CN202510468628.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The surface impurities detection efficiency of traditional copper plates is low, and manual identification is inaccurate. It is easy to sample the impurity area during the sampling process, which affects the detection effect. The sampling at the fixed position of the drill bit leads to inaccurate results, making the optimal path impossible to plan, affecting the sampling efficiency.
The RGB camera is used to collect impurities images of copper plate surfaces, feature extraction and decision-making control are performed through deep reinforcement learning models, attention enhancement modules and reinforcement learning decision networks are used to generate drill bit action space and state value scalars, and the sampling path is optimized.
It realizes efficient and accurate identification of impurities on the surface of copper plate and drill bit sampling, avoids impurity areas, improves sampling efficiency, reduces error detection rate, and reduces equipment wear.
Smart Images

Figure CN120318193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent sampling method for identifying impurities on the surface of copper plates and decision control, belonging to the technical fields of computer vision and metal material quality inspection. Background Art
[0002] During the production and processing of copper plates, surface impurity detection is a key link in quality control. Traditional surface impurity detection is usually manual detection, which has problems such as low impurity identification efficiency and inaccurate manual detection results. At the same time, the traditional sampling process is prone to sampling samples containing impurities under the premise of manual detection, unable to accurately avoid the impurity areas of the copper plates, affecting the subsequent detection effect. At the same time, when performing multi-hole sampling, sampling at a fixed position of the drill bit will cause surface impurities to enter the sample, resulting in inaccurate detection results, and it is also impossible to plan the optimal path, affecting the sampling efficiency. Therefore, how to achieve high-efficiency and high-precision identification of impurities on the surface of copper plates and how to make the drill bit sampling generate a path to avoid impurity positions for intelligent sampling have become an important research direction in the current industrial automation field. Summary of the Invention
[0003] In order to overcome the problems in the background art, the purpose of the present invention is to provide an intelligent sampling method for identifying impurities on the surface of copper plates and decision control.
[0004] In order to achieve the above purpose, the present invention is realized through the following technical solutions:
[0005] An intelligent sampling method for identifying impurities on the surface of copper plates and decision control includes the following steps:
[0006] (1) Using an RGB camera to sample the impurities on the surface of the copper plate to be detected to obtain RGB image data;
[0007] (2) Sequentially performing normalization, downsampling, and noise filtering on the RGB image data to obtain a preprocessed image;
[0008] (3) Establishing a deep reinforcement learning model for identifying impurities on the surface of copper plates and decision control, using the preprocessed image as input, performing feature extraction in a multi-layer neural network, and then using an attention enhancement module and a reinforcement learning decision network to obtain the drill bit action space and the state value scalar V(s); the drill bit action space includes the direction of drill bit movement, the movement step of the drill bit, and the sampling instruction of the drill bit;
[0009] (4) Training the deep reinforcement learning model for identifying impurities on the surface of copper plates and decision control with a data set.
[0010] More preferably, the feature extraction specifically includes: inputting the preprocessed image, and performing feature extraction through an initial convolutional layer and three residual modules to obtain image features.
[0011] The specific layers are as follows:
[0012] The first layer is the initial convolutional layer, with an input shape of (512, 512, 3) and an output shape of (256, 256, 64).
[0013] The second layer is residual block 1, with an input shape of (256, 256, 64) and an output shape of (128, 128, 128).
[0014] The third layer is residual block 2, with an input shape of (128, 128, 128) and an output shape of (64, 64, 256).
[0015] The fourth layer is residual block 3, with an input shape of (64, 64, 256) and an output shape of (32, 32, 512).
[0016] More preferably, the attention enhancement module includes a channel attention mechanism and a spatial attention mechanism; the feature image is passed through the channel attention mechanism to obtain a channel weight vector Mc, and through the spatial attention mechanism to obtain a spatial weight map Ms; the input feature image is weighted by channel features and spatial features based on the channel weight vector Mc and the spatial weight map Ms to obtain an output image.
[0017] More preferably, the reinforcement learning decision network includes a feature fusion layer, a policy network, and a value network; specifically, it includes the following steps:
[0018] S3-1: The output image obtained by the attention enhancement module is input into the feature fusion layer, flattened into a 524288-dimensional vector, and the image features are compressed to 1024 dimensions using a fully connected layer, and are concatenated with a 256-dimensional environmental state vector to form a 1280-dimensional joint feature;
[0019] S3-2: The 1280-dimensional joint feature passes through the policy network to obtain a 128-dimensional drill bit action space. The encoding method is to map the multi-dimensional action combination to 64 decisions, and expand it to 128 dimensions to be compatible with future expansions;
[0020] S3-3: The 1280-dimensional joint feature passes through the value network to obtain a state value scalar V(s);
[0021] S3-4: The policy network and the value network adopt a cooperative training mechanism. The policy network generates a drill bit action space and interacts with the environment to generate experiences; the value network calculates the current state value scalar V(s) and the next state value scalar V(s′), which are used to calculate the advantage function and guide the policy gradient update.
[0022] The drill bit movement directions include 8 directions: up, down, left, right, upper left, lower left, upper right, and lower right. The drill bit movement step size includes 3 gears, and the sampling instructions for the drill bit include start and stop.
[0023] The upper left, lower left, upper right, and lower right are respectively at an angle of 45° to the left and right horizontal directions.
[0024] More preferably, step (4) specifically includes:
[0025] Train the deep reinforcement learning model for copper plate surface impurity recognition and decision control on a dataset. The dataset is 3000 - 5000 groups of copper plate RGB images collected. The training specifically includes: using the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and 100 training epochs.
[0026] Advantages of the present invention: First, the present invention collects impurity images on the surface of the copper plate to be detected through an RGB camera. Compared with manual recognition, the recognition efficiency is higher. At the same time, the present invention designs a deep reinforcement learning model for copper plate surface impurity recognition and decision control. By inputting the RGB image, the model can directly output the drill bit action space (movement direction, movement step, whether to sample) and the state value scalar. During the process of image feature extraction, an attention enhancement module is used to suppress background noise and increase the response weight of the impurity area. The detection accuracy of small impurities is improved through spatial-channel dual attention. At the same time, the policy network and the value network adopt a collaborative training mechanism. The policy network generates actions to interact with the environment to generate experiences, and the value network calculates the current state value and the next state value to guide the policy gradient update, realizing parameter sharing and independent output layers. Through the above model, the present invention realizes the rapid recognition and positioning of impurities on the copper plate surface. At the same time, according to the position of impurities on the copper plate surface, decision control is performed on the drill bit to generate an action space (movement direction, movement step, whether to sample), and value evaluation is carried out simultaneously. The state value continuously guides the policy gradient update to make the action space more accurate. The present invention solves the problem of low efficiency of traditional manual recognition. At the same time, combined with the recognition of the surface impurity position to control the drill bit for uniform sampling at non-impurity points, it effectively prevents problems such as impurity sampling, inaccurate sample detection, etc. At the same time, it continuously optimizes and updates the drill bit sampling path, effectively improving the sampling efficiency, reducing the false detection rate, and reducing the wear degree of the equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic flow chart of the intelligent sampling method for copper plate surface impurity recognition and decision control based on deep reinforcement learning in the present invention.
[0028] Figure 2 It is a schematic structural diagram of the RGB camera in the present invention for collecting the position information of impurities on the copper plate surface and the drill bit sampling.
[0029] Figure 3 It is a schematic structural diagram of the deep reinforcement learning model for copper plate surface impurity recognition and decision control in the present invention. Detailed Implementation Modes
[0030] The present invention will be further described in detail below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto.
[0031] Embodiment 1
[0032] As Figures 1 - 3 , an intelligent sampling method for impurity recognition and decision control on the surface of copper plates based on deep reinforcement learning includes the following steps:
[0033] (1) Acquisition of original image data. As Figure 2 shown in the device, an RGB camera is used to sample the impurities on the surface of the copper plate to be detected in a specified area to obtain an RGB image. The RGB image format is a resolution of 2048×2048, ensuring that the image resolution of RGB meets the requirements of subsequent processing.
[0034] (2) Preprocessing the RGB image obtained in step (1): Normalize, downsample, and filter the noise of the RGB image. The normalized RGB image (2048, 2048, 3) is adjusted to the network input size (512, 512, 3) through downsampling to reduce the amount of calculation, and then noise filtering is performed to eliminate imaging noise through Gaussian filtering or median filtering. The obtained image remains in the format of (512, 512, 3).
[0035] (3) Establish a deep reinforcement learning model for impurity recognition and decision control on the surface of copper plates. Take the image (512, 512, 3) obtained in step (2) as the input, perform feature extraction in a multi-layer neural network, use an attention enhancement module and a reinforcement learning decision network, and finally obtain the drill action space (Direction, Step, Instruction) and the state value scalar (V(s)). Direction is the moving direction of the drill, Step is the moving step of the drill, and Instrution is the sampling instruction of the drill;
[0036] Specifically, it includes:
[0037] S3.1. Feature extraction network: Input the image (512, 512, 3) obtained in step (2), and pass through an initial convolutional layer and three residual modules to obtain features (32, 32, 512). The specific structures of each layer are as follows:
[0038] The first layer is the initial convolutional layer, which includes a convolutional layer. The convolutional kernel size of the convolutional layer is 7×7, the stride is 2, the padding is same, the input shape is (512, 512, 3), and the output shape is (256, 256, 64); the output of the convolutional layer is normalized to make the data distribution more stable, thus alleviating the problems of vanishing gradients and exploding gradients; a ReLU activation function is added after the convolutional layer to introduce non-linearity so that the neural network can learn complex patterns, and at the same time, it can alleviate the problem of vanishing gradients and improve the training effect of the deep network.
[0039] The second layer is Residual Block 1, which includes Convolutional Layer 1, Convolutional Layer 2, and a max pooling layer; the input shape is (256, 256, 64), and the output shape is (128, 128, 128); first is Convolutional Layer 1, which can extract local features of the data; the convolutional kernel size is 3×3, the stride is 1, the padding is same, the input shape is (256, 256, 64), and the output shape is (256, 256, 128). The output of Convolutional Layer 1 is normalized to make the data distribution more stable, thus alleviating the problems of vanishing gradients and exploding gradients. A ReLU activation function is added after the convolutional layer to introduce non-linearity so that the neural network can learn complex patterns, and at the same time, it can alleviate the problem of vanishing gradients and improve the training effect of the deep network.
[0040] Secondly is Convolutional Layer 2, the convolutional kernel size is 3×3, the stride is 1, the padding is same, the input shape is (256, 256, 128), and the output shape is (256, 256, 128). The output of Convolutional Layer 2 is normalized to make the data distribution more stable, thus alleviating the problems of vanishing gradients and exploding gradients. A ReLU activation function is added after Convolutional Layer 2 to introduce non-linearity so that the neural network can learn complex patterns, and at the same time, it can alleviate the problem of vanishing gradients and improve the training effect of the deep network; a skip connection is added between the output of the initial convolutional layer and the output of Convolutional Layer 2 to form a standard residual block; a residual network is introduced to achieve residual learning through skip connections to solve the problems of vanishing gradients and degradation in deep networks; the number of input and output channels is inconsistent, and a 1x1 convolution is used to adjust the input shape; the formula for the standard residual block is: v = F(x, W i ) + u, where u is the input, F(x, W i ) is the output of the two convolutional layers (residual mapping), v is the final output, x is the input variable, and W i are the learnable parameters in the residual function F, and i is the index of the parameters.
[0041] Finally, there is the max pooling layer. The pooling layer can reduce the spatial dimension of the feature map, reduce the computational amount, and prevent overfitting. The pooling kernel size is 2×2, the stride is 2, the input shape is (256, 256, 128), and the output shape is (128, 128, 128).
[0042] The third layer is Residual Block 2, which includes Convolution Layer 1, Convolution Layer 2, and the max pooling layer. The input shape is (128, 128, 128), and the output shape is (64, 64, 256):
[0043] First, there is Convolution Layer 1. The convolution layer can extract local features of the data. The convolution kernel size is 3×3, the stride is 1, the padding is same, the input shape is (128, 128, 128), and the output shape is (128, 128, 256). The output of Convolution Layer 1 is normalized to make the data distribution more stable, thus alleviating the problems of gradient vanishing and gradient explosion. After the convolution layer, the ReLU activation function is added to introduce non-linearity so that the neural network can learn complex patterns and at the same time alleviate the gradient vanishing problem and improve the training effect of the deep network.
[0044] Secondly, there is Convolution Layer 2. The convolution kernel size is 3×3, the stride is 1, the padding is same, the input shape is (128, 128, 256), and the output shape is (128, 128, 256). The output of Convolution Layer 2 is normalized to make the data distribution more stable, thus alleviating the problems of gradient vanishing and gradient explosion. After the convolution layer, the ReLU activation function is added to introduce non-linearity so that the neural network can learn complex patterns and at the same time alleviate the gradient vanishing problem and improve the training effect of the deep network. A skip connection is added between the output of Residual Block 1 and the output of Convolution Layer 2 to form a standard residual block. The residual network is introduced, and residual learning is achieved through skip connections, which can solve the problems of gradient vanishing and degradation in deep networks. Since the number of input and output channels is inconsistent, a 1x1 convolution is used to adjust the input shape. The formula for the standard residual block is: v = F(x, W i ) + u, where u is the input, and F(x, W i ) is the output of the two convolution layers (residual mapping), and v is the final output.
[0045] Finally, there is the max pooling layer. The pooling layer can reduce the spatial dimension of the feature map, reduce the computational amount, and prevent overfitting. The pooling kernel size is 2×2, the stride is 2, the input shape is (128, 128, 256), and the output shape is (64, 64, 256).
[0046] The fourth layer is Residual Block 3, which includes Convolution Layer 1, Convolution Layer 2, and the max pooling layer. The input shape is (64, 64, 256), and the output shape is (32, 32, 512):
[0047] First is the convolutional layer 1. The convolutional layer can extract local features of the data. The size of the convolutional kernel is 3×3, the stride is 1, the padding is same, the input shape is (64, 64, 256), and the output shape is (64, 64, 512). Normalize the output of the convolutional layer to make the data distribution more stable, thus alleviating the problems of vanishing gradients and exploding gradients. Add the ReLU activation function after the convolutional layer to introduce non-linearity, enabling the neural network to learn complex patterns and alleviating the vanishing gradient problem, improving the training effect of the deep network.
[0048] Second is the convolutional layer 2. The size of the convolutional kernel is 3×3, the stride is 1, the padding is same, the input shape is (64, 64, 512), and the output shape is (64, 64, 512). Normalize the output of the convolutional layer to make the data distribution more stable, thus alleviating the problems of vanishing gradients and exploding gradients. Add the ReLU activation function after the convolutional layer to introduce non-linearity, enabling the neural network to learn complex patterns and alleviating the vanishing gradient problem, improving the training effect of the deep network. A skip connection is added between the output of the residual block 2 and the output of the convolutional layer 2 to form a standard residual block. Introduce the residual network. Through the skip connection, residual learning can be achieved, which can solve the problems of vanishing gradients and degradation in deep networks. Since the number of input and output channels is inconsistent, use a 1x1 convolution to adjust the input shape. The formula for the standard residual block is: v = F(x, W i ) + u, where u is the input, and F(x, W i ) is the output of two convolutional layers (residual mapping), and v is the final output.
[0049] Finally is the max pooling layer. The pooling layer can reduce the spatial dimension of the feature map, reduce the amount of calculation, and prevent overfitting. The size of the pooling kernel is 2×2, the stride is 2, the input shape is (64, 64, 512), and the output shape is (32, 32, 512).
[0050] S3.2. Attention enhancement module. The attention enhancement module is divided into channel attention and spatial attention. The (32, 32, 512) obtained from S3.1 passes through the channel attention and spatial attention of the attention enhancement module respectively to obtain the channel weight vector Mc and the spatial weight map Ms, and the input (32, 32, 512) is weighted by channel and spatial features to obtain the output (32, 32, 512). Specifically, it includes:
[0051] S3.2.1. Channel attention mechanism, including one layer of global average pooling layer and two layers of fully connected layers. The input shape is (32, 32, 512), and the output shape is (1, 1, 512). Adjust the channel dimension of the feature map to obtain the channel weight vector Mc. The specific structure of each layer is as follows:
[0052] The first layer is the global average pooling layer with a pooling kernel size of 32×32, a stride of 2, an input shape of (32, 32, 512), and an output shape of (1, 1, 512).
[0053] The second layer is the fully connected layer 1 with a weight shape of (32, 512), a bias of (32), an input shape of (1, 1, 512), and an output shape of (1, 1, 32), which reduces the dimension to extract the non-linear relationship between channels; a ReLU activation function is added to introduce non-linear features.
[0054] The third layer is the fully connected layer 2 with a weight shape of (32, 512), a bias of (32), an input shape of (1, 1, 32), and an output shape of (1, 1, 512), which restores the original channel dimension; a Sigmoid activation function is added to generate a channel weight vector Mc ∈ [0, 1]∧512.
[0055] S3.2.2. Spatial attention mechanism, including a cross-channel max pooling layer, a cross-channel average pooling layer, and a convolutional layer, with an input shape of (32, 32, 512) and an output shape of (32, 32, 1), highlighting the impurity region features and generating a spatial weight map Ms; the specific structures of each layer are as follows:
[0056] The first layer is the cross-channel max pooling layer with a pooling kernel size of 1×1×512, a stride of 1×1×512, an input shape of (32, 32, 512), and an output shape of (32, 32, 1), extracting the maximum response value at each spatial position.
[0057] The second layer is the cross-channel average pooling layer with a pooling kernel size of 1×1×512, a stride of 1×1×512, an input shape of (32, 32, 512), and an output shape of (32, 32, 1), calculating the average response value at each spatial position.
[0058] The outputs of the above two layers are concatenated along the channel dimension, with an output shape of (32×32×2), combining the two spatial statistical features.
[0059] The third layer is the convolutional layer with a convolutional kernel size of 7×7, a stride of 3, padding same, an input shape of (32, 32, 2), and an output shape of (32, 32, 1), generating a spatial weight map Ms ∈ [0, 1]∧(32×32).
[0060] S3.2.3. Feature fusion, which is divided into channel weighting and spatial weighting; specifically:
[0061] Channel weighting: Assign weights to each channel to enhance the contribution of important channels. Broadcast the channel weight vector Mc ∈ [0,1] ∧ 512 to the same spatial dimension as the original feature map (32,32,512), multiply channel by channel, and the output size is (32,32,512). The features of important channels are amplified, and redundant channels are suppressed.
[0062] Spatial weighting: Assign weights to each spatial position to highlight the impurity region. Broadcast the spatial weight map Ms ∈ [0,1] ∧ (32×32) to all channel dimensions, multiply the eigenvalue of each spatial position of the channel-weighted result (32,32,512) by the corresponding spatial weight map Ms ∈ [0,1] ∧ (32×32), and the final output size is (32,32,512). The features of the impurity region are significantly enhanced, and the background is weakened.
[0063] Through the above process, by combining the channel and spatial weighting results, the retention of dual attention information is achieved.
[0064] S3.3. The reinforcement learning decision network includes a feature fusion layer, a policy network, and a value network; specifically including:
[0065] S3.3.1. Feature fusion layer, input the (32,32,512) obtained from S3.2, flatten it into a 524288-dimensional vector, use a fully connected layer to compress the image features from 524288 dimensions to 1024 dimensions, and at the same time concatenate it with a 256-dimensional environmental state vector to form a 1280-dimensional joint feature.
[0066] S3.3.2. Policy network, including three fully connected layers; the 1280-dimensional joint feature obtained from S3.3.1 passes through the policy network to obtain a 128-dimensional drill bit action space. The drill bit action space includes the drill bit movement direction (8 directions: up, down, left, right, upper left, lower left, upper right, lower right), the drill bit movement step size (3 gears: slow, medium, fast), and the sampling instruction of the drill bit (2 instructions: start, stop). The encoding method is to map the multi-dimensional action combination to 64 (8×3×2) kinds of decisions and expand it to 128 dimensions to be compatible with future expansions.
[0067] The specific structure of each layer is as follows:
[0068] The first layer is the fully connected layer 1, the weight shape is (512,1280), the bias is (512), the input shape is (1280), and the output shape is (512).
[0069] The second layer is the fully connected layer 1, the weight shape is (256,512), the bias is (256), the input shape is (512), and the output shape is (256).
[0070] The third layer is the fully-connected layer 1, with a weight shape of (32, 256), a bias of (128), an input shape of (256), and an output shape of (128).
[0071] S3.3.3. Value network, which shares features with the policy network. The 1280-dimensional joint features obtained from S3.3.1 are used by the value network to obtain the current state value scalar V(s), evaluating the long-term benefits of the current state. The specific structure of each layer is as follows:
[0072] The first layer is the fully-connected layer 1, with a weight shape of (256, 1280), a bias of (256), an input shape of (1280), and an output shape of (256).
[0073] The second layer is the fully-connected layer 1, with a weight shape of (64, 256), a bias of (64), an input shape of (128), and an output shape of (64).
[0074] The third layer is the fully-connected layer 1, with a weight shape of (1, 64), a bias of (1), an input shape of (128), and an output shape of (128).
[0075] S3.3.4. Co-training mechanism. The policy network generates the drill bit action space. After the drill bit executes the action and interacts with the environment, rewards (obtained from the reward function) and a new state (new spatial state) are generated. The value network calculates the current state value V(s) and the next state value V(s'), which are used to calculate the advantage function and guide the policy gradient update, achieving parameter sharing and independence. Parameter sharing is reflected in shared feature extraction, reducing computational costs and ensuring consistent understanding of the state by the two networks. Independence is reflected in the independent output layer. The policy network optimizes the action distribution, and the value network focuses on value evaluation. Specifically, it includes:
[0076] (1) Policy network: Generates the drill bit action space (such as moving direction, moving step size, whether to sample) based on the current copper plate surface image state (such as impurity distribution, drill bit position).
[0077] The steps for the drill bit to execute the action and interact with the environment to generate rewards include:
[0078] Step 1. Record the events during the training process
[0079] 1. Whether successfully avoiding impurity sampling: Record whether the impurity area is successfully avoided during each sampling. If successfully avoided, mark it as "Yes", otherwise mark it as "No".
[0080] 2. Sampling point distribution: Check whether the sampling points are evenly distributed. It can be judged by calculating indicators such as the distance between sampling points and the distribution density. If it meets a certain uniform distribution standard (for example, the distance between adjacent sampling points is within a certain range and the distribution coverage meets the expectation), it is marked as "even", otherwise it is marked as "uneven".
[0081] 3. Whether entering the impurity area: Monitor in real time whether the drill bit enters the impurity area. If it enters, it is marked as "yes", otherwise it is marked as "no".
[0082] 4. Whether the path is repeated: Record the action path of the agent, and judge whether there is a repeated path by comparing the front and back paths. If there is a repeated path, it is marked as "yes", otherwise it is marked as "no".
[0083] Step 2. Calculate the reward value of a single action
[0084] The reward function is designed as:
[0085] Step 3. Accumulate the total reward
[0086] Accumulate the reward value r of each action during the training process t to obtain the total reward R at the end of the training:
[0087]
[0088] where T is the total number of actions during the training process.
[0089] R > 0: It indicates that the agent performs well during the training process, can better complete tasks such as avoiding impurity sampling and maintaining uniform distribution of sampling points, and rarely enters the impurity area and has repeated paths. The specific value of R can be further analyzed. The larger the value, the better the performance.
[0090] R = 0: It means that the agent's performance is average, and the completion of each task cancels each other out, without obvious advantages or disadvantages. At this time, it can be further analyzed which task rewards cancel each other out to find the direction for improvement.
[0091] R < 0: It shows that there are many problems with the agent during the training process, such as frequently entering the impurity area, having many repeated paths, or failing to effectively avoid impurity sampling and maintain uniform distribution of sampling points.
[0092] (2) Value network: Scalarize and evaluate the value of the current state (i.e., V(s)) to predict the future cumulative reward (such as the long-term benefit of successfully avoiding impurity sampling).
[0093] The value network guides the policy gradient update by calculating the advantage function, and the formula is:
[0094] A(s,a) = Q(s,a) - V(s)
[0095] Where: A(s,a) is the advantage function, Q(s,a) is the action value function, and it is estimated by the temporal difference (TD) error:
[0096] δ = r + γV(s′) - V(s)
[0097] Where: δ is the temporal difference error, r is the immediate reward, γ is the discount factor, and V(s′) is the next state value.
[0098] S4. Train the model established in S3 with a dataset. The dataset is 3000 groups of copper plate RGB images collected. Use the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and 100 training epochs for training. Through training, the total reward function value R is gradually increased, thereby improving the accuracy of the model.
[0099] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. An intelligent sampling method for identifying impurities and making decision control on the surface of copper plates, characterized in that: It includes the following steps: (1) Use an RGB camera to sample the impurities on the surface of the copper plate to be detected to obtain RGB image data; (2) Normalize, downsample, and filter the noise of the RGB image data in sequence to obtain a preprocessed image; (3) Establish a deep reinforcement learning model for impurity recognition and decision control on the surface of the copper plate. Take the preprocessed image as the input, extract features in a multi-layer neural network, and then use an attention enhancement module and a reinforcement learning decision network to obtain the drill bit action space and the state value scalar V(s); the drill bit action space includes the direction of drill bit movement, the movement step of the drill bit, and the sampling instruction of the drill bit; (4) Train the deep reinforcement learning model for impurity recognition and decision control on the surface of the copper plate with a dataset.
2. The intelligent sampling method for identifying impurities on the copper plate surface and decision control according to claim 1, characterized in that: The feature extraction specifically includes: inputting the preprocessed image, and performing feature extraction through an initial convolutional layer and three residual modules to obtain a feature image.
3. The intelligent sampling method for identifying and decision-making control of impurities on the surface of copper plates according to claim 2, characterized in that: The attention enhancement module includes a channel attention and a spatial attention mechanism; passing the feature image through the channel attention mechanism to obtain a channel weight vector Mc, and passing it through the spatial attention mechanism to obtain a spatial weight map Ms; based on the channel weight vector Mc and the spatial weight map Ms, perform channel feature weighting and spatial feature weighting on the input feature image to obtain an output image.
4. The intelligent sampling method for identifying impurities and making decision control on the surface of copper plates according to claim 3, wherein: The channel attention mechanism includes a layer of global average pooling layer and two layers of fully connected layers.
5. The intelligent sampling method for identifying and decision-making control of impurities on the copper plate surface according to claim 3, characterized in that: The spatial attention mechanism includes a cross-channel maximum pooling layer, a cross-channel average pooling layer, and a convolutional layer.
6. The intelligent sampling method for identifying and decision-making control of impurities on the copper plate surface according to claim 3, characterized in that: The reinforcement learning decision network includes a feature fusion layer, a policy network, and a value network; it specifically includes the following steps: S3-1: Flatten the output image features obtained by the attention enhancement mechanism into a 524288-dimensional vector and input it into the feature fusion layer. Use a fully connected layer to compress the image features to 1024 dimensions, and at the same time concatenate them with a 256-dimensional environmental state vector to form a 1280-dimensional joint feature; S3-2: The 1280-dimensional joint feature passes through the policy network to obtain a 128-dimensional drill bit action space. The encoding method is to map the multi-dimensional action combination to 64 decisions and expand it to 128 dimensions; S3-3: The 1280-dimensional joint feature passes through the value network to obtain the state value scalar V(s); S3-4: The policy network and the value network adopt a cooperative training mechanism. The policy network generates the drill bit action space and interacts with the environment to generate experiences; the value network calculates the current state value scalar V(s) and the next state value scalar V(s′) for calculating the advantage function to guide the policy gradient update.
7. The intelligent sampling method for identifying impurities and making decision control on the surface of copper plates according to claim 1, characterized in that: In the step (4), the dataset is 3000 - 5000 groups of copper plate RGB images collected; the training specifically includes: using the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and a training epoch of 100.