Intelligent PLC processing method based on visual deep reinforcement learning
By integrating the visual depth reinforcement learning module in the PLC control system, real-time monitoring and adjustment of processing parameters, the limitations of traditional PLC control methods in handling complex processing tasks are solved, and processing accuracy and production efficiency are significantly improved, and cost and scrap rate are reduced.
Patent Information
- Application Number
- CN202510011253.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional PLC control methods have limitations in dealing with complex machining tasks, especially in dealing with systematic errors such as concentricity errors in cutting knife processing, which lacks real-time monitoring and adaptive adjustment capabilities, resulting in reduced product quality and inefficient production efficiency.
The intelligent PLC machining method based on visual depth reinforcement learning is adopted, and through the integrated vision module and deep learning algorithm, key parameters in the processing process are monitored and dynamically adjusted in real time, tool motion strategies are optimized, and concentricity errors are reduced.
The processing accuracy is significantly improved, and the concentricity error is reduced from 25μm to less than 2μm, reducing manual intervention, improving production efficiency and safety, reducing labor costs and waste rates, and reducing energy consumption and maintenance costs.
Smart Images

Figure CN120032228A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an intelligent PLC processing method based on visual deep reinforcement learning, and belongs to the technical field of artificial intelligence. Background Art
[0002] In the field of intelligent manufacturing, industrial automation technology is experiencing rapid development, especially in precision machining and automated production environments. Traditional programmable logic controller (PLC) methods have played an important role in automation control, relying on pre-set rules to perform tasks. However, these methods have limitations in handling complex machining tasks, especially in dealing with systematic errors in the machining process, such as concentricity errors in wedge machining.
[0003] Concentricity error will affect the life of the splitter and processing accuracy, resulting in reduced product quality. Secondly, traditional processing methods lack the ability to monitor the processing process in real time and adjust it adaptively, making it difficult to achieve efficient processing. Manual intervention and debugging processes are time-consuming and labor-intensive, reducing production efficiency. In addition, the high labor costs and scrap rates of traditional processing methods have led to high production costs and reduced corporate competitiveness.
[0004] In recent years, the development of deep reinforcement learning algorithms has provided new possibilities for solving this problem. Deep reinforcement learning algorithms allow optimal strategies to be learned through interaction with the environment, thereby achieving optimized control in a dynamic environment. In addition, advances in computer vision technology, especially the application of the OpenCV library, have made image processing and feature extraction more efficient and accurate, which is essential for achieving precise processing control.
[0005] In the field of industrial automation, the Siemens S7-1200 series PLC system is a widely used platform that supports data communication with external computing systems through Python code embedding, such as using the Modbus protocol for data transmission. This provides a technical basis for integrating deep reinforcement learning models, enabling the PLC system to make intelligent decisions based on real-time data.
[0006] Although existing intelligent manufacturing technologies have achieved certain results in improving production efficiency and machining accuracy, there is still room for improvement in dealing with complex machining tasks and systematic errors. This method deeply integrates visual perception, deep reinforcement learning and PLC control technology, monitors and dynamically adjusts errors in the machining process in real time, adaptively adjusts machining parameters and strategies, reduces manual intervention, reduces production costs, improves production efficiency and safety performance, and provides a new solution for the field of CNC machining. Summary of the invention
[0007] The purpose of the present invention is to address the concentricity error problem in the hard material processing process in the field of precision manufacturing, and propose an intelligent PLC processing method based on deep reinforcement learning. This method integrates advanced visual modules and deep learning algorithms to achieve real-time monitoring and dynamic adjustment of key parameters in the processing process, thereby maximizing the accuracy and efficiency of the processing process. The proposed solution can solve the limitations of traditional PLC control methods in handling complex processing tasks, especially the real-time adjustment capability of systematic errors, and can significantly improve processing quality, reduce tool wear, and improve production efficiency.
[0008] The technical method adopted by the present invention to solve the technical problem is: an intelligent PLC processing method based on visual deep reinforcement learning, which includes the following steps:
[0009] Step 1: Build a visual module;
[0010] Step 1-1: Select and install a camera that can capture the target image;
[0011] Step 1-2: Image preprocessing to extract relevant information;
[0012] Step 1-3: Use image processing algorithms to extract features and calculate errors;
[0013] Step 2: Train the deep reinforcement learning model;
[0014] Step 2-1: Define state and action space;
[0015] Step 2-2: Design a reward mechanism and optimize behavioral strategies;
[0016] Step 2-3: Build a deep neural network;
[0017] Step 2-4: Train the model;
[0018] Step 3: Integrate intelligent PLC control module;
[0019] Step 3-1: Select the PLC system and embed Python code;
[0020] Step 3-2: Use PID control algorithm to close the loop output.
[0021] Furthermore, step 1 of the present invention includes: constructing a visual module, and the specific calculation steps are:
[0022] Step 1-1: Select an industrial camera with high resolution and high frame rate to ensure that it can clearly capture the images of the workpiece and the tool. Install the camera in a suitable position, such as above or on the side of the machine tool, to ensure that it can cover the entire processing area.
[0023] Step 1-2: Convert the color image to a grayscale image to reduce the amount of calculation, using the formula:
[0024] I gray =0.2989R+0.5870G+0.1140B
[0025] Where R, G and B represent the red, green and blue channel values of the color image respectively;
[0026] Then, to remove image noise and improve image quality, a Gaussian filter is used to perform convolution on the image to smooth the image and remove noise. The formula of the Gaussian filter is:
[0027]
[0028] Where x and y are the coordinates of the pixel in the image, which are used to determine the value of the Gaussian function at that point, and σ is the standard deviation of the Gaussian kernel, which determines the width of the Gaussian function, that is, the degree of blurring of the filter;
[0029] Convert the grayscale image to a black and white image for subsequent processing. Use the threshold T to convert the grayscale image to a binary image, where pixels with values greater than T are set to white, otherwise they are set to black. The selection of the threshold requires calculating the cumulative number of pixels for each grayscale value first:
[0030]
[0031] Where h(j) is the number of pixels with gray value j;
[0032] Then for each possible threshold, the between-class variance between foreground and background is calculated:
[0033]
[0034] where w 0 (t) and w 1 (t) are the weights (number of pixels) of the background below and foreground above the threshold t, μ 0 (t) and μ 1 (t) are the average grayscale values of the background and foreground, respectively, and w(t) is the total number of pixels;
[0035] Finally, the optimal threshold t is obtained * :
[0036]
[0037] In order to extract edge information in the image for identifying the position of the tool and workpiece, the gradient strength and direction of each pixel in the image need to be calculated;
[0038] The gradient strength can be calculated by the following formula;
[0039]
[0040] Among them, I x and I y are the gradients of the image in the x and y directions respectively;
[0041] The gradient direction can be calculated by the following formula:
[0042]
[0043] For each pixel, check whether its gradient strength is the local maximum in its gradient direction. If not, set the gradient strength of the point to 0. This step can eliminate the "burrs" and "double edges" of the edge.
[0044] Using two thresholds T low and T high To determine the edge, if the gradient strength of a pixel is greater than T high , it is considered an edge point; if it is less than T low , it is considered not an edge point; if low and T high , it is considered an edge point only when it is connected to an edge point with a gradient strength greater than Thigh:
[0045]
[0046] Among them, connected means that the point is connected to a point with a gradient strength greater than T high The edge points are connected, T low and T high The choice of depends on the characteristics of the image and the required edge detection effect;
[0047] Step 1-3: Use image processing algorithms to extract features such as tool position, angle, and workpiece concentricity deviation, and use contour detection algorithms to identify the contours of the tool and workpiece;
[0048] The contour detection algorithm can be implemented using the findContours function in the OpenCV library, and its geometric parameters such as area, distance, and angle are calculated for error feedback;
[0049] Calculate the area of the contour using the formula:
[0050]
[0051] Among them, (x 1 ,y 1 ),(x 2 ,y 2 ),...,(x n,y n ) are the coordinates of the points on the contour;
[0052] And use the formula The distance between two points is calculated, and these geometric parameters will be used as input to the deep reinforcement learning module to evaluate the machining error.
[0053] Furthermore, the step 2 of the present invention includes: training a deep reinforcement learning model, including:
[0054] Step 2-1: Define the state vector s, which will be used as the input of the deep reinforcement learning module to evaluate the processing state:
[0055] s=[x,y,θ,v,f]
[0056] Where x and y are the position coordinates of the tool, θ is the angle of the tool, v is the speed of the tool, f is the feed rate, and w is the material property;
[0057] Define the action space a:
[0058] a=[Δx,Δy,Δθ,Δv]
[0059] Where Δx and Δy are the changes in position, Δθ is the change in angle, Δv is the change in velocity, Δf is the change in feed rate, and the action space defines all possible actions that the agent can take;
[0060] Step 2-2: Design a reward mechanism. When the error e is less than the preset threshold T, a positive reward R is given. + The formula is:
[0061] R + =k 1 (Te)
[0062] When the error e is less than the preset threshold T, a negative reward R - The formula is:
[0063] R - =-k 2 (eT)
[0064] When the knife can maintain the ideal angle for a long time, give continuous reward R C , stability can be defined as the time t that the tool maintains the ideal angle, and its formula is:
[0065] R C =k 3 t
[0066] Among them, k1 is the positive reward coefficient, k2 is the negative reward coefficient, and k3 is the continuous reward coefficient, which is used to adjust the size of the reward value;
[0067] Therefore, the total reward R can be expressed as:
[0068] R=R + +R - +R C
[0069] Step 2-3: Build a deep neural network. For the input layer, receive the state vector and use the fully connected layer to map the state vector to the hidden layer. The formula of the fully connected layer is y=Wx+b, where W is the weight matrix, b is the bias vector, x is the input vector, and y is the output vector.
[0070] For the hidden layer, the ReLU activation function f(x)=max(0,x) is used to limit the output of the hidden layer to non-negative values, improve the nonlinear ability of the model, and accelerate the training process of the neural network and improve the generalization ability of the model. Finally, the action probability distribution is output, and the softmax function is used to convert the output of the hidden layer into a probability distribution. The formula of the softmax function is:
[0071]
[0072] where z a is the logit value of action a, and n is the dimension of the action space;
[0073] Step 2-4: Collect training data to prepare the training model, including state, action, reward, and next state. Use a simulation environment or actual processing process to collect data, store experience in an experience pool, and randomly sample data from it for training, using uniform sampling or priority sampling.
[0074] The experience replay mechanism can break the correlation between data, improve the generalization ability of the model, and use priority experience replay to improve training efficiency, where the priority of each experience is determined by its TD error (Temporal Difference Error):
[0075]
[0076] Where r is the immediate reward from state s to state s′, γ is the time discount factor, which usually takes a value between 0 and 1.
[0077] Q(s,a) is the Q value of taking action a in the current state s. is the maximum Q value of all possible actions in the next state s′, so the priority of each experience is expressed as:
[0078] Priority=|δ+ε
[0079] Where ∈ is a small positive number used to avoid priority 0;
[0080] In the experience replay pool, a combination of uniform sampling and priority sampling is used to select training data to balance exploration and utilization;
[0081]
[0082] Where α is the priority weight, and N is the size of the experience replay pool;
[0083] The Adam optimizer is used to update the model parameters. The optimizer can adaptively adjust the learning rate to improve the training efficiency. The formula of the gradient descent algorithm is:
[0084]
[0085] Where θ is the model parameter, α is the learning rate, and L(θ) is the loss function;
[0086] For the loss function, we define it as:
[0087]
[0088] where y i is the true value, is the predicted value.
[0089] Furthermore, the step 3 of the present invention includes: integrating an intelligent PLC control module, including:
[0090] Step 3-1: Select the Siemens S7-1200 series, which has powerful computing power and rich communication interfaces. The PLC system is the core of the intelligent control module, responsible for receiving control instructions from the deep reinforcement learning module and controlling the movement of the machine tool. Embedding Python code can be implemented through the programming interface of the PLC system, such as using the TIA Portal software of the Siemens S7-1200 series for programming;
[0091] Step 3-2: Define the error function e(t) = target value - current value, which is used to measure the deviation between the current processing state and the desired target. The target value is the ideal position or angle of the tool, and the current value is the actual position or angle of the tool. The output u(t) of the PID controller is calculated by the following formula:
[0092]
[0093] Among them, K p The proportional coefficient directly affects the size of the error. The larger the error, the stronger the control effect. i The integral coefficient is proportional to the accumulated time of the error and is used to eliminate the steady-state error; Kd The differential coefficient is proportional to the rate of change of the error and is used to predict the future trend of the error and make adjustments in advance.
[0094] Beneficial effects:
[0095] 1. The present invention optimizes the tool motion strategy and achieves a breakthrough in significantly reducing the concentricity error from 25μm to less than 2μm. Further real machine experiments verified the high-precision capability of the model. After 500 rounds of training, the concentricity error was stabilized at 1.5μm. This achievement is unprecedented in the existing technology and significantly improves the processing accuracy.
[0096] 2. In terms of high-efficiency machining, the present invention significantly reduces manual intervention through automated control, thereby improving machining efficiency. On actual equipment, the model can quickly adapt and adjust the tool motion path and parameters, so that the concentricity error is gradually reduced. Experimental data show that the response speed and adjustment efficiency of the model in actual application are much higher than traditional methods, showing good performance and reliability.
[0097] 3. By implementing the present invention, enterprises can achieve 20%-30% savings in labor costs, reduce scrap rates by more than 30%, and reduce energy consumption by 10%-20%. At the same time, maintenance costs are also significantly reduced due to the low maintenance requirements and long service life of automated equipment. These comprehensive benefits significantly improve production efficiency and bring significant economic advantages to enterprises.
[0098] 4. The method of the present invention improves production safety through real-time monitoring and early warning functions. Through visual feedback and deep reinforcement learning, the system can predict and avoid potential processing errors. The safety test results show that the present invention has obvious advantages in reducing safety accidents during processing.
[0099] 5. The present invention is an intelligent PLC processing method based on deep reinforcement learning, which aims to improve processing accuracy and efficiency, and effectively solve the concentricity error problem in chopping knife processing, so as to realize the intelligentization, automation and efficiency of the processing process. The present invention combines advanced visual perception technology, deep reinforcement learning algorithm and PLC control technology, and provides a new solution for the field of CNC processing, which has broad application prospects and important technical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 The figure is an overall block diagram of the method proposed in the present invention.
[0101] Figure 2 Schematic diagram of the vision module that preprocesses the camera-captured image to accurately display the tool position.
[0102] Figure 3 A diagram showing the fluctuating total reward value as the model learns and optimizes the tool motion strategy.
[0103] Figure 4 A schematic diagram showing the changes in the model's centralization loss during training. DETAILED DESCRIPTION
[0104] The invention is further described in detail below in conjunction with the accompanying drawings.
[0105] like Figure 1 As shown, the present invention provides an intelligent PLC processing method based on visual deep reinforcement learning, which includes the following steps.
[0106] Step 1: Build a visual module.
[0107] Step 1.1, select a high-resolution, high-frame-rate industrial camera to ensure that the images of the workpiece and tool can be captured clearly. Install the camera in a suitable position, such as above or on the side of the machine tool, to ensure that the entire processing area can be covered;
[0108] Step 1.2, convert the color image to grayscale image to reduce the amount of calculation. Use the formula:
[0109] I gray =0.2989R+0.5870G+0.1140B
[0110] Where R, G, and B represent the red, green, and blue channel values of the color image, respectively.
[0111] Next, remove the image noise and improve the image quality. Use the Gaussian filter to perform convolution on the image to smooth the image and remove the noise. The formula of the Gaussian filter is:
[0112]
[0113] Where x and y are the coordinates of the pixel in the image, which are used to determine the value of the Gaussian function at that point. σ is the standard deviation of the Gaussian kernel, which determines the width of the Gaussian function, that is, the blurriness of the filter.
[0114] Convert grayscale images to black and white images for subsequent processing. Use threshold T to convert grayscale images into binary images, where pixels with values greater than T are set to white, otherwise set to black. The selection of the threshold requires calculating the cumulative number of pixels for each grayscale value:
[0115]
[0116] where h(j) is the number of pixels with gray value j.
[0117] Then for each possible threshold, the between-class variance between foreground and background is calculated:
[0118]
[0119] where w 0 (t) and w 1 (t) are the weights (number of pixels) of the background below and foreground above the threshold t, respectively. 0 (t) and μ 1 (t) are the average grayscale values of the background and foreground, respectively, and w(t) is the total number of pixels.
[0120] Finally, the optimal threshold t is obtained * :
[0121]
[0122] In order to extract edge information in the image and identify the position of the tool and workpiece, the gradient strength and direction of each pixel in the image must be calculated.
[0123] The gradient strength can be calculated by the following formula;
[0124]
[0125] Among them, I x and I y are the gradients of the image in the x and y directions, respectively.
[0126] The gradient direction can be calculated by the following formula:
[0127]
[0128] For each pixel, check whether its gradient strength is the local maximum in its gradient direction. If not, set the gradient strength of the point to 0. This step can eliminate the "burrs" and "double edges" of the edge.
[0129] Using two thresholds T low and T high To determine the edge. If the gradient strength of a pixel is greater than T high , it is considered an edge point; if it is less than T low , it is considered not an edge point; if low and T high If it is connected to an edge point with a gradient strength greater than Thigh, it is considered an edge point.
[0130]
[0131] Among them, connected means that the point is connected to a point with a gradient strength greater than T high The edge points are connected. low and T high The choice depends on the characteristics of the image and the desired edge detection effect.
[0132] Step 1.3, use image processing algorithms to extract features such as tool position, angle, and workpiece concentricity deviation. Use contour detection algorithms to identify the contours of the tool and workpiece.
[0133] The contour detection algorithm can be implemented using the findContours function in the OpenCV library and calculate its geometric parameters such as area, distance, and angle for error feedback.
[0134] Calculate the area of the contour using the formula:
[0135]
[0136] Among them, (x 1 ,y 1 ),(x 2 ,y 2 ),...,(x n ,y n ) are the coordinates of the points on the contour.
[0137] And use the formula Calculate the distance between two points. These geometric parameters will be used as input to the deep reinforcement learning module to evaluate the machining error.
[0138] Step 2: Train the deep reinforcement learning model.
[0139] Step 2.1, define the state vector s, which will be used as the input of the deep reinforcement learning module to evaluate the processing state:
[0140] s=[x,y,θ,v,f]
[0141] Where x and y are the position coordinates of the tool, θ is the angle of the tool, v is the speed of the tool, f is the feed rate, and w is the material property.
[0142] Define the action space a:
[0143] a=[Δx,Δy,Δθ,Δv]
[0144] Where Δx and Δy are the changes in position, Δθ is the change in angle, Δv is the change in velocity, and Δf is the change in feed rate. The action space defines all possible actions that the agent can take.
[0145] Step 2.2, design a reward mechanism. When the error e is less than the preset threshold T, a positive reward R is given. + The formula is:
[0146] R + =k 1 (Te)
[0147] When the error e is less than the preset threshold T, a negative reward R - The formula is:
[0148] R - =-k 2 (eT)
[0149] When the knife can maintain the ideal angle for a long time, give continuous reward R C , stability can be defined as the time t that the tool maintains the ideal angle, and its formula is:
[0150] R C =k 3 t
[0151] Among them, k1 is the positive reward coefficient, k2 is the negative reward coefficient, and k3 is the continuous reward coefficient, which is used to adjust the size of the reward value.
[0152] Therefore, the total reward R can be expressed as:
[0153] R=R + +R - +R C
[0154] Step 2.3, build a deep neural network. For the input layer, receive the state vector and use the fully connected layer to map the state vector to the hidden layer. The formula of the fully connected layer is y=Wx+b, where W is the weight matrix, b is the bias vector, x is the input vector, and y is the output vector.
[0155] For the hidden layer, the ReLU activation function f(x)=max(0,x) is used to limit the output of the hidden layer to non-negative values, thereby improving the nonlinear ability of the model. The ReLU activation function can accelerate the training process of the neural network and improve the generalization ability of the model. Finally, the action probability distribution is output, and the softmax function is used to convert the output of the hidden layer into a probability distribution. The formula of the softmax function is:
[0156]
[0157] where z a is the logit value of action a, and n is the dimension of the action space.
[0158] Step 2.4, collect training data to prepare the training model, including state, action, reward and next state, and use simulation environment or actual processing to collect data. Store the experience in the experience pool and randomly sample data from it for training, using uniform sampling or priority sampling.
[0159] The experience replay mechanism can break the correlation between data and improve the generalization ability of the model. Use priority experience replay to improve training efficiency, where the priority of each experience is determined by its TD error (Temporal Difference Error):
[0160]
[0161] Where r is the immediate reward from state s to state s′. γ is the time discount factor, which usually takes a value between 0 and 1.
[0162] Q(s,a) is the Q value of taking action a in the current state s. is the maximum Q value of all possible actions in the next state s′, so the priority of each experience is expressed as:
[0163] Priority=|δ|+ε
[0164] where ∈ is a small positive number used to avoid priority 0.
[0165] In the experience replay pool, a combination of uniform sampling and priority sampling is used to select training data to balance exploration and exploitation.
[0166]
[0167] Where α is the priority weight and N is the size of the experience replay pool.
[0168] The Adam optimizer is used to update the model parameters. The optimizer can adaptively adjust the learning rate to improve the training efficiency. The formula of the gradient descent algorithm is:
[0169]
[0170] Where θ is the model parameter, α is the learning rate, and L(θ) is the loss function.
[0171] For the loss function, we define it as:
[0172]
[0173] where y i is the true value, is the predicted value.
[0174] Step 3: Integrate the intelligent PLC control module.
[0175] Step 3.1, select the Siemens S7-1200 series, which has powerful computing power and rich communication interfaces. The PLC system is the core of the intelligent control module, responsible for receiving control instructions from the deep reinforcement learning module and controlling the movement of the machine tool. Embedding Python code can be implemented through the programming interface of the PLC system, such as programming using the TIA Portal software of the Siemens S7-1200 series.
[0176] Step 3.2, define the error function e(t) = target value - current value, which is used to measure the deviation between the current processing state and the desired target. The target value is the ideal position or angle of the tool, and the current value is the actual position or angle of the tool. The output u(t) of the PID controller is calculated by the following formula:
[0177]
[0178] Among them, K p The proportional coefficient directly affects the size of the error. The larger the error, the stronger the control effect. i The integral coefficient is proportional to the accumulated time of the error and is used to eliminate the steady-state error; K d The differential coefficient is proportional to the rate of change of the error and is used to predict the future trend of the error and make adjustments in advance.
[0179] The performance effect of the present invention can be further illustrated by the following simulation, which specifically includes the following:
[0180] 1. Simulation hardware conditions
[0181] The simulation experiment of the present invention is carried out on the simulation platform of Python 3.8 and Pytorch 11.7. The computer CPU model is E5-2680 v4, the number is 5, and the GPU model is NVIDIA Geforce RTX 3080. The video memory is 12.6GB.
[0182] 2. Simulation conditions
[0183] The image resolution is 128*128. The tool motion path, angle, speed, and concentricity error are simulated in the simulation environment. 50 rounds of training are performed to optimize the model performance and reduce the concentricity error.
[0184] 3. Simulation content
[0185] Figure 2The image shows the workpiece and tool captured by the vision module. The black dots indicate the position of the tool, while the gray background represents the processing area. This image is part of the vision module, which is used to monitor and extract key features such as tool position, angle, and workpiece concentricity deviation in real time. The accurate extraction of these features is the key to achieving high-precision machining. Through visual feedback, the model can adjust the tool motion path and parameters in real time, reduce manual intervention, and improve production efficiency.
[0186] Figure 3 The key performance indicators of the present invention are demonstrated, recording how the model gradually improves its strategy to reduce concentricity error after 50 rounds of training in a simulated environment. The fluctuations and improvements in the reward values reflect the model's progress in learning and optimizing the tool motion strategy to reduce concentricity error. The reward values in the figure reach a peak around the 10th and 20th rounds, which not only indicates that the model has found an effective strategy at these stages, but also reflects that the model can significantly reduce the concentricity error to within 2μm driven by the deep reinforcement learning algorithm. This achievement reflects the innovation of the present invention in improving machining accuracy, and also demonstrates its potential to reduce tool wear and improve production efficiency in practical applications.
[0187] Figure 4 The change of the centralization loss of the model in the present invention during the training process is shown, which is closely related to the error feedback mechanism in the deep reinforcement learning model training process mentioned in the present invention. The downward trend of the loss value shows that the model has made significant progress in reducing the concentricity error, which is consistent with the method mentioned in the present invention of adaptively adjusting the processing parameters and strategies by real-time monitoring and dynamic adjustment of errors in the processing process. The continuous reduction of this loss reflects the technical advantage of this method in improving processing accuracy, that is, significantly reducing tool wear and improving production efficiency through automated control, while reducing maintenance costs and energy consumption, bringing significant economic advantages to enterprises.
[0188] Based on the above simulation results and analysis, the PLC processing method based on visual deep reinforcement learning proposed in the present invention realizes the precise control of concentricity error in the CNC machining process by integrating visual perception technology and deep reinforcement learning algorithm. Its advantage is that it significantly improves the machining accuracy, reducing the error from 25μm to within 2μm, and reduces manual intervention through automated control, thereby improving production efficiency and safety, while reducing labor costs and scrap rates, reducing energy consumption and maintenance costs, and bringing significant economic advantages to enterprises.
Claims
1. An intelligent PLC processing method based on visual deep reinforcement learning, characterized in that: The method comprises the following steps: Step 1: Build a visual module; Step 1-1: Select and install a camera that can capture the target image; Step 1-2: Image preprocessing to extract relevant information; Step 1-3: Use image processing algorithms to extract features and calculate errors; Step 2: Train the deep reinforcement learning model; Step 2-1: Define state and action space; Step 2-2: Design a reward mechanism and optimize behavioral strategies; Step 2-3: Build a deep neural network; Step 2-4: Train the model; Step 3: Integrate intelligent PLC control module; Step 3-1: Select the PLC system and embed Python code; Step 3-2: Use PID control algorithm to close the loop output.
2. According to claim 1, an intelligent PLC processing method based on deep reinforcement learning is characterized in that: The step 1 includes: constructing a visual module, and the specific calculation steps are: Step 1-1: Select an industrial camera with high resolution and high frame rate to ensure that it can clearly capture the images of the workpiece and the tool. Install the camera in a suitable position, such as above or on the side of the machine tool, to ensure that it can cover the entire processing area. Step 1-2: Convert the color image to a grayscale image to reduce the amount of calculation, using the formula: I gray =0.2989R+0.5870G+0.1140B Where R, G and B represent the red, green and blue channel values of the color image respectively; Then, to remove image noise and improve image quality, a Gaussian filter is used to perform convolution on the image to smooth the image and remove noise. The formula of the Gaussian filter is: Where x and y are the coordinates of the pixel in the image, which are used to determine the value of the Gaussian function at that point, and σ is the standard deviation of the Gaussian kernel, which determines the width of the Gaussian function, that is, the degree of blurring of the filter; Convert the grayscale image to a black and white image for subsequent processing. Use the threshold T to convert the grayscale image to a binary image, where pixels with values greater than T are set to white, otherwise they are set to black. The selection of the threshold requires calculating the cumulative number of pixels for each grayscale value first: Where h(j) is the number of pixels with gray value j; Then for each possible threshold, the between-class variance between foreground and background is calculated: Where w0(t) and w1(t) are the weights (number of pixels) of the background below and foreground above the threshold t, μ0(t) and μ1(t) are the average grayscale values of the background and foreground, respectively, and w(t) is the total number of pixels; Finally, the optimal threshold t is obtained * : In order to extract edge information in the image for identifying the position of the tool and workpiece, the gradient strength and direction of each pixel in the image need to be calculated; The gradient strength can be calculated by the following formula; Among them, I x and I y are the gradients of the image in the x and y directions respectively; The gradient direction can be calculated by the following formula: For each pixel, check whether its gradient strength is the local maximum in its gradient direction. If not, set the gradient strength of the point to 0. This step can eliminate the "burrs" and "double edges" of the edge. Using two thresholds T low and T high To determine the edge, if the gradient strength of a pixel is greater than T high , it is considered an edge point; if it is less than T low , it is considered not an edge point; if low and T high , it is considered an edge point only when it is connected to an edge point with a gradient strength greater than Thigh: Among them, connected means that the point is connected to a point with a gradient strength greater than T high The edge points are connected, T low and T high The choice of depends on the characteristics of the image and the required edge detection effect; Step 1-3: Use image processing algorithms to extract features such as tool position, angle, and workpiece concentricity deviation, and use contour detection algorithms to identify the contours of the tool and workpiece; The contour detection algorithm can be implemented using the findContours function in the OpenCV library, and its geometric parameters such as area, distance, and angle are calculated for error feedback; Calculate the area of the contour using the formula: Among them, (x1,y1),(x2,y2),...,(x n ,y n ) are the coordinates of the points on the contour; And use the formula The distance between two points is calculated, and these geometric parameters will be used as input to the deep reinforcement learning module to evaluate the machining error.
3. According to the intelligent PLC processing method based on deep reinforcement learning described in claim 1, it is characterized in that: The step 2 includes: training a deep reinforcement learning model, including: Step 2-1: Define the state vector s, which will be used as the input of the deep reinforcement learning module to evaluate the processing state: s=[x,y,θ,v,f] Where x and y are the position coordinates of the tool, θ is the angle of the tool, v is the speed of the tool, f is the feed rate, and w is the material property; Define the action space a: a=[Δx,Δy,Δθ,Δv] Where Δx and Δy are the changes in position, Δθ is the change in angle, Δv is the change in velocity, Δf is the change in feed rate, and the action space defines all possible actions that the agent can take; Step 2-2: Design a reward mechanism. When the error e is less than the preset threshold T, a positive reward R is given. + The formula is: R + =k1(T-e) When the error e is less than the preset threshold T, a negative reward R - The formula is: R - =-k2(e-T) When the knife can maintain the ideal angle for a long time, give continuous reward R C , stability can be defined as the time t that the tool maintains the ideal angle, and its formula is: R C =k3t Among them, k1 is the positive reward coefficient, k2 is the negative reward coefficient, and k3 is the continuous reward coefficient, which is used to adjust the size of the reward value; Therefore, the total reward R can be expressed as: R=R + +R - +R C Step 2-3: Build a deep neural network. For the input layer, receive the state vector and use the fully connected layer to map the state vector to the hidden layer. The formula of the fully connected layer is y=Wx+b, where W is the weight matrix, b is the bias vector, x is the input vector, and y is the output vector. For the hidden layer, the ReLU activation function f(x)=max(0,x) is used to limit the output of the hidden layer to non-negative values, improve the nonlinear ability of the model, and accelerate the training process of the neural network and improve the generalization ability of the model. Finally, the action probability distribution is output, and the softmax function is used to convert the output of the hidden layer into a probability distribution. The formula of the softmax function is: where z a is the logit value of action a, and n is the dimension of the action space; Step 2-4: Collect training data to prepare the training model, including state, action, reward, and next state. Use a simulation environment or actual processing process to collect data, store experience in an experience pool, and randomly sample data from it for training, using uniform sampling or priority sampling. The experience replay mechanism can break the correlation between data, improve the generalization ability of the model, and use priority experience replay to improve training efficiency, where the priority of each experience is determined by its TD error (Temporal Difference Error): Where r is the immediate reward from state s to state s′, γ is the time discount factor, which usually takes a value between 0 and 1. Q(s,a) is the Q value of taking action a in the current state s. is the maximum Q value of all possible actions in the next state s′, so the priority of each experience is expressed as: Priority=|δ+ε Where ∈ is a small positive number used to avoid priority 0; In the experience replay pool, a combination of uniform sampling and priority sampling is used to select training data to balance exploration and utilization; Where α is the priority weight, and N is the size of the experience replay pool; The Adam optimizer is used to update the model parameters. The optimizer can adaptively adjust the learning rate to improve the training efficiency. The formula of the gradient descent algorithm is: Where θ is the model parameter, α is the learning rate, and L(θ) is the loss function; For the loss function, we define it as: where y i is the true value, is the predicted value.
4. The intelligent PLC processing method based on deep reinforcement learning described in claim 1 is characterized in that: The step 3 includes: integrating an intelligent PLC control module, including: Step 3-1: Select the Siemens S7-1200 series, which has powerful computing power and rich communication interfaces. The PLC system is the core of the intelligent control module, responsible for receiving control instructions from the deep reinforcement learning module and controlling the movement of the machine tool. Embedding Python code can be implemented through the programming interface of the PLC system, such as using the TIA Portal software of the Siemens S7-1200 series for programming; Step 3-2: Define the error function e(t) = target value - current value, which is used to measure the deviation between the current processing state and the desired target. The target value is the ideal position or angle of the tool, and the current value is the actual position or angle of the tool. The output u(t) of the PID controller is calculated by the following formula: Among them, K p The proportional coefficient directly affects the size of the error. The larger the error, the stronger the control effect. i The integral coefficient is proportional to the accumulated time of the error and is used to eliminate the steady-state error; K d The differential coefficient is proportional to the rate of change of the error and is used to predict the future trend of the error and make adjustments in advance.