Intelligent production line CCD defect detection method based on reinforcement learning
By applying a smart detection method based on reinforcement learning in CCD defect detection, combined with Dreamer and MAP-Elites algorithm, the problems of poor adaptability and low detection efficiency of traditional methods in complex environments are solved, and high-precision, adaptive and efficient defect detection are achieved.
Patent Information
- Application Number
- CN202510668946.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional CCD defect detection methods have problems such as poor adaptability, low detection accuracy and low efficiency in complex production environments, especially when facing diverse defect types and changing production conditions.
Using the intelligent production line CCD defect detection method based on reinforcement learning, combined with reinforcement learning, image processing technology and multimodal data fusion, the combination of Dreamer algorithm model and MAP-Elites algorithm is used to construct an adaptive detection strategy and optimize the detection process.
It improves the adaptability, detection accuracy and efficiency of the detection system, significantly reduces manual intervention, and ensures the consistency and reliability of product quality.
Smart Images

Figure CN120198424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision defect detection, and particularly to an intelligent production line CCD defect detection method based on reinforcement learning. Background Art
[0002] With the rapid development of industrial automation, intelligent manufacturing plays an increasingly important role in improving production efficiency and product quality. In modern production lines, defect detection is a key link to ensure product quality, especially in industries with high-precision requirements, such as electronic products, automotive parts, and high-end machinery manufacturing. Traditional defect detection methods usually rely on manual inspection or rule-based automated detection methods. However, these methods usually have relatively serious limitations, especially in terms of detection accuracy, efficiency, and adaptability, and are difficult to meet the requirements of modern production lines.
[0003] Currently, defect detection technology based on CCD (Charge Coupled Device) cameras has been widely applied to various production lines. CCD cameras capture image information on the surface of an object through image acquisition technology, and then extract features through image processing algorithms to judge defects. However, traditional CCD defect detection methods rely on fixed image processing algorithms and preset rules, such as edge detection, texture analysis, color distribution, etc. Although these methods can identify defects to a certain extent, they are often affected by various factors in practical applications, such as changes in lighting, differences in the surface materials of objects, and the complexity of defect types, which makes it difficult to guarantee the accuracy and stability of detection results.
[0004] In addition, existing detection methods usually do not have sufficient flexibility and adaptability. In the actual production process, changes in the production environment and products lead to the complexity of detection tasks. For example, there may be slight surface differences between different batches of products, and it is very difficult for traditional algorithms to accurately identify them in this case. At the same time, existing algorithms often rely on manually set rules and thresholds, lacking the ability to effectively learn and adapt to new types of defects. The lack of flexible and intelligent detection methods has led to an increase in manual intervention, which not only reduces production efficiency but also affects the overall quality of products.
[0005] With the rapid development of artificial intelligence technologies, especially deep learning and reinforcement learning, more and more research has begun to apply these advanced technologies to the field of defect detection. Deep learning technology, by establishing deep neural networks, can learn high-level feature representations from a large amount of data and thus performs excellently in image recognition tasks. However, despite the remarkable achievements of deep learning technology in image classification and detection, its application in the field of defect detection still faces some problems. First, deep learning models usually require a large amount of labeled data for training, and the training process has high requirements for computing resources. Second, the training of deep learning models requires a large amount of labeled data, which may not be easily obtained in some production environments. In addition, deep learning algorithms are usually a static training process, lacking adaptability and continuous optimization capabilities, which may limit their application in actual production lines for a constantly changing production environment.
[0006] In contrast, reinforcement learning, as an intelligent decision-making algorithm, can learn the optimal strategy through interaction with the environment, which gives it greater advantages in dealing with complex dynamic systems. Reinforcement learning enables the agent to optimize its behavior through continuous trial and error through a reward feedback mechanism. In the field of defect detection, the advantage of reinforcement learning is that it can adaptively adjust the detection strategy according to different detection tasks and production environments, improving the accuracy and flexibility of the system. For example, reinforcement learning can continuously optimize the detection strategy based on historical data, dynamically adapt to different defect types and complex environments, thereby improving the detection efficiency and accuracy.
[0007] However, the existing defect detection methods based on reinforcement learning still have certain defects. First, although reinforcement learning has strong adaptability, it still faces the problems of low model training efficiency and scarce sample data in practical applications. Traditional reinforcement learning algorithms require a large number of trial-and-error processes, which may lead to an overly long training process for the real-time detection requirements of the production line, thus affecting real-time performance. Second, reinforcement learning still relies on manually designed features in feature extraction, such as the edges and textures of images, and the manual design of these features often fails to fully capture the complex information in the images, especially in complex production environments, making it difficult to ensure that the detection system can adapt to various different types of defects.
[0008] Therefore, how to provide an intelligent production line CCD defect detection method based on reinforcement learning is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0009] An object of the present invention is to propose an intelligent production line CCD defect detection method based on reinforcement learning. The present invention makes full use of reinforcement learning, image processing technology, and multi-modal data fusion, and details an intelligent defect detection algorithm optimized based on reinforcement learning. By adopting a method combining reinforcement learning with CCD image features, the present invention realizes the automatic optimization of defect detection tasks in complex production environments and can continuously adjust detection strategies according to changes in the environment and products. The method of the present invention has the advantages of high self-adaptability, high detection accuracy, and strong flexibility, can significantly improve the detection efficiency of the production line, reduce manual intervention, and effectively improve product quality.
[0010] An intelligent production line CCD defect detection method based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1. Collect image data through a CCD camera and perform preprocessing; S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of the image, and encode the features to form an image feature matrix; S3. Build a Dreamer algorithm model, perform sequence modeling on the image feature matrix, predict the potential state trajectory and the corresponding reward sequence, and generate an initial policy sample set; S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set; S5. Construct a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and perform defect determination on the potential state trajectory according to the policy action, and output the defect classification result; S6. Collect the difference information between the defect classification result and the true defect annotation result, and update the potential state network parameters and reward prediction network parameters in the Dreamer algorithm model; S7. Repredict the potential state trajectory and the corresponding reward sequence as a new round of initial policy sample set, and repeat steps S4-S6 until the policy performance meets the preset determination condition, and output the CCD defect detection result.
[0011] Optionally, the image data includes resolution, color, spatial information, temporal information, noise information, and annotation information.
[0012] Optionally, the preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.
[0013] Optionally, the S2 specifically includes: S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image: ; ; Among them, represents the edge intensity of the image at the coordinate point , and respectively represent the horizontal and vertical gradients of the image at , represents the pixel value of the image at ; S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image: ; Among them, represents the texture intensity of the image at the coordinate point , is the weighting coefficient with a window size of , is the pixel value of the image at ; S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image: ; Among them, represents the spatial feature of the image at the coordinate point , represents the number of neighborhood points to be calculated, is the coordinate of the neighborhood points in the image, represents the pixel value of the image at the neighborhood point ; S24. Combine the edge features , texture features and spatial features to form an image feature matrix .
[0014] Optionally, the image feature matrix is formed by weighted combination of the edge features , texture features and spatial features to form a matrix that comprehensively represents the image features: ; Among them, represents the comprehensive feature value of the image at the coordinate point , , and are the weighted coefficients of the edge feature, texture feature, and spatial feature respectively, , and are the eigenvalue of the edge, texture, and spatial feature respectively.
[0015] Optionally, the S3 specifically includes: S31. Construct a Dreamer algorithm model, and perform sequence modeling on the image feature matrix to calculate the latent state trajectory at time step : ; where represents the latent state at time step , represents the latent state at time step , is the image feature matrix at time step , is the Hadamard product, representing element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters , represents the high-order feature extraction function of the latent state, and is the hyperparameter of the Dreamer algorithm model; S32. Based on the latent state trajectory and the action taken at the current moment, calculate the reward sequence at time step : ; where represents the reward sequence at time step , represents the action taken at time step , is the reward prediction function of the Dreamer algorithm model, represents the parameter of the reward prediction function, is the weight coefficient, representing the influence degree of the latent state on the current reward prediction, is the image feature matrix at time step , is the hyperparameter for controlling the long-term influence of the state sequence, represents the compensation coefficient related to the latent state, used to adjust the influence of the past state on the current reward, represents the policy function based on the latent state , and respectively represent the impacts of weighted image features and policy features is the maximum number of steps of the time series, representing the time span of the whole process; S33. Generate an initial policy sample set through the latent state trajectory and the reward sequence: ; Among them, represents the initial policy sample set, including the time steps from 1 to all latent states at the moment of , the corresponding actions and the predicted rewards .
[0016] Optionally, the Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision module. The latent state network generates a latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts future rewards based on the latent state and action inputs, combines the historical image feature sequence, and feeds back to the learning process for optimization. The sequence modeling module uses the edge features, texture features, and spatial features of the image to fuse multi-dimensional feature information into the Dreamer algorithm model to ensure the generalization ability of the model in complex production environments. The state-action decision module generates the corresponding action and the predicted reward , combines the policy selection under each latent state, and generates the corresponding action through a policy optimization method to achieve defect determination and action execution.
[0017] Optionally, the specific steps of S4 include: S41. Use the MAP-Elites algorithm to construct a feature space and map the initial policy sample set to each sub-region in the feature space. The feature space is defined by the following formula: ; Among them, represents the th sub-region, represents the latent state at the time step , is the action taken at the time step , represents the reward sequence at the time step , is the All potential states in the sub-region, is the action set corresponding to the th sub-region, is the reward sequence set of the th sub-region; S42. Evaluate the policy samples in each sub-region and calculate the performance value of the policy samples: ; wherein, represents the reward sequence of the policy sample at the th moment in the sub-region , represents the number of policy samples corresponding to the reward sequence set in the th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps in the time series, representing the time span of the whole process; S43. Select the policy samples with the best performance from each sub-region and generate the optimal policy set: ; wherein, is the optimal policy set, containing the policy samples with the best performance in all sub-regions, is the number of sub-regions in the feature space, represents selecting the policy sample with the longest reward sequence in the sub-region.
[0018] Optionally, the MAP-Elites algorithm is an optimization method based on feature space partitioning. The optimization method maps the policy sample set to each sub-region in the feature space, and by evaluating the performance values of the policy samples in each sub-region, selects the policy samples with the best performance in each sub-region, thereby generating the optimal policy set and using it for further policy optimization and iteration.
[0019] Optionally, the S5 specifically includes: S51. Construct a parallel training environment according to the optimal policy set and asynchronously evaluate and update each policy sample using the A3C algorithm ; wherein, represents the updated potential state, is the step size factor related to the time step , represents the potential state at the time step , Denote the parameters of the Dreamer algorithm model The constructed non - linear mapping function is the time step The image feature matrix at the moment is the maximum number of steps of the time series, representing the time span of the whole process; S52. Update the policy value for each policy sample, and the policy value is updated by the following formula: ; where is the policy value of the latent state , is the reward signal is the discount factor is the learning rate is the latent state at the next time step; S53. Determine the defects of the latent state trajectory according to the policy action and output the defect classification result: ; where represents the defect classification result is the weight coefficient related to the category and the feature , represents the number of features represents the number of categories is the normalization factor of the image feature matrix is the weighted sum of each category represents selecting the category corresponding to the maximum value as the defect classification result
[0020] The beneficial effects of the present invention are as follows: First of all, by combining reinforcement learning and CCD image processing technology, the present invention solves the problem of poor adaptability of traditional defect detection methods in complex production environments. Traditional methods rely on fixed rules and manually set features, and cannot effectively cope with diverse defect types and changing production conditions. However, through the reinforcement learning algorithm, the defect detection system of the present invention can adaptively adjust the detection strategy, automatically optimize the detection process according to different production batches and environmental conditions, thereby improving the robustness of the system in dynamic environments.
[0021] Secondly, by introducing the combination of the Dreamer algorithm model and the MAP-Elites algorithm, the defect detection process of the present invention not only has strong intelligent capabilities but also higher learning efficiency. Through the reinforcement learning model for sequence modeling and policy optimization of the image feature matrix, different types of defects can be quickly and accurately identified, while avoiding the problems of heavy dependence on manual adjustment and excessive rule setting in traditional methods. This method enables the system to complete the learning and optimization of new defect types in the shortest time, improving the detection accuracy and efficiency.
[0022] Finally, the present invention combines reinforcement learning with image processing, which can effectively reduce manual intervention and significantly improve the detection speed while ensuring the detection accuracy. Traditional defect detection methods often require manual monitoring and intervention, resulting in reduced production efficiency. Through the automated defect detection method of the present invention, efficient and accurate full-automated defect classification and judgment can be achieved, further improving the working efficiency of the production line, reducing human errors, and ensuring the consistency and reliability of product quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention; Figure 2 is a schematic diagram of the construction of the image feature matrix and the generation of the initial policy sample set of an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention; Figure 3 is a schematic diagram of policy evaluation optimization and defect determination of an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0025] Refer to Figures 1 - 3 , an intelligent production line CCD defect detection method based on reinforcement learning, includes the following steps: S1. Collect image data through a CCD camera and perform preprocessing; S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of the image, and encode the features to form an image feature matrix; S3. Build a Dreamer algorithm model to perform sequence modeling on the image feature matrix, predict the latent state trajectory and the corresponding reward sequence, and generate an initial policy sample set; S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set; S5. Construct a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and determine defects for the latent state trajectory according to the policy actions, and output the defect classification results; S6. Collect the difference information between the defect classification results and the true defect annotation results, and update the latent state network parameters and reward prediction network parameters in the Dreamer algorithm model; S7. Re-predict the latent state trajectory and the corresponding reward sequence as a new round of initial policy sample set, repeat steps S4 - S6 until the policy performance meets the preset determination conditions, and output the CCD defect detection results.
[0026] In the present invention, by using a CCD camera to collect image data and perform preprocessing, the accuracy and real-time performance of defect detection can be effectively improved. By extracting edge, texture, and spatial features from the image data and combining the training and optimization of the reinforcement learning model, the present invention can adaptively adjust the detection strategy according to the changes in the actual production environment, so as to ensure the detection effect under different working conditions. This method significantly improves the intelligent level of the production line, reduces manual intervention, and improves production efficiency.
[0027] In this embodiment, the image data includes resolution, color, spatial information, temporal information, noise information, and annotation information.
[0028] The present invention makes full use of the multi-dimensional information of the image data. In addition to the conventional resolution and color, it also combines spatial information, temporal information, noise information, and annotation information. By integrating these data, the defect detection method of the present invention can more accurately capture the key features in the image, improve the accuracy and adaptability of the detection method, especially in a complex production environment.
[0029] In this embodiment, the preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.
[0030] In the preprocessing stage, the present invention performs grayscale conversion, edge enhancement, noise suppression, and size normalization on the image, thereby ensuring that the image data input into the subsequent algorithms has higher quality and consistency. This process effectively removes noise and irrelevant interference signals, improves the effect of image feature extraction, and provides more accurate data support for subsequent reinforcement learning and defect determination.
[0031] In this embodiment, S2 specifically includes: S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image: ; ; Wherein, represents the edge intensity of the image at the coordinate point , and respectively represent the gradients of the image in the horizontal and vertical directions at , represents the pixel value of the image at ; S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image: ; Wherein, represents the texture intensity of the image at the coordinate point , is the weighting coefficient with a window size of , is the pixel value of the image at ; S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image: ; Wherein, represents the spatial feature of the image at the coordinate point , represents the number of neighborhood points to be calculated, is the coordinate of the neighborhood point in the image, represents the pixel value of the image at the neighborhood point ; S24. Combine the edge feature , texture feature and spatial feature to form an image feature matrix .
[0032] The feature extraction method of the present invention further improves the recognition ability of subtle differences in images by performing edge detection, texture analysis, and spatial feature analysis on the images. By independently calculating and weighted combining different features of each frame of the image, an image feature matrix is constructed. This integration method of multi-dimensional features effectively enhances the expression ability of image information, making defect detection more precise and accurate.
[0033] In this embodiment, the image feature matrix is formed by weighted combining the edge feature , the texture feature and the spatial feature ; wherein, represents the comprehensive feature value of the image at the coordinate point , , and are the weighted coefficients of the edge feature, the texture feature, and the spatial feature respectively, , and are the feature values of the edge, texture, and spatial features respectively.
[0034] By weighted combining the image feature matrix, the present invention forms a matrix that comprehensively represents the image features, further improving the effectiveness of feature extraction. The weighted coefficients are dynamically adjusted according to the importance of different features, so that the important information in the image can be more prominently represented, thereby improving the accuracy of defect detection, especially in the case of complex images and large background noise.
[0035] In this embodiment, S3 specifically includes: S31. Construct a Dreamer algorithm model, and perform sequence modeling on the image feature matrix to calculate the potential state trajectory at time step : ; wherein, represents the potential state at time step , represents the potential state at time step , is the image feature matrix at time step , is the Hadamard product, indicating element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters . A high-order feature extraction function representing a latent state, and is a hyperparameter of the Dreamer algorithm model; S32. Calculate the reward sequence at time step based on the latent state trajectory and the action taken at the current moment: ; Among them, represents the reward sequence at time step , represents the action taken at time step , is the reward prediction function of the Dreamer algorithm model, represents the parameters of the reward prediction function, is the weight coefficient, indicating the influence degree of the latent state on the current reward prediction, is the image feature matrix at time step , is a hyperparameter used to control the long-term influence of the state sequence, represents the compensation coefficient related to the latent state, used to adjust the influence of past states on the current reward, represents the policy function based on the latent state , and respectively represent the influence of the weighted image features and policy features, is the maximum number of steps of the time series, representing the time span of the whole process; S33. Generate an initial policy sample set through the latent state trajectory and the reward sequence: ; Among them, represents the initial policy sample set, including all latent states at time steps from 1 to , the corresponding actions and the predicted rewards . .
[0036] By constructing the Dreamer algorithm model, the present invention performs sequence modeling on the image feature matrix, can capture the latent state trajectory in the image, and predict the corresponding reward sequence. Through the process of reinforcement learning, the model can optimize the policy in continuous feedback and learning, thereby improving the accuracy of defect detection and the adaptive ability of the system. This method has strong dynamic adjustment ability and can quickly learn and adapt to new defect types in the actual production process.
[0037] In this embodiment, the Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision-making module. The latent state network generates a latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts future rewards based on the latent state and action inputs, combines the historical image feature sequence, and feeds back to the learning process for optimization. The sequence modeling module utilizes the edge features, texture features, and spatial features of the image to fuse multi-dimensional feature information into the Dreamer algorithm model to ensure the generalization ability of the model in complex production environments. The state-action decision-making module generates corresponding actions through a policy optimization method according to the current latent state and the predicted reward , and combines the policy selection under each latent state to generate corresponding actions to achieve defect determination and action execution.
[0038] The present invention improves the adaptive ability and intelligent level of the defect detection method through the Dreamer algorithm model. The combination of the latent state network, the reward prediction network, the sequence modeling module, and the state-action decision-making module realizes the dynamic generation of the latent state trajectory and the real-time optimization of the strategy. The fusion of image features ensures accurate detection in complex environments. The state-action decision-making module generates the optimal detection action according to the latent state and the predicted reward, improving the detection accuracy and automation degree.
[0039] In this embodiment, S4 specifically includes: S41. Construct a feature space using the MAP-Elites algorithm and map the initial policy sample set to each sub-region in the feature space. The feature space is defined by the following formula: ; where represents the th sub-region, represents the latent state at time step , is the action taken at time step , represents the reward sequence at time step , is all the latent states in the th sub-region, is the action set corresponding to the th sub-region, is the reward sequence set of the th sub-region; S42. Evaluate the policy samples in each sub-region and calculate the performance value of the policy samples: ; wherein, represents the reward sequence of the policy sample at the -th moment in the sub-region , represents the number of policy samples corresponding to the reward sequence set in the -th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps of the time series, representing the time span of the whole process; S43. Select the policy samples with the best performance from each sub-region and generate an optimal policy set: ; wherein, is the optimal policy set, containing the policy samples with the best performance in all sub-regions, is the number of sub-regions of the feature space, represents selecting the policy sample with the longest reward sequence in the sub-region.
[0040] In the present invention, by introducing the MAP-Elites algorithm, the initial policy sample set is mapped to different sub-regions in the feature space, and by evaluating the performance of the policy samples in each sub-region, the policy samples with the best performance are selected. This optimization process not only improves the detection accuracy of the model, but also enhances its generalization ability in different production environments.
[0041] In this embodiment, the MAP-Elites algorithm is an optimization method based on the partition of the feature space. The optimization method maps the policy sample set to each sub-region in the feature space, and by evaluating the performance value of the policy samples in each sub-region, selects the policy samples with the best performance in each sub-region, so as to generate an optimal policy set and use it for further policy optimization and iteration.
[0042] In the present invention, the feature space is optimized and partitioned by the MAP-Elites algorithm, and the policy samples with the best performance are selected in each sub-region. This method can identify the most effective policies in the complex feature space, improve the detection ability and learning efficiency of the model, and at the same time reduce the computational burden in the training process, further improving the working efficiency of the production line.
[0043] In this embodiment, the specific steps of S5 include: S51. According to the optimal policy set Build a parallel training environment and use the A3C algorithm to asynchronously evaluate and update each policy sample ; Among them, represents the updated latent state, is the step size factor related to the time step ; represents the latent state at the time step ; represents the non - linear mapping function constructed by the Dreamer algorithm model parameters ; is the image feature matrix at the time step ; is the maximum number of steps in the time series, representing the time span of the whole process; S52. Update the policy value for each policy sample. The policy value is updated through the following formula: ; Among them, is the policy value of the latent state ; is the reward signal, is the discount factor, is the learning rate, is the latent state of the next time step; S53. Determine the defects of the latent state trajectory according to the policy action and output the defect classification result: ; Among them, represents the defect classification result, is the weight coefficient related to the category and the feature ; represents the number of features, represents the number of categories, is the normalization factor of the image feature matrix, is the weighted sum of each category, represents selecting the category corresponding to the maximum value as the defect classification result.
[0044] By building a parallel training environment and asynchronously evaluating and updating each policy based on the A3C algorithm, the training efficiency and real - time performance of the model are significantly improved. Through this asynchronous training method, the policy can be updated efficiently, improving the adaptability and decision - making ability of the method of the present invention, and ensuring high - efficiency and accurate defect detection performance in a rapidly changing production environment.
[0045] Example 1: To verify the feasibility of the present invention in implementation, the present invention is applied to the CCD defect detection task on a certain intelligent manufacturing production line. This production line is mainly used for manufacturing the casings of electronic products. During the manufacturing process, various minor defects such as scratches, stains, and bubbles may appear on the surface of the casings of electronic products, and these defects will directly affect the appearance and performance of the products. Traditional defect detection methods mainly rely on manual inspection or simple image processing techniques, and cannot effectively meet the requirements of high-precision and high-speed automated production. Especially when faced with changes in the production environment and the complexity of the product surface, the detection accuracy and efficiency of traditional methods are inadequate. To improve the efficiency and accuracy of defect detection, the present invention proposes an intelligent CCD defect detection method based on reinforcement learning. By extracting, fusing, and optimizing the multi-dimensional features of images, and combining reinforcement learning for adaptive detection, many problems of traditional methods in intelligent production are successfully solved.
[0046] In this scenario application, first, a CCD camera is used to collect real-time images of each electronic casing on the production line. Each collected image will first undergo preprocessing steps such as grayscale conversion, edge enhancement, and noise suppression to ensure that the input data has high quality and consistency. Next, image processing algorithms are used to extract the edge, texture, and spatial features of each frame of the image. These feature information are encoded as a feature matrix and input into the Dreamer algorithm model based on reinforcement learning for further analysis.
[0047] In the Dreamer algorithm model, the feature matrix of each image frame is analyzed through sequence modeling, the potential state trajectory is calculated, and the probability of defect occurrence is predicted. Through the MAP-Elites algorithm, these policy samples are mapped to each sub-region in the feature space, and the corresponding detection strategy is selected according to the optimal strategy in each sub-region. Finally, parallel training is performed based on the A3C algorithm, and the policy is updated after each training to ensure that the model can adaptively adjust the detection scheme to cope with defects of different types and in different environments.
[0048] To ensure the effectiveness of the present invention, before detection, the system was fully trained using approximately 5000 sample images. These sample images include various types of defects such as minor scratches, bubbles, stains, etc., as well as normal casing images. During the training process, the system can gradually improve the detection accuracy by continuously adjusting parameters and optimizing strategies. After training, the system can perform real-time detection on each electronic casing on the production line and automatically reject unqualified products according to the detection results.
[0049] In practical applications, we monitored the detection process and collected data. During the detection process, the system can process approximately 25 frames of images per second, and the detection speed is about 30 times higher than that of traditional manual detection. For each frame of image, the system can complete the defect determination within 2 seconds. We compared the results of manual detection and automatic system detection and found that the accuracy of automatic detection reached 95%, the false detection rate was 5%, and the missed detection rate was 3%. In contrast, the accuracy of manual detection was only 85%, the false detection rate was 10%, and the missed detection rate was 8%. From the data, it can be seen that automated detection not only improves the detection efficiency but also significantly improves the detection accuracy and reduces the interference of human factors.
[0050] In addition, when the system faces different production environments and product changes, it can quickly adapt and optimize the detection strategy. For example, in different batches of products, the surface material, color, and lighting conditions may vary. Traditional detection methods require re-adjusting parameters or setting new rules. However, the detection method based on reinforcement learning of the present invention can dynamically adjust the detection strategy according to the image features of each batch of products, avoiding the need for manual intervention while ensuring high precision.
[0051] Through the application of this embodiment, the effectiveness of the present invention in the actual production environment is verified. Especially in improving the detection accuracy, production efficiency, and reducing manual intervention, remarkable results have been achieved. The present invention has not only been widely applied in existing production lines but also provides an innovative solution for future intelligent manufacturing systems, with broad application prospects.
[0052] Table 1 Comparison data of the performance of the defect detection system ; From the above data table, we can see that the automatic detection system of the present invention shows excellent detection effects in each batch. Compared with manual detection, the accuracy has been significantly improved, and the false detection rate and missed detection rate have been significantly reduced, demonstrating the superiority of this technology in practical applications.
[0053] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. An intelligent production line CCD defect detection method based on reinforcement learning, characterized in that It includes the following steps: S1. Collect image data through a CCD camera and perform preprocessing; S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of the image, and encode the features to form an image feature matrix; S3. Build a Dreamer algorithm model, perform sequence modeling on the image feature matrix, and predict the latent state trajectory and the corresponding reward sequence to generate an initial policy sample set; S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set; S5. Construct a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and determine defects on the latent state trajectory according to the policy actions, and output the defect classification result; S6. Collect the difference information between the defect classification result and the true defect annotation result, and update the latent state network parameters and reward prediction network parameters in the Dreamer algorithm model; S7. Re-predict the latent state trajectory and the corresponding reward sequence as a new round of initial policy sample set, repeat steps S4 - S6 until the policy performance meets the preset determination condition, and output the CCD defect detection result.
2. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, wherein The image data includes resolution, color, spatial information, temporal information, noise information, and annotation information.
3. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, wherein The preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.
4. A method for defect detection of CCD in an intelligent production line based on reinforcement learning according to claim 1, characterized in that, The specific content of S2 includes: S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image; ; ; Among them, represents the edge intensity of the image at the coordinate point ; and respectively represent the gradients of the image in the horizontal and vertical directions at ; represents the pixel value of the image at . S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image; ; Among them, represents the texture intensity of the image at the coordinate point , is the weighting coefficient with a window size of , is the pixel value of the image at . S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image; ; Among them, represents the spatial feature of the image at the coordinate point ; represents the number of neighborhood points to be calculated, are the coordinates of the neighborhood points in the image, represents the pixel value of the image at the neighborhood point ; S24. Combine the edge feature , texture feature and spatial feature to form an image feature matrix .
5. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 4, wherein, The described image feature matrix By performing weighted combination on edge features , texture features and spatial features to form a matrix that comprehensively represents image features: ; Among them, represents the comprehensive eigenvalue of the image at the coordinate point . , and are the weighting coefficients of the edge feature, texture feature, and spatial feature respectively, , and are the eigenvalues of the edge, texture, and spatial features respectively.
6. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, characterized in that The specific content of S3 includes: S31. Build a Dreamer algorithm model and perform sequence modeling on the image feature matrix to calculate the latent state trajectory of the time step : ; Among them, represents the latent state at time step moment, represents the latent state at time step moment, is the image feature matrix at time step moment, is the Hadamard product, representing element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters ; represents the high-order feature extraction function of the latent state, and is the hyperparameter of the Dreamer algorithm model; S32. Calculate the reward sequence at the time step based on the latent state trajectory and the action taken at the current moment at the moment: ; Among them, represents the reward sequence at time step moment, represents the action taken at time step moment, is the reward prediction function of the Dreamer algorithm model, represents the parameters of the reward prediction function, is the weight coefficient, indicating the influence degree of the latent state on the current reward prediction, is the image feature matrix at time step moment, is a hyperparameter used to control the long-term influence of the state sequence, represents the compensation coefficient related to the latent state, used to adjust the influence of past states on the current reward, represents the policy function based on the latent state ; and respectively represent the influence of the weighted image features and policy features, is the maximum number of steps of the time series, representing the time span of the whole process; S33. Generate an initial policy sample set through the latent state trajectory and the reward sequence; ; Among them, represents the initial policy sample set, including time steps from 1 to all potential states at the moment , corresponding actions and predicted rewards .
7. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 6, wherein, The Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision module. The latent state network generates a latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts future rewards based on latent states and action inputs, combines historical image feature sequences, and feeds back to the learning process for optimization. The sequence modeling module utilizes the edge features, texture features, and spatial features of images to fuse multi-dimensional feature information into the Dreamer algorithm model, ensuring the generalization ability of the model in complex production environments. The state-action decision module generates corresponding actions through a policy optimization method according to the current latent state and predicted rewards , in combination with the policy selection under each latent state, to achieve defect determination and action execution.
8. A method for defect detection of CCD in an intelligent production line based on reinforcement learning according to claim 1, characterized in that, The specific content of S4 includes: S41. Use the MAP-Elites algorithm to construct a feature space and map the initial policy sample set to each sub-region in the feature space, where the feature space is defined by the following formula: ; Among them, represents the th sub-region, represents the potential state at time step moment, is the action taken at time step moment, represents the reward sequence at time step moment, is the th all potential states in the sub-region, is the action set corresponding to the th sub-region, is the reward sequence set of the th sub-region; S42. Evaluate the policy samples in each sub-region and calculate the performance value of the policy samples; ; Among them, represents the reward sequence of the policy sample at the th moment within the sub-region, represents the number of policy samples corresponding to the reward sequence set in the th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps of the time series, representing the time span of the whole process; S43. Select the optimal policy sample from each sub-region and generate an optimal policy set; ; Among them, is the set of optimal strategies, which contains the optimal strategy samples in all sub-regions. is the number of sub-regions in the feature space. represents the reward sequence in the selected sub-region with the longest strategy sample.
9. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 8, characterized in that, The MAP-Elites algorithm is an optimization method based on feature space partitioning. The optimization method maps the policy sample set to each sub-region in the feature space, and by evaluating the performance value of the policy samples in each sub-region, selects the optimal policy sample in each sub-region, thereby generating an optimal policy set for further policy optimization and iteration.
10. A method for defect detection of CCD in an intelligent production line based on reinforcement learning according to claim 1, characterized in that, The specific content of S5 includes: S51. According to the optimal policy set Construct a parallel training environment, and use the A3C algorithm to asynchronously evaluate and update each policy sample ; Among them, represents the updated potential state, is the step size factor related to the time step and represents the step size factor related to the time step represents the potential state at the time step moment, represents the non - linear mapping function constructed by the Dreamer algorithm model parameters and represents the non - linear mapping function constructed by the Dreamer algorithm model parameters is the image feature matrix at the time step moment, is the maximum number of steps in the time series, representing the time span of the whole process; S52. Update the policy value of each policy sample, and the policy value is updated through the following formula: ; Among them, is the policy value in the latent state , is the reward signal, is the discount factor, is the learning rate, is the latent state at the next time step; S53. Defect determination is performed on the potential state trajectory according to the policy action, and the defect classification result is output: ; Among them, represents the defect classification result, is the weight coefficient related to the category and the feature , represents the number of features, represents the number of categories, is the normalization factor of the image feature matrix, is the weighted sum of each category, represents selecting the category corresponding to the maximum value as the defect classification result.
Citation Information
Patent Citations
Weld joint surface defect intelligent detection method based on deep reinforcement learning
CN116205877A
Typical defect detection method and system for target object in industrial video image
CN118470013A
Precise component surface detection method and system based on laser parallel control
CN119260164A