An intelligent production line CCD defect detection method based on reinforcement learning

By combining reinforcement learning and CCD image features, using Dreamer, MAP-Elites and A3C algorithms to optimize defect detection, the problem of poor adaptability in traditional methods is solved, efficient and accurate automated defect detection is achieved, and the degree of intelligence of the production line is improved.

CN120198424BActive Publication Date: 2025-07-22KUNSHAN ZHONGLIANXIN PRECISION MACHINERY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510668946.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-22
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing CCD defect detection methods lack flexibility and adaptability in complex production environments, making it difficult to deal with diverse defect types and changes in production conditions, resulting in insufficiency of detection and increased dependence on manual intervention.

Method used

Using a method of combining reinforcement learning with CCD image features, we optimize defect detection through Dreamer algorithm and MAP-Elites algorithm, build a parallel training environment, use the A3C algorithm for asynchronous evaluation and update, adaptively adjust the detection strategy, and combine the edge, texture and spatial features of the image for multi-dimensional feature extraction and fusion.

Benefits of technology

It improves the accuracy and efficiency of defect detection, reduces manual intervention, ensures the robustness and efficiency of the detection system in a dynamic environment, and realizes high-precision automated defect classification and judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198424B_ABST
    Figure CN120198424B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent production line CCD defect detection method based on reinforcement learning, which comprises the following steps: S1, collecting image data through a CCD camera and preprocessing; S2, extracting image edge, texture and spatial features, and encoding to construct a feature matrix; S3, constructing a Dreamer algorithm model, and predicting potential state trajectories and corresponding reward sequences; S4, constructing a feature space by using the MAP-Elites algorithm, and mapping an initial policy sample set to each sub-region; S5, constructing a parallel training environment, asynchronously evaluating and updating the policy based on the A3C algorithm, and performing defect determination on the potential state trajectories; S6, collecting defect difference information and updating network parameters; S7, re-performing prediction, and repeating steps S4-S6 until a preset condition is met and the detection result is output. The present invention improves the detection accuracy and efficiency, reduces manual intervention, and improves the production automation level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision defect detection, and particularly to an intelligent production line CCD defect detection method based on reinforcement learning. Background Art

[0002] With the rapid development of industrial automation, intelligent manufacturing is playing an increasingly important role in improving production efficiency and product quality. In modern production lines, defect detection is a key link to ensure product quality, especially in industries with high-precision requirements, such as electronic products, automotive parts, and high-end machinery manufacturing. Traditional defect detection methods usually rely on manual inspection or rule-based automated detection methods. However, these methods usually have relatively serious limitations, especially in terms of detection accuracy, efficiency, and adaptability, and it is difficult to meet the requirements of modern production lines.

[0003] At present, defect detection technologies based on CCD (Charge Coupled Device) cameras have been widely used in various production lines. The CCD camera captures the image information of the object surface through image acquisition technology, and then extracts features through image processing algorithms to judge defects. However, traditional CCD defect detection methods rely on fixed image processing algorithms and preset rules, such as edge detection, texture analysis, color distribution, etc. Although these methods can identify defects to a certain extent, they are often affected by various factors in practical applications, such as changes in lighting, differences in the surface materials of objects, and the complexity of defect types, which makes it difficult to guarantee the accuracy and stability of the detection results.

[0004] In addition, existing detection methods usually do not have sufficient flexibility and adaptability. In the actual production process, the changes in the production environment and products lead to the complexity of the detection task. For example, there may be slight surface differences between different batches of products, and it is very difficult for traditional algorithms to accurately identify in this case. At the same time, existing algorithms often rely on manually set rules and thresholds, lacking the effective learning and adaptation ability for new defect types. The lack of flexible and intelligent detection methods leads to an increase in manual intervention, which not only reduces production efficiency but also affects the overall quality of products.

[0005] With the rapid development of artificial intelligence technologies, especially deep learning and reinforcement learning, more and more research has begun to apply these advanced technologies to the field of defect detection. Deep learning technologies can learn high-level feature representations from a large amount of data by building deep neural networks, and thus perform excellently in image recognition tasks. However, despite the remarkable achievements of deep learning technologies in image classification and detection, their application in the field of defect detection still faces some problems. First, deep learning models usually require a large amount of labeled data for training, and the training process has high requirements for computing resources. Second, the training of deep learning models requires a large amount of labeled data, which may not be easily obtained in some production environments. In addition, deep learning algorithms are usually a static training process, lacking adaptability and continuous optimization capabilities, which may limit their application in actual production lines for the constantly changing production environment.

[0006] In contrast, as an intelligent decision-making algorithm, reinforcement learning can learn the optimal strategy through interaction with the environment, which gives it greater advantages in dealing with complex dynamic systems. Reinforcement learning enables the agent to optimize its own behavior through continuous trial and error through the reward feedback mechanism. In the field of defect detection, the advantage of reinforcement learning is that it can adaptively adjust the detection strategy according to different detection tasks and production environments, improving the accuracy and flexibility of the system. For example, reinforcement learning can continuously optimize the detection strategy based on historical data, dynamically adapt to different defect types and complex environments, thereby improving the detection efficiency and accuracy.

[0007] However, the existing defect detection methods based on reinforcement learning still have certain defects. First, although reinforcement learning has strong adaptability, it still faces the problems of low model training efficiency and scarce sample data in practical applications. Traditional reinforcement learning algorithms require a large number of trial-and-error processes, which may lead to an overly long training process for the real-time detection requirements of the production line, thus affecting real-time performance. Second, reinforcement learning still relies on manually designed features in feature extraction, such as the edges and textures of images, and the manual design of these features often fails to fully capture the complex information in the image, especially in complex production environments, making it difficult to ensure that the detection system can adapt to various different types of defects.

[0008] Therefore, how to provide an intelligent production line CCD defect detection method based on reinforcement learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0009] An object of the present invention is to propose an intelligent production line CCD defect detection method based on reinforcement learning. The present invention makes full use of reinforcement learning, image processing technology, and multi-modal data fusion, and details an intelligent defect detection algorithm optimized based on reinforcement learning. By adopting a method combining reinforcement learning with CCD image features, the present invention realizes the automatic optimization of defect detection tasks in complex production environments and can continuously adjust detection strategies according to changes in the environment and products. The method of the present invention has the advantages of high self-adaptability, high detection accuracy, and strong flexibility, can significantly improve the detection efficiency of the production line, reduce manual intervention, and effectively improve product quality.

[0010] An intelligent production line CCD defect detection method based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0011] S1. Collect image data through a CCD camera and perform preprocessing;

[0012] S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of the image, and encode the features to form an image feature matrix;

[0013] S3. Build a Dreamer algorithm model, perform sequence modeling on the image feature matrix, and predict the potential state trajectory and the corresponding reward sequence to generate an initial policy sample set;

[0014] S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set;

[0015] S5. Build a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and perform defect determination on the potential state trajectory according to the policy action, and output the defect classification result;

[0016] S6. Collect the difference information between the defect classification result and the true defect annotation result, and update the potential state network parameters and reward prediction network parameters in the Dreamer algorithm model;

[0017] S7. Repredict the potential state trajectory and the corresponding reward sequence as a new round of initial policy sample set, repeat steps S4-S6 until the policy performance meets the preset determination condition, and output the CCD defect detection result.

[0018] Optionally, the image data includes resolution, color, spatial information, timing information, noise information, and annotation information.

[0019] Optionally, the preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.

[0020] Optionally, S2 specifically includes:

[0021] S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image:

[0022] ;

[0023] ;

[0024] where represents the edge intensity of the image at the coordinate point , and respectively represent the horizontal and vertical gradients of the image at , represents the pixel value of the image at ;

[0025] S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image:

[0026] ;

[0027] where represents the texture intensity of the image at the coordinate point , is the weighting coefficient with a window size of , is the pixel value of the image at ;

[0028] S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image:

[0029] ;

[0030] where represents the spatial feature of the image at the coordinate point , represents the number of neighborhood points to be calculated, is the coordinate of the neighborhood point in the image, represents the pixel value of the image at the neighborhood point ;

[0031] S24. Combine the edge features , texture features and spatial features to form an image feature matrix .

[0032] Optionally, the image feature matrix is formed by weighted combination of edge features , texture features and spatial features to form a matrix that comprehensively represents image features:

[0033] ;

[0034] Among them, represents the comprehensive feature value of the image at the coordinate point , , and are the weighted coefficients of edge features, texture features and spatial features respectively, , and are the feature values of edge, texture and spatial features respectively.

[0035] Optionally, S3 specifically includes:

[0036] S31. Construct a Dreamer algorithm model, perform sequence modeling on the image feature matrix , and calculate the potential state trajectory at time step :

[0037] ;

[0038] Among them, represents the potential state at time step , represents the potential state at time step , is the image feature matrix at time step , is the Hadamard product, indicating element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters , represents the high-order feature extraction function of the potential state, and is a hyperparameter of the Dreamer algorithm model;

[0039] S32. Based on the potential state trajectory and the action taken at the current moment, calculate the reward sequence at time step :

[0040] ;

[0041] Among them, represents at time step The reward sequence at a moment, indicating the time step The action taken at the moment, is the reward prediction function of the Dreamer algorithm model, indicating the parameters of the reward prediction function, is the weight coefficient, indicating the influence degree of the latent state on the current reward prediction, is the time step The image feature matrix at the moment, is a hyperparameter used to control the long-term influence of the state sequence, indicating the compensation coefficient related to the latent state, used to adjust the influence of past states on the current reward, indicating based on the latent state The policy function, and respectively represent the influence of the weighted image features and policy features, is the maximum number of steps of the time series, indicating the time span of the whole process;

[0042] S33. Generate the initial policy sample set through the latent state trajectory and the reward sequence:

[0043] ;

[0044] Among them, indicates the initial policy sample set, including all latent states at the time steps from 1 to the moment , the corresponding actions and the predicted rewards .

[0045] Optionally, the Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision module. The latent state network generates the latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts the future reward through the input of the latent state and action, combines the historical image feature sequence, and feeds it back into the learning process for optimization. The sequence modeling module uses the edge features, texture features, and spatial features of the image to fuse multi-dimensional feature information into the Dreamer algorithm model to ensure the generalization ability of the model in complex production environments. The state-action decision module generates the corresponding action according to the current latent state and the predicted reward , to achieve defect determination and execution of actions.

[0046] Optionally, the S4 specifically includes:

[0047] S41. Use the MAP-Elites algorithm to construct a feature space and map the initial policy sample set to each sub-region in the feature space. The feature space is defined by the following formula:

[0048] ;

[0049] where represents the th sub-region, represents the latent state at time step , is the action taken at time step , represents the reward sequence at time step , is the th all latent states in the sub-region, is the action set corresponding to the th sub-region, is the reward sequence set of the th sub-region;

[0050] S42. Evaluate the policy samples in each sub-region and calculate the performance value of the policy samples:

[0051] ;

[0052] where represents the reward sequence of the policy sample at the th time step in the sub-region , represents the number of policy samples corresponding to the reward sequence set in the th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps in the time series, representing the time span of the whole process;

[0053] S43. Select the policy samples with the best performance from each sub-region and generate an optimal policy set:

[0054] ;

[0055] where is the optimal policy set, containing the policy samples with the best performance in all sub-regions, is the number of sub-regions in the feature space, represents the reward sequence in the selected sub-region the longest policy sample.

[0056] Optionally, the MAP-Elites algorithm is an optimization method based on feature space partitioning. The optimization method maps the policy sample set to each sub-region in the feature space, and by evaluating the performance values of the policy samples in each sub-region, selects the policy sample with the best performance in each sub-region, thereby generating an optimal policy set and using it for further policy optimization and iteration.

[0057] Optionally, the S5 specifically includes:

[0058] S51. According to the optimal policy set construct a parallel training environment and use the A3C algorithm to asynchronously evaluate and update each policy sample

[0059] ;

[0060] wherein, represents the updated latent state, is the step factor related to the time step , represents the latent state at the time step , represents the non-linear mapping function constructed by the Dreamer algorithm model parameters , is the image feature matrix at the time step , is the maximum number of steps in the time series, representing the time span of the whole process;

[0061] S52. Update the policy value for each policy sample, and the policy value is updated through the following formula:

[0062] ;

[0063] wherein, is the policy value of the latent state , is the reward signal, is the discount factor, is the learning rate, is the latent state at the next time step;

[0064] S53. Determine the defects of the latent state trajectory according to the policy action and output the defect classification result:

[0065] ;

[0066] Among them, represents the defect classification result, is the weight coefficient related to the category and the feature related, represents the number of features, represents the number of categories, is the normalization factor of the image feature matrix, is the weighted sum for each category, represents selecting the category corresponding to the maximum value as the defect classification result.

[0067] The beneficial effects of the present invention are as follows:

[0068] First of all, by combining reinforcement learning and CCD image processing technology, the present invention solves the problem of poor adaptability of traditional defect detection methods in complex production environments. Traditional methods rely on fixed rules and manually set features and cannot effectively handle diverse defect types and changing production conditions. However, through the reinforcement learning algorithm, the defect detection system of the present invention can adaptively adjust the detection strategy, automatically optimize the detection process according to different production batches and environmental conditions, thereby improving the robustness of the system in a dynamic environment.

[0069] Secondly, by introducing the combination of the Dreamer algorithm model and the MAP-Elites algorithm, the defect detection process of the present invention not only has strong intelligent capabilities but also has higher learning efficiency. Through the reinforcement learning model for sequence modeling and policy optimization of the image feature matrix, different types of defects can be quickly and accurately identified, while avoiding the problems of relying heavily on manual adjustment and excessive rule setting in traditional methods. This method enables the system to complete the learning and optimization of new defect types in the shortest time, improving the detection accuracy and efficiency.

[0070] Finally, the present invention adopts the combination of reinforcement learning and image processing, which can effectively reduce manual intervention and greatly improve the detection speed while ensuring the detection accuracy. Traditional defect detection methods often require manual monitoring and intervention, resulting in reduced production efficiency. Through the automated defect detection method of the present invention, efficient and accurate full-automated defect classification and judgment can be achieved, further improving the working efficiency of the production line, reducing human errors, and ensuring the consistency and reliability of product quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0072] Figure 1 The flowchart of an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention;

[0073] Figure 2 The schematic diagram of image feature matrix construction and initial policy sample set generation for an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention;

[0074] Figure 3 The schematic diagram of policy evaluation optimization and defect determination for an intelligent production line CCD defect detection method based on reinforcement learning proposed by the present invention. Specific implementation manners

[0075] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0076] Refer to Figures 1 - 3 , an intelligent production line CCD defect detection method based on reinforcement learning, comprising the following steps:

[0077] S1. Collect image data through a CCD camera and perform preprocessing;

[0078] S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of image, and encode the features to form an image feature matrix;

[0079] S3. Build a Dreamer algorithm model, perform sequence modeling on the image feature matrix, and predict the potential state trajectory and the corresponding reward sequence to generate an initial policy sample set;

[0080] S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set;

[0081] S5. Construct a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and perform defect determination on the potential state trajectory according to the policy action, and output the defect classification result;

[0082] S6. Collect the difference information between the defect classification result and the true defect annotation result, and update the potential state network parameters and reward prediction network parameters in the Dreamer algorithm model;

[0083] S7. Re-predict the potential state trajectory and the corresponding reward sequence as a new round of initial policy sample set, repeat steps S4 - S6 until the policy performance meets the preset determination condition, and output the CCD defect detection result.

[0084] By collecting and preprocessing image data using a CCD camera, the present invention can effectively improve the accuracy and real-time performance of defect detection. By extracting edge, texture, and spatial features from the image data and combining the training and optimization of the reinforcement learning model, the present invention can adaptively adjust the detection strategy according to the changes in the actual production environment, thereby ensuring the detection effect under different working conditions. This method significantly improves the intelligence level of the production line, reduces manual intervention, and improves production efficiency.

[0085] In this embodiment, the image data includes resolution, color, spatial information, temporal information, noise information, and annotation information.

[0086] The present invention comprehensively utilizes the multi-dimensional information of the image data. In addition to the conventional resolution and color, it also combines spatial information, temporal information, noise information, and annotation information. By integrating these data, the defect detection method of the present invention can more accurately capture the key features in the image, improve the accuracy and adaptability of the detection method, especially in complex production environments.

[0087] In this embodiment, the preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.

[0088] The present invention performs grayscale conversion, edge enhancement, noise suppression, and size normalization on the image in the preprocessing stage, thereby ensuring that the image data input to the subsequent algorithm has higher quality and consistency. This process effectively removes noise and irrelevant interference signals, improves the effect of image feature extraction, and provides more accurate data support for subsequent reinforcement learning and defect determination.

[0089] In this embodiment, S2 specifically includes:

[0090] S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image:

[0091] ;

[0092] ;

[0093] Wherein, represents the edge intensity of the image at the coordinate point , and respectively represent the image at The gradients in the horizontal and vertical directions at represent the pixel value of the image at ;

[0094] S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image:

[0095] ;

[0096] Among them, represents the texture intensity of the image at the coordinate point ; is the weighting coefficient with a window size of ; is the pixel value of the image at ;

[0097] S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image:

[0098] ;

[0099] Among them, represents the spatial feature of the image at the coordinate point ; represents the number of neighborhood points to be calculated, are the coordinates of the neighborhood points in the image, represents the pixel value of the image at the neighborhood point ;

[0100] S24. Combine the edge feature , texture feature and spatial feature to form an image feature matrix .

[0101] The feature extraction method of the present invention further improves the recognition ability of subtle differences in the image by performing edge detection, texture analysis, and spatial feature analysis on the image. By independently calculating and weighted combining different features of each frame of the image to form an image feature matrix, this integration method of multi-dimensional features effectively enhances the expression ability of image information, making defect detection more refined and accurate.

[0102] In this embodiment, the image feature matrix is formed by weighted combining the edge feature , texture feature and spatial feature to form a matrix comprehensively representing the image features:

[0103] ;

[0104] Among them, represents the comprehensive eigenvalue of the image at the coordinate point . , and are the weighted coefficients of edge feature, texture feature and spatial feature respectively, , and are the eigenvalue of edge, texture and spatial feature respectively.

[0105] In the present invention, through the weighted combination of the image feature matrix, a matrix for comprehensively representing the image features is formed, further improving the effectiveness of feature extraction. The weighted coefficients are dynamically adjusted according to the importance of different features, so that the important information in the image can be more prominently represented, thereby improving the accuracy of defect detection, especially in the case of complex images and large background noise.

[0106] In this embodiment, S3 specifically includes:

[0107] S31. Construct a Dreamer algorithm model and perform sequence modeling on the image feature matrix to calculate the potential state trajectory at time step :

[0108] ;

[0109] Among them, represents the potential state at time step , represents the potential state at time step , is the image feature matrix at time step , is the Hadamard product, indicating element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters , represents the high-order feature extraction function of the potential state, and is the hyperparameter of the Dreamer algorithm model;

[0110] S32. Based on the potential state trajectory and the action taken at the current moment, calculate the reward sequence at time step :

[0111] ;

[0112] Among them, represents the reward sequence at time step , Represents the time step The action taken at the moment, is the reward prediction function of the Dreamer algorithm model, represents the parameters of the reward prediction function, is the weight coefficient, indicating the influence degree of the latent state on the current reward prediction, is the time step the image feature matrix at the moment, is a hyperparameter used to control the long-term influence of the state sequence, represents the compensation coefficient related to the latent state, used to adjust the influence of past states on the current reward, represents based on the latent state of the policy function, and respectively represent the influence of the weighted image features and policy features, is the maximum number of steps of the time series, representing the time span of the whole process;

[0113] S33. Generate the initial policy sample set through the latent state trajectory and the reward sequence:

[0114] ;

[0115] Among them, represents the initial policy sample set, including the time step from 1 to all the latent states at the moment the corresponding actions and the predicted rewards .

[0116] The present invention constructs a Dreamer algorithm model to perform sequence modeling on the image feature matrix, which can capture the latent state trajectory in the image and predict the corresponding reward sequence. Through the process of reinforcement learning, the model can optimize the policy in continuous feedback and learning, thereby improving the accuracy of defect detection and the adaptive ability of the system. This method has strong dynamic adjustment ability and can quickly learn and adapt to new defect types in the actual production process.

[0117] In this embodiment, the Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision module. The latent state network generates a latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts future rewards based on the latent state and action inputs, combines the historical image feature sequence, and feeds back to the learning process for optimization. The sequence modeling module utilizes the edge features, texture features, and spatial features of the image to fuse multi-dimensional feature information into the Dreamer algorithm model, ensuring the generalization ability of the model in complex production environments. The state-action decision module generates corresponding actions through a policy optimization method according to the current latent state and the predicted reward , combines the policy selection under each latent state, and generates corresponding actions to achieve defect determination and action execution.

[0118] The present invention improves the adaptive ability and intelligent level of the defect detection method through the Dreamer algorithm model. The combination of the latent state network, the reward prediction network, the sequence modeling module, and the state-action decision module realizes the dynamic generation of the latent state trajectory and the real-time optimization of the strategy. The fusion of image features ensures accurate detection in complex environments. The state-action decision module generates optimal detection actions according to the latent state and the predicted reward, improving the detection accuracy and automation degree.

[0119] In this embodiment, S4 specifically includes:

[0120] S41. Construct a feature space using the MAP-Elites algorithm and map the initial policy sample set to each sub-region in the feature space. The feature space is defined by the following formula:

[0121] ;

[0122] where represents the th sub-region, represents the latent state at time step , is the action taken at time step , represents the reward sequence at time step , is the th sub-region's all latent states, is the action set corresponding to the th sub-region, is the set of reward sequences for the th sub-region;

[0123] S42. Evaluate the policy samples in each sub-region and calculate the performance value of the policy samples:

[0124] ;

[0125] where represents the reward sequence of the policy sample at the th moment in the sub-region , represents the number of policy samples corresponding to the set of reward sequences in the th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps in the time series, representing the time span of the whole process;

[0126] S43. Select the policy samples with the best performance from each sub-region and generate the optimal policy set:

[0127] ;

[0128] where is the optimal policy set, containing the policy samples with the best performance in all sub-regions, is the number of sub-regions in the feature space, represents selecting the policy sample with the longest reward sequence in the sub-region.

[0129] The present invention maps the initial policy sample set to different sub-regions in the feature space by introducing the MAP-Elites algorithm, and screens out the policy samples with the best performance by evaluating the performance of the policy samples in each sub-region. This optimization process not only improves the detection accuracy of the model, but also enhances its generalization ability in different production environments.

[0130] In this embodiment, the MAP-Elites algorithm is an optimization method based on feature space partitioning. The optimization method maps the policy sample set to each sub-region in the feature space, and selects the policy samples with the best performance in each sub-region by evaluating the performance values of the policy samples in each sub-region, so as to generate the optimal policy set and use it for further policy optimization and iteration.

[0131] The present invention optimizes the division of the feature space through the MAP-Elites algorithm and selects the optimal strategy samples in each sub-region. This method can identify the most effective strategies in a complex feature space, improve the detection ability and learning efficiency of the model, and at the same time reduce the computational burden in the training process, further improving the working efficiency of the production line.

[0132] In this embodiment, S5 specifically includes:

[0133] S51. Construct a parallel training environment according to the optimal policy set and asynchronously evaluate and update each policy sample using the A3C algorithm

[0134] ;

[0135] wherein, represents the updated latent state, is the step size factor related to the time step , represents the latent state at the time step , represents the non-linear mapping function constructed by the Dreamer algorithm model parameters , is the image feature matrix at the time step , is the maximum number of steps in the time series, representing the time span of the whole process;

[0136] S52. Update the policy value for each policy sample, and the policy value is updated through the following formula:

[0137] ;

[0138] wherein, is the policy value of the latent state , is the reward signal, is the discount factor, is the learning rate, is the latent state at the next time step;

[0139] S53. Determine the defects of the latent state trajectory according to the policy action and output the defect classification result:

[0140] ;

[0141] wherein, represents the defect classification result, is related to the category and the feature The relevant weight coefficient represents the number of features represents the number of categories is the normalization factor of the image feature matrix is the weighted sum for each category represents selecting the category corresponding to the maximum value as the defect classification result

[0142] The present invention significantly improves the training efficiency and real-time performance of the model by constructing a parallel training environment and asynchronously evaluating and updating each policy based on the A3C algorithm. Through this asynchronous training method, the policy can be updated efficiently, improving the adaptability and decision-making ability of the method of the present invention, and ensuring high-efficiency and accurate defect detection performance in a rapidly changing production environment.

[0143] Example 1:

[0144] To verify the feasibility of the present invention in implementation, the present invention is applied to the CCD defect detection task on a certain intelligent manufacturing production line. This production line is mainly used for manufacturing the outer shells of electronic products. During the manufacturing process, various micro-defects such as scratches, stains, and bubbles may appear on the surface of the outer shells of electronic products, which will directly affect the appearance and performance of the products. Traditional defect detection methods mainly rely on manual inspection or simple image processing techniques, and cannot effectively meet the requirements of high-precision and high-speed automated production. Especially when facing changes in the production environment and the complexity of the product surface, the detection accuracy and efficiency of traditional methods are insufficient. To improve the efficiency and accuracy of defect detection, the present invention proposes an intelligent CCD defect detection method based on reinforcement learning. By extracting, fusing, and optimizing the multi-dimensional features of images and combining reinforcement learning for adaptive detection, many problems of traditional methods in intelligent production are successfully solved.

[0145] In this scenario application, first, a CCD camera is used to collect real-time images of each electronic outer shell on the production line. Each collected image will first go through preprocessing steps such as grayscale conversion, edge enhancement, and noise suppression to ensure that the input data has high quality and consistency. Next, edge, texture, and spatial features of each frame of the image are extracted through image processing algorithms. These feature information is encoded into a feature matrix and input into the Dreamer algorithm model based on reinforcement learning for further analysis.

[0146] In the Dreamer algorithm model, the feature matrix of each image frame is analyzed through sequence modeling, the latent state trajectory is calculated, and the probability of defect occurrence is predicted. Through the MAP-Elites algorithm, these policy samples are mapped to each sub-region in the feature space, and the corresponding detection strategy is selected according to the optimal strategy in each sub-region. Finally, parallel training is carried out based on the A3C algorithm, and the policy is updated after each training to ensure that the model can adaptively adjust the detection scheme to cope with defects in different types and environments.

[0147] To ensure the effectiveness of the present invention, before detection, the system was fully trained using approximately 5000 sample images. These sample images included various types of defects, such as tiny scratches, bubbles, stains, etc., as well as normal housing images. During the training process, the system could gradually improve the detection accuracy by continuously adjusting parameters and optimizing strategies. After training, the system could perform real-time detection on each electronic housing on the production line and automatically reject unqualified products according to the detection results.

[0148] In practical applications, we monitored and collected data during the detection process. During the detection process, the system could process approximately 25 frames of images per second, and the detection speed was increased by approximately 30 times compared to traditional manual detection. For each frame of image, the system could complete defect determination within 2 seconds. We compared the results of manual detection and automatic detection by the system and found that the accuracy of automatic detection reached 95%, the false detection rate was 5%, and the missed detection rate was 3%. In contrast, the accuracy of manual detection was only 85%, the false detection rate was 10%, and the missed detection rate was 8%. From the data, it can be seen that automated detection not only improves the detection efficiency but also significantly improves the detection accuracy and reduces the interference of human factors.

[0149] In addition, when facing different production environments and product changes, the system can quickly adapt and optimize the detection strategy. For example, in different batches of products, the surface material, color, and lighting conditions may be different. Traditional detection methods require re-adjusting parameters or setting new rules. However, the detection method based on reinforcement learning of the present invention can dynamically adjust the detection strategy according to the image features of each batch of products, avoiding the need for manual intervention while ensuring high precision.

[0150] Through the application of this embodiment, the effectiveness of the present invention in the actual production environment is verified. Especially in improving the detection accuracy, production efficiency, and reducing manual intervention, remarkable results have been achieved. The present invention has not only been widely applied in existing production lines but also provides an innovative solution for future intelligent manufacturing systems, with broad application prospects.

[0151] Table 1 Comparison data of the performance of the defect detection system

[0152] ;

[0153] From the above data table, we can see that the automatic detection system of the present invention has shown excellent detection effects in each batch. Compared with manual detection, the accuracy rate has been significantly improved, and the false detection rate and missed detection rate have been significantly reduced, demonstrating the superiority of this technology in practical applications.

[0154] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.

Claims

1. An intelligent production line CCD defect detection method based on reinforcement learning, characterized in that It includes the following steps: S1. Collect image data through a CCD camera and perform preprocessing; S2. Based on the preprocessed image data, extract the edge features, texture features, and spatial features of each frame of the image, and encode the features to form an image feature matrix; S3. Build a Dreamer algorithm model, perform sequence modeling on the image feature matrix, predict the latent state trajectory and the corresponding reward sequence, and generate an initial policy sample set; The Dreamer algorithm model is composed of a neural network architecture based on reinforcement learning. The neural network architecture includes a latent state network, a reward prediction network, a sequence modeling module, and a state-action decision module. The latent state network generates a latent state trajectory through sequence modeling and is dynamically updated through the Dreamer framework in reinforcement learning. The reward prediction network predicts future rewards based on latent states and action inputs, combines historical image feature sequences, and feeds back to the learning process for optimization. The sequence modeling module utilizes the edge features, texture features, and spatial features of images to fuse multi-dimensional feature information into the Dreamer algorithm model, ensuring the generalization ability of the model in complex production environments. The state-action decision module generates corresponding actions through a policy optimization method according to the current latent state and predicted rewards , combines the policy selection under each latent state, and realizes defect determination and action execution; S4. Use the MAP-Elites algorithm to construct a feature space, map the initial policy sample set to each sub-region in the feature space, and retain the optimal policy sample in each sub-region to form an optimal policy set; The MAP-Elites algorithm is an optimization method based on feature space partitioning. The optimization method maps the policy sample set to each sub-region in the feature space, and by evaluating the performance values of the policy samples in each sub-region, selects the optimal policy sample in each sub-region, thereby generating an optimal policy set and using it for further policy optimization and iteration; S5. Construct a parallel training environment according to the optimal policy set, asynchronously evaluate and update each policy based on the A3C algorithm, and determine defects for the latent state trajectory according to the policy actions, and output the defect classification result; S6. Collect the difference information between the defect classification result and the true defect annotation result, and update the latent state network parameters and reward prediction network parameters in the Dreamer algorithm model; S7. Re-predict the latent state trajectory and the corresponding reward sequence as a new round of initial policy sample set, repeat steps S4 - S6 until the policy performance meets the preset determination condition, and output the CCD defect detection result.

2. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, wherein The image data includes resolution, color, spatial information, temporal information, noise information, and annotation information.

3. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, characterized in that The preprocessing includes image grayscale conversion, edge enhancement, noise suppression, and size normalization.

4. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, wherein, The specific content of S2 includes: S21. Perform edge detection on the preprocessed image data and calculate the edge features of each frame of the image; ; ; Among them, represents the edge intensity of the image at the coordinate point , and respectively represent the gradients of the image in the horizontal and vertical directions at , represents the pixel value of the image at . S22. Perform texture analysis on the preprocessed image data and calculate the texture features of each frame of the image; ; Among them, represents the texture intensity of the image at the coordinate point , is the weighting coefficient with a window size of , is the pixel value of the image at . S23. Perform spatial feature analysis on the preprocessed image data and calculate the spatial features of each frame of the image; ; Among them, represents the spatial feature of the image at the coordinate point ; represents the number of neighborhood points to be calculated, are the coordinates of the neighborhood points in the image, represents the pixel value of the image at the neighborhood point ; S24. Combine the edge feature , texture feature and spatial feature to form an image feature matrix .

5. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 4, characterized in that, The image feature matrix By performing weighted combination on edge features , texture features and spatial features to form a matrix that comprehensively represents image features: ; Among them, represents the comprehensive eigenvalue of the image at the coordinate point . , and are the weighting coefficients of the edge feature, texture feature, and spatial feature respectively, , and are the eigenvalue of the edge, texture, and spatial feature respectively.

6. The intelligent production line CCD defect detection method based on reinforcement learning according to claim 1, wherein The specific content of S3 includes: S31. Build a Dreamer algorithm model and perform sequence modeling on the image feature matrix to calculate the latent state trajectory at time step : ; Among them, represents the latent state at time step moment, represents the latent state at time step moment, is the image feature matrix at time step moment, is the Hadamard product, indicating element-wise multiplication, represents the non-linear mapping function constructed by the Dreamer algorithm model parameters and represents the high-order feature extraction function of the latent state, and is the hyperparameter of the Dreamer algorithm model; S32. Calculate the reward sequence at the time step based on the latent state trajectory and the action taken at the current moment at the moment: ; Among them, represents the reward sequence at time step moment, represents the action taken at time step moment, is the reward prediction function of the Dreamer algorithm model, represents the parameters of the reward prediction function, is the weight coefficient, indicating the influence degree of the latent state on the current reward prediction, is the image feature matrix at time step moment, is a hyperparameter used to control the long-term influence of the state sequence, represents the compensation coefficient related to the latent state, used to adjust the influence of the past state on the current reward, represents the policy function based on the latent state and and respectively represent the influence of the weighted image feature and the policy feature, is the maximum number of steps of the time series, indicating the time span of the whole process; S33. Generate an initial policy sample set through the latent state trajectory and the reward sequence; ; Among them, represents the initial policy sample set, including time steps from 1 to all potential states at the corresponding actions and predicted rewards .

7. A CCD defect detection method for an intelligent production line based on reinforcement learning according to claim 1, characterized in that The specific content of S4 includes: S41. Use the MAP-Elites algorithm to construct a feature space and map the initial policy sample set to each sub-region in the feature space, where the feature space is defined by the following formula: ; Among them, represents the th sub-region, represents the potential state at time step moment, is the action taken at time step moment, represents the reward sequence at time step moment, is the set of all potential states in the th sub-region, is the set of actions corresponding to the th sub-region, is the set of reward sequences for the th sub-region; S42. Evaluate the policy samples in each sub-region and calculate the performance values of the policy samples; ; Among them, represents the reward sequence of the policy sample at the th moment within the sub-region, represents the number of policy samples corresponding to the reward sequence set in the th sub-region, represents the average performance of the policy in the sub-region , is the maximum number of steps of the time series, representing the time span of the whole process; S43. Select the optimal policy sample from each sub-region and generate an optimal policy set; ; Among them, is the set of optimal strategies, which contains the optimal strategy samples in all sub-regions. is the number of sub-regions in the feature space. represents selecting the strategy sample with the longest reward sequence in the sub-region.

8. A method for defect detection of CCD in an intelligent production line based on reinforcement learning according to claim 1, characterized in that, The specific content of S5 includes: S51. According to the optimal policy set Construct a parallel training environment, and use the A3C algorithm to asynchronously evaluate and update each policy sample ; Among them, represents the updated potential state, is the step size factor related to the time step and represents the potential state at the time step moment, represents the non - linear mapping function constructed by the Dreamer algorithm model parameters and is the image feature matrix at the time step moment, is the maximum number of steps in the time series, representing the time span of the whole process; S52. Update the policy value for each policy sample, and the policy value is updated through the following formula: ; Among them, is the policy value in the potential state , is the reward signal is the discount factor is the learning rate is the potential state at the next time step; S53. Determine defects based on strategic actions Perform defect determination on the potential state trajectory and output the defect classification result: ; Among them, represents the defect classification result, is the weight coefficient related to the category and the feature , represents the number of features, represents the number of categories, is the normalization factor of the image feature matrix, is the weighted sum of each category, represents selecting the category corresponding to the maximum value as the defect classification result.

Citation Information

Patent Citations

  • Typical defect detection method and system for target object in industrial video image

    CN118470013A

  • Precise component surface detection method and system based on laser parallel control

    CN119260164A