Self-adaptive industrial metal surface defect detection method based on reinforcement learning

By applying an adaptive detection method based on reinforcement learning in industrial metal surface detection and combining deep convolutional neural network for image detection, the problem of degradation of detection accuracy in complex production environments is solved, and the detection effect of high accuracy, robustness and adaptability is achieved.

CN120031833APending Publication Date: 2025-05-23HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510109389.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing industrial metal surface detection methods are difficult to adapt to in complex production environments, and there are problems such as decreasing detection accuracy and excessive calculation pressure due to excessive model parameters.

Method used

Adaptive detection method based on reinforcement learning is adopted, and image enhancement strategies are adaptively selected through reinforcement learning models, combined with deep convolutional neural networks for image detection, so as to realize joint training and optimization of the model.

Benefits of technology

It improves the detection accuracy and robustness of the detection model, reduces the model training time, enhances the adaptability and versatility of the model in different environments, and improves production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031833A_ABST
    Figure CN120031833A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive industrial metal surface defect detection method based on reinforcement learning. Comprising the following steps: firstly, constructing an image feature vector set and an enhanced feature vector set according to an obtained industrial metal surface original image; then, in combination with the environment information of each image, carrying out image enhancement training on the reinforcement learning model by utilizing the industrial metal surface preprocessing image set to obtain a reinforcement learning model after preliminary training; then, training the image detection model by using the industrial metal surface image enhancement set to obtain a preliminarily trained image detection model; carrying out joint training to obtain a trained industrial metal surface defect detection model; and finally, preprocessing an industrial metal surface original image to be detected, inputting the image into the trained industrial metal surface defect detection model, and outputting a defect detection result by the model. According to the invention, industrial images in complex scenes can be processed and metal surface detection results can be output, so that factory efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a metal surface detection method in the field of industrial metal surface detection, and in particular to an adaptive industrial metal surface defect detection method based on reinforcement learning. Background Art

[0002] There is a pain point of abnormal maintenance of production lines in intelligent manufacturing. Good operation of equipment often means a higher input-output ratio and lower maintenance costs. To keep the equipment running, it needs to be maintained. Existing industrial detection methods can monitor equipment abnormalities in real time, so that plans can be made before equipment abnormalities occur, so as to reasonably arrange production and maintenance time and prevent normal production from being affected by sudden equipment maintenance.

[0003] Current industrial detection methods are divided into two categories: traditional detection methods and deep learning detection methods.

[0004] The methods used in traditional detection methods include threshold segmentation, edge detection, and region growing. The core is to use image features that can be directly calculated, such as illumination and noise, to segment the defective parts in the image, and further process and analyze them. Most traditional detection methods are directly based on mathematical and optical technologies. Since they do not involve model training, they have the highest reliability and stability. However, traditional detection methods rely on artificially defined image rules and features to achieve segmentation. The versatility and robustness of the solutions are low, and the detection efficiency is difficult to meet industrial needs.

[0005] Compared with manual or traditional optical detection methods, deep learning solutions have advantages in efficiency and accuracy. The current research directions of industrial surface detection in deep learning detection methods mainly include segmentation-detection methods and direct detection methods. Among them, the segmentation-detection method divides the image into several regions using a candidate box generator, and then classifies and detects different regions. Representative ones are convolutional neural networks CNN and CNN variant networks, which use the feature extraction ability of CNN to extract image features and select the region with the highest detection demand after regional classification, and use traditional detection methods to detect it. Direct detection methods tend to directly use convolutional neural networks to identify on images to achieve end-to-end detection. In this part, the YOLO algorithm and its variants and SSD algorithms are representative. Direct detection methods can effectively improve detection accuracy and detection speed, but there are still problems such as large number of parameters, difficulty in obtaining training data, low model generalization ability and robustness, and low practicality in actual deployment. At the same time, the real production environment of the factory is very complex, and general surface detection methods are difficult to adapt to complex environments. Therefore, it is necessary to provide a metal surface detection method that can adapt to the complex production environment of the factory. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention proposes an adaptive industrial metal surface defect detection method based on reinforcement learning. The method proposed by the present invention can be applied to a variety of production environments, can greatly improve the adaptability in different environments, and avoid excessive calculation pressure caused by too high a number of model parameters.

[0007] The technical solution of the present invention is as follows:

[0008] 1. An adaptive industrial metal surface defect detection method based on reinforcement learning

[0009] Step 1: After preprocessing several original images of industrial metal surfaces, an image feature vector set is obtained; and after image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained;

[0010] Step 2: Combine the environmental information of each image and use the industrial metal surface preprocessing image set to train the reinforcement learning model on image enhancement to obtain the reinforcement learning model after preliminary training;

[0011] Step 3: After training the image detection model using the industrial metal surface image enhancement set, a preliminarily trained image detection model is obtained;

[0012] Step 4: Connect the reinforcement learning model and the image detection model after preliminary training to form an industrial metal surface defect detection model, and then use the industrial metal surface preprocessing image set to train the industrial metal surface defect detection model to obtain a trained industrial metal surface defect detection model;

[0013] Step 5: After preprocessing, the original image of the industrial metal surface to be inspected is input into the trained industrial metal surface defect detection model, and the model outputs the defect detection result.

[0014] In step 1, a plurality of original images of industrial metal surfaces are preprocessed to obtain an image feature vector set, specifically:

[0015] S1: Acquire several original images of industrial metal surfaces;

[0016] S2: performing cropping processing on each original image of the industrial metal surface to obtain a cropped metal surface image;

[0017] S3: converting the cropped metal surface image into a metal surface grayscale image;

[0018] S4: performing normalization processing on the metal surface grayscale image to obtain a normalized metal surface image;

[0019] S5: extracting the features of the normalized metal surface image, and forming an image feature vector from all the features;

[0020] S6: Repeat S2-S5, traverse and process the remaining industrial metal surface original images, obtain corresponding image feature vectors respectively, and thus obtain an image feature vector set.

[0021] In step 1, after image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained, which is specifically:

[0022] After performing image enhancement on each image feature vector in the image feature vector set, a corresponding plurality of enhanced feature vectors are obtained. The enhanced feature vector set is composed of all the image feature vectors themselves and a plurality of enhanced feature vectors corresponding to each image feature vector.

[0023] The step 2 is specifically as follows:

[0024] Step 2.1: Construct several image enhancement tools and generate tool paths, and combine the image enhancement tools and tool path graphs into the action space of the reinforcement learning model;

[0025] Step 2.2: In each training, the reinforcement learning model selects an image feature vector from the image feature vector set as the current state, combines the environmental information corresponding to the current state to enhance the tool and tool path in the action space and generate a tool path diagram, uses the tool path diagram to enhance the current state, obtains the state after image enhancement, calculates the reward function according to the current state and the state after image enhancement, and then adjusts the model parameters and the weights of each node in the tool path diagram; repeats the training multiple times until the training is completed, and obtains the reinforcement learning model after preliminary training.

[0026] In step 2.2, the reward function satisfies the following formula:

[0027]

[0028] Among them, is indicates whether the image enhancement tool in the action space is used, p indicates the weight of the image enhancement tool, and P(L loss ) represents the model performance gain under L2 loss, L loss Represents the loss value of the model output image, L ori Represents the standard loss without selecting any enhancement tool.

[0029] The environmental information of each image is specifically the complexity of the environment, and the complexity of the environment is determined by the number of potential interference factors existing in the production workshop where the metal parts are located.

[0030] The image enhancement tool set includes an image denoising tool, a ghost repair tool, a blur repair tool, and a contrast enhancement tool.

[0031] The image detection model is a deep convolutional neural network.

[0032] In step 4, during the training of the industrial metal surface defect detection model using the industrial metal surface preprocessing image set, a recurrent neural network based on an attention mechanism is used to weight the data in the experience pool of the reinforcement learning model according to its reward value and store it in the reward pool.

[0033] 2. A computer device

[0034] The device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the adaptive industrial metal surface defect detection method based on reinforcement learning are implemented.

[0035] The beneficial effects of the present invention are:

[0036] 1. The present invention utilizes reinforcement learning to perform image enhancement training, which can automatically select tool paths, improve the efficiency and quality of image enhancement, and further improve the detection accuracy of the defect detection model.

[0037] 2. The present invention jointly trains the reinforcement learning model and the image detection model, which effectively reduces the time required for model training and can adapt to image features in different environments, and has good versatility and generalizability.

[0038] 3. The method proposed in the present invention can effectively improve the detection accuracy of metal surface defects, especially in complex industrial environments, and exhibits strong robustness and adaptability.

[0039] 4. The present invention can process industrial metal surface images in complex scenarios, providing a more intelligent and automated solution for metal surface defect detection, significantly improving the factory's production efficiency and product quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The present invention is a flow chart of the method.

[0041] Figure 2 Schematic diagram of the recurrent neural network structure based on the attention mechanism.

[0042] Figure 3 Schematic diagram of the three network channel structure of the decision block.

[0043] Figure 4 This is an example diagram of the image enhancement process. DETAILED DESCRIPTION

[0044] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0045] The present invention proposes an adaptive industrial metal surface defect detection method based on reinforcement learning, such as Figure 1 As shown, the method comprises the following steps:

[0046] Step 1: After preprocessing several original images of industrial metal surfaces, an image feature vector set is obtained; and after image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained;

[0047] In step 1, several original images of industrial metal surfaces are preprocessed to obtain image feature vector sets, specifically:

[0048] S1: Use a high-resolution industrial camera or sensor to take real-time photos of industrial metal surfaces and obtain several original images of industrial metal surfaces. The images are colorful, clear, and cover the area to be detected. The original images of industrial metal surfaces cover all possible image distortions to ensure the generalization ability of the model.

[0049] S2: Crop each original image of the industrial metal surface to obtain a cropped metal surface image, that is, remove irrelevant background and noise areas and retain the effective area of ​​the metal surface to reduce data interference and improve the accuracy and efficiency of subsequent processing;

[0050] S3: Convert the cropped metal surface image into a metal surface grayscale image, which can reduce the data dimension, highlight the texture and defect features of the metal surface, reduce the computational complexity, and retain key information;

[0051] S4: performing normalization processing on the grayscale image of the metal surface to obtain a normalized metal surface image. The normalization processing is used to adjust the brightness and contrast of the image to meet the input requirements of the model and eliminate the influence caused by changes in lighting conditions.

[0052] S5: extracting features of the normalized metal surface image (including edge, texture, and trace features) using an image feature extraction method, and forming an image feature vector from all the features;

[0053] S6: Repeat S2-S5, traverse and process the remaining industrial metal surface original images, obtain corresponding image feature vectors respectively, and thus obtain an image feature vector set.

[0054] After image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained, which is specifically:

[0055] After performing image enhancement on each image feature vector in the image feature vector set, a corresponding plurality of enhanced feature vectors are obtained. The enhanced feature vector set is composed of all the image feature vectors themselves and a plurality of enhanced feature vectors corresponding to each image feature vector.

[0056] Step 2: Combine the environmental information of each image and use the industrial metal surface preprocessing image set to train the reinforcement learning model on image enhancement to obtain the reinforcement learning model after preliminary training;

[0057] Step 2 is as follows:

[0058] Step 2.1: Construct several image enhancement tools and generate tool paths (specifically, verify through experiments and use the paths with better effects as tool paths). The image enhancement tools and tool path diagrams form the action space of the reinforcement learning model; the actions of the reinforcement learning model include tool selection and path selection. The state space includes each image feature vector and environmental parameters in the image feature vector set. The image enhancement tool is a small neural network tool trained using the PolyU, Urban 100, and Flare7K data sets to deal with Gaussian noise, Gaussian blur, and image artifacts. The specific construction process is to train small neural network tools to form decision blocks, and further use the decision blocks to construct the tool path diagram required for explorer training. The loss function Loss of this stage of training satisfies the following formula:

[0059] Loss = 1 / n(yf(State)) 2 +ρ*∑(f(State)-State) 2

[0060] Among them, n represents the number of network nodes, y represents the original input, State represents the image state, f(State) represents the image output of the tool node, and ρ represents the loss weight of the network channel in the entire decision block.

[0061] The first half of the formula is the mean square error (MSE) of the image output, which obtains the average of the square of the difference between the model prediction value f(x) and the sample true value y as a relatively stable quality estimate of the model. The loss function in the second half of the formula is a correction parameter, which is used to align the output state dimensions between different tool path diagram nodes.

[0062] The loss parameter in the image enhancement tool training is set to 0.2. At the same time, for ease of processing, the training image is adjusted to a pixel size of 64*64, and 32 image samples are trained in each batch. In the model training of this article, in order to make the model more robust, Adam is selected as the optimizer. The learning rate is set to 8*10 -3 , and it will be reduced by 1 / 3 after each round of training.

[0063] Step 2.2: In each training, the reinforcement learning model selects an image feature vector from the image feature vector set as the current state, combines the environmental information corresponding to the current state, and generates a tool path diagram in the form of a directed graph from the image enhancement tools and tool paths in the action space. The tool path diagram contains multiple nodes. A single image enhancement tool or a combination of multiple image enhancement tools is used as a node in the tool path diagram, providing a rich selection for the action space of the reinforcement learning model. The tool path between the image enhancement tools is the connection relationship between the two nodes. The tool path diagram is used to enhance the current state and obtain the state after image enhancement. The reward function is calculated based on the current state and the state after image enhancement, and then the model parameters and the weights of each node in the tool path diagram are adjusted; the training is repeated many times until the training is completed, and the reinforcement learning model after preliminary training is obtained. The trained reinforcement learning model can adaptively select the optimal image enhancement tool and path, thereby completing the adaptive image enhancement of the image feature vector.

[0064] The task of the reinforcement learning model needs to consider two aspects: image enhancement tool selection and path selection. Therefore, the reward function satisfies the following formula:

[0065]

[0066] Among them, is indicates whether the image enhancement tool in the action space is used, and p indicates the weight of the image enhancement tool, which is updated through model learning. The higher the weight, the more important the tool is in the current environment. P(L loss ) represents the model performance gain under L2 loss, L loss Represents the loss value of the model output image, L ori Represents the preset standard loss without selecting any enhancement tool.

[0067] The -∑is*p in the formula belongs to path selection, which means that when the explorer decides to use a more complex tool subgraph for image enhancement, the model will get a smaller reward value. loss )、L loss , L ori The three together constitute the reward for image enhancement tool selection, which belongs to image enhancement tool selection. loss When L is smaller, the image loss is lower and the difficulty of image enhancement is also reduced. loss When it decreases, the reward value will decrease, and the reinforcement learning model tends to make decisions on image enhancement tools with less computational pressure, which is specifically manifested as image enhancement tools with fewer layers in the decision block. These two parts constitute the reward function of the reinforcement learning model. i,The reinforcement learning model can use the reward function to calculate the reward obtained by the decision, thereby obtaining the mapping between the image enhancement requirements and the image enhancement tools.

[0068] The reward function proposed in this invention comprehensively considers the quality improvement and computational cost after image enhancement, and aims to encourage the reinforcement learning model to reduce the consumption of computing resources while improving image quality, thereby ensuring the real-time and high efficiency of the system.

[0069] The environmental information of each image is specifically the complexity of the environment. The complexity of the environment is determined by the number of potential interference factors in the production workshop where the metal parts are located. The potential interference factors include lighting conditions, noise levels, metal surface materials, equipment vibrations and other factors.

[0070] The image enhancement toolset includes image denoising tools, ghost repair tools, blur repair tools, and contrast enhancement tools. In the specific implementation, various image enhancement tools are pre-designed and trained, and each tool is optimized for specific image quality issues.

[0071] The reinforcement learning model includes the following processes:

[0072] Value function and strategy update: Based on the principle of reinforcement learning, the value network is used to evaluate the pros and cons of the current strategy, guide the update of the strategy network, and through iterative optimization, the model gradually learns the optimal image enhancement strategy;

[0073] Environmental interaction and reward feedback: After applying the image enhancement strategy, the model obtains the enhanced image quality index and calculates the reward value as the basis for model optimization;

[0074] Experience pool and experience replay: The state, action, reward, and next state quadruple generated during the interaction is stored in the experience pool. The experience replay mechanism is used to train the model in batches to improve training efficiency and stability.

[0075] Model convergence and verification: Through multiple iterations of training, the model strategy is gradually optimized until convergence. The model is evaluated using a validation set to ensure that the model's image enhancement capabilities meet actual application requirements.

[0076] The present invention adopts a reinforcement learning model to decide the optimal enhancement tool from the action space and store the tuple {image state S, predicted image state S' at the next time step, decision action a, reward r}, and finally jointly trains multiple convolutional networks to extract the enhanced image features output by the explorer and output the industrial detection results. The specific implementation method is:

[0077] Parameters are initialized, and the cv library of Python is used to implement image cutting, black and white processing, and vectorization before inputting into the model.

[0078] Construct a tool path diagram. The decision block is the tool decision unit of the tool path diagram and is a collection of image enhancement tools. Figure 3 As shown in the figure, the decision block has three network channels. Each channel is a small neural network composed of a fully connected FC layer and a convolutional neural layer. According to the different network structures, they correspond to three types of image enhancement tools: image denoising, image ghosting, and image blur repair. The reinforcement learning model can use the policy network in the actor to adaptively select the appropriate tool map path for each time step image and the tools in each decision block. The data changes of the image through the decision block are: Where D i+1 and D i Respectively represent the output and input of the i-th decision block, n ij is an upsampling mask for path subgraph selection, d k is a vector used to assist in calculating the direction of the subgraph. is the input D i The function for feature transformation, ⊙ represents the element-wise multiplication of two vector matrices, n ij is an upsampling selection code for selecting tools in the path graph, R is the reward in the reinforcement learning model, and a i represents the actions or decisions made by the model, and N represents the number of decision blocks.

[0079] It should be noted that in different graph nodes, the parameters of the corresponding tool nodes may be different, but the decision parameters of the reinforcement learning model are consistent. The path is the tool quantity decision unit of the tool path graph.

[0080] Compared with the one-dimensional tool chain implemented by a single one-hot vector commonly used in reinforcement learning, the two-dimensional path graph can transform the decision on the number of tools into the decision on the path subgraph, thus covering a more complex image enhancement environment at a lower cost.

[0081] The path subgraph is a path that responds to image enhancement requirements of different scales and contains different numbers of decision blocks. The explorer can adaptively select the tool path according to the needs of image enhancement, that is, the explorer's action space contains two one-hot vectors: {tool decision, route decision}.

[0082] The tool path graph contains the input image feature X in , output the enhanced image Y out In the path graph, when the data passes through the nth decision block of the i-th path, it will perform an enhancement in the decision block and output the enhanced image features.

[0083] The decision network undertakes the decision-making task of the explorer and makes decisions on the types and number of tools. In the deep deterministic policy gradient algorithm, there are two types of networks: Actor and Critic. The Actor network belongs to the deterministic policy network and directly outputs continuous actions based on the input state s. The Critic network is used to evaluate the decisions made by the policy network. The two are coupled to form a decision network. The Critic network follows the following Bellman equation:

[0084]

[0085] Among them, Q(s,a) is the reward data that the agent can obtain for its decision action in the current state s, that is, whether the image enhancement tool can maximize the enhancement of the image, and r(s,a) is the immediate reward obtained by the image enhancement tool in the current state and action. It represents the expected value of all possible states s, and is usually used to represent the average performance of the model under different states. Q is the reward value that the agent can obtain when performing action a in the current state s, indicating that the agent chooses the optimal action path in the current state, and a′ selects the action that can maximize future rewards from all possible actions. s′ is the next state obtained by sampling. The reward data of subsequent time steps is predicted using a neural network with parameters, which improves the accuracy of Q value calculation. The Actor is responsible for maintaining the policy network, while the Critic is responsible for calculating r (value) based on the state transmitted by the policy network and returning it. The image of the environment will input the image state S to the Actor t , the policy network in the Actor selects a path subgraph from the action space as an action to interact with the environment. The environment returns the state S of the next time step t+1 , transmitted to the Critic network via the Actor and the reward value r of the state is calculated using the Q network. The experience pool processes the data sent by the Actor into {S t , S t+1 , a, r} and store them for batch extraction of training actors and critics. As the decision network gradually obtains better image parameters, the rewards of its decision strategy will also be accumulated by the model.

[0086] The role of the experience replay buffer is to store the four-tuple {S t , S t+1, a, r}, and adopt random and other algorithms to batch sample experience for training. By introducing the Attention-based RNN module in the experience pool, weighted storage of experience data is realized, focusing on the high-weight part of the experience to improve the efficiency of experience storage and training, while reducing the computing pressure. First, the data transmitted by the Actor is organized into tuples and the softmax activation function is used to assign a retention probability σ to the tuple, so that the data probability σ is retained and combined into the input sequence of the encoder. Softmax ensures that the sum of all retention probabilities in the same tuple is equal to 1, preventing errors in the data proportion in the sequence. is the kth input sequence, For the input experience data sequence, the weight of the attention mechanism input at time t is added to the sequence before entering the encoder This attention weight will adaptively extract the features of the driving input sequence and input them into the RNN module of the encoder. Specifically, it is the attention weight given by the temporal attention mechanism. To select the hidden state of the time series encoder. Adaptive feature extraction of input sequence is input to the encoder and updates its hidden state S t-1 In the processing module, the short-term storage state ψ of the RNN module is updated for each time step of the input sequence t This is used to further update the hidden state. The update strategy of the storage state. After completing the weight addition, the data enters the decoder part. In this part, the time series encoder and decoder are connected to each other through the context vector and decode the output.

[0087] Step 3: After training the image detection model using the industrial metal surface image enhancement set, a preliminarily trained image detection model is obtained;

[0088] The image detection model is used to output the enhanced image in the front module as the industrial detection result. The image detection model mainly includes the convolutional neural network conv2D, the fully connected layer FC and the maximum pooling layer MaxPooling. The image detection model uses multiple convolutional layers to realize image feature extraction and defect detection, and finally outputs the detection results through the FC fully connected layer. Its core is to use the high-dimensional feature extraction capability of the convolutional layer to achieve high-precision industrial surface detection.

[0089] The L2 loss function is used in the image detection model training, as shown in the formula:

[0090] L 2 =∑(r i -f(wS i +b)) 2

[0091] Among them, L 2represents the L2 loss value, r i represents the reward calculated by the image state, w represents the matrix weight parameter of the network, b represents the network bias value, S i Indicates the image state, f(wS i +b) represents the reward estimated based on the image state. The loss function of this step is obtained by calculating the sum of the squares of the difference between the two, and the parameters of the network are updated accordingly.

[0092] During the training process at this stage, the loss parameter of the model is set to 0.01, and 32 image samples are trained in each batch. In order to make the model more robust, Adam is also selected as the optimizer. The learning rate is set to 2*10 -3 , and it will be reduced by 1 / 3 after each round of training.

[0093] Step 4: Connect the reinforcement learning model and the image detection model after preliminary training to form an industrial metal surface defect detection model, and then use the industrial metal surface preprocessing image set to train the industrial metal surface defect detection model, that is, jointly train the reinforcement learning model and the image detection model, and finally obtain a trained industrial metal surface defect detection model; the joint training is specifically that during the training process, the parameters of the two models influence and optimize each other, and the output of the enhanced model directly affects the performance of defect detection; design a comprehensive loss function, taking into account the image enhancement quality and defect detection performance, to ensure that the optimization goals of the two models are consistent; introduce an experience pool based on the attention mechanism to improve the efficiency and effect of model training. Through joint training, the present invention makes the output of the image enhancement model more suitable for the input requirements of the defect detection model, and improves the accuracy and robustness of the overall detection system.

[0094] In the joint training part, the model will synchronously update the parameters of the explorer, image enhancement tool, and image detection model in each patch to strengthen the coupling of the model and improve the model performance under high-pressure environments. The image detection model will use its loss function to further update its own parameters; the image enhancement tool will use the zeroed loss function to jointly train its own parameters. This change is to focus on training the image detection model, ignoring the intermediate loss of the decision block, so as to calculate the image state change through the loss function and return the parameter update; the explorer still uses the policy algorithm.

[0095] In step 4, during the training of the industrial metal surface defect detection model using the industrial metal surface preprocessing image set, the data in the experience pool of the reinforcement learning model is weighted according to its reward value using a recurrent neural network based on an attention mechanism and stored in the reward pool. The network structure diagram of the recurrent neural network based on the attention mechanism is shown in FIG. Figure 2As shown. This allows experiences with higher weights to be used more frequently in training, highlighting experiences that contribute more to model optimization. By introducing the attention mechanism, the present invention enables the model to learn effective strategies more quickly, reduce training time, and improve the convergence speed and performance of the model; while using the attention mechanism, measures are taken to prevent the model from over-relying on a small amount of high-reward experience data, thereby maintaining the generalization ability of the model.

[0096] Step 5: Deploy the model in the actual production environment to ensure the compatibility and stability of hardware and software. Obtain the original image of the industrial metal surface to be inspected, pre-process the original image of the industrial metal surface to be inspected, and then input it into the trained industrial metal surface defect detection model, and the model outputs the defect detection result. The image enhancement process is as follows Figure 4 As shown, Figure 4 (a) is the original image of the industrial metal surface to be detected. Figure 4 (a) is an industrial metal surface image in image enhancement. Figure 4 (c) is an industrial metal surface image after image enhancement. During the production process, the image data of the metal surface is collected in real time, and the model quickly preprocesses, enhances the image, and detects defects to meet the real-time requirements of the production line. The model adaptively adjusts the image enhancement strategy according to the actual production environment and image quality to ensure that a high level of detection performance can be maintained in different environments. The detection results are output in real time, including information such as the location, type, and severity of the defects, for reference by production line operators, and the results are fed back to the production control system for corresponding adjustments and early warnings.

[0097] The image detection model is a deep convolutional neural network, including multiple convolutional layers, pooling layers and fully connected layers, which can effectively extract the deep features of the image and accurately locate and classify defects. The network parameters of the deep convolutional neural network include the number of filters and kernel size of the convolutional layer, the window size of the pooling layer, the number of neurons in the fully connected layer, the type of activation function, and the optimizer selection.

[0098] Step 6: Continue to collect new image data and defect cases in the production process, enrich the data set, cover more environments and defect types, and then continuously optimize and update the trained industrial metal surface defect detection model. Specifically, adjust the model's hyperparameters, such as learning rate, network depth, loss function weight, etc., to optimize model performance; according to production needs, add new image enhancement tools and detection functions, such as dedicated detection modules for specific defect types, to improve the functionality and practicality of the system, thereby maintaining the model's adaptability to new environments and new defects.

[0099] Through the above steps, the present invention provides an adaptive industrial metal surface defect detection method based on reinforcement learning, which uses the reinforcement learning model to adaptively select the optimal image enhancement strategy, and solves the problem of reduced detection accuracy in industrial detection due to complex production environment, unstable lighting conditions, noise interference, etc. Through joint training with the defect detection model, the coupling degree and detection performance of the model are improved, and real-time and high-precision detection of industrial metal surface defects can be achieved in a complex and changeable production environment.

[0100] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modification and change made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. An adaptive industrial metal surface defect detection method based on reinforcement learning, characterized in that: The following steps are involved: Step 1: After preprocessing several original images of industrial metal surfaces, an image feature vector set is obtained; and after image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained; Step 2: Combine the environmental information of each image and use the industrial metal surface preprocessing image set to train the reinforcement learning model on image enhancement to obtain the reinforcement learning model after preliminary training; Step 3: After training the image detection model using the industrial metal surface image enhancement set, a preliminarily trained image detection model is obtained; Step 4: Connect the reinforcement learning model and the image detection model after preliminary training to form an industrial metal surface defect detection model, and then use the industrial metal surface preprocessing image set to train the industrial metal surface defect detection model to obtain a trained industrial metal surface defect detection model; Step 5: After preprocessing, the original image of the industrial metal surface to be inspected is input into the trained industrial metal surface defect detection model, and the model outputs the defect detection result.

2. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 1, characterized in that: In step 1, a plurality of original images of industrial metal surfaces are preprocessed to obtain an image feature vector set, specifically: S1: Acquire several original images of industrial metal surfaces; S2: performing cropping processing on each original image of the industrial metal surface to obtain a cropped metal surface image; S3: converting the cropped metal surface image into a metal surface grayscale image; S4: performing normalization processing on the metal surface grayscale image to obtain a normalized metal surface image; S5: extracting the features of the normalized metal surface image, and forming an image feature vector from all the features; S6: Repeat S2-S5, traverse and process the remaining industrial metal surface original images, obtain corresponding image feature vectors respectively, and thus obtain an image feature vector set.

3. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 1 is characterized in that: In step 1, after image enhancement is performed on the image feature vector set, an enhanced feature vector set is obtained, which is specifically: After performing image enhancement on each image feature vector in the image feature vector set, a corresponding plurality of enhanced feature vectors are obtained. The enhanced feature vector set is composed of all the image feature vectors themselves and a plurality of enhanced feature vectors corresponding to each image feature vector.

4. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 1 is characterized in that: The step 2 is specifically as follows: Step 2.1: Construct several image enhancement tools and generate tool paths, and combine the image enhancement tools and tool path graphs into the action space of the reinforcement learning model; Step 2.2: In each training, the reinforcement learning model selects an image feature vector from the image feature vector set as the current state, and combines the environmental information corresponding to the current state to enhance the tool and tool path from the action space and generate a tool path graph. The tool path graph is used to enhance the current state to obtain the state after image enhancement. The reward function is calculated based on the current state and the state after image enhancement, and then the model parameters and the weights of each node in the tool path graph are adjusted. Repeat the training several times until the training is completed to obtain the reinforcement learning model after preliminary training.

5. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 4 is characterized in that: In step 2.2, the reward function satisfies the following formula: Among them, is indicates whether the image enhancement tool in the action space is used, p indicates the weight of the image enhancement tool, and P(L loss ) represents the model performance gain under L2 loss, L loss Represents the loss value of the model output image, L ori Represents the standard loss without selecting any enhancement tool.

6. The adaptive industrial metal surface defect detection method based on reinforcement learning according to claim 1 is characterized in that: The environmental information of each image is specifically the complexity of the environment, and the complexity of the environment is determined by the number of potential interference factors existing in the production workshop where the metal parts are located.

7. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 4 is characterized in that: The image enhancement tool set includes an image denoising tool, a ghost repair tool, a blur repair tool, and a contrast enhancement tool.

8. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 1, characterized in that: The image detection model is a deep convolutional neural network.

9. The method for adaptive industrial metal surface defect detection based on reinforcement learning according to claim 1, characterized in that: In step 4, during the training of the industrial metal surface defect detection model using the industrial metal surface preprocessing image set, a recurrent neural network based on an attention mechanism is used to weight the data in the experience pool of the reinforcement learning model according to its reward value and store it in the reward pool.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the adaptive industrial metal surface defect detection method based on reinforcement learning as described in any one of claims 1 to 9 are implemented.