Shaving board target detection system and detection method based on deep reinforcement learning

Through deep reinforcement learning method, the initial candidate rectangular box in particleboard object detection is transformed in shape, which solves the problems of large computing overhead and limited adaptability in the prior art, and achieves efficient and accurate object detection.

CN120164017AActive Publication Date: 2025-06-17NANJING FORESTRY UNIV

Patent Information

Application Number
CN202510211623.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing object detection algorithm has high computational overhead and limited adaptability in particleboard object detection.

Method used

Deep reinforcement learning method is used to continuously transform the initial candidate rectangle box to complete the target defect position detection. This method designs a dataset loader and a framework for deep reinforcement learning object detection model by defining the state space and action space of reinforcement learning agents, and trains a deep reinforcement learning object detection model by combining object detection evaluation indicators and reward functions.

Benefits of technology

It reduces the computational complexity, improves adaptive detection capabilities, has high detection accuracy, and is suitable for targets of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164017A_ABST
    Figure CN120164017A_ABST
Patent Text Reader

Abstract

The invention discloses a shaving board target detection system and detection method based on deep reinforcement learning. The method comprises the following steps: acquiring a shaving board defect sample image data set; defining the state and action space of the Agent; designing a data set loader, and building a framework of a deep reinforcement learning target detection model; designing a target detection evaluation index and a reward function; and training a deep reinforcement learning target detection model by using the divided target detection data set, and selecting an optimal deep reinforcement learning target detection model in combination with a target detection evaluation index. According to the method, deep reinforcement learning is used, continuous morphological transformation is carried out on an initial candidate rectangular frame, and target defect position detection is completed; the method can act on the defect position detection task of the shaving board image, the calculation complexity is reduced, the self-adaptive detection capability is high, and the detection accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent forestry and computer technology, and particularly to the fields of particleboard defect detection, deep reinforcement learning, and machine vision. Specifically, it is a particleboard target detection system and detection method based on deep reinforcement learning. Background Art

[0002] Particleboard is a type of artificial board made from wood chips and adhesives through hot pressing, and is widely used in furniture manufacturing and other fields. Quality inspection during its production process is one of the important application scenarios of target detection technology. Target detection is an important research direction in the field of machine vision and has been widely applied in many aspects such as the mechanical industry, medical field, military field, and agricultural and forestry product fields. Establishing a perfect target detection model and improving the accuracy of target detection tasks are the key research contents at present.

[0003] Target detection algorithms have continuously improved in detection performance from the early R-CNN to the later two-stage detectors such as Fast R-CNN, and single-stage detectors such as YOLO and SSD. However, the existing single-stage detectors and two-stage detectors at present increase the network structure and parameter calculation amount at the cost of continuously improving the detection performance. Therefore, there are problems such as complex network structure and large computational overhead, and the adaptive ability in dynamic scenarios is limited.

[0004] Reinforcement learning learns the optimal strategy through continuous interaction between the agent and the environment, and its decision-making mechanism is highly similar to the human visual attention system. Introducing reinforcement learning into the particleboard target detection task can enable the detector to dynamically adjust the search strategy according to scene changes, which helps to reduce the computational complexity and improve the system robustness.

[0005] In summary, there is a need to provide a particleboard target detection system and detection method to realize the detection of the defect positions on the particleboard and further solve the problems such as large computational overhead and limited adaptive ability existing when the existing target detection algorithms are used for particleboard target detection. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a particleboard target detection system and detection method based on deep reinforcement learning for the problems of large computational overhead and limited adaptive ability of the existing target detection algorithms. This method uses deep reinforcement learning to perform continuous morphological transformations on the initial candidate rectangular frames to complete the detection of target defect positions. This method can be applied to the defect position detection task of particleboard images, reduces the computational complexity, can adapt to targets of different sizes, has strong adaptive detection ability and high detection accuracy.

[0007] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0008] A particleboard target detection method based on deep reinforcement learning, comprising:

[0009] Step 1: Obtain a particleboard defect sample image dataset, including the original sample images and the corresponding target defect category labels and target defect position labels of the original images; divide the original sample images and the target defect category labels into a training set and a test set as the target classification dataset for training a classification model, and divide the original sample images and the target defect position labels into a training set and a test set as the target detection dataset for training a deep reinforcement learning target detection model; the target defect categories include large particles, glue spots, sand leakage, scratches, and dust spots;

[0010] Step 2: Define the state space and action space of the reinforcement learning Agent; the state space consists of global features, local features, and historical action sequences to assist the Agent in making decisions; the action space includes translation operations, scaling operations, and termination actions, which act on the Mask for positioning the target; the Mask represents a rectangular box;

[0011] Step 3: Design a dataset loader and build the framework of the deep reinforcement learning target detection model;

[0012] Step 4: Design the target detection evaluation index and the reward function;

[0013] Step 5: Use the divided target classification dataset to train the classification model to obtain a trained classification model; use the divided target detection dataset to train the deep reinforcement learning target detection model, and combine the target detection evaluation index to select the best deep reinforcement learning target detection model as the particleboard target detection model.

[0014] As a further improved technical solution of the present invention, the specific content of Step 2 includes:

[0015] 2.1 Define the state space of the reinforcement learning Agent: The state space consists of one-dimensional arrays of three parts, specifically including global features, local features, and historical action sequences;

[0016] 2.2 Design an image feature extractor: The image feature extractor is based on the ResNet101 network. After removing its last fully connected layer, it is used as the backbone network, and the pre-trained weight parameters are loaded, and the network is set to the evaluation mode;

[0017] 2.3 Global features: The original image is first pre-processed, that is, it is scaled to a fixed size using a scaling operation, and then normalized using a preset mean and standard deviation; the pre-processed image passes through the image feature extractor in Step 2.2 to output global features;

[0018] 2.4 Local features: The image content within the Mask first undergoes the preprocessing operations of scaling and normalization in Step 2.3. The preprocessed image then passes through the image feature extractor in Step 2.2 to output local features;

[0019] 2.5 Historical action sequence: Each action is encoded using one-hot encoding. That is, in the action space dimension, the position of the executed action is marked as 1, and the remaining positions are 0; The sequence stores the encoding of the most recent n actions vertically, with the new action encoding at the top and the historical encodings shifting downwards in sequence; The initial historical action sequence is an all-1 array;

[0020] A i =[0,1,0,0,…,0], H=[A1,A2,…,A 10 (1);

[0021] where A i is the encoding of the i-th action, and H is the historical action sequence;

[0022] 2.6 Concatenate the global features in Step 2.3, the local features in Step 2.4, and the historical action sequence in Step 2.5 along the length direction. After standardization, the state S of the Agent is obtained;

[0023] 2.7 Define the action space of the reinforcement learning Agent: The action space includes translation operations, scaling operations, and termination actions; Each time the Agent takes an action, it samples an action from the action space according to the probability distribution;

[0024] 2.8 The scaling operations include zooming in on all four sides, zooming out on all four sides, horizontal zooming in, horizontal zooming out, vertical zooming in, and vertical zooming out; The change in width and height for each action is α times the corresponding width and height of the current Mask (0 ≤ α ≤ 1);

[0025]

[0026] where the calculation formulas for scaling on all four sides, horizontal scaling, and vertical scaling are shown from left to right. (x1,y1) is the coordinate of the upper left corner of the Mask, (x2,y2) is the coordinate of the lower right corner of the Mask, (x′1,y′1) is the coordinate of the upper left corner of the updated Mask, (x′2,y′2) is the coordinate of the lower right corner of the updated Mask, and W and H are the width and height of the Mask respectively;

[0027] 2.9 The translation operations include horizontal movement and vertical movement. Vertical movement includes upward translation and downward translation, and horizontal movement includes leftward translation and rightward translation; The change in width and height for each action is α times the corresponding width and height of the current Mask (0 ≤ α ≤ 1);

[0028]

[0029] Among them, from left to right are the calculation formulas for horizontal movement and vertical movement operations respectively;

[0030] 2.10. The termination action indicates that the Agent ends the iteration of the current image with the current Mask as the final positioning box.

[0031] As a further improved technical solution of the present invention, step 3 specifically includes:

[0032] 3.1. Design a dataset loader to read sample image data, extract relevant parameters from the target defect category label and the target defect position label, and arrange them in a unified format: [class, x t1 , x t2 , y t1 , y t2 , where class is the value of the category to which the target in the image belongs, and (x t1 , y t1 ) is the upper left corner coordinate of the true positioning box recorded by the target defect position label, and (x t2 , y t2 ) is the lower right corner coordinate of the true positioning box recorded by the target defect position label;

[0033] 3.2. Build the framework of the deep reinforcement learning object detection model, including a behavior decision network and a value evaluation network. Among them, the behavior decision network is used to perform action policy selection and baseline value estimation, and the value evaluation network is used to be responsible for the evaluation of the state value function;

[0034] 3.3. The behavior decision network consists of a backbone network, an action head, and a baseline value head; the backbone network consists of a linear layer, two linear modules, and a root mean square normalization layer;

[0035] 3.4. The linear module consists of a root mean square normalization layer, a linear layer, a quadratic ReLU activation function, a linear layer, and a dropout layer;

[0036] ReLUSquared(x) = (ReLU(x)) 2 (4);

[0037] Among them, ReLUSquared is the quadratic ReLU activation function;

[0038] 3.5. The state S in step 2.6 is processed by the first linear layer of the backbone network in the behavior decision network, and the output feature A1 is obtained; the feature A1 is processed by the first linear module, and the output feature A2 is obtained; the residual connection combines A1 + A2 = A3; after being processed by the second linear module, the output feature A4 is obtained; the residual connection combines A3 + A4 = A5; and then through the root mean square normalization layer, the normalized feature A6 is obtained;

[0039] 3.6. The action head includes a linear layer, a quadratic ReLU activation function layer, and a linear layer; the feature A6 in step 3.5 is processed by the action head to obtain the feature A7, and then the action probability distribution Action is output through the SoftMax layer prob ;

[0040] 3.7. The baseline value head includes a linear layer, a quadratic ReLU activation function layer, and a linear layer; the feature A6 in step 3.5 is processed by the baseline value head to obtain the baseline value output A8;

[0041] 3.8. The value evaluation network consists of a backbone network and a state value head; the backbone network consists of a linear layer, six linear modules as described in step 3.4, and a root mean square normalization layer;

[0042] 3.9. The state S in step 2.6 is processed by the first linear layer of the backbone network in the value evaluation network to output the feature B1; the feature B1 is processed by the first linear module to output the feature B2; the residual connection combines B1 + B2 = B3; it is processed by the second linear module to output the feature B4; the residual connection combines B3 + B4 = B5; it is processed by the third linear module to output the feature B6; the residual connection combines B5 + B6 = B7; it is processed by the fourth linear module to output the feature B8; the residual connection combines B7 + B8 = B9; it is processed by the fifth linear module to output the feature B 10 ; the residual connection combines B9 + B 10 = B 11 ; it is processed by the sixth linear module to output the feature B 12 ; the residual connection combines B 11 + B 12 = B 13 ; finally, through the root mean square normalization layer, the normalized feature B is obtained 14 ;

[0043] 3.10. The state value head is a linear layer; the feature B in step 3.9 14 is processed by the state value head to obtain the state value output B 15 .

[0044] As a further improved technical solution of the present invention, step 4 specifically includes:

[0045] 4.1. The object detection evaluation metrics adopt multi-class mean average precision mAP, recall Recall, and intersection over union IoU;

[0046]

[0047] where C is the total number of classes, APi Denote the average precision of the \(i\)-th class, where \(i\) is the class index; \(TP\) represents the number of correctly detected targets, and \(FN\) represents the number of undetected targets; \(A\) and \(B\) represent two bounding boxes; \(Area\) intersection is the intersection area of the two bounding boxes; \(Area\) union is the union area of the two bounding boxes; \(Area\) A , \(Area\) B are the areas of bounding boxes \(A\) and \(B\) respectively;

[0048] 4.2. The reward function is divided into normal action reward and termination action reward, and they are gradually superimposed;

[0049] 4.3. Normal actions include scaling actions and translation actions. When these actions are selected, first check whether the Mask exceeds the range of the original image. If it does, a reward of \(-r\) will be given b is given;

[0050] 4.4. After the normal action is executed, if it exceeds the range of the original image, the Mask will be restricted within the correct range by the constraint conditions;

[0051]

[0052] Among them, \((x_1,y_1)\) are the coordinates of the upper left corner of the Mask, \((x_2,y_2)\) are the coordinates of the lower right corner of the Mask, \((x_1',y_1')\) are the coordinates of the upper left corner of the updated Mask, \((x_2',y_2')\) are the coordinates of the lower right corner of the updated Mask, \(\max(a,b)\) and \(\min(a,b)\) are used to take the maximum and minimum values of the array \((a,b)\) respectively;

[0053] 4.5. Calculate the IOU between the Mask after the normal action and the true Mask of the target defect position label through the formula in step 4.1; The IOU reward function is as follows:

[0054]

[0055] Among them, \(i\) is the index of the \(i\)-th action, \(1,2,3,\cdots\); \(\beta\) u is the weight of the positive reward, \(\beta\) d is the weight of the negative reward. The initial IOU is calculated from the initial Mask and the true Mask, and the size of the initial Mask is the size of the original image;

[0056] 4.6. When the termination action is selected, first set the termination flag \(done = 1\), and give a step reward according to the IOU calculated by the previous action;

[0057]

[0058] 4.7. If the Agent selects the termination action in the first action selection, a reward of -r b is given.

[0059] As a further improved technical solution of the present invention, step 5 specifically includes:

[0060] 5.1. Train the classification model VGG16 using the target classification dataset in step 1;

[0061] Load the sample original image data through step 3.1 and extract the information in the target defect category label and the target defect position label. Steps 2, 3, and 4 build the framework of the complete deep reinforcement learning target detection model, and initialize the experience pool, which is used to store past action information;

[0062] 5.2. Select a sample original image, initialize the Mask size to the size of the entire image, and RC Vec stores the coordinates of the upper left corner and the lower right corner of the Mask. The historical action sequence H is a full-1 array with n rows and 10 columns, and the total reward R all is assigned 0, and the termination flag done is also 0;

[0063] 5.3. Calculate the initial IOU0 based on the IOU formula in step 4.1; adopt the feature extraction method in step 2 to extract the global feature, local feature, and historical action sequence respectively, and splice them into the initial state S;

[0064] 5.4. Use the state S as the input of the behavior decision network and the value evaluation network, and output the action probability distribution and the state value v; then the Agent randomly samples an action a based on the action probability distribution and calculates the logarithmic probability a of the selected action log ;

[0065] 5.5. Construct a one-hot vector according to the encoding method in step 2.5 and update the historical action sequence H;

[0066] 5.6. According to steps 2.7 to 2.10, execute the action a, update the Mask and RC Vec , and recalculate the IOU. Then, according to the reward function in steps 4.2 to 4.7, calculate the reward r and the termination flag done;

[0067] 5.7. Accumulate the single-step reward r and update the total reward R all ;

[0068] 5.8. Extract the updated image features as local features according to the feature extraction method in step 2, and combine them with the global features of the original image extracted in step 5.3 and the updated historical action sequence to construct the next state S_;

[0069] 5.9, (S, a, r, done, a log , v) is put into the experience pool as a complete action process;

[0070] 5.10. When the Agent has executed enough actions, the past experience and the next state S - will be used to train the Agent, that is, to train the deep reinforcement learning object detection model, and the experience pool is cleared;

[0071] 5.11. If the termination flag done = 1 or the upper limit of the number of iterations per single image is reached, the iterative detection of the current image is exited; after the updated local image undergoes the preprocessing operation in step 2.3, it is input into the pre-trained classification model VGG16 to identify the target defect category; then the next image is selected from the training set, and it jumps to step 5.2 to continue the loop; otherwise, the next state S - is used as the current state, and it jumps to step 5.4 to continue the loop;

[0072] 5.12. When all the images in the training set have been traversed; similarly, the images in the test set are traversed by the method in steps 5.2 to 5.11, where the training processes in steps 5.9 and 5.10 are skipped; after the test set is traversed, the mAP and Recall are calculated according to step 4.1;

[0073] 5.13. Steps 5.2 to 5.12 will be looped Eposide times, and the best detection model parameters are selected according to the object detection evaluation metrics to obtain the best deep reinforcement learning object detection model.

[0074] To achieve the above technical objectives, another technical solution adopted by the present invention is:

[0075] A particleboard object detection system based on deep reinforcement learning, comprising:

[0076] An object detection dataset processing module, which is used to obtain a particleboard defect sample image dataset. The particleboard defect sample image dataset includes sample original images and the corresponding target defect category labels and target defect position labels of the original images; it is also used to check whether the original images in the particleboard defect sample image dataset match the target defect category labels and target defect position labels, and it is also used to divide the sample original images and the target defect category labels into a training set and a test set as the target classification dataset for training the classification model, and divide the sample original images and the target defect position labels into a training set and a test set as the target detection dataset for training the deep reinforcement learning object detection model;

[0077] A state and action space definition module, which is used to construct the state space and action space of the reinforcement learning Agent;

[0078] The reinforcement learning algorithm construction module is used to design a data set loader and construct the framework of a deep reinforcement learning object detection model;

[0079] The metric and reward module is used to select object detection evaluation metrics and design a reward function;

[0080] The network training module is used to train a classification model using the divided object classification data set, and within the specified number of iterations, train the deep reinforcement learning object detection model using the divided object detection data set, update the network parameters, and select the best network parameters according to the object detection evaluation metrics to obtain the best deep reinforcement learning object detection model.

[0081] The beneficial effects of the present invention are as follows:

[0082] (1) For the particleboard object detection task, the present invention proposes a novel particleboard object detection method and detection system based on deep reinforcement learning, and completes the object detection task through continuous morphological transformation operations on the positioning frame (rectangular frame). Currently, existing single-stage detectors and two-stage detectors increase the detection performance at the cost of increasing and widening the network structure and increasing the parameter calculation amount. To address these problems, the particleboard object detection method using deep reinforcement learning is adopted. Its algorithm backbone network consists of only multiple fully connected layers, greatly reducing the parameters and the calculation amount. Moreover, by adjusting the morphological change amount and the number of iterations, it can adapt to objects of different sizes and achieve the effect of precise positioning.

[0083] (2) The present invention optimizes and improves the network architecture of the Agent in reinforcement learning. The linear module is composed of multiple linear fully connected layers, a dropout layer, a root mean square normalization layer, and a quadratic ReLU activation function. The network structure is clear and efficient. Among them, the root mean square normalization layer improves the training stability, the quadratic ReLU activation function enhances the non-linear expression ability of the network, and the dropout layer is used to prevent overfitting. Compared with traditional one-stage and two-stage detectors, its overall structure is more lightweight, has high computational efficiency, and has good generalization performance.

[0084] (3) The present invention proposes a state calculation method for multi-feature fusion, which fuses global features, local features, and historical action sequences. The global features are used to guide the Agent to explore the correct position, the local features are used to guide the Agent to refine the Mask, and the historical action sequences are used to optimize the action selection process of the Agent.

[0085] (4) The present invention proposes a multi - angle and phased reward function, which balances the single - step reward and the end - point reward. When taking actions, based on the change in IOU, the Agent is rewarded, which helps to locate the target faster; when the target is successfully detected, a relatively large reward value is given, which is beneficial for the Agent to clarify the target task and accelerate convergence; when the Mask after the action exceeds the range of the original image, a negative reward value is given, indicating that this action is a dangerous action, thus gradually avoiding it.

[0086] (5) The present invention can be applied to the defect (target) detection task of particleboard images (the types of defects include: large particles, glue spots, sand leakage, scratches, dust spots), can accurately detect the defect positions, has strong adaptive detection ability and high detection accuracy. The present invention is applicable to single - target detection, that is, there is one target on a particleboard image. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 It is a flow chart of the particleboard target detection method for deep reinforcement learning.

[0088] Figure 2 It is a training flow chart of the particleboard target detection method for deep reinforcement learning

[0089] Figure 3 It is a schematic diagram of the original image and the region image in the state space.

[0090] Figure 4 It is a schematic diagram of the actions in the action space.

[0091] Figure 5 It is a network framework diagram of the reinforcement learning algorithm.

[0092] Figure 6 It is a schematic diagram of particleboard defect detection by the method of the present invention.

[0093] Figure 7 It is a diagram of the detection results of other four kinds of defects. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0094] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings:

[0095] A particleboard target detection method for deep reinforcement learning, as Figure 1 shown, includes:

[0096] Step 1: Obtain a dataset of particleboard defect sample images, including the original sample images and the corresponding target defect category labels and target defect location labels; divide the original sample images and the target defect category labels into a training set and a test set according to a specified ratio, as the target classification dataset for training the classification model, and divide the original sample images and the target defect location labels into a training set and a test set according to a specified ratio, as the target detection dataset for training the deep reinforcement learning object detection model; the target defect categories include large particles, glue spots, sand leakage, scratches, and dust spots.

[0097] Step 2: Define the state space and action space of the reinforcement learning Agent (in reinforcement learning, an Agent is a "decision maker" that learns the optimal behavior strategy by continuously interacting with the environment); the state space consists of global features, local features, and the historical action sequence, which assist the Agent in making decisions; the action space includes translation operations, scaling operations, and termination actions, which act on the Mask for locating the target; the Mask is a rectangular box used to cover the target to be detected.

[0098] Step 3: Design a dataset loader and build the framework of the deep reinforcement learning object detection model.

[0099] Step 4: Combine the task objectives and design the object detection evaluation metrics and the reward function.

[0100] Step 5: Use the divided target classification dataset to train the classification model to obtain a trained classification model; use the divided target detection dataset to train the deep reinforcement learning object detection model, and combine the object detection evaluation metrics to select the best deep reinforcement learning object detection model as the particleboard object detection model.

[0101] The specific content of Step 1 includes:

[0102] 1.1. Obtain a dataset of particleboard defect sample images, including the original sample images, the target defect category labels, and the target defect location labels, check whether the labels correspond to the images, and ensure the accuracy of the label candidate boxes. 1.2. Randomly divide the dataset into a training set and a test set according to a specified ratio, keep the images and labels of each sample paired completely, and ensure that there is no data loss or incorrect matching.

[0103] The specific content of Step 2 includes:

[0104] 2.1. Define the state space of the reinforcement learning Agent: The state space consists of one-dimensional arrays of three parts, specifically including global features (such as Figure 3 the original image features shown), local features (such as Figure 3 the image features within the rectangular box Mask shown, that is, the regional image), and the historical action sequence.

[0105] 2.2. Design an image feature extractor: The image feature extractor is based on the ResNet101 network. After removing its last fully connected layer, it is used as the backbone network, and the pre-trained weight parameters are loaded, and the network is set to the evaluation mode.

[0106] 2.3. Global features: The original image is first preprocessed, that is, it is scaled to a fixed size using a scaling operation, and then normalized using the preset mean and standard deviation; the preprocessed image passes through the image feature extractor in step 2.2 to output global features.

[0107] 2.4. Local features: The image content within the Mask first undergoes the preprocessing operations of scaling and normalization in step 2.3. The preprocessed image passes through the image feature extractor in step 2.2 to output local features.

[0108] 2.5. Historical action sequence: Each action is encoded using one-hot encoding, that is, in the dimension of the action space, the position of the executed action is marked as 1, and the rest are 0; the sequence stores the encoding of the most recent n actions vertically, with the new action encoding at the top and the historical encodings moving down in sequence; the initial historical action sequence is an all-1 array;

[0109] A i =[0,1,0,0,…,0],H=[A1,A2,…,A 10 (1);

[0110] where, A i is the encoding of the i-th action, and H is the historical action sequence.

[0111] 2.6. Concatenate the global features in step 2.3, the local features in step 2.4, and the historical action sequence in step 2.5 along the length direction. After standardization, the state S of the Agent is obtained.

[0112] 2.7. Define the action space of the reinforcement learning Agent: The action space includes translation, scaling, and termination actions as shown in Figure 4 ; each time the Agent takes an action, it samples an action from the action space according to the probability distribution.

[0113] 2.8. As shown in Figure 4 , the scaling operation includes zooming in on all four sides, zooming out on all four sides, horizontal zooming in, horizontal zooming out, vertical zooming in, and vertical zooming out; the change in width and height for each action is α times the width and height corresponding to the current Mask (0 ≤ α ≤ 1);

[0114]

[0115] Among them, from left to right are the calculation formulas for the four-sided scaling, horizontal scaling, and vertical scaling operations. (x1, y1) is the coordinate of the upper left corner of the Mask, (x2, y2) is the coordinate of the lower right corner of the Mask, (x′1, y′1) is the coordinate of the upper left corner of the Mask after update, (x′2, y′2) is the coordinate of the lower right corner of the Mask after update, and W and H are the width and height of the Mask respectively.

[0116] 2.9. As Figure 4 shown, the translation operations include upward translation, downward translation, leftward translation, and rightward translation; the amount of change in width and height for each action is α times the corresponding width and height of the current Mask (0 ≤ α ≤ 1);

[0117]

[0118] Among them, from left to right are the calculation formulas for the horizontal movement (including leftward translation and rightward translation) and vertical movement (including upward translation and downward translation) operations; the other parameters are the same as those in step 2.8.

[0119] 2.10. The termination action indicates that the Agent ends the iteration of the current image with the current Mask as the final positioning box.

[0120] The specific steps of step 3 include:

[0121] 3.1. Design a dataset loader to read the sample image data, extract relevant parameters from the target defect category label and the target defect location label, and arrange them in a unified format: [class, x t1 , x t2 , y t1 , y t2 , where class is the value of the category to which the target in the image belongs, (x t1 , y t1 ) is the coordinate of the upper left corner of the true positioning box (true Mask) recorded by the target defect location label, and (x t2 , y t2 ) is the coordinate of the lower right corner of the true positioning box (true Mask) recorded by the target defect location label.

[0122] 3.2. Build a deep reinforcement learning framework, that is, the framework of the deep reinforcement learning object detection model, including a behavior decision network and a value evaluation network. Among them, the behavior decision network is used to perform action policy selection and baseline value estimation, and the value evaluation network is used to be responsible for the evaluation of the state value function.

[0123] 3.3. As Figure 5 shown, the behavior decision network consists of a backbone network, an action head, and a baseline value head; the backbone network consists of a linear layer, two linear modules, and a root mean square normalization layer.

[0124] 3.4. The linear module consists of a root mean square normalization layer, a linear layer, a quadratic ReLU activation function, a linear layer, and a dropout layer;

[0125] ReLUSquared(x) = (ReLU(x)) 2 (4);

[0126] where ReLUSquared is the quadratic ReLU activation function.

[0127] 3.5. The state S in step 2.6 is processed by the first linear layer of the backbone network in the action decision network, and the output feature A1 is obtained; the feature A1 is processed by the first linear module, and the output feature A2 is obtained; the residual connection combines A1 + A2 = A3; after being processed by the second linear module, the output feature A4 is obtained; the residual connection combines A3 + A4 = A5; and then through the root mean square normalization layer, the normalized feature A6 is obtained.

[0128] 3.6. The action head consists of a linear layer, a quadratic ReLU activation function layer, and a linear layer; the feature A6 in step 3.5 is processed by the action head to obtain the feature A7, and then through the SoftMax layer, the action probability distribution Action is output prob .

[0129] 3.7. The baseline value head consists of a linear layer, a quadratic ReLU activation function layer, and a linear layer; the feature A6 in step 3.5 is processed by the baseline value head to obtain the baseline value output A8.

[0130] 3.8. As Figure 5 shown, the value evaluation network consists of a backbone network and a state value head; the backbone network consists of a linear layer, six linear modules as described in step 3.4, and a root mean square normalization layer.

[0131] 3.9. The state S in step 2.6 is processed by the first linear layer of the backbone network in the value evaluation network, and the output feature B1 is obtained; the feature B1 is processed by the first linear module, and the output feature B2 is obtained; the residual connection combines B1 + B2 = B3; after being processed by the second linear module, the output feature B4 is obtained; the residual connection combines B3 + B4 = B5; after being processed by the third linear module, the output feature B6 is obtained; the residual connection combines B5 + B6 = B7; after being processed by the fourth linear module, the output feature B8 is obtained; the residual connection combines B7 + B8 = B9; after being processed by the fifth linear module, the output feature B 10 ; the residual connection combines B9 + B 10 = B 11 ; after being processed by the sixth linear module, the output feature B 12 ; the residual connection combines B 11+B 12 = B 13 ; Finally, through the root mean square normalization layer, the normalized feature B is obtained 14 .

[0132] 3.10. The state value head is a linear layer; the feature B in step 3.9 14 is processed by the state value head to obtain the state value output B 15 .

[0133] The specific steps of step 4 include:

[0134] 4.1. The target detection evaluation metrics adopt multi-class mean average precision mAP, recall Recall, and intersection over union IoU;

[0135]

[0136] Among them, C is the total number of classes, and AP i represents the mean average precision of the i-th class, where i is the class index; TP represents the number of correctly detected targets, and FN is the number of undetected targets; A and B represent two bounding boxes, specifically referring to the true Mask (rectangular box) recorded by the target defect position label and the Mask (rectangular box) calculated by the model; Area intersection is the intersection area of the two bounding boxes; Area union is the union area of the two bounding boxes; Area A , Area B are the areas of bounding boxes A and B respectively.

[0137] 4.2. The reward function needs to be set according to the task objective and is one of the cores of the reinforcement learning algorithm; the reward function is divided into normal action rewards and termination action rewards, and they are gradually superimposed.

[0138] 4.3. Normal actions include scaling actions and translation actions. When these actions are selected, first, it will be checked whether the Mask exceeds the range of the original image. If it does, a reward of -r b will be given.

[0139] 4.4. After the normal action is executed, if it exceeds the range of the original image, the Mask will be restricted within the correct range by the constraint conditions;

[0140]

[0141] Among them, (x1, y1) are the coordinates of the upper left corner of the Mask, (x2, y2) are the coordinates of the lower right corner of the Mask, (x′1, y′1) are the coordinates of the upper left corner of the Mask after update, (x′2, y′2) are the coordinates of the lower right corner of the Mask after update, max(a, b) and min(a, b) are used to obtain the maximum and minimum values of the array (a, b) respectively.

[0142] 4.5. After normal operation, the IOU between the Mask and the true Mask of the target defect position label is calculated through the formula in step 4.1; the IOU reward function is as follows:

[0143]

[0144] Among them, i is the index of the i-th action, 1, 2, 3...; β u is the weight of the positive reward, β d is the weight of the negative reward. The initial IOU is calculated from the initial Mask and the true Mask, and the size of the initial Mask is the size of the original image.

[0145] 4.6. When the termination action is selected, first set the termination flag done = 1, and give a stepped reward according to the IOU calculated by the previous action;

[0146]

[0147] 4.7. If the Agent selects the termination action when choosing the action for the first time, then give a reward of -r b .

[0148] As Figure 2 shown, step 5 specifically includes:

[0149] 5.1. Use the target classification dataset in step 1 to train the classification model VGG16. The specific training method of the classification model VGG16 adopts the existing technology.

[0150] Load the sample original image data through step 3.1 and extract the information in the target defect category label and the target defect position label. Steps 2, 3, and 4 build the framework of the complete deep reinforcement learning target detection model, and initialize the experience pool, which is used to store the past action information.

[0151] 5.2. Select a sample original image, initialize the Mask size to the size of the entire image, and RC Vec stores the coordinates of the upper left corner and the lower right corner of the Mask. The historical action sequence H is a full 1 array with n rows and 10 columns, and the total reward R all is assigned 0, and the termination flag done is also 0.

[0152] 5.3. Calculate the initial IOU0 based on the IOU formula in step 4.1; adopt the feature extraction method in step 2 to extract the global feature, local feature, and historical action sequence respectively, and splice them into the initial state S.

[0153] 5.4. Use the state S as the input of the action decision network and the value evaluation network, and output the action probability distribution and the state value v; then the Agent randomly samples an action a based on the action probability distribution and calculates the log probability of the selected action a log .

[0154] 5.5. Construct a one-hot vector according to the encoding method in step 2.5 and update the historical action sequence H.

[0155] 5.6. According to steps 2.7 to 2.10, execute the action a, update Mask and RC Vec , and recalculate the IOU, and then calculate the reward r and the termination flag done according to the reward function in steps 4.2 to 4.7.

[0156] 5.7. Accumulate the single-step reward r and update the total reward R all . These rewards are calculated and accumulated for each action of each image. After switching images, the total reward is reset to 0.

[0157] 5.8. According to the feature extraction method in step 2, extract the features of the updated image (the image within the updated Mask) as the local feature, and construct the next state S_ together with the global feature of the original image extracted in step 5.3 and the updated historical action sequence. The global feature refers to the feature of the original image, which is extracted only once at the beginning, and the local feature refers to the feature of the updated image.

[0158] 5.9. (S, a, r, done, a log , v) is put into the experience pool as a complete action process.

[0159] 5.10. When the Agent has executed enough actions, the past experience and the next state S - will be used to train the Agent, that is, to train the deep reinforcement learning object detection model, and the experience pool is cleared. Among them, the reward represents the quality of the actions made by the algorithm, and its role is to give a reference to the reinforcement learning algorithm and let it update in the direction of increasing the total reward. The reward is used to train the model in step 5.10; the baseline value output in step 3.7 and the state value output in step 3.10 are used in step 5.10 to calculate the loss, train the model, and update the parameters.

[0160] 5.11. If the termination flag done = 1 or the upper limit of the number of iterations for a single image is reached, exit the iterative detection of the current image; after the updated local image (the image within the updated Mask) undergoes the preprocessing operation in step 2.3, it is input into the pre-trained classification model VGG16 to identify the target defect category.

[0161] Then select the next image from the training set, jump to step 5.2, and continue the loop. Otherwise, use the following state S - as the current state, jump to step 5.4, and continue the loop.

[0162] 5.12. When all images in the training set have been traversed; similarly, traverse the images in the test set using the methods in steps 5.2 - 5.8 and 5.11; after the test set is traversed, calculate mAP and Recall according to step 4.1.

[0163] 5.13. Steps 5.2 to 5.12 will be looped Eposide times. According to the object detection evaluation metrics (mAP and Receall), select the best detection model parameters (parameters within the best behavior decision network and value evaluation network) to obtain the best deep reinforcement learning object detection model.

[0164] During actual prediction, the original image to be measured is processed in the same way as in steps 5.2 - 5.8 (during the process, the state S is input into the behavior decision network and value evaluation network of the best deep reinforcement learning object detection model) and step 5.11, and the finally updated local image is output, denoted as the target defect finally detected. After the updated local image undergoes the preprocessing operation in step 2.3, it is input into the pre-trained classification model VGG16 to identify the target defect category.

[0165] Figure 6 It is a schematic diagram of the process of detecting the target defect (adhesive spot) of the particleboard original image by the method of the present invention. Figure 7 It is a schematic diagram of the detection results of the other four defects (the target defects in the images within the red Mask from left to right are dust spots, sand leakage, scratches, and large particles). As Figure 6 - Figure 7 can be seen, the method of the present invention can be applied to the defect (target) detection task of particleboard images (the defect types include: large particles, adhesive spots, sand leakage, scratches, dust spots), can accurately detect the defect positions, has strong adaptive detection ability and high detection accuracy. The present invention is applicable to single-object detection, that is, there is one target on a particleboard image.

[0166] This embodiment also provides a particleboard object detection system based on deep reinforcement learning, including:

[0167] The target detection dataset processing module is used to obtain the particleboard defect sample image dataset, which includes the original sample images and the corresponding target defect category labels and target defect location labels of the original images; it is also used to check whether the original images in the particleboard defect sample image dataset match the target defect category labels and target defect location labels, and is also used to divide the original sample images and the target defect category labels into training sets and test sets as the target classification dataset for training the classification model, and divide the original sample images and the target defect location labels into training sets and test sets as the target detection dataset for training the deep reinforcement learning target detection model;

[0168] The state and action space definition module is used to construct the state space and action space of the reinforcement learning Agent;

[0169] The reinforcement learning algorithm construction module is used to design a dataset loader and construct the framework of the deep reinforcement learning target detection model;

[0170] The metric and reward module is used to select the target detection evaluation metrics and design the reward function;

[0171] The network training module is used to train the classification model using the divided target classification dataset, and within the specified number of iterations, use the divided target detection dataset to train the deep reinforcement learning target detection model, update the network parameters, and select the best network parameters according to the target detection evaluation metrics to obtain the best deep reinforcement learning target detection model.

[0172] The protection scope of the present invention includes but is not limited to the above embodiments. The protection scope of the present invention is subject to the claims, and any substitutions, deformations, and improvements that are easily conceivable by those skilled in the art to this technology fall within the protection scope of the present invention.

Claims

1. A particleboard target detection method based on deep reinforcement learning, characterized in that: include: Step 1: Obtain a particleboard defect sample image dataset, including the sample original image and the target defect category label and target defect position label corresponding to the original image; The sample original images and target defect category labels are divided into training sets and test sets as the target classification dataset for training the classification model; the sample original images and target defect position labels are divided into training sets and test sets as the target detection dataset for training the deep reinforcement learning target detection model; Step 2: Define the state space and action space of the reinforcement learning agent; the state space consists of global features, local features and historical action sequences to assist the agent in making decisions; the action space includes translation operations, scaling operations and termination operations, acting on the mask of the positioning target; Step 3: Design a dataset loader and build a framework for the deep reinforcement learning object detection model; Step 4: Design target detection evaluation indicators and reward functions; Step 5: Use the divided target classification data set to train the classification model to obtain a trained classification model; use the divided target detection data set to train the deep reinforcement learning target detection model, and combine the target detection evaluation indicators to select the best deep reinforcement learning target detection model as the particleboard target detection model.

2. The particleboard target detection method based on deep reinforcement learning according to claim 1, characterized in that: The step 2 specifically includes: 2.

1. Define the state space of the reinforcement learning agent: The state space consists of three one-dimensional arrays, including global features, local features, and historical action sequences; 2.

2. Design image feature extractor: The image feature extractor is based on the ResNet101 network. After removing its last fully connected layer, it is used as the backbone network, the pre-trained weight parameters are loaded, and the network is set to evaluation mode; 2.

3. Global features: The original image is first preprocessed, that is, it is scaled to a fixed size and then normalized using a preset mean and standard deviation. The preprocessed image is passed through the image feature extractor in step 2.2 to output global features. 2.

4. Local features: The image content in the Mask is first preprocessed by scaling and normalization in step 2.

3. The preprocessed image is passed through the image feature extractor in step 2.2 to output local features. 2.

5. Historical action sequence: Each action is one-hot encoded, that is, in the action space dimension, the executed action position is marked as 1, and the rest are 0; the sequence vertically stores the latest n action codes, with the new action code at the top and the historical codes moving down in sequence; the initial historical action sequence is an array of all 1s; A i =[0,1,0,0,…,0],H=[A1,A2,…,A 10 ](1); Among them, A i is the encoding of the ith action, H is the historical action sequence; 2.

6. The global features of step 2.3, the local features of step 2.4 and the historical action sequence of step 2.5 are concatenated in length direction and standardized to obtain the state S of the Agent. 2.

7. Define the action space of the reinforcement learning agent: The action space includes translation operations, scaling operations, and termination actions. Each time the agent takes an action, it samples an action from the action space according to the probability distribution. 2.

8. Zooming operations include zooming in and out, zooming in and out horizontally, zooming in and out horizontally, zooming in and out vertically, and zooming in and out vertically. The width and height change of each action is α times the width and height of the current Mask, 0≤α≤1. Among them, from left to right are the calculation formulas for the four-way scaling, horizontal scaling, and vertical scaling operations. (x1, y1) is the coordinate of the upper left corner of the Mask, (x2, y2) is the coordinate of the lower right corner of the Mask, (x1 ′ ,y1 ′ ) is the coordinate of the upper left corner after the Mask is updated, (x ′ 2,y2 ′ ) is the coordinate of the lower right corner of the Mask after update, W and H are the width and height of the Mask respectively; 2.

9. Translation operations include horizontal and vertical movements. Vertical movements include upward and downward movements. Horizontal movements include left and right movements. The width and height change of each movement is α times the width and height of the current Mask, 0≤α≤1. Among them, from left to right are the calculation formulas for horizontal movement and vertical movement operations; 2.

10. The termination action means that the Agent ends the iteration of the current image with the current Mask as the final positioning box.

3. The particleboard target detection method based on deep reinforcement learning according to claim 1, characterized in that: The step 3 specifically includes: 3.

1. Design a dataset loader to read sample image data, extract relevant parameters from the target defect category label and target defect location label, and arrange them in a unified format: [class, x t1 ,x t2 ,y t1 ,y t2 ], class is the class value of the target in the image, (x t1 ,y t1 ) is the upper left corner coordinate of the real positioning frame recorded by the target defect position label, (x t2 ,y t2 ) is the lower right corner coordinate of the real positioning frame recorded by the target defect position label; 3.

2. Build a framework for deep reinforcement learning target detection model, including a behavior decision network and a value evaluation network. The behavior decision network is used to perform action strategy selection and baseline value estimation, and the value evaluation network is responsible for the evaluation of the state value function. 3.

3. The behavior decision network consists of a backbone network, an action head, and a baseline value head. The backbone network in the behavior decision network consists of a linear layer, two linear modules, and a root mean square normalization layer. 3.4, the linear module consists of a root mean square normalization layer, a linear layer, a quadratic ReLU activation function, a linear layer and a random dropout layer; ReLUSquared(x)=(ReLU(x)) 2 (4); Among them, ReLUSquared is the quadratic ReLU activation function; 3.

5. The state S in step 2.6 is processed by the first linear layer of the backbone network in the behavior decision network, and the output feature A1 is output; the feature A1 is processed by the first linear module, and the output feature A2 is output; the residual connection merges A1+A2=A3; the residual connection merges A3+A4=A5; and the normalized feature A6 is obtained by the RMS normalization layer; 3.

6. The action head contains a linear layer, a quadratic ReLU activation function layer and a linear layer. Feature A6 in step 3.5 is processed by the action head to obtain feature A7, which is then passed through the SoftMax layer to output the action probability distribution Action. prob ; 3.

7. The baseline value head includes a linear layer, a quadratic ReLU activation function layer and a linear layer. The feature A6 in step 3.5 is processed by the baseline value head to obtain the baseline value output A8. 3.

8. The value evaluation network consists of a backbone network and a state value head; the backbone network in the value evaluation network consists of a linear layer, six linear modules described in step 3.4, and a root mean square normalization layer; 3.

9. The state S of step 2.6 is processed by the first linear layer of the backbone network in the value evaluation network, and outputs feature B1; feature B1 is processed by the first linear module, and outputs feature B2; residual connection merges B1+B2=B3; after processing by the second linear module, output feature B4; residual connection merges B3+B4=B5; after processing by the third linear module, output feature B6; residual connection merges B5+B6=B7; after processing by the fourth linear module, output feature B8; residual connection merges B7+B8=B9; after processing by the fifth linear module, output feature B 10 ; Residual connection merge B9+B 10 =B 11 ; After being processed by the sixth linear module, the output feature B 12 ; Residual connection merge B 11 +B 12 =B 13 ; Finally, after the RMS normalization layer, the normalized feature B is obtained 14 ; 3.10, the state value head is a linear layer; feature B in step 3.9 14 After the state value header is processed, the state value output B is obtained 15 .

4. The particleboard target detection method based on deep reinforcement learning according to claim 1, characterized in that: The step 4 specifically includes: 4.

1. The target detection evaluation indicators include multi-category average precision mAP, recall rate Recall and intersection over union (IoU); Among them, C is the total number of categories, AP i represents the average precision of the i-th category, i is the category index; TP represents the number of correctly detected targets, and FN represents the number of undetected targets; A and B represent two bounding boxes; Area intersection is the intersection area of ​​the two bounding boxes; Area union Area is the union area of ​​the two bounding boxes; A ,Area B are the areas of bounding boxes A and B respectively; 4.

2. The reward function is divided into normal action reward and termination action reward, and they are gradually superimposed; 4.

3. Normal actions include zooming and panning. When these actions are selected, the mask will first be checked to see if it exceeds the original image range. If it exceeds, -r will be given. b Rewards; 4.

4. After the normal action is executed, if it exceeds the range of the original image, the Mask will be restricted to the correct range by the constraints; Among them, (x1, y1) is the coordinate of the upper left corner of the Mask, (x2, y2) is the coordinate of the lower right corner of the Mask, (x1 ′ ,y1 ′ ) is the coordinate of the upper left corner after the Mask is updated, (x ′ 2,y2 ′ ) is the coordinate of the lower right corner after the Mask is updated, max(a,b) and min(a,b) are used to obtain the maximum and minimum values ​​of the array (a,b) respectively; 4.

5. The IOU of the Mask after normal action and the real Mask of the target defect position label is calculated by the formula in step 4.1; the IOU reward function is as follows: Where i is the index of the i-th action, 1, 2, 3...; β u is the weight of positive reward, β d is the weight of the negative reward. The initial IOU is calculated by the initial Mask and the true Mask. The initial Mask size is the original image size. 4.

6. When the termination action is selected, first set the termination flag done = 1, and give a step-by-step reward based on the IOU calculated for the previous action; 4.

7. If the agent chooses the termination action when it first chooses an action, it is given -r b Reward.

5. The particleboard target detection method based on deep reinforcement learning according to claim 1, characterized in that: The step 5 specifically includes: 5.

1. Use the target classification dataset in step 1 to train the classification model VGG16; Through step 3.1, load the sample original image data and extract the information in the target defect category label and the target defect location label. Steps 2, 3, and 4 build a complete framework of the deep reinforcement learning target detection model and initialize the experience pool, which is used to store past action information. 5.

2. Select a sample original image and initialize the Mask size to the size of the entire image. RC Vec Store the coordinates of the upper left corner and the lower right corner of the Mask. The historical action sequence H is an array of n rows and 10 columns with all 1s. The total reward R all The value is assigned to 0, and the termination flag done is also 0; 5.

3. Calculate the initial IOU0 based on the IOU formula in step 4.1; use the feature extraction method in step 2 to extract global features, local features and historical action sequences respectively, and splice them into the initial state S; 5.

4. The state S is used as the input of the behavior decision network and the value evaluation network, and the output is the action probability distribution and the state value v; then the agent randomly samples the action a based on the action probability distribution and calculates the logarithmic probability a of the selected action log ; 5.

5. Construct a one-hot vector according to the encoding method of step 2.5 and update the historical action sequence H; 5.

6. According to steps 2.7 to 2.10, perform action a to update Mask and RC Vec , and recalculate the IOU, and then calculate the reward r and the termination flag done according to the reward function from step 4.2 to step 4.7; 5.

7. Accumulate single-step rewards r and update total rewards R all ; 5.

8. According to the feature extraction method in step 2, extract the updated image features as local features, and combine them with the global features of the original image extracted in step 5.3 and the updated historical action sequence to construct the next state S_; 5.9、(S,a,r,done,a log ,v) is put into the experience pool as a complete action process; 5.

10. When the agent performs enough actions, past experience and the next state S - It will be used to train the Agent, that is, to train the deep reinforcement learning target detection model and clear the experience pool; 5.

11. If the termination flag done = 1 or the upper limit of the number of iterations for a single image is reached, the iterative detection of the current image is exited; the updated local image is preprocessed in step 2.3 and input into the pre-trained classification model VGG16 to identify the target defect category; then the next image is selected from the training set, and the process jumps to step 5.2 to continue the cycle; otherwise, the next state S is used. - As the current state, jump to step 5.4 and continue the loop; 5.

12. When all images in the training set have been traversed, traverse the images in the test set in the same way as steps 5.2-5.8 and 5.

11. After traversing the images in the test set, calculate mAP and Recall according to step 4.

1. 5.

13. Steps 5.2 to 5.12 will be cycled Eposide times, and the optimal detection model parameters will be selected according to the target detection evaluation index to obtain the optimal deep reinforcement learning target detection model.

6. A particleboard target detection system based on deep reinforcement learning, characterized in that: include: The target detection data set processing module is used to obtain a particleboard defect sample image data set, which includes a sample original image and a target defect category label and a target defect position label corresponding to the original image; it is also used to check whether the original image in the particleboard defect sample image data set matches the target defect category label and the target defect position label, and is also used to divide the sample original image and the target defect category label into a training set and a test set as a target classification data set for training a classification model, and divide the sample original image and the target defect position label into a training set and a test set as a target detection data set for training a deep reinforcement learning target detection model; State and action space definition module, used to construct the state space and action space of reinforcement learning agent; A reinforcement learning algorithm building module, which is used to design a dataset loader and build a framework for deep reinforcement learning target detection models; The indicator and reward module is used to select target detection evaluation indicators and design reward functions; The network training module is used to train the classification model using the divided target classification data set, and to train the deep reinforcement learning target detection model using the divided target detection data set within a specified number of iterations, update the network parameters, and select the optimal network parameters based on the target detection evaluation index to obtain the optimal deep reinforcement learning target detection model.

Citation Information

Patent Citations

  • Shaving board surface defect detection method self-adaptive to board thickness

    CN114511503A

Cited By

  • Weld joint intelligent defect detection model training method, detection method and electronic equipment

    CN122199516A

  • Weld intelligent defect detection model training method, detection method and electronic equipment

    CN122199516B