Model training method and device, computer equipment and computer readable storage medium

By performing specific model training methods on deep learning models, the prediction results of the initial defect interception strategy model and the initial defect interception value model are adjusted, which solves the problems of large error and low accuracy in defect product interception, and achieves higher interception prediction accuracy.

CN120198747APending Publication Date: 2025-06-24SHENZHEN SMARTMORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200549.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When detecting defective products, traditional deep learning models cannot accurately determine whether an interception operation is required based on the defect situation and degree of defects, resulting in large interception errors and low accuracy of product defects.

Method used

Through a model training method, the product defect training image and corresponding training defect features are input into the initial defect interception strategy model and the initial defect interception value model for processing, output the predicted interception label and predicted interception value, adjust the model until the convergence conditions are met, and the target defect interception strategy model and the target defect interception value model are obtained.

Benefits of technology

It improves the accuracy of the model's interception prediction of defective products, reduces interception errors, and improves the accuracy of product detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198747A_ABST
    Figure CN120198747A_ABST
Patent Text Reader

Abstract

The invention relates to a model training method and device, computer equipment and a computer readable storage medium. The method comprises the following steps: respectively obtaining a product defect training image, a training defect feature and a training interception label; inputting the training defect features into the initial defect interception strategy model, and outputting prediction interception labels corresponding to the training defect features; inputting the training defect features into an initial defect interception value model, and outputting predicted interception values corresponding to the training defect features; determining a training interception value based on the prediction interception label, the training interception label and the training defect feature; adjusting the initial defect interception strategy model based on the predicted interception value, the training interception value and the predicted interception label; and adjusting the initial defect interception value model based on the predicted interception value and the training interception value until a convergence condition is met to obtain a target defect interception strategy model and a target defect interception value model. By adopting the method, the accuracy of intercepting and predicting the defective product by the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a model training method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] In recent years, with the rapid development of deep learning models, deep learning models can be applied to the field of appearance detection. Usually, a deep learning model is used to detect defects in products and intercept the detected defective products.

[0003] However, traditional deep learning models uniformly intercept the detected defective products and cannot accurately determine whether an interception operation is required based on the defect situation (such as complex defect features) and the degree of defect (such as minor defects), resulting in large interception errors and low accuracy for product defects. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a model training method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of the model's interception prediction for defective products.

[0005] In a first aspect, the present application provides a model training method, including:

[0006] Obtaining product defect training images, training defect features corresponding to the product defect training images, and training interception labels corresponding to the training defect features respectively;

[0007] Inputting the training defect features into an initial defect interception policy model for processing, and outputting a predicted interception label corresponding to the training defect features;

[0008] Inputting the training defect features into an initial defect interception value model for processing, and outputting a predicted interception value corresponding to the training defect features;

[0009] Determining the training interception value based on the predicted interception label, the training interception label, and the training defect features;

[0010] Adjusting the initial defect interception policy model respectively based on the predicted interception value, the training interception value, and the predicted interception label; adjusting the initial defect interception value model based on the predicted interception value and the training interception value until the convergence condition is met, and obtaining a target defect interception policy model and a target defect interception value model.

[0011] In a second aspect, the present application provides a model training apparatus, including:

[0012] A data acquisition module, configured to respectively obtain product defect training images, training defect features corresponding to the product defect training images, and training interception labels corresponding to the training defect features;

[0013] An interception prediction module, configured to input the training defect features into an initial defect interception policy model for processing, and output prediction interception labels corresponding to the training defect features;

[0014] A value prediction module, configured to input the training defect features into an initial defect interception value model for processing, and output prediction interception values corresponding to the training defect features;

[0015] A value calculation module, configured to determine the training interception value based on the prediction interception labels, the training interception labels, and the training defect features;

[0016] A model update module, configured to respectively adjust the initial defect interception policy model based on the prediction interception value, the training interception value, and the prediction interception labels; and adjust the initial defect interception value model based on the prediction interception value and the training interception value until the convergence condition is met, so as to obtain a target defect interception policy model and a target defect interception value model.

[0017] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method are implemented.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0019] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the above method are implemented.

[0020] The above model training method, device, computer device, computer-readable storage medium, and computer program product input training defect features into an initial defect interception policy model and an initial defect interception value model respectively to obtain a predicted interception label and a predicted interception value corresponding to the training defect features, representing the interception action of whether to intercept the defective product to which the training defect features belong and the predicted interception value; then determine the training interception value according to the predicted interception label, the training interception label, and the training defect features, representing the actual generated interception value determined by the real interception action and the predicted interception action; then use the predicted interception value and the training interception value to adjust the initial defect interception value model, and use the predicted interception value, the training interception value, and the predicted interception label to adjust the initial defect interception policy model, which can, during the process of model iteration, participate the predicted interception value output by the simultaneously trained initial defect interception value model in the next training of the initial defect interception policy model, ensuring the accuracy of the predicted interception value output by the initial defect interception value model, thereby improving the training accuracy of the initial defect interception policy model and the interception prediction accuracy of the target defect interception policy model for defective products. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. is an application environment diagram of a model training method provided by an embodiment of the present application;

[0022] Figure 2 FIG. is a flowchart of a model training method provided by an embodiment of the present application;

[0023] Figure 3 FIG. is a flowchart of an industrial defect detection based on reinforcement learning provided by an embodiment of the present application;

[0024] Figure 4 FIG. is a structural block diagram of a model training device provided by an embodiment of the present application;

[0025] Figure 5 FIG. is an internal structure diagram of a computer device provided by an embodiment of the present application;

[0026] Figure 6 FIG. is another internal structure diagram of a computer device provided by an embodiment of the present application;

[0027] Figure 7 FIG. is an internal structure diagram of a computer-readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0029] The model training method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a communication network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0030] As Figure 2 shown, the embodiments of the present application provide a model training method. Taking the method applied to the Figure 1 terminal 102 or the server 104 in it as an example for illustration. It can be understood that the computer device can include at least one of the terminal and the server. The method includes the following steps:

[0031] S202, respectively obtain product defect training images, training defect features corresponding to the product defect training images, and training interception labels corresponding to the training defect features.

[0032] Among them, the product defect training image refers to an image collected from a defective product and used as a training image for extracting defect features. The training defect feature is a defect feature extracted from the product defect in the product defect training image and used as a feature for model training. The training defect feature can refer to the attribute description information of the product defect, such as the length, depth, position coordinates, etc. of the product defect. The training interception label refers to the information on whether to intercept the product defect in the product defect training image.

[0033] Specifically, the server obtains each product defect training image sent by the terminal. Each product defect training image includes at least one product defect and the corresponding interception label for the product defect. The interception label can be information indicating whether to intercept or not, which is determined in advance according to the actual demand by judging the severity of the product defect and then marked according to the severity. Generally, for product defects judged to have a higher severity, the marked labels can be "intercept", "NG", etc., and for product defects judged to have a lower severity, the marked labels can be "do not intercept", "OK", etc.

[0034] After the server obtains each product defect training image, it extracts the defect features corresponding to the product defects from the product defect training image as training defect features, and takes the interception label corresponding to the product defect in the product defect training image as the training interception label corresponding to the training defect features.

[0035] S204. Input the training defect features into the initial defect interception policy model for processing, and output the predicted interception label corresponding to the training defect features.

[0036] Among them, the initial defect interception policy model refers to the defect interception policy model to be trained. The defect interception policy model is a decision-making model that predicts the probability of the product defect to which the input defect features belong under each interception policy. The interception policy can be understood as the interception action taken according to the defect features, such as the probability of interception and the probability of non-interception. Generally, the interception action with a higher probability value is used as the predicted interception label output by the defect interception policy model.

[0037] Specifically, the server randomly initializes the decision model network to obtain the initial defect interception policy model, which includes a feature extraction layer and a classification layer. The server inputs the training defect features into the initial defect interception policy model. The feature extraction layer extracts vector features from the training defect features, and the classification layer outputs the probabilities of different types of interception actions 0 and 1. 0 indicates non-interception, 1 indicates interception, and then the interception action 0 or 1 with a higher probability value is used as the training interception label.

[0038] S206. Input the training defect features into the initial defect interception value model for processing, and output the predicted interception value corresponding to the training defect features.

[0039] Among them, the initial defect interception value model refers to the defect interception value model to be trained. The defect interception value model is a prediction model that evaluates the interception value of a defect feature based on the input defect feature. The interception value is a representation value indicating the quality of taking a certain interception action according to the defect feature, which can be understood as the feedback information on whether a certain interception action meets the actual interception expectation. For example, when the interception action is good, it means that the interception action meets the actual interception expectation, that is, the interception value of the interception action is positive feedback information; when the interception action is bad, it means that the interception action does not meet the actual interception expectation, that is, the interception value of the interception action is negative feedback information. The predicted interception value can be understood as the feedback information when taking a certain interception action predicted by the initial defect interception value model according to the defect feature. For example, the predicted interception value is the positive feedback information or negative feedback information predicted by the initial defect interception value model according to the defect feature. Specifically, the server randomly initializes the prediction model network to obtain the initial defect interception value model, and the initial defect interception value model includes a feature extraction layer and a fully connected layer. The server inputs the training defect feature into the initial defect interception value model, extracts the vector feature of the training defect feature through the feature extraction layer, and outputs the predicted interception value corresponding to the training defect feature through the fully connected layer.

[0040] In some embodiments, the initial defect interception policy model can share the feature extraction layer with the initial defect interception value model, while the classification heads corresponding to the initial defect interception policy model and the initial defect interception value model are different. The server inputs the training defect feature into the feature extraction network for feature vector extraction, and then outputs the training interception label through the classification head (which can be a SoftMax classification head) corresponding to the initial defect interception policy model, and outputs the training interception value through the classification head (fully connected layer) corresponding to the initial defect interception value model.

[0041] S208. Determine the training interception value based on the predicted interception label, the training interception label, and the training defect feature.

[0042] Among them, the training interception value can be understood as the real feedback information on whether the interception action corresponding to the predicted interception label meets the true interception expectation corresponding to the training interception label.

[0043] Specifically, after the server obtains the predicted interception label through the initial defect interception policy model, it compares the predicted interception label with the training interception label, obtains the corresponding weighting parameter according to the comparison result, and calculates based on the weighting parameter and the training defect feature to obtain the training interception value.

[0044] S210. Adjust the initial defect interception policy model respectively based on the predicted interception value, the training interception value, and the predicted interception label; adjust the initial defect interception value model based on the predicted interception value and the training interception value until the convergence condition is met, and obtain the target defect interception policy model and the target defect interception value model.

[0045] Specifically, the server calculates the policy model loss corresponding to the initial defect interception policy model according to the difference between the training interception value and the predicted interception value and the predicted interception label, and updates the initial defect interception policy model according to the policy model loss to obtain an updated defect interception policy model; and synchronously calculates the value model loss corresponding to the initial defect interception value model according to the difference between the training interception value and the predicted interception value, and updates the initial defect interception value model according to the value model loss to obtain an updated defect interception value model.

[0046] Then, take the updated defect interception policy model as the initial defect interception policy model, take the updated defect interception value model as the initial defect interception value model, and return to the step of inputting the training defect feature into the initial defect interception policy model to obtain the predicted interception label corresponding to the training defect feature, and execute until the convergence condition is met to obtain the target defect interception policy model and the target defect interception value model.

[0047] It can be seen that in the embodiments of the present application, by inputting the training defect feature into the initial defect interception policy model and the initial defect interception value model respectively, the predicted interception label and the predicted interception value corresponding to the training defect feature are obtained, indicating the interception action predicted for the defective product to which the training defect feature belongs and the predicted interception value; then, according to the predicted interception label, the training interception label, and the training defect feature, the training interception value is determined, indicating the actual generated interception value determined by the real interception action and the predicted interception action; then, the initial defect interception value model is adjusted using the predicted interception value and the training interception value, and the initial defect interception policy model is adjusted using the predicted interception value, the training interception value, and the predicted interception label, so that during the model iteration process, the predicted interception value output by the synchronously trained initial defect interception value model can participate in the next training of the initial defect interception policy model, ensuring the accuracy of the predicted interception value output by the initial defect interception value model, thereby improving the training accuracy of the initial defect interception policy model and the interception prediction accuracy of the target defect interception policy model for defective products.

[0048] In some embodiments, S208. Determining the training interception value based on the predicted interception label, the training interception label, and the training defect feature includes:

[0049] Obtain the weighted parameters corresponding to each feature dimension in the training defect features; weight the training defect features based on the weighted parameters to obtain weighted training defect features;

[0050] When the predicted interception label is consistent with the training interception label, calculate the positive interception value based on the weighted training defect features to obtain the training interception value; or,

[0051] When the predicted interception label is inconsistent with the training interception label, calculate the negative interception value based on the weighted training defect features to obtain the training interception value.

[0052] Among them, the positive interception value is the interception value greater than zero, representing a better positive result or expectation, that is, the interception action meets the actual interception expectation. The negative interception value is the interception value less than zero, representing a worse negative result or expectation, that is, the interception action does not meet the actual interception expectation. The feature dimension refers to the attribute dimension of each feature in the training defect features, such as dimensions like length, width, area, etc.

[0053] Specifically, the server obtains the weighted parameters corresponding to each feature dimension in the training defect features, and weights the defect features of the corresponding feature dimensions in the training defect features according to the weighted parameters to obtain weighted training defect features. Generally, the weighted parameter of a feature dimension is positively correlated with the importance degree of that feature dimension, that is, the higher the importance degree of the feature dimension with a larger weighted parameter. The weighted parameter can be a weight parameter set according to the feature importance degrees of each feature dimension after judging the feature importance degrees of each feature dimension according to actual needs.

[0054] Then the server compares the training interception label with the predicted interception label. When the predicted interception label is the same as the training interception label, it means that the predicted interception label output by the initial defect interception policy model is the same as the training interception label, that is, the interception action corresponding to the predicted interception label meets the actual interception expectation corresponding to the training interception label. Calculate the positive interception value (i.e., positive feedback information) according to the weighted training defect, and obtain the training interception value when the predicted interception label is the same as the training interception label. Among them, the positive interception value is used to give positive rewards to the initial defect interception policy model, so that the initial defect interception policy model can perform positive learning based on the currently predicted correct interception label. Or, when the predicted interception label is different from the training interception label, it means that the predicted interception label output by the initial defect interception policy model is different from the training interception label, that is, the interception action corresponding to the predicted interception label does not meet the actual interception expectation corresponding to the training interception label. Calculate the negative interception value (i.e., negative feedback information) according to the weighted training defect, and obtain the training interception value when the predicted interception label is different from the training interception label. Among them, the negative interception value is used to give negative rewards to the initial defect interception policy model, so that the initial defect interception policy model can perform reverse learning based on the currently predicted incorrect interception label. The training interception value can be calculated using a reward function, and the reward function is used to feedback the pros and cons of the interception action to help the model learn. The reward function is shown in formula (1):

[0055] (1)

[0056] Where Rt is the training interception value (which can be understood as the model reward); St represents the training defect feature; yt represents the training interception label; at represents the predicted interception label; λ is an adjustment coefficient used to control the influence degree of the defect feature on the model reward; is a weighting function (i.e., weighting parameter) used to calculate the increase and decrease effect of important defect features on the model reward, and can be determined according to the actual importance degree of each training feature. Formula (1) means that if the interception action predicted by the model is the same as the actual label (interception label), a positive reward is given, and if the interception action predicted by the model is different from the actual label (interception label), a negative reward is given.

[0057] It can be seen that in this embodiment, by weighting important training defect features and calculating the corresponding model reward (training interception value) according to the comparison result of the training label and the predicted label, the model can learn important defect features and improve the interception prediction accuracy of product defects.

[0058] In some embodiments, S210, based on the predicted interception value, the training interception value, and the predicted interception label, adjust the initial defect interception policy model until the convergence condition is met to obtain the target defect interception policy model, including:

[0059] Calculate the value difference between the predicted interception value and the training interception value;

[0060] When the predicted interception label is within the preset label value range, determine the policy model loss corresponding to the initial defect interception policy model based on the value difference and the predicted interception label; or,

[0061] When the predicted interception label is not within the preset label value range, determine the policy model loss corresponding to the initial defect interception policy model based on the preset label value range and the value difference;

[0062] Update the initial defect interception policy model based on the policy model loss until the convergence condition is met, and obtain the target defect interception policy model.

[0063] Wherein, the preset label value range is a range used to limit the size of the policy model loss determined according to the predicted interception label. The policy model loss is loss information used to update the initial defect interception policy model. The range values corresponding to the preset label value range include a range upper limit value and a range lower limit value. The value difference refers to the difference between the predicted interception value and the training interception value.

[0064] Specifically, after the server obtains the predicted interception label and the predicted interception value, it calculates the value difference between the predicted interception label and the predicted interception value. The calculation of the value difference is shown in formula (2):

[0065] (2)

[0066] Wherein, is the value difference; is the training interception value; is the predicted interception value.

[0067] Then update the initial defect interception policy model and the initial defect interception value model. Updating the initial defect interception policy model includes: obtaining the preset label value range, determining whether the training interception label is within the preset label value range. If it is, calculate the policy model loss corresponding to the initial defect interception policy model according to the value difference and the predicted interception label; if the predicted interception label is greater than the range upper limit value of the preset label value range, calculate the policy model loss corresponding to the initial defect interception policy model according to the value difference and the range upper limit value; if the predicted interception label is less than the range lower limit value of the preset label value range, calculate the policy model loss corresponding to the initial defect interception policy model according to the value difference and the range lower limit value, and update the initial defect interception policy model according to the policy model loss until the convergence condition is met, and obtain the target defect interception policy model. The calculation of the policy model loss is shown in formula (3).

[0068] (3)

[0069] Wherein, is the loss of the policy model; is the value difference; is the predicted interception label; represents restricting the elements in the array to the specified minimum value and the maximum value . ([[]]END]] , ) is the preset label value range, is the lower limit value of the range, is the upper limit value of the range. is the clipping amplitude, which can be determined according to the adaptation degree of the training scenario and the model loss. For example, if the loss is always large, then reduce . is used to constrain the updated interception policy to be close to the current interception policy to prevent instability caused by excessive updates and can limit the amplitude of model updates. Among them, formula (3) means that when the predicted interception label is within the preset label value range, the loss calculated using the predicted interception label is the same as the loss calculated using the range value of the preset label value range, and the loss calculated using the predicted interception label can be optionally used as the loss of the policy model; when the predicted interception label is outside the preset label value range, the loss calculated using the range value of the preset label value range is always less than the loss calculated using the predicted interception label, so the loss calculated using the range value of the preset label value range is selected as the loss of the policy model.

[0070] It can be seen that in this embodiment, by restricting the amplitude of model updates through the preset label value range, the update accuracy of the initial defect interception policy model is ensured.

[0071] In some embodiments, in S210, based on the predicted interception value and the training interception value, the initial defect interception value model is adjusted until the convergence condition is met to obtain the target defect interception value model, including:

[0072] Calculating the value difference between the predicted interception value and the training interception value;

[0073] Based on the value difference, determining the value model loss corresponding to the initial defect interception value model;

[0074] Updating the initial defect interception value model based on the value model loss until the convergence condition is met to obtain the target defect interception value model.

[0075] Wherein, the value model loss is the loss information used to update the initial defect interception value model.

[0076] Specifically, the server updates the initial defect interception value model. It can calculate the value difference between the training interception value and the predicted interception value, then calculate the square value of the value difference to obtain the value model loss corresponding to the initial defect interception value model, and update the initial defect interception value model according to the value model loss until the convergence condition is met to obtain the target defect interception value model. The calculation of the value model loss is shown in formula (4):

[0077] (4)

[0078] Wherein, is the value model loss; is the training interception value; is the predicted interception value.

[0079] It can be seen that in this embodiment, updating the initial defect interception value model according to the value model loss ensures the accuracy of the defect interception value model for the predicted interception value.

[0080] In some embodiments, in S202, obtaining the product defect training image, the training defect feature corresponding to the product defect training image, and the training interception label corresponding to the training defect feature respectively includes:

[0081] Obtain the product defect training image, input the product defect training image into the defect detection model for processing, and output the defect type and defect area in the product defect training image;

[0082] Obtain the target feature dimension corresponding to the defect type, and extract the training defect feature corresponding to the target feature dimension and the training interception label corresponding to the training defect feature from the defect area.

[0083] Wherein, the target feature dimension refers to the feature attribute of the feature to be extracted determined according to the defect type, and different defect types correspond to different feature dimensions.

[0084] Specifically, the server obtains the product defect training image. The product defect training image can be a pre-determined image with product defects. Then, the product defect training image is input into the defect detection model. The defect detection model identifies the defect area with defects in the product defect training image and the defect type of the defect, such as defect types like scratches and different colors. Then, the corresponding target feature dimension is determined according to the defect type, and the corresponding defect feature is extracted from the defect area according to the target feature dimension as the training defect feature corresponding to the product defect training image, and the manual label corresponding to the product defect is obtained from the product defect training image as the training interception label corresponding to the training defect feature.

[0085] In some embodiments, for an input workpiece image to be detected, the defect type and defect location of each product defect are output through a defect detection model. For the image information within each defect location, training defect features are extracted. For example, if the contour coordinate values of a defect in the input image are contours, the following basic defect features can be constructed based on the contour area of the defect within this contour range, including: the area area corresponding to contours, the length rect_h of the minimum bounding rectangle, the width rect_w of the minimum bounding rectangle, the skeleton length seleton of contours, the center point coordinates (c_x, c_y) of contours, and other features. Among them, if the defect type is a different-color type, the defect features also include the average gray value gray_avg within the contour area, etc. The defect feature St of a defect with a different-color type can be expressed as:

[0086] .

[0087] It can be seen that in this embodiment, by determining the corresponding target feature dimension according to the defect type and obtaining the training defect features according to the target feature dimension, the detection accuracy of product defects of this defect type can be ensured.

[0088] In some embodiments, the model training method further includes:

[0089] Obtain the current product defect image corresponding to the target product;

[0090] Extract the current defect features corresponding to the current product defect image;

[0091] Input the current defect features into the target defect interception strategy model for processing, and output the predicted interception label corresponding to the current defect features;

[0092] Based on the predicted interception label corresponding to the current defect features, determine the defect detection result corresponding to the target product.

[0093] Among them, the target product refers to the product for which product defect detection is currently required. The current product defect image refers to the image of the target product for which the interception strategy is currently required to be predicted.

[0094] Specifically, after the server finishes training the initial defect interception policy model, it obtains the target defect interception policy model. Then, it obtains the current product defect image corresponding to the target product. It can obtain the product image corresponding to the target product, input the product image into the defect detection model, and perform defect detection on the product image through the defect detection model. If no product defect is detected in the product image, it is determined that the target product is a normal product; if a product defect is detected in the product image, the product image is used as the current product defect image corresponding to the target product, and the current defect feature corresponding to the current product defect image is extracted through the defect detection model. Then, the current defect feature is input into the target defect interception policy model for processing, and the predicted interception label corresponding to the current defect feature is output. When the predicted interception label corresponding to the current defect feature is "do not intercept", it indicates that the product defect in the target product is relatively minor (i.e., the defect severity is low and can be ignored), and it is determined that the defect detection result of the target product is a normal product; when the predicted interception label corresponding to the current defect feature is "intercept", it indicates that the product defect in the target product is relatively serious (i.e., the defect severity is high and cannot be ignored), and it is determined that the defect detection result of the target product is an abnormal product.

[0095] It can be seen that in this embodiment, by using the target defect interception policy model to predict the interception label of product defects, the detection accuracy of product defects can be guaranteed.

[0096] In some embodiments, as Figure 3 shown. A flowchart of industrial defect detection based on reinforcement learning is provided. The initial defect interception policy model and the initial defect interception value model are trained through a reinforcement learning mechanism. Reinforcement learning refers to a machine learning method in which the model learns the optimal policy through trial and error in the interaction with the environment (Environment). The model executes actions (Action) in the environment, such as the interception policy output by the defect interception policy model, and receives feedback, that is, a reward (Reward), according to the result of the action. The reward signal is used to guide the model to adjust its policy to maximize the long-term cumulative reward.

[0097] In reinforcement learning, it is necessary to define the reinforcement learning environment, including the state space, action space, and reward function. Among them, the state space refers to the training defect features, and the feature dimension of the training defect features can be determined according to the defect type, expressed as , where t represents the t-th time step in reinforcement learning; the action space includes two predefined actions, respectively representing whether to intercept, that is, the interception action at the t-th time step {0, 1}, where 0 indicates that the defect should not be intercepted, and 1 indicates that the defect should be intercepted; the reward function is used to feedback the quality of the predicted interception action and help the model learn.

[0098] Then, the interception strategy of product defects is optimized through the policy network (i.e., the defect interception policy model) and the value network (i.e., the defect interception value model). Among them, the policy network is used to output the probability of taking each action. For each state (i.e., the defect characteristics of each defect), the policy network outputs a probability distribution for the probability of selecting the interception action 0 or action 1, denoted as . The value network is used to evaluate the "value" of the current state (i.e., defect characteristics), that is, the expected total return of the model in this state, to better measure the long-term benefits brought by the interception action, denoted as . The two networks of the policy network and the value network share the feature extraction layer, only the final classification heads are different. The policy network uses the SoftMax classification head and finally outputs the probability values of the two categories of 0 and 1. The value network uses the fully connected layer and finally outputs the predicted interception value .

[0099] The model training through reinforcement learning includes:

[0100] Extract the feature list: Obtain the original defect sample image, which includes the workpiece area and the real defects in the workpiece area; input the original defect sample image into the deep defect detection model, detect the defect location through the deep defect detection model, then crop out the defect area, and extract the defect characteristics according to the defect area to generate the feature list.

[0101] Initialize the model parameters: Randomly initialize the policy network and the value network.

[0102] Data collection: Input the feature list into the policy network for sampling and output the predicted interception action. Through multiple samplings of the policy network in the environment (state space, action space, and reward function), accumulate a batch of state-action-reward sequences, which can be used as trajectories, denoted as , where t is the time step. During the training process, reinforcement learning will repeatedly collect new trajectories for updating the model.

[0103] Training process: Use the policy network to generate a batch of state-action-reward trajectories, and based on the predicted interception value output by the value network, calculate the advantage function of each state under the current interception policy to measure the quality of the interception action in the current state; the advantage function can be to calculate the estimated return of the value network (i.e., the predicted interception value ) and the real return (i.e., the training interception value ) the value difference between. Calculate the loss function of reinforcement learning, including the policy model loss and the value model loss. Gradient update: Based on the loss function, perform backpropagation and update on the parameters of the policy network and the value network. Repeat update: After multiple rounds of sampling and updating, use the updated policy network for subsequent sampling.

[0104] The reinforcement learning model uses the trained policy network to predict the interception decision for each defect sample. For a defect sample to be detected, its feature vector is input into the policy network for feature input: the action probability is output through the policy network, and the action with a higher probability is selected as the decision of the model, that is, whether to intercept the defect.

[0105] It can be seen that in this embodiment, the interception strategy is automatically adjusted based on the actual interception situation through the reward function, enabling the policy model to dynamically adapt to different defect categories and different defect features, and improving the accuracy of interception prediction for product defects.

[0106] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0107] Based on the same inventive concept, an embodiment of the present application also provides a model training device. The implementation solution for solving problems provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model training device provided below can refer to the limitations on the model training method in the above text, and will not be repeated here.

[0108] As Figure 4 shown, an embodiment of the present application provides a model training device 400, including:

[0109] A data acquisition module 402, configured to respectively obtain product defect training images, training defect features corresponding to the product defect training images, and training interception labels corresponding to the training defect features;

[0110] An interception prediction module 404, configured to input the training defect features into an initial defect interception policy model for processing, and output a predicted interception label corresponding to the training defect features;

[0111] A value prediction module 406, configured to input the training defect features into an initial defect interception value model for processing, and output a predicted interception value corresponding to the training defect features;

[0112] A value calculation module 408 for determining a training interception value based on a predicted interception label, a training interception label, and training defect features;

[0113] A model update module 410 for adjusting an initial defect interception policy model respectively based on a predicted interception value, a training interception value, and a predicted interception label; and adjusting an initial defect interception value model based on the predicted interception value and the training interception value until a convergence condition is satisfied, so as to obtain a target defect interception policy model and a target defect interception value model.

[0114] In some embodiments, in terms of determining a training interception value based on a predicted interception label, a training interception label, and training defect features, the value calculation module 408 is specifically configured to:

[0115] Obtain weighted parameters corresponding to respective feature dimensions in the training defect features; weight the training defect features based on the weighted parameters to obtain weighted training defect features;

[0116] When the predicted interception label is consistent with the training interception label, calculate a positive interception value based on the weighted training defect features to obtain the training interception value; or,

[0117] When the predicted interception label is inconsistent with the training interception label, calculate a negative interception value based on the weighted training defect features to obtain the training interception value.

[0118] In some embodiments, in terms of adjusting an initial defect interception policy model based on a predicted interception value, a training interception value, and a predicted interception label until a convergence condition is satisfied to obtain a target defect interception policy model, the model update module 410 is specifically configured to:

[0119] Calculate a value difference between the predicted interception value and the training interception value;

[0120] When the predicted interception label is within a preset label value range, determine a policy model loss corresponding to the initial defect interception policy model based on the value difference and the predicted interception label; or,

[0121] When the predicted interception label is not within the preset label value range, obtain a policy model loss corresponding to the initial defect interception policy model based on the preset label value range and the value difference;

[0122] Update the initial defect interception policy model based on the policy model loss until a convergence condition is satisfied to obtain a target defect interception policy model.

[0123] In some embodiments, in terms of adjusting an initial defect interception value model based on a predicted interception value and a training interception value until a convergence condition is satisfied to obtain a target defect interception value model, the model update module 410 is specifically configured to:

[0124] Calculate the value difference between the predicted interception value and the training interception value;

[0125] Based on the value difference, determine the value model loss corresponding to the initial defect interception value model;

[0126] Update the initial defect interception value model based on the value model loss until the convergence condition is met to obtain the target defect interception value model.

[0127] In some embodiments, in terms of separately obtaining the product defect training images, the training defect features corresponding to the product defect training images, and the training interception labels corresponding to the training defect features, the data acquisition module 402 is specifically configured to:

[0128] Obtain the product defect training images, input the product defect training images into a defect detection model for processing, and output the defect types and defect regions in the product defect training images;

[0129] Obtain the target feature dimension corresponding to the defect type, and extract the training defect features corresponding to the target feature dimension and the training interception labels corresponding to the training defect features from the defect regions.

[0130] In some embodiments, the model training device 400 further includes a model application module, and this model application module is used for:

[0131] Obtain the current product defect image corresponding to the target product; extract the current defect features corresponding to the current product defect image; input the current defect features into the target defect interception policy model for processing, and output the predicted interception labels corresponding to the current defect features; based on the predicted interception labels corresponding to the current defect features, determine the defect detection result corresponding to the target product.

[0132] Each module in the above model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0133] In some embodiments, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to model training, etc. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the steps in the above-mentioned model training method.

[0134] In some embodiments, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 6 shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the above-mentioned model training method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen; the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0135] Those skilled in the art can understand that Figure 5 or Figure 6The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0136] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0137] In some embodiments, as Figure 7 shown, an internal structure diagram of a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0138] In some embodiments, a computer program product is provided. The computer program product includes a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0139] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0142] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A model training method, characterized in that: include: Respectively obtaining a product defect training image, a training defect feature corresponding to the product defect training image, and a training interception label corresponding to the training defect feature; Input the training defect features into the initial defect interception strategy model for processing, and output the predicted interception label corresponding to the training defect features; Inputting the training defect feature into the initial defect interception value model for processing, and outputting the predicted interception value corresponding to the training defect feature; Determining a training interception value based on the predicted interception label, the training interception label, and the training defect feature; The initial defect interception strategy model is adjusted based on the predicted interception value, the training interception value and the predicted interception label respectively; the initial defect interception value model is adjusted based on the predicted interception value and the training interception value until the convergence condition is met, thereby obtaining a target defect interception strategy model and a target defect interception value model.

2. The method according to claim 1, characterized in that The determining of the training interception value based on the predicted interception label, the training interception label, and the training defect feature includes: Obtaining weighting parameters corresponding to each feature dimension in the training defect feature; weighting the training defect feature based on the weighting parameters to obtain a weighted training defect feature; When the predicted interception label is consistent with the training interception label, a forward interception value is calculated based on the weighted training defect feature to obtain a training interception value; or, When the predicted interception label is inconsistent with the training interception label, a reverse interception value is calculated based on the weighted training defect feature to obtain a training interception value.

3. The method according to claim 1, characterized in that: The adjusting the initial defect interception strategy model based on the predicted interception value, the training interception value and the predicted interception label until a convergence condition is met to obtain a target defect interception strategy model includes: calculating a value difference between the predicted intercept value and the trained intercept value; When the predicted interception label is within a preset label value range, determining the strategy model loss corresponding to the initial defect interception strategy model based on the value difference and the predicted interception label; or, When the predicted interception label is not within the preset label value range, determining a strategy model loss corresponding to the initial defect interception strategy model based on the preset label value range and the value difference; Based on the strategy model loss, the initial defect interception strategy model is updated until a convergence condition is met to obtain a target defect interception strategy model.

4. The method according to claim 1, characterized in that: The adjusting the initial defect interception value model based on the predicted interception value and the training interception value until a convergence condition is met to obtain a target defect interception value model includes: calculating a value difference between the predicted intercept value and the trained intercept value; Based on the value difference, determining a value model loss corresponding to the initial defect interception value model; The initial defect interception value model is updated based on the value model loss until a convergence condition is met to obtain a target defect interception value model.

5. The method according to claim 1, characterized in that The obtaining of the product defect training image, the training defect feature corresponding to the product defect training image, and the training interception label corresponding to the training defect feature respectively includes: Obtaining a product defect training image, inputting the product defect training image into a defect detection model for processing, and outputting a defect type and defect area in the product defect training image; A target feature dimension corresponding to the defect type is obtained, and a training defect feature corresponding to the target feature dimension and a training interception label corresponding to the training defect feature are extracted from the defect area.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining a current product defect image corresponding to the target product; Extracting current defect features corresponding to the current product defect image; Inputting the current defect feature into the target defect interception strategy model for processing, and outputting a predicted interception label corresponding to the current defect feature; Based on the predicted interception label corresponding to the current defect feature, a defect detection result corresponding to the target product is determined.

7. A model training device, characterized in that: include: A data acquisition module, used to respectively obtain a product defect training image, a training defect feature corresponding to the product defect training image, and a training interception label corresponding to the training defect feature; An interception prediction module, used to input the training defect features into the initial defect interception strategy model for processing, and output a predicted interception label corresponding to the training defect features; A value prediction module, used for inputting the training defect features into the initial defect interception value model for processing, and outputting the predicted interception value corresponding to the training defect features; A value calculation module, used to determine the training interception value based on the predicted interception label, the training interception label, and the training defect feature; A model updating module is used to adjust the initial defect interception strategy model based on the predicted interception value, the training interception value and the predicted interception label respectively; based on the predicted interception value and the training interception value, adjust the initial defect interception value model until the convergence condition is met, thereby obtaining a target defect interception strategy model and a target defect interception value model.

8. The device according to claim 7, characterized in that In determining the training interception value based on the predicted interception label, the training interception label, and the training defect feature, the value calculation module is specifically used to: Obtaining weighting parameters corresponding to each feature dimension in the training defect feature, and weighting the training defect feature based on the weighting parameters to obtain a weighted training defect feature; When the predicted interception label is consistent with the training interception label, a positive interception value is calculated based on the weighted training defect feature to obtain a training interception value; or, When the predicted interception label is inconsistent with the training interception label, a reverse interception value is calculated based on the weighted training defect feature to obtain a training interception value.

9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.