Model training method and device, computer equipment and computer readable storage medium

By introducing multiple attention modules into the image segmentation model and adjusting the cue image and annotation mask based on error loss, the problem of low training efficiency in traditional methods is solved, achieving efficient training and improved accuracy of the image segmentation model.

CN120612335BActive Publication Date: 2025-11-04SIMOU GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511127447.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-04
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In the training process of traditional image segmentation models, the use of L1 error loss results in low training efficiency and long training time.

Method used

An image segmentation model employing multiple attention modules is proposed. Each module corresponds to a cue image and a cue annotation mask. By determining the error loss between the predicted annotation mask and the trained annotation mask, the cue image and cue annotation mask of each attention module are adjusted to obtain the target image segmentation model.

Benefits of technology

Multiple cue images and multiple cue annotation masks can be adjusted with a single training iteration, reducing the number of training iterations for the image segmentation model and improving training efficiency and model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612335B_ABST
    Figure CN120612335B_ABST
Patent Text Reader

Abstract

The application relates to a model training method and device, computer equipment and a computer readable storage medium. The method is applied to an image segmentation model comprising a plurality of attention modules, each attention module corresponding to a prompt image and a prompt annotation mask. The method comprises: obtaining a training image and a training annotation mask corresponding to the training image; inputting the training image into an initial image segmentation model for processing to output a predicted annotation mask; determining an error loss between the predicted annotation mask and the training annotation mask based on a predicted annotation of a pixel point in the predicted annotation mask, a training annotation of a pixel point in the training annotation mask and a reference annotation corresponding to various pixel types; and adjusting the prompt image and the prompt annotation mask corresponding to each attention module based on the error loss to obtain a target image segmentation model. The method can improve the efficiency of image segmentation model training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automation, and in particular to a model training method and device, computer equipment and a computer readable storage medium. BACKGROUND

[0002] With the development of automation technology, the automatic detection of product defects is realized. By inputting the image of the product to be detected into a trained image segmentation model, a labeled mask can be obtained, and the labeled mask can be used to determine the defect detection result of the product to be detected.

[0003] In the training process of the image segmentation model in the prior art, the L1 error loss (Least Absolute Deviations Loss, L1 Loss for short) is used to adjust the parameters of the image segmentation model, which results in a long training time of the image segmentation model and low efficiency of the image segmentation model training. SUMMARY

[0004] Therefore, it is necessary to provide a model training method, device, computer equipment and computer readable storage medium to improve the efficiency of image segmentation model training.

[0005] In a first aspect, the present application provides a model training method applied to an image segmentation model comprising a plurality of attention modules, each of the attention modules corresponding to a prompt image and a prompt labeled mask, the method comprising:

[0006] obtaining a training image and a training labeled mask corresponding to the training image;

[0007] inputting the training image into an initial image segmentation model for processing to output a predicted labeled mask;

[0008] determining an error loss between the predicted labeled mask and the training labeled mask based on a predicted label of a pixel point in the predicted labeled mask, a training label of the pixel point in the training labeled mask and a reference label corresponding to various pixel types;

[0009] adjusting the prompt image and the prompt labeled mask corresponding to each attention module based on the error loss to obtain a target image segmentation model.

[0010] In a second aspect, the present application further provides a model training device applied to an image segmentation model comprising a plurality of attention modules, each of the attention modules corresponding to a prompt image and a prompt labeled mask, the device comprising:

[0011] an obtaining module configured to obtain a training image and a training labeled mask corresponding to the training image;

[0012] The processing module is configured to input the training image into the initial image segmentation model for processing, and output a predicted annotation mask.

[0013] The determining module is configured to determine an error loss between the predicted annotation mask and the training annotation mask based on the predicted annotation of the pixel point in the predicted annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to various pixel types.

[0014] The training module is configured to adjust the prompt image and the prompt annotation mask corresponding to each attention module based on the error loss, to obtain a target image segmentation model.

[0015] In a third aspect, the present application provides a computer device, which comprises a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method when executing the computer program.

[0016] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method.

[0017] In a fifth aspect, the present application further provides a computer program product, which comprises a computer program. The computer program is executed by a processor to implement the steps in the above method.

[0018] The above model training method, device, computer device and computer readable storage medium can determine the predicted annotation mask corresponding to the training image through the initial image segmentation model, determine the error loss between the predicted annotation mask and the training annotation mask based on the predicted annotation of the pixel point in the predicted annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to various pixel types, adjust the prompt image and the prompt annotation mask corresponding to each attention module using the error loss, and obtain a target image segmentation model. Through one training, multiple prompt images and multiple prompt annotation masks can be adjusted. Compared with adjusting only one prompt image and one prompt annotation mask each time, the training times of the image segmentation model are reduced, thereby shortening the training time of the image segmentation model and improving the efficiency of the image segmentation model training. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 An application environment diagram of a model training method provided by an embodiment of the present application;

[0020] Figure 2 A flowchart of a model training method provided by an embodiment of the present application;

[0021] Figure 3 A schematic diagram of an image segmentation model provided by an embodiment of the present application;

[0022] Figure 4 A flowchart of an error loss determination step provided for an embodiment of the present application is shown in the figure.

[0023] Figure 5 A flowchart of another model training method provided for an embodiment of the present application is shown in the figure.

[0024] Figure 6 A structural block diagram of a model training device provided for an embodiment of the present application is shown in the figure.

[0025] Figure 7 An internal structure diagram of a computer device provided for an embodiment of the present application is shown in the figure.

[0026] Figure 8 An internal structure diagram of a computer readable storage medium provided for an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical scheme and advantages of the present application clearer, further detailed description will be made to the present application in combination with the figures and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0028] The model training method provided by the embodiments of the present application can be applied in the application environment as shown in the figure. Figure 1 The terminal 102 communicates with the server 104 through a communication network. The data storage system can store the data required to be processed by the server 104. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be realized by an independent server or a server cluster composed of multiple servers.

[0029] As shown in the figure, Figure 2 The present application provides a model training method applied to an image segmentation model comprising multiple attention modules, each attention module corresponding to one prompt image and one prompt annotation mask. Taking the case of applying the method to a computer device for illustration, the method comprises the following steps:

[0030] Step 202, obtaining a training image and a training annotation mask corresponding to the training image.

[0031] The image segmentation model refers to a neural network model used to determine the annotation mask. The image segmentation model includes but is not limited to an encoder and a decoder. The encoder is composed of multiple attention modules. The encoder is used to extract semantic features, that is, the encoder converts the input information into high-dimensional semantic features. The decoder is used to restore the high-dimensional semantic features to a pixel-level segmentation image, and the segmentation image can be annotated with a mask. The attention module is the core constituent unit of the encoder, which is used to establish a connection between different regions of the image and capture the context semantics. The input of the attention module is a prompt image, a prompt annotation mask, an input image, and a blank annotation mask. For the first attention module in the encoder, the input image is a training image. For other attention modules in the encoder, the input image is a feature image output by the previous adjacent attention module. The output of the attention module is a feature image. The prompt image and the prompt annotation mask corresponding to the attention module are part of the image segmentation model. It can be understood that for a trained target image segmentation model, the target prompt image and the target prompt annotation mask corresponding to each attention module in the target image segmentation model are fixed. The size of the prompt image and the prompt annotation mask corresponding to one attention module is the same. The sizes of the prompt images corresponding to different attention modules can be different. The image segmentation model can be SegGPT (Segmenting Everything in Context, prompt image segmentation model). SegGPT is a general image segmentation model based on the Transformer architecture. For example, as shown in the schematic diagram of the image segmentation model in Figure 3 The image segmentation model is composed of an encoder and a decoder. The encoder is composed of n attention modules. The input of the first attention module includes a prompt image 1 and a prompt annotation mask 1. The input of the second attention module includes a prompt image 2 and a prompt annotation mask 2. The input of the nth attention module includes a prompt image n and a prompt annotation mask n.

[0032] The prompt image refers to an example image used to guide the image segmentation model to learn the segmentation target. The annotation mask refers to a mask image representing the pixel type of a pixel point in an image. The annotation mask can be a binary mask image, a multi-value mask image, or a color mask image. In the case of including one defect type, the annotation mask can be a binary mask image or a color mask image. For example, in the binary mask image, the pixel point with a pixel value of 1 is a defect pixel point, and the pixel point with a pixel value of 0 is a normal pixel point. Alternatively, in the color mask image, the pixel point with a pixel value of (255, 255, 255) is a defect pixel point, and the pixel point with a pixel value of (0, 0, 0) is a normal pixel point. In the case of including multiple defect types, the prompt annotation mask can be a multi-value mask image or a color mask image. For example, in the multi-value mask image, the pixel point with a pixel value of 1 is a scratch pixel point, the pixel point with a pixel value of 2 is an indentation pixel point, and the pixel point with a pixel value of 0 is a normal pixel point. Alternatively, in the color mask image, the pixel point with a pixel value of (255, 255, 255) is a scratch pixel point, the pixel point with a pixel value of (100, 105, 150) is an indentation pixel point, and the pixel point with a pixel value of (0, 0, 0) is a normal pixel point. The prompt annotation mask refers to the annotation mask corresponding to the prompt image. The annotation mask in the following description is described by taking a color mask image as an example. The color mask image is an RGB image, and each pixel point in the annotation mask corresponds to a red channel brightness value (R value), a green channel brightness value (G value), and a blue channel brightness value (B value). That is, each pixel point in the annotation mask corresponds to a pixel value of (R value, G value, B value). The training annotation mask refers to the annotation mask corresponding to the training image. The blank annotation mask refers to an annotation mask in which the pixel values of the pixel points are all preset values.

[0033] The training image refers to an image used to train the image segmentation model. The training image can include a detection object. The detection object can be a product, for example, the detection object is an industrial product of an automated production plastic part or a glass product. The number of training images is a plurality, and each training image includes at least one detection object. The detection objects included in the plurality of training images can be the same type of detection object. For example, the detection object is a mobile phone shell, and each training image includes at least one mobile phone shell. The mobile phone shells included in different training images can have different defects. The detection objects included in the plurality of training images can be different types of objects but have similar defects. For example, both glass and the protective film of the explosion-proof valve can have defects such as scratches, cracks, and incompleteness. Therefore, the detection objects included in the plurality of training images can be at least one of glass and the protective film. The training image can be an RGB image.

[0034] For example, the computer device obtains a training image and a training annotation mask corresponding to the training image from a sample set.

[0035] Step 204, input the training image into the initial image segmentation model for processing, and output a predicted annotation mask.

[0036] The initial image segmentation model refers to an image segmentation model that needs to be trained. The initial image segmentation model can be an untrained image segmentation model or an image segmentation model that has been trained but still needs to be further trained. The predicted annotation mask refers to a mask image corresponding to the training image.

[0037] For example, the computer device inputs the training image and the blank annotation mask into the initial image segmentation model for processing, and outputs the predicted annotation mask.

[0038] Step 206, based on the predicted annotation of the pixel points in the predicted annotation mask, the training annotation of the pixel points in the training annotation mask, and the reference annotation corresponding to various pixel types, determine the error loss between the predicted annotation mask and the training annotation mask.

[0039] The predicted annotation refers to the pixel value of the pixel point in the predicted annotation mask. The training annotation refers to the pixel value of the pixel point in the training annotation mask. The pixel type refers to the type of the pixel point, which can be normal pixel points and defect pixel points. Defect pixel points can be divided into many different defect pixel points, for example, pixel types include scratch pixel points, broken pixel points, indentation pixel points, and normal pixel points. The reference annotation refers to the pixel value corresponding to the pixel type, for example, the reference annotation of the scratch pixel point is (100, 105, 150), and the pixel point with pixel value (100, 105, 150) is a scratch pixel point. The error loss refers to a value representing the difference between the predicted annotation mask and the training annotation mask. The larger the error loss, the lower the accuracy of the image segmentation model; the smaller the error loss, the higher the accuracy of the image segmentation model.

[0040] For example, the computer device obtains the reference annotation corresponding to various pixel types preset in advance, determines the type error corresponding to each pixel type based on the predicted annotation of the pixel points in the predicted annotation mask, the training annotation of the pixel points in the training annotation mask, and the reference annotation corresponding to various pixel types, and determines the error loss between the predicted annotation mask and the training annotation mask based on the type error corresponding to various pixel types.

[0041] Step 208, based on the error loss, adjust the prompt image and the prompt annotation mask corresponding to each attention module to obtain a target image segmentation model.

[0042] The target image segmentation model refers to a trained image segmentation model.

[0043] Exemplarily, the computer device adjusts the prompt image and the prompt annotation mask corresponding to each attention module based on the error loss, obtains an updated prompt image and an updated prompt annotation mask corresponding to each attention module, and repeats the steps 202 to 208 until a training stop condition is met, to obtain a target prompt image and a target prompt annotation mask corresponding to each attention module, i.e., to obtain a target image segmentation model. The training stop condition can be that the number of training times is equal to a preset number, the error loss is less than or equal to a loss threshold, etc. The training stop condition can be set according to actual needs, which is not limited here.

[0044] In the above model training method, the prediction annotation mask of the training image is determined by the initial image segmentation model, the error loss between the prediction annotation mask and the training annotation mask is determined based on the prediction annotation of the pixel point in the prediction annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to various pixel types, the prompt image and the prompt annotation mask corresponding to each attention module are adjusted using the error loss, and the target image segmentation model is obtained. Through one training, multiple prompt images and multiple prompt annotation masks can be adjusted. Compared with adjusting only one prompt image and one prompt annotation mask each time, the number of training times of the image segmentation model is reduced, thereby shortening the training time of the image segmentation model and improving the efficiency of the image segmentation model training.

[0045] In some embodiments, as shown in Figure 4 The error loss between the prediction annotation mask and the training annotation mask is determined based on the prediction annotation of the pixel point in the prediction annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to various pixel types, including:

[0046] In step 402, for each pixel point in the prediction annotation mask, the prediction probability of the pixel point belonging to various pixel types is determined based on the prediction annotation of the pixel point and the reference annotation corresponding to various pixel types.

[0047] The prediction probability refers to the probability of the pixel point belonging to the pixel type.

[0048] Exemplarily, the computer device calculates the difference value between the prediction annotation of the pixel point and the reference annotation corresponding to the pixel type, determines the prediction probability of the pixel point belonging to the pixel type based on the difference value, and repeats the above steps to obtain the prediction probability of the pixel point belonging to various pixel types.

[0049] At step 404, for each pixel type, a mask image corresponding to the pixel type is determined based on the training labels of the pixels in the training label mask; the mask value of a pixel in the mask image is the first identifier or the second identifier, the first identifier indicates that the corresponding training label of the pixel in the training label mask is the same as the reference label corresponding to the pixel type, and the second identifier indicates that the corresponding training label of the pixel in the training label mask is different from the reference label corresponding to the pixel type.

[0050] The mask image refers to a mask indicating whether a pixel in the training image belongs to a pixel type, i.e., the mask value of a pixel in the mask image has only two values, the mask value of the pixel is the first identifier or the second identifier. The first identifier indicates that the pixel belongs to the pixel type, and the second identifier indicates that the pixel does not belong to the pixel type. The first identifier can be 1, and the second identifier can be 0.

[0051] For example, for each pixel type, the computer device obtains the reference label corresponding to the pixel type, compares the training label of each pixel in the training image with the reference label corresponding to the pixel type, and determines the mask value of the pixel as the first identifier if the training label is the same as the reference label, or as the second identifier if the training label is different from the reference label; and determines the mask image corresponding to the pixel type based on the mask values of the pixels.

[0052] At step 406, the type error corresponding to the pixel type is determined based on the predicted probability that a pixel in the predicted label mask belongs to the pixel type and the mask image corresponding to the pixel type.

[0053] The type error refers to the cumulative error of all pixels in the training image that are of a certain pixel type.

[0054] At step 408, the error loss between the predicted label mask and the training label mask is determined based on the type errors corresponding to various pixel types.

[0055] For example, the computer device performs weighted summation on the type errors corresponding to various pixel types to obtain the error loss between the predicted label mask and the training label mask.

[0056] In the embodiment, the type error corresponding to the pixel type is determined by predicting the prediction probability of the pixel point in the prediction annotation mask belonging to the pixel type and the mask image corresponding to the pixel type, that is, the type error corresponding to each pixel type is determined, and then the error loss between the prediction annotation mask and the training annotation mask is determined based on the type error corresponding to each pixel type. Compared with directly determining the error loss between the prediction annotation mask and the training annotation mask, the regions composed of pixel points belonging to different pixel types in the training image are fully considered, and the accuracy of the error loss is improved. Moreover, the prompt image and the prompt annotation mask are adjusted using the error loss, the detection capability of the target image segmentation model for small objects to be detected is improved, the performance of the target image segmentation model for image data with unbalanced pixel categories is optimized, and thus the accuracy of the annotation mask output by the target image segmentation model is improved.

[0057] In some embodiments, the prediction probability of the pixel point belonging to each pixel type is determined based on the prediction annotation of the pixel point and the reference annotation corresponding to each pixel type, including:

[0058] For each pixel type, a difference value between the prediction annotation of the pixel point and the reference annotation corresponding to the pixel type is determined.

[0059] From the difference values corresponding to each pixel type, a maximum difference value is obtained.

[0060] A ratio between the maximum difference value and the difference value corresponding to the pixel type is calculated to obtain a difference ratio value corresponding to the pixel type.

[0061] The difference ratio values corresponding to each pixel type are normalized to obtain the prediction probability of the pixel point belonging to each pixel type.

[0062] The difference value refers to a numerical value representing the difference between the prediction annotation of the pixel point and the reference annotation corresponding to the pixel type. The difference value can be represented by the straight-line distance between the prediction annotation of the pixel point and the reference annotation corresponding to the pixel type. Normalization refers to the process of converting data to a unified numerical range according to certain rules. Different methods can be used for normalization, which are not limited herein, for example, the softmax function can be used for normalization.

[0063] Exemplarily, for each pixel type, the computer device calculates a linear distance between the predicted label of the pixel point and the reference label corresponding to the pixel type, determines the linear distance as a difference value between the predicted label of the pixel point and the reference label corresponding to the pixel type, and repeats the above steps, so as to obtain the difference values corresponding to various pixel types of the pixel point; filters the maximum difference value from the difference values corresponding to various pixel types of the pixel point; for each pixel type, calculates a ratio between the maximum difference value and the difference value corresponding to the pixel type, to obtain a difference ratio value corresponding to the pixel type, and repeats the above steps, so as to obtain the difference ratio values corresponding to various pixel types; and performs normalization processing on the difference ratio values corresponding to various pixel types, to obtain the predicted probabilities of the pixel point belonging to various pixel types.

[0064] For example, the predicted label of the pixel point (x, y) in the prediction label mask is (R, G, B), the reference label of the pixel type 1 is (R1, G1, B1), the reference label of the pixel type 2 is (R2, G2, B2), and the reference label of the pixel type 3 is (R3, G3, B3), the difference value between the predicted label of the pixel point (x, y) and the reference label corresponding to the pixel type 1 is , the difference value between the predicted label of the pixel point (x, y) and the reference label corresponding to the pixel type 2 is , and the difference value between the predicted label of the pixel point (x, y) and the reference label corresponding to the pixel type 3 is The maximum difference value in D1, D2 and D3 is Dmax; the difference ratio value of the pixel type 1 is p1=Dmax / D1, the difference ratio value of the pixel type 2 is p2=Dmax / D2, and the difference ratio value of the pixel type 3 is p3=Dmax / D3; the predicted probability of the pixel point (x, y) belonging to the pixel type 1 is Q1=p1 / (p1+p2+p3), the predicted probability of the pixel point (x, y) belonging to the pixel type 2 is Q2=p2 / (p1+p2+p3), and the predicted probability of the pixel point (x, y) belonging to the pixel type 3 is Q3=p3 / (p1+p2+p3).

[0065] In the embodiment, the difference value between the predicted label of the pixel point and the reference label corresponding to the pixel type is determined, the ratio between the maximum difference value and the difference value corresponding to the pixel type is determined, to obtain the difference ratio value corresponding to the pixel type, the greater the difference ratio value represents the smaller the difference between the predicted label and the reference label, the smaller the difference ratio value represents the greater the difference between the predicted label and the reference label, that is, the difference ratio value is inversely proportional to the difference value, and the difference ratio value corresponding to the pixel type is proportional to the probability of the pixel point belonging to the pixel type, so that the predicted probabilities of the pixel point belonging to various pixel types can be obtained by performing normalization processing on the difference ratio values corresponding to various pixel types, the sum of the predicted probabilities of the pixel point belonging to various pixel types is equal to 1, and the basis data for subsequent determination of error loss is provided.

[0066] In some embodiments, based on the predicted probability that a pixel in the predicted annotation mask belongs to a pixel type and the mask image corresponding to the pixel type, a type error corresponding to the pixel type is determined, including:

[0067] For each pixel in the predicted annotation mask, a product of the predicted probability that the pixel belongs to a pixel type and the mask value corresponding to the pixel in the mask image is calculated to obtain a target probability that the pixel belongs to the pixel type;

[0068] The target probabilities that each pixel belongs to the pixel type are accumulated to obtain a target statistical probability;

[0069] The predicted probabilities that each pixel belongs to the pixel type are accumulated to obtain a predicted statistical probability;

[0070] The mask values corresponding to each pixel in the mask image are accumulated to obtain a mask statistical value;

[0071] Based on the target statistical probability, the predicted statistical probability, and the mask statistical value, a type error corresponding to the pixel type is determined.

[0072] The target probability is equal to zero or the predicted probability, i.e., when the mask value is zero, the target probability is equal to zero; when the mask value is 1, the target probability is equal to the predicted probability. The target statistical probability refers to the sum of the target probabilities that all pixels in the training image belong to the same pixel type. The predicted statistical value refers to the sum of the predicted probabilities that all pixels in the training image belong to the same pixel type. The mask statistical value refers to the sum of the mask values of all pixels in the mask image corresponding to the pixel type.

[0073] For example, for each pixel in the predicted annotation mask, the computer device calculates a product of the predicted probability that the pixel belongs to a pixel type and the mask value corresponding to the pixel in the mask image to obtain a target probability that the pixel belongs to the pixel type, adds the target probabilities that each pixel belongs to the pixel type to obtain a target statistical probability, adds the predicted probabilities that each pixel belongs to the pixel type to obtain a predicted statistical probability, adds the mask values corresponding to each pixel in the mask image to obtain a mask statistical value, and determines a type error corresponding to the pixel type based on the target statistical probability, the predicted statistical probability, and the mask statistical value.

[0074] In some embodiments, the calculation formula of the type error is as follows:

[0075] Formula (1)

[0076] wherein, is a type error of a pixel type; i is a pixel point identifier, the pixel points in the training image and the pixel points in the predicted annotation mask are one-to-one corresponding, the pixel point identifier i can be the i th pixel point in the training image and the i th pixel point in the predicted annotation mask; N is a total number of pixel points, the total number of pixel points in the training image and the total number of pixel points in the predicted annotation mask are the same, N can be the total number of pixel points in the training image or the total number of pixel points in the predicted annotation mask; is a predicted probability that the pixel point belongs to the pixel type; is a mask value of the pixel point in the predicted annotation mask, is 0 or 1, 1 represents that the pixel point belongs to the pixel type, and 0 represents that the pixel point does not belong to the pixel type; is a target statistical probability; is a predicted statistical probability; is a mask statistical value.

[0077] In the embodiment, the type error corresponding to each pixel type is calculated respectively, thereby providing accurate basic data for subsequent determination of the error loss between the predicted annotation mask and the training annotation mask.

[0078] In some embodiments, based on the type error corresponding to each pixel type, the error loss between the predicted annotation mask and the training annotation mask is determined, including:

[0079] For each pixel type, the pixel points in the training annotation mask that are the same as the reference annotation corresponding to the pixel type are counted to obtain a total number of pixel points corresponding to the pixel type;

[0080] Based on the total number of pixel points corresponding to the pixel type, a correction coefficient corresponding to the pixel type is determined; the correction coefficient corresponding to the pixel type is in an inverse proportional relationship with the total number of pixel points corresponding to the pixel type;

[0081] Based on the preset weight value corresponding to the pixel type and the correction coefficient corresponding to the pixel type, a target weighting coefficient corresponding to the pixel type is determined;

[0082] Based on the target weighting coefficient corresponding to each pixel type, the type errors corresponding to the various pixel types are weighted and summed to obtain the error loss between the predicted annotation mask and the training annotation mask.

[0083] The total number of pixel points refers to the number of pixel points in the training annotation mask with the same mask value of the reference annotation. The correction coefficient refers to a coefficient for correcting the preset weight value. The preset weight value refers to a weight value corresponding to the pixel type and set in advance. The preset weight value can be set according to the importance of the pixel type, and the preset weight value is in a proportional relationship with the importance of the pixel type. The target weighting coefficient refers to the product of the preset weight value and the corresponding correction coefficient.

[0084] Exemplarily, for each pixel type, the computer device counts the pixel points in the training annotation mask corresponding to the pixel type which are the same as the reference annotation, obtains the total number of pixel points corresponding to the pixel type, substitutes the total number of pixel points corresponding to the pixel type into the objective function in which the correction coefficient is inversely proportional to the total number of pixel points, obtains the correction coefficient corresponding to the pixel type, multiplies the preset weight value corresponding to the pixel type by the correction coefficient corresponding to the pixel type, obtains the target weighting coefficient corresponding to the pixel type, and repeats the above process to obtain the target weighting coefficient corresponding to each pixel type; and the type errors corresponding to various pixel types are weighted and summed based on the target weighting coefficients corresponding to the various pixel types to obtain the error loss between the prediction annotation mask and the training annotation mask.

[0085] In the embodiment, the correction coefficient corresponding to the pixel type is inversely proportional to the total number of pixel points corresponding to the pixel type, which means that the more the total number of pixel points corresponding to the pixel type, the smaller the correction coefficient corresponding to the pixel type. By multiplying the preset weight value corresponding to the pixel type by the correction coefficient, the target weighting coefficient corresponding to the pixel type is reduced. For the pixel type with a large total number, the type error corresponding to the pixel type is large. By multiplying the large type error by the small target weighting coefficient, the proportion of the product of the type error and the target weighting coefficient in the error loss is reduced, so that the error loss fully reflects the errors corresponding to various pixel types. Using the error loss to adjust the initial image segmentation model optimizes the performance of the target image segmentation model on the image data with an unbalanced pixel class, that is, improves the accuracy of the target image segmentation model in determining the annotation mask of the image data with an unbalanced pixel type.

[0086] In some embodiments, the training image is input into the initial image segmentation model for processing to output a prediction annotation mask, including:

[0087] For each attention module, the prompt image corresponding to the attention module and the input image are merged to obtain a first image; when the attention module is the first attention module, the input image is the training image, otherwise, the input image is the feature image output by the previous adjacent attention module of the attention module;

[0088] The prompt annotation mask corresponding to the attention module and the blank annotation mask are merged to obtain a second image;

[0089] The first image and the second image are input into the attention module to obtain a feature image output by the attention module;

[0090] Until the feature image output by the last attention module, the feature image output by the last attention module is input into the decoding module for processing to output a prediction annotation mask.

[0091] The first image refers to the image obtained by merging the prompt image and the input image. The prompt image and the input image can be merged in the height direction or in the width direction. The second image refers to the image obtained by merging the prompt annotation mask and the blank annotation mask. The prompt annotation mask and the blank annotation mask can be merged in the height direction or in the width direction. The merging manner of the prompt annotation mask and the blank annotation mask is the same as the merging manner of the prompt image and the input image. The feature image refers to the image obtained by the attention module performing semantic feature extraction on the first image and the second image.

[0092] For example, for each attention module in the encoder, for the first attention module in the encoder, the computer device merges the prompt image corresponding to the attention module and the training image to obtain a first image, and merges the prompt annotation mask corresponding to the attention module and the blank annotation mask to obtain a second image, and inputs the first image and the second image into the attention module to obtain a feature image output by the attention module; for each of the second attention module to the last attention module in the encoder, the computer device merges the prompt image corresponding to the attention module and the feature image output by the previous adjacent attention module to obtain a first image, and merges the prompt annotation mask corresponding to the attention module and the blank annotation mask to obtain a second image, and inputs the first image and the second image into the attention module to obtain a feature image output by the attention module; until the last attention module outputs a feature image, and the feature image output by the last attention module is input into the decoding module for processing to obtain a predicted annotation mask output by the decoding module.

[0093] In this embodiment, each attention module in the encoder performs image segmentation on the training image or the feature image by referring to the corresponding prompt image and the prompt annotation mask, thereby improving the accuracy of the feature image output by the last attention module. The decoding module decodes the feature image output by the last attention module to obtain a predicted annotation mask, thereby improving the accuracy of the predicted annotation mask.

[0094] In some embodiments, the model training method further comprises:

[0095] inputting the obtained target image into the target image segmentation model for processing to output a target annotation mask;

[0096] Based on the target annotation mask and the reference annotation corresponding to various pixel types, the defect type corresponding to the object to be detected in the target image is determined.

[0097] The target image refers to an image for which the annotation mask needs to be determined. The defect type refers to the type of defect of the object to be detected. The target annotation mask refers to the annotation mask corresponding to the target image.

[0098] Exemplarily, the computer device compares the target annotation in the target annotation mask with the reference annotation corresponding to the pixel type. If there is at least one target annotation identical to the reference annotation, the defect type corresponding to the pixel type of the reference annotation is determined as the defect type corresponding to the object to be detected in the target image. If there is no target annotation identical to the reference annotation, the object to be detected in the target image is determined as a normal object. The defect type corresponding to the pixel type refers to the type of defect corresponding to the pixel type. For example, if the pixel type is a crack pixel, the defect type corresponding to the pixel type is a breakage. Or, if the pixel type is a scratch pixel, the defect type corresponding to the pixel type is a scratch. The normal object refers to an object without defects.

[0099] In this embodiment, the target image segmentation model is used to determine the target annotation mask corresponding to the target image, which improves the accuracy of the target annotation mask. The target annotation mask and the reference annotations corresponding to various pixel types are used to determine the defect type corresponding to the object to be detected in the target image, thereby improving the accuracy of the defect type determination.

[0100] In some exemplary embodiments, the image segmentation model includes an encoder and a decoder. The encoder is composed of multiple attention modules, each attention module corresponding to one prompt image and one prompt annotation mask. The encoder is used to extract semantic features, and the decoder is used to restore the high-dimensional semantic features to a pixel-level segmentation image. The flowchart of the model training method is shown in FIG. 1. Figure 5 As shown in FIG. 1, the method comprises the following steps:

[0101] Step 502, determining the prediction annotation mask corresponding to the training image.

[0102] The computer device acquires training images and corresponding training annotation masks from the sample set. For the first attention module in the encoder, the computer device merges the prompt image and the training image corresponding to the attention module to obtain a first image, and merges the prompt annotation mask and the blank annotation mask corresponding to the attention module to obtain a second image. The first and second images are input into the attention module to obtain the feature image output by the attention module. For each attention module from the second to the last attention module in the encoder, the computer device merges the prompt image corresponding to the attention module with the feature image output by the previous adjacent attention module to obtain a first image, and merges the prompt annotation mask and the blank annotation mask corresponding to the attention module to obtain a second image. The first and second images are input into the attention module to obtain the feature image output by the attention module. This process continues until the last attention module outputs a feature image, which is then input into the decoding module to obtain the prediction annotation mask.

[0103] Step 504: Determine the predicted probability of each pixel in the training image belonging to various pixel types.

[0104] For each pixel type, the computer calculates the straight-line distance between the predicted label of the pixel and the baseline label corresponding to the pixel type. This straight-line distance is defined as the difference between the predicted label of the pixel and the baseline label corresponding to the pixel type. Repeating the above steps yields the difference values ​​corresponding to various pixel types for the aforementioned pixel. The difference values ​​corresponding to various pixel types for the aforementioned pixel are compared to obtain the maximum difference value. For each pixel type, the ratio between the maximum difference value and the difference value corresponding to the pixel type is determined to obtain the difference ratio corresponding to the pixel type. Repeating the above steps yields the difference ratio corresponding to various pixel types. The difference ratio corresponding to various pixel types is normalized to obtain the predicted probability of the pixel belonging to each pixel type.

[0105] Step 506: Determine the type error corresponding to each pixel type.

[0106] For each pixel type, the computer device obtains the reference label corresponding to the pixel type. For each pixel point in the training image, the training label of the pixel point is compared with the reference label corresponding to the pixel type. If the training label is the same as the reference label, the mask value of the pixel point is determined as the first identifier. If the training label is different from the reference label, the mask value of the pixel point is determined as the second identifier. Based on the mask values of the pixel points, the mask image corresponding to the pixel type is determined. Based on the prediction probability of the pixel point in the prediction label mask belonging to the pixel type and the mask image corresponding to the pixel type, the type error corresponding to the pixel type is determined using formula (1). The above steps are repeated to determine the type error corresponding to each pixel type.

[0107] In step 508, the error loss between the prediction label mask and the training label mask is determined.

[0108] The computer device performs weighted summation on the type errors corresponding to various pixel types to obtain the error loss between the prediction label mask and the training label mask.

[0109] In step 510, the prompt image and the prompt label mask corresponding to each attention module are adjusted based on the error loss.

[0110] The computer device adjusts the prompt image and the prompt label mask corresponding to each attention module based on the error loss to obtain the updated prompt image and the updated prompt label mask corresponding to each attention module.

[0111] Steps 502 to 510 are a training process. The above steps 502 to 510 are repeated until the training stop condition is met to obtain the target prompt image and the target prompt label mask corresponding to each attention module, i.e., to obtain the target image segmentation model.

[0112] In the aforementioned model training method, a predicted annotation mask for the training image is determined through an initial image segmentation model. Based on the predicted annotations of pixels in the predicted annotation mask, the training annotations of pixels in the training annotation mask, and the baseline annotations corresponding to various pixel types, an error loss is determined between the predicted annotation mask and the training annotation mask. This error loss is used to adjust the cue image and cue annotation mask corresponding to each attention module, resulting in the target image segmentation model. Multiple cue images and multiple cue annotation masks can be adjusted in a single training iteration, compared to adjusting only one cue image and one cue annotation mask per training iteration. This shortens the number of training iterations for the image segmentation model, thereby reducing the training time and improving the training efficiency. Furthermore, using this error loss to adjust the cue image and cue annotation mask optimizes the performance of the target image segmentation model on image data with imbalanced pixel types, thus improving the accuracy of the target image segmentation model in determining the annotation mask for image data with imbalanced pixel types. Furthermore, by using this error loss to adjust the cue image and cue annotation mask, the detection capability of the target image segmentation model for smaller objects to be detected is improved, and the performance of the target image segmentation model for image data with unbalanced pixel categories is optimized, thereby improving the accuracy of the annotation mask output by the target image segmentation model.

[0113] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0114] Based on the same inventive concept, this application also provides a model training apparatus. The solution provided by this apparatus is similar to the solution described in the above method. Therefore, the specific limitations of one or more model training apparatus embodiments provided below can be found in the limitations of the model training method above, and will not be repeated here.

[0115] like Figure 6 As shown, this application embodiment provides a model training device 600, including:

[0116] The acquisition module 602 is used to acquire the training image and the training annotation mask corresponding to the training image;

[0117] The processing module 604 is configured to input the training image into the initial image segmentation model for processing, and output a predicted annotation mask.

[0118] The determining module 606 is configured to determine an error loss between the predicted annotation mask and the training annotation mask based on the predicted annotation of the pixel point in the predicted annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to each pixel type.

[0119] The training module 608 is configured to adjust the prompt image and the prompt annotation mask corresponding to each attention module based on the error loss, to obtain a target image segmentation model.

[0120] In some embodiments, in the aspect of determining the error loss between the predicted annotation mask and the training annotation mask based on the predicted annotation of the pixel point in the predicted annotation mask, the training annotation of the pixel point in the training annotation mask, and the reference annotation corresponding to each pixel type, the determining module 606 is specifically configured to:

[0121] For each pixel point in the predicted annotation mask, determine a predicted probability that the pixel point belongs to each pixel type based on the predicted annotation of the pixel point and the reference annotation corresponding to each pixel type.

[0122] For each pixel type, determine a mask image corresponding to the pixel type based on the training annotation of each pixel point in the training annotation mask; the mask value of the pixel point in the mask image is a first identifier or a second identifier, the first identifier represents that the corresponding training annotation of the pixel point in the training annotation mask is the same as the reference annotation corresponding to the pixel type, and the second identifier represents that the corresponding training annotation of the pixel point in the training annotation mask is different from the reference annotation corresponding to the pixel type.

[0123] Determine a type error corresponding to the pixel type based on the predicted probability that the pixel point in the predicted annotation mask belongs to the pixel type and the mask image corresponding to the pixel type.

[0124] Determine the error loss between the predicted annotation mask and the training annotation mask based on the type error corresponding to each pixel type.

[0125] In some embodiments, in the aspect of determining the predicted probability that the pixel point belongs to each pixel type based on the predicted annotation of the pixel point and the reference annotation corresponding to each pixel type, the determining module 606 is specifically configured to:

[0126] For each pixel type, determine a difference value between the predicted annotation of the pixel point and the reference annotation corresponding to the pixel type.

[0127] From the difference values corresponding to each pixel type, filter to obtain a maximum difference value.

[0128] a ratio between the maximum difference value and the difference value corresponding to the pixel type is calculated to obtain a difference ratio corresponding to the pixel type;

[0129] The difference ratios corresponding to various pixel types are normalized to obtain a prediction probability of the pixel point belonging to various pixel types.

[0130] In some embodiments, in terms of determining the type error corresponding to the pixel type based on the prediction probability of the pixel point belonging to the pixel type in the prediction annotation mask and the mask image corresponding to the pixel type, the determining module 606 is specifically configured to:

[0131] For each pixel point in the prediction annotation mask, a product of the prediction probability of the pixel point belonging to the pixel type and the mask value corresponding to the pixel point in the mask image is calculated to obtain a target probability of the pixel point belonging to the pixel type;

[0132] The target probabilities of the various pixel points belonging to the pixel type are accumulated to obtain a target statistical probability;

[0133] The prediction probabilities of the various pixel points belonging to the pixel type are accumulated to obtain a prediction statistical probability;

[0134] The mask values corresponding to the various pixel points in the mask image are accumulated to obtain a mask statistical value;

[0135] Based on the target statistical probability, the prediction statistical probability, and the mask statistical value, the type error corresponding to the pixel type is determined.

[0136] In some embodiments, in terms of determining the error loss between the prediction annotation mask and the training annotation mask based on the type errors corresponding to various pixel types, the determining module 606 is specifically configured to:

[0137] For each pixel type, the pixel points in the training annotation mask that are the same as the reference annotation corresponding to the pixel type are counted to obtain a total number of pixel points corresponding to the pixel type;

[0138] Based on the total number of pixel points corresponding to the pixel type, a correction coefficient corresponding to the pixel type is determined; the correction coefficient corresponding to the pixel type is inversely proportional to the total number of pixel points corresponding to the pixel type;

[0139] Based on the preset weight corresponding to the pixel type and the correction coefficient corresponding to the pixel type, a target weighting coefficient corresponding to the pixel type is determined;

[0140] Based on the target weighting coefficients corresponding to various pixel types, the type errors corresponding to various pixel types are weighted and summed to obtain the error loss between the prediction annotation mask and the training annotation mask.

[0141] In some embodiments, in the aspect of inputting the training image into the initial image segmentation model for processing and outputting the predicted annotation mask, the processing module 604 is specifically configured to:

[0142] For each attention module, the prompt image corresponding to the attention module and the input image are merged to obtain a first image; when the attention module is the first attention module, the input image is the training image, otherwise, the input image is the feature image output by the previous adjacent attention module of the attention module;

[0143] The prompt annotation mask corresponding to the attention module and the blank annotation mask are merged to obtain a second image;

[0144] The first image and the second image are input into the attention module to obtain the feature image output by the attention module;

[0145] Until the feature image output by the last attention module, the feature image output by the last attention module is input into the decoding module for processing to output the predicted annotation mask.

[0146] In some embodiments, the model training apparatus 600 further comprises a defect detection module, and the defect detection module is configured to:

[0147] Input the obtained target image into the target image segmentation model for processing to output a target annotation mask;

[0148] Based on the target annotation mask and the reference annotation corresponding to various pixel types, the defect type corresponding to the object to be detected in the target image is determined.

[0149] Each module in the above model training apparatus can be realized by software, hardware and combinations thereof in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations of the above modules by the processor.

[0150] In some embodiments, a computer device is provided, which can be a terminal, and the internal structure diagram thereof can be as shown in Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through the system bus, the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control ability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the external terminal in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to realize the steps in the above model training method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0151] Those skilled in the art can understand that Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0152] In some embodiments, a computer device is provided, which includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the steps in each of the above method embodiments.

[0153] In some embodiments, as Figure 8 A computer readable storage medium 800 is provided, which stores a computer program 802, and the computer program 802 is executed by a processor to realize the steps in each of the above method embodiments.

[0154] In some embodiments, a computer program product is provided, which includes a computer program, and the computer program is executed by a processor to realize the steps in each of the above method embodiments.

[0155] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.

[0156] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of each method can be included. Any reference to a memory, database or other medium used in the embodiments provided by the present application can include at least one of a non-volatile and volatile memory. The non-volatile memory can include a read-only memory (Read-Only Memory, ROM), a magnetic tape, a floppy disk, a flash memory, an optical storage, a high-density embedded non-volatile memory, a resistive memory (ReRAM), a magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), a ferroelectric memory (Ferroelectric Random Access Memory, FRAM), a phase change memory (Phase Change Memory, PCM), a graphene memory, etc. The volatile memory can include a random access memory (Random Access Memory, RAM) or an external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0157] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present disclosure.

[0158] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the patent scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A model training method, characterized in that, The method is applied to an image segmentation model comprising a plurality of attention modules, each of the attention modules corresponding to a prompt image and a prompt annotation mask, and the method comprises: obtaining a training image and a training annotation mask corresponding to the training image; for each of the attention modules, merging the prompt image corresponding to the attention module and an input image to obtain a first image; when the attention module is a first attention module, the input image is the training image, otherwise, the input image is a feature image output by a previous adjacent attention module of the attention module; merging the prompt annotation mask corresponding to the attention module and a blank annotation mask to obtain a second image; inputting the first image and the second image into the attention module to obtain a feature image output by the attention module; until a feature image output by a last attention module, inputting the feature image output by the last attention module into a decoding module for processing to output a predicted annotation mask; for each pixel point in the predicted annotation mask, determining a predicted probability that the pixel point belongs to each of various pixel types based on a predicted annotation of the pixel point and a reference annotation corresponding to each of the pixel types; for each of the pixel types, determining a mask image corresponding to the pixel type based on training annotations of each of the pixel points in the training annotation mask; a mask value of a pixel point in the mask image is a first identifier or a second identifier, the first identifier indicating that the training annotation corresponding to the pixel point in the training annotation mask is the same as a reference annotation corresponding to the pixel type, and the second identifier indicating that the training annotation corresponding to the pixel point in the training annotation mask is different from the reference annotation corresponding to the pixel type; determining a type error corresponding to the pixel type based on the predicted probability that the pixel point in the predicted annotation mask belongs to the pixel type and the mask image corresponding to the pixel type; determining an error loss between the predicted annotation mask and the training annotation mask based on the type error corresponding to each of the pixel types; based on the error loss, adjusting the prompt image and the prompt annotation mask corresponding to each of the attention modules to obtain a target image segmentation model.

2. The method of claim 1, wherein, The method for determining, for each of the pixel types, a mask image corresponding to the pixel type based on training annotations of each of the pixel points in the training annotation mask comprises: for each of the pixel types, obtaining a reference annotation corresponding to the pixel type; for each of the pixel points in the training image, comparing the training annotation of the pixel point with the reference annotation corresponding to the pixel type; if the training annotation is the same as the reference annotation, determining a mask value of the pixel point as the first identifier; or if the training annotation is different from the reference annotation, determining the mask value of the pixel point as the second identifier; determining the mask image corresponding to the pixel type based on the mask value of each of the pixel points.

3. The method of claim 1, wherein, The prediction label of the pixel point and the reference label corresponding to each pixel type are used to determine a prediction probability of the pixel point belonging to each pixel type, including: For each pixel type, a difference value between the prediction label of the pixel point and the reference label corresponding to the pixel type is determined; The maximum difference value is obtained by screening from the difference values corresponding to each pixel type; A difference ratio value corresponding to the pixel type is calculated by calculating a ratio between the maximum difference value and the difference value corresponding to the pixel type; The difference ratio values corresponding to each pixel type are normalized to obtain the prediction probability of the pixel point belonging to each pixel type.

4. The method of claim 1, wherein, The prediction probability of the pixel point belonging to the pixel type in the prediction label mask and the mask image corresponding to the pixel type are used to determine a type error corresponding to the pixel type, including: For each pixel point in the prediction label mask, a target probability of the pixel point belonging to the pixel type is obtained by calculating a product of the prediction probability of the pixel point belonging to the pixel type and a mask value corresponding to the pixel point in the mask image; A target statistical probability is obtained by accumulating the target probabilities of each pixel point belonging to the pixel type; A prediction statistical probability is obtained by accumulating the prediction probabilities of each pixel point belonging to the pixel type; A mask statistical value is obtained by accumulating the mask values corresponding to each pixel point in the mask image; The type error corresponding to the pixel type is determined based on the target statistical probability, the prediction statistical probability, and the mask statistical value.

5. The method of claim 1, wherein, The type errors corresponding to each pixel type are used to determine an error loss between the prediction label mask and the training label mask, including: For each pixel type, a total number of pixel points corresponding to the pixel type is obtained by counting the pixel points in the training label mask that have the same reference label as the pixel type; A correction coefficient corresponding to the pixel type is determined based on the total number of pixel points corresponding to the pixel type; the correction coefficient corresponding to the pixel type is inversely proportional to the total number of pixel points corresponding to the pixel type; A target weighting coefficient corresponding to the pixel type is determined based on a preset weight value corresponding to the pixel type and the correction coefficient corresponding to the pixel type; The type errors corresponding to each pixel type are weighted and summed based on the target weighting coefficients corresponding to each pixel type to obtain the error loss between the prediction label mask and the training label mask.

6. The method of claim 1, wherein, The prompt image and the prompt label mask corresponding to each attention module are adjusted based on the error loss to obtain a target image segmentation model, including: The prompt image and the prompt label mask corresponding to each attention module are adjusted based on the error loss to obtain an updated prompt image and an updated prompt label mask corresponding to each attention module; The step of obtaining the training image and the training annotation mask corresponding to the training image is repeatedly performed until a training stop condition is met, to obtain a target prompt image and a target prompt annotation mask corresponding to each attention module, and to obtain a target image segmentation model.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: inputting the obtained target image into the target image segmentation model for processing, and outputting a target annotation mask; based on the target annotation mask and the reference annotation corresponding to each of the pixel types, determining a defect type corresponding to an object to be detected in the target image.

8. A model training apparatus, comprising: The application is applied to an image segmentation model comprising a plurality of attention modules, each of the attention modules corresponding to a prompt image and a prompt annotation mask, and the device comprises: an obtaining module configured to obtain a training image and a training annotation mask corresponding to the training image; a processing module configured to, for each of the attention modules, combine the prompt image corresponding to the attention module and an input image to obtain a first image; when the attention module is the first attention module, the input image is the training image; otherwise, the input image is a feature image output by a previous adjacent attention module of the attention module; combine the prompt annotation mask corresponding to the attention module and a blank annotation mask to obtain a second image; input the first image and the second image into the attention module to obtain a feature image output by the attention module; and input the feature image output by the last attention module into a decoding module for processing to output a predicted annotation mask; a determining module configured to, for each pixel point in the predicted annotation mask, determine a predicted probability that the pixel point belongs to each of the pixel types based on a predicted annotation of the pixel point and a reference annotation corresponding to each of the pixel types; for each of the pixel types, determine a mask image corresponding to the pixel type based on a training annotation of each pixel point in the training annotation mask; the mask value of a pixel point in the mask image is a first identifier or a second identifier, the first identifier indicating that the training annotation of the pixel point in the training annotation mask is the same as the reference annotation corresponding to the pixel type, and the second identifier indicating that the training annotation of the pixel point in the training annotation mask is different from the reference annotation corresponding to the pixel type; determine a type error corresponding to the pixel type based on the predicted probability that the pixel point in the predicted annotation mask belongs to the pixel type and the mask image corresponding to the pixel type; and determine an error loss between the predicted annotation mask and the training annotation mask based on the type errors corresponding to each of the pixel types. a training module configured to adjust the prompt image and the prompt annotation mask corresponding to each of the attention modules based on the error loss to obtain a target image segmentation model. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image segmentation method and device, electronic equipment and storage medium

    CN118505722A

  • Network training method, image segmentation method, apparatus, device, medium, and product

    WO2022127071A1