Defect data set auxiliary labeling method based on visual large model
Through the defect dataset assisted labeling method based on visual large model, the problem of low manual labeling efficiency in CT image defect detection is solved, efficient and accurate defect dataset labeling is achieved, and the cost and time of manual participation is reduced.
Patent Information
- Application Number
- CN202510324137.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-27
AI Technical Summary
Deep learning models have high time cost and inefficiency problems in the manual annotation process in CT image defect detection, resulting in inefficient data set annotation.
The defect dataset assisted labeling method based on visual big model is adopted. By cropping CT slice images, manually labeling a small amount of data, generating point prompts and box prompts, and inputting the image dataset and prompt information into the visual big model to perform defect prediction and label optimization, and finally obtaining high-quality defect labeling datasets.
It reduces the cost and time of manual labeling, improves the efficiency of data set labeling, ensures the accuracy and consistency of labeling results, and provides efficient and accurate technical support for the materials field and artificial intelligence field.
Smart Images

Figure CN120219344A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of CT image processing and artificial intelligence, and particularly to a method for assisting in annotating a defect data set based on a large vision model. Background Art
[0002] The rapid development of modern industry has put forward higher requirements for the reliability and quality of materials. In the fields of aerospace, nuclear industry, precision manufacturing, energy, and construction, minute defects (such as cracks, holes, inclusions, delaminations, etc.) inside materials may directly lead to a decline in product performance or catastrophic accidents. As a core technology for ensuring the integrity of materials and structures, non-destructive testing has been widely applied in industrial production. And the non-destructive testing technology based on Computed Tomography (CT) has become a key means for detecting material defects due to its high resolution, full three-dimensional imaging ability, and intuitive display ability of the internal structure of materials.
[0003] In recent years, the rapid development of deep learning technology has brought important breakthroughs to the field of CT non-destructive testing. Traditional CT detection methods have deficiencies in processing complex data and identifying diverse defects, while the defect detection technology based on deep learning is gradually changing this situation with its excellent feature extraction ability and adaptability. Compared with traditional methods, deep learning can automatically extract deep features from CT images, achieving a high degree of automation and precision in the detection process. This technological transformation not only improves the accuracy of detection results but also significantly reduces the need for manual intervention, enabling the intelligent development of material defect detection to enter a new stage.
[0004] However, the training of deep learning models relies on large-scale high-quality data sets. In practice, the construction of data sets often faces significant challenges. Manually annotating images is extremely time-consuming and inefficient, and at the same time, it places high requirements on the professional knowledge of annotators. To ensure the accuracy and consistency of annotations, multiple annotators need to annotate the same batch of data, and then comparison and correction are required. This not only increases the labor cost but also makes the model training inefficient. Summary of the Invention
[0005] To solve the problems of high time cost and low efficiency in the process of manually annotating a CT image defect data set, the present invention provides a method for assisting in annotating a defect data set based on a large vision model, which makes full use of the robustness of the large vision model, effectively optimizes the annotation process of the defect data set, reduces the cost of manual annotation, improves the efficiency of data set annotation, and provides efficient and accurate technical support for the material field and the artificial intelligence field.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A method for assisting in annotating a defect data set based on a large vision model, comprising:
[0008] Step 1, obtain a plurality of CT slice images containing defects, perform cropping processing on the CT slice images to obtain an image data set;
[0009] Step 2, manually annotate a small amount of data in the image data set to obtain manually annotated images, and generate point prompts and box prompts;
[0010] Step 3, input the image data set, the point prompts, and the box prompts into a large vision model to obtain defect prediction images;
[0011] Step 4, set a target threshold, quantitatively compare the defect prediction images with the manually annotated images, and determine whether the comparison result meets the target threshold. If the target threshold is met, obtain the optimal large vision model; if the target threshold is not met, return to Step 2, increase the point prompts until the optimal large vision model is obtained;
[0012] Step 5, input the remaining data in the image data set into the optimal large vision model, and combine the point prompt information to obtain a large number of target defect annotation data sets.
[0013] Optionally, generating the point prompts and the box prompts includes:
[0014] Randomly select a point coordinate from the defect positions annotated in the manually annotated image as the point prompt, calculate the position points of the defect annotation boundary, determine the diagonal coordinates according to the position points, and obtain the box prompt corresponding to the diagonal coordinates.
[0015] Optionally, the large vision model includes:
[0016] An image encoder for capturing the global and local features of the images in the image data set to obtain image embedding vectors;
[0017] A prompt encoder for encoding the point prompts and the box prompts into high-dimensional vectors;
[0018] An image decoder for fusing the image embedding vectors and the high-dimensional vectors, focusing on the target region of the fused image through an attention mechanism, and obtaining the defect prediction images.
[0019] Optionally, the image encoder includes:
[0020] A position encoding unit for dividing the images in the image data set into small blocks of a target size, each small block undergoing a linear transformation, being mapped into an embedding vector of a fixed dimension, and adding position encoding to each small block;
[0021] A number of self-attention mechanisms and multi-layer perception units are used to utilize the embedding vectors and the positional encoding, focus on different positions in the image, capture different spatial information, and extract the image embedding vectors.
[0022] Optionally, the prompt encoder includes:
[0023] A prompt encoding unit is used to multiply the spatial coordinates in the point prompt and the box prompt by a vector of a Gaussian distribution to generate a positional encoding, and at the same time add a one-dimensional description vector to obtain the high-dimensional vector; the one-dimensional description vector is: a one-dimensional vector describing the current point state.
[0024] Optionally, the image decoder includes:
[0025] A number of Transformer units are used to fuse the image embedding vectors and the high-dimensional vectors;
[0026] A number of transposed convolution units are used to increase the resolution of the fused image to a target value and process it through a Sigmoid activation function to obtain the defect prediction image.
[0027] Optionally, after obtaining the defect prediction image, it includes:
[0028] Perform regional connectivity denoising on the defect prediction image, remove the connected regions that do not conform to the target features, and generate a denoised image.
[0029] Optionally, obtaining the comparison result includes:
[0030] Quantitatively compare the denoised image with the manually annotated image, calculate the Dice coefficient between each denoised image and the manually annotated image, and calculate the average value of the Dice coefficients, and use the average value as the comparison result.
[0031] Optionally, calculating the Dice coefficient between each denoised image and the manually annotated image includes:
[0032]
[0033] where X′ out represents the predicted image of the post-processed vision large model, X label represents the manually annotated segmentation image, and i represents the serial number of the slice image.
[0034] Optionally, calculating the average value of the Dice coefficients includes:
[0035]
[0036] Among them, i represents the serial number of the sliced image, and Dice i represents the Dice coefficient of the i-th slice, and m represents the total number of all sliced images.
[0037] The beneficial effects of the present invention are as follows:
[0038] This method provides an auxiliary annotation method for defect datasets based on large vision models. By integrating advanced model capabilities with actual annotation requirements, it reduces the intensity of manual participation and time costs while ensuring the accuracy and consistency of annotation results. In addition, the improvement of annotation efficiency not only meets the needs of constructing large-scale defect datasets but also provides more reliable data support for scientific research and engineering applications in the field of materials. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a flowchart of an auxiliary annotation method for defect datasets based on large vision models according to an embodiment of the present invention;
[0041] Figure 2 It is a schematic diagram of image annotation of a large vision model according to an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of a material defect dataset according to an embodiment of the present invention; among them, (a) is a CT sliced image of a material with crack defects, and (b) is an image of material cracks predicted by a large vision model. Detailed Embodiments
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0045] With the emergence and development of artificial intelligence and deep learning technologies, the object surface defect detection technology based on vision technology has been greatly improved. The following is a comprehensive elaboration on the training of defect detection models:
[0046] 101. Obtain a training image set and perform data annotation on the training image set to obtain annotation information;
[0047] 102. Decode the annotation information to obtain a pixel-level annotation image and a target box annotation image;
[0048] 103. Construct a training set and a test set based on the pixel-level annotation image, the target box annotation image, and the training image set;
[0049] 104. Construct a defect detection model and a target loss function. The defect detection model includes at least a backbone segmentation sub-model, an auxiliary segmentation sub-model, and a convolutional variant sub-model;
[0050] 105. Train the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model according to the training set and the test set, and obtain the optimal parameters of the defect detection model according to the target loss function;
[0051] 106. Save the optimal parameters of the defect detection model.
[0052] In the embodiment of the present application, by performing target box annotation and pixel-level annotation on the training image set, a training set and a test set can be obtained. Furthermore, the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model can be trained through the training set and the test set. Furthermore, through the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model, the amount of training data and data annotation in the training process can be reduced.
[0053] In the embodiment of the present application, the defect detection model includes a backbone segmentation sub-model, an auxiliary segmentation sub-model, and a convolutional variant sub-model. Among them, the auxiliary segmentation sub-model is used to output an initial segmentation image and a target box, and the backbone segmentation sub-model is used to perform precise segmentation on the image based on the initial segmentation image and the target box output by the auxiliary segmentation sub-model and generate a target semantic segmentation image. The convolutional variant sub-model refines the target semantic segmentation image generated by the backbone segmentation sub-model.
[0054] In the embodiment of the present application, the target loss function can be a cross loss function.
[0055] As an optional implementation manner, the auxiliary segmentation sub-model includes a target box encoder, a segmentation decoder, and a segmentation encoder. The target box encoder takes the training set and the test set as inputs and outputs a three-dimensional binary feature tensor and a first three-dimensional feature of the training set and the test set. The segmentation encoder takes the training set and the test set as inputs and outputs a second three-dimensional feature of the training set and the test set. Furthermore, a second prediction result can be obtained according to the three-dimensional binary feature tensor, the first three-dimensional feature, and the second three-dimensional feature.
[0056] It should be noted that the second prediction result represents the initial segmentation image and target bounding box of the picture.
[0057] In this optional embodiment, further, the auxiliary segmentation sub-model further includes a first convolutional layer and a first activation function. Among them, the three-dimensional binary feature tensor and the first three-dimensional feature can be encoded according to the first convolutional layer and the first activation function to obtain an attention map, and then the attention map is element-wise multiplied with the second three-dimensional feature to obtain a feature map. Furthermore, the feature map can be used as the input of the segmentation decoder, so that the segmentation decoder outputs the second prediction result.
[0058] It can be seen that in this optional embodiment, the auxiliary segmentation sub-model includes a target bounding box encoder, a segmentation decoder, a segmentation encoder, a first convolutional layer, and a first activation function. In this way, an attention map can be obtained by outputting the three-dimensional binary feature tensor and the first three-dimensional feature through the target bounding box encoder, and then the second three-dimensional feature output by the segmentation encoder is element-wise multiplied to obtain a feature map, and then the second prediction result is obtained according to the feature map.
[0059] As an optional embodiment, the convolutional variant sub-model includes two second convolutional layers. Among them, one second convolutional layer is a 3x3 convolutional network, and the other second convolutional layer is a 3x3 convolutional network as an activation function. The second convolutional layer is used to perform a convolutional operation on the probability map (target semantic segmentation image) output by the backbone segmentation sub-model and the probability map (initial segmentation image) output by the auxiliary segmentation sub-model, so as to refine the target semantic segmentation image generated by the backbone segmentation sub-model.
[0060] Through the second convolutional layer, the convolutional variant sub-model can refine the segmentations output by the auxiliary segmentation sub-model and the backbone segmentation sub-model, and can improve the segmentation accuracy.
[0061] As an optional embodiment, after step 103: constructing a training set and a test set according to the pixel-level annotation image, the target bounding box annotation image, and the training picture set, and before step 104: training the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model according to the target loss function, the training set, and the test set and obtaining the optimal parameters of the defect detection model according to the target loss function, the defect detection model training method further includes the step of: performing data augmentation on the training set to expand the training set.
[0062] In this optional embodiment, performing data augmentation on the training set can expand the training set, increase the density of training data, and avoid overfitting of training data, thereby further improving the pre-defect recognition accuracy of the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model.
[0063] As an alternative implementation, step 105: Train the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model based on the training set and the test set, and obtain the optimal parameters of the defect detection model according to the target loss function, including the sub-steps:
[0064] Perform pixel mask prediction on the training set and the test set according to the backbone segmentation sub-model to obtain a first prediction result;
[0065] Perform target box prediction on the training set and the test set according to the auxiliary segmentation sub-model to obtain a second prediction result;
[0066] Perform self-correction on the first prediction result and the second prediction result according to the convolutional variant sub-model to obtain a corrected first prediction result and a corrected second prediction result;
[0067] Calculate the loss of the corrected first prediction result and the corrected second prediction result according to the target loss function to obtain the optimal parameters of the defect detection model.
[0068] In this alternative implementation, it is possible to perform pixel mask prediction on the training set and the test set according to the backbone segmentation sub-model to obtain a first prediction result, and it is possible to perform self-correction on the first prediction result and the second prediction result according to the convolutional variant sub-model to obtain a corrected first prediction result and a corrected second prediction result. Furthermore, perform self-correction on the first prediction result and the second prediction result according to the convolutional variant sub-model to obtain a corrected first prediction result and a corrected second prediction result, so as to calculate the loss of the corrected first prediction result and the corrected second prediction result according to the target loss function to obtain the optimal parameters of the defect detection model.
[0069] As an alternative implementation, the annotation information includes pixel-level annotation information and target box annotation information;
[0070] And, step 101: Obtain a training image set and perform data annotation on the training image set, including the sub-steps:
[0071] Divide the training image set into a first image subset and a second image subset according to a preset ratio;
[0072] Perform pixel-level annotation on the images in the first image subset to obtain pixel-level annotation information;
[0073] Perform target box annotation on the images in the second image subset to obtain target box annotation information.
[0074] In this alternative implementation, by dividing the training image set into a first image subset and a second image subset, pixel-level annotation and target box annotation can be performed on the first image subset and the second image subset respectively.
[0075] The following is a comprehensive elaboration on defect detection:
[0076] 201. Obtain the optimal parameters of the defect detection model;
[0077] 202. Predict the imaging picture of the object to be detected according to the defect detection model and the optimal parameters of the defect detection model to obtain the defect prediction result;
[0078] 203. Generate connected components according to the defect prediction result;
[0079] 204. Calculate the minimum bounding rectangle of the connected components to obtain the picture character box of the imaging picture;
[0080] 205. Perform pixel cutting and defect classification on the picture character box of the imaging picture to obtain the defect category.
[0081] In the embodiment of the present application, by importing the optimal parameters into the defect detection model of the first aspect, the defect detection of the object to be detected can be performed through the defect detection model. Then, based on the defect detection result, the minimum bounding rectangle of the connected components is calculated to obtain the picture character box of the imaging picture. Furthermore, pixel cutting and defect classification are performed according to the picture character box to obtain the defect category of the object to be detected.
[0082] The defect detection model training device includes:
[0083] The first acquisition module 301 is used to acquire the training picture set and perform data annotation on the training picture set to obtain the annotation information;
[0084] The decoding module 302 is used to decode the annotation information to obtain the pixel-level annotation image and the target box annotation image;
[0085] The first construction module 303 is used to construct the training set and the test set according to the pixel-level annotation image, the target box annotation image, and the training picture set;
[0086] The second construction module 304 is used to construct the defect detection model and the target loss function. The defect detection model at least includes a backbone segmentation sub-model, an auxiliary segmentation sub-model, and a convolutional variant sub-model;
[0087] The training module 305 is used to train the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model according to the training set and the test set, and obtain the optimal parameters of the defect detection model according to the target loss function;
[0088] The saving module 306 is used to save the optimal parameters of the defect detection model.
[0089] The device according to the embodiment of the present application can obtain a training set and a test set by performing object box annotation and pixel-level annotation on a training image set through executing a defect detection model training method, and then can train a backbone segmentation sub-model, an auxiliary segmentation sub-model, and a convolutional variant sub-model through the training set and the test set. Furthermore, the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model can reduce the amount of training data and data annotation in the training process.
[0090] In the embodiment of the present application, the defect detection model includes a backbone segmentation sub-model, an auxiliary segmentation sub-model, and a convolutional variant sub-model. Among them, the auxiliary segmentation sub-model is used to output an initial segmentation image and an object box, and the backbone segmentation sub-model is used to perform precise segmentation on the image based on the initial segmentation image and the object box output by the auxiliary segmentation sub-model and generate an object semantic segmentation image. The convolutional variant sub-model refines the object semantic segmentation image generated by the backbone segmentation sub-model.
[0091] In the embodiment of the present application, the target loss function can be a cross loss function.
[0092] As an optional implementation manner, the auxiliary segmentation sub-model includes an object box encoder, a segmentation decoder, and a segmentation encoder. The object box encoder takes the training set and the test set as inputs and outputs a three-dimensional binary feature tensor and a first three-dimensional feature of the training set and the test set. The segmentation encoder takes the training set and the test set as inputs and outputs a second three-dimensional feature of the training set and the test set. Furthermore, a second prediction result can be obtained according to the three-dimensional binary feature tensor, the first three-dimensional feature, and the second three-dimensional feature.
[0093] It should be noted that the second prediction result represents the initial segmentation image and the object box of the picture.
[0094] In this optional implementation manner, further optionally, the auxiliary segmentation sub-model further includes a first convolutional layer and a first activation function. The three-dimensional binary feature tensor and the first three-dimensional feature can be encoded according to the first convolutional layer and the first activation function to obtain an attention map, and then the attention map is multiplied element-wise with the second three-dimensional feature to obtain a feature map. Furthermore, the feature map can be used as the input of the segmentation decoder, so that the segmentation decoder outputs the second prediction result.
[0095] It can be seen that in this optional implementation manner, the auxiliary segmentation sub-model includes an object box encoder, a segmentation decoder, a segmentation encoder, a first convolutional layer, and a first activation function. In this way, an attention map can be obtained by outputting a three-dimensional binary feature tensor and a first three-dimensional feature through the object box encoder, and then the second three-dimensional feature output by the segmentation encoder is multiplied element-wise to obtain a feature map. Furthermore, the second prediction result is obtained according to the feature map.
[0096] As an alternative implementation, the convolutional variant sub-model includes two second convolutional layers. One of the second convolutional layers is a 3x3 convolutional network, and the other second convolutional layer is a 3x3 convolutional network serving as an activation function. The second convolutional layer is used to perform convolutional operations on the probability map (target semantic segmentation image) output by the backbone segmentation sub-model and the probability map (initial segmentation image) output by the auxiliary segmentation sub-model, thereby refining the target semantic segmentation image generated by the backbone segmentation sub-model.
[0097] It can be seen that through the second convolutional layer, the convolutional variant sub-model can refine the segmentations output by the auxiliary segmentation sub-model and the backbone segmentation sub-model, and can improve the segmentation accuracy.
[0098] As an alternative implementation, the defect detection model training device further includes:
[0099] A data augmentation module for augmenting the training set to expand the training set.
[0100] In this alternative implementation, augmenting the training set can expand the training set, increase the density of training data, and avoid overfitting of training data, thereby further improving the pre-defect recognition accuracy of the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model.
[0101] As an alternative implementation, the specific manner in which the training module 305 trains the backbone segmentation sub-model, the auxiliary segmentation sub-model, and the convolutional variant sub-model according to the training set and the test set and obtains the optimal parameters of the defect detection model is as follows:
[0102] Perform pixel mask prediction on the training set and the test set according to the backbone segmentation sub-model to obtain a first prediction result;
[0103] Perform target box prediction on the training set and the test set according to the auxiliary segmentation sub-model to obtain a second prediction result;
[0104] Perform self-correction on the first prediction result and the second prediction result according to the convolutional variant sub-model to obtain a corrected first prediction result and a corrected second prediction result;
[0105] Perform loss calculation on the corrected first prediction result and the corrected second prediction result according to the target loss function to obtain the optimal parameters of the defect detection model.
[0106] In this optional embodiment, the backbone segmentation sub-model can be used to perform pixel mask prediction on the training set and the test set to obtain a first prediction result, and the convolutional variant sub-model can be used to perform self-correction on the first prediction result and the second prediction result to obtain a corrected first prediction result and a corrected second prediction result. Furthermore, the convolutional variant sub-model can be used to perform self-correction on the first prediction result and the second prediction result to obtain a corrected first prediction result and a corrected second prediction result. Then, the objective loss function can be used to calculate the loss of the corrected first prediction result and the corrected second prediction result to obtain the optimal parameters of the defect detection model.
[0107] As an optional embodiment, the annotation information includes pixel-level annotation information and bounding box annotation information.
[0108] Moreover, the specific manner in which the first acquisition module 301 executes to acquire the training image set and perform data annotation on the training image set is as follows:
[0109] The training image set is divided into a first image subset and a second image subset according to a preset ratio.
[0110] Perform pixel-level annotation on the images in the first image subset to obtain pixel-level annotation information.
[0111] Perform bounding box annotation on the images in the second image subset to obtain bounding box annotation information.
[0112] In this optional embodiment, by dividing the training image set into a first image subset and a second image subset, pixel-level annotation and bounding box annotation can be respectively performed on the first image subset and the second image subset.
[0113] The defect detection device includes:
[0114] A second acquisition module 401, configured to acquire the optimal parameters of the defect detection model.
[0115] A prediction module 402, configured to perform prediction on the imaging picture of the object to be detected according to the defect detection model and the optimal parameters of the defect detection model to obtain a defect prediction result.
[0116] A generation module 403, configured to generate a connected component according to the defect prediction result.
[0117] A calculation module 404, configured to calculate the minimum bounding rectangle of the connected component to obtain a picture character box of the imaging picture.
[0118] A classification module 405, configured to perform pixel cutting and defect classification on the picture character box of the imaging picture to obtain a defect category.
[0119] In the embodiments of the present application, by importing the optimal parameters into the defect detection model of the first aspect, it is possible to detect defects in the object to be detected through the defect detection model, and then calculate the minimum circumscribed rectangle of the connected region based on the defect detection result to obtain the picture character frame of the imaging picture, and then perform pixel cutting and defect classification according to the picture character frame to obtain the defect category of the object to be detected.
[0120] As Figure 1 shown, this embodiment discloses a method for assisting in annotating a defect data set based on a vision large model, including: Step 1, obtaining a plurality of CT slice images containing defects, performing cropping processing on the CT slice images to obtain an image data set; Step 2, manually annotating a small amount of data in the image data set to obtain manually annotated images, and generating point prompts and box prompts; Step 3, inputting the image data set, point prompts, and box prompts into the vision large model to obtain defect prediction images;
[0121] Step 4, setting a target threshold, quantitatively comparing the defect prediction image with the manually annotated image, and determining whether the comparison result meets the target threshold. If the target threshold is met, the optimal vision large model is obtained; if the target threshold is not met, return to Step 2 to increase the point prompts until the optimal vision large model is obtained; Step 5, inputting the remaining data in the image data set into the optimal vision large model, and combining the point prompt information to obtain a large number of target defect annotation data sets.
[0122] Further, generating the point prompts and box prompts includes: randomly selecting a point coordinate from the defect positions annotated in the manually annotated image as the point prompt, calculating the position points of the defect annotation boundary, determining the diagonal coordinates according to the position points, and obtaining the box prompt corresponding to the diagonal coordinates.
[0123] Specifically:
[0124] In order to effectively process a large number of CT slice images of materials containing defects, first perform cropping processing on the images, and crop the continuous CT image data of materials containing defects into multiple images X image . Then randomly divide the multiple groups of data into two parts: the first part of the small amount of data is manually annotated, and the second part of the large amount of data is predicted by the model.
[0125] Then, use data annotation software to manually annotate the first part of the small amount of image data, annotating different types of defects. After the annotation is completed, set the pixel value of the background area to 0 and the pixel value of the defect area to 1, and save the annotated image X label .
[0126] Generate point prompts and box prompts based on the labeled material defect images. Randomly select a point coordinate (x, y) from the labeled defect positions as the point prompt information; calculate the position points of the defect annotation boundary: the left boundary a1, the right boundary a2, the upper boundary b1, and the lower boundary b2, and determine the diagonal coordinates of the box prompt through the boundary points: (a1, b1), (a2, b2).
[0127] Further, the vision large model includes: an image encoder for capturing the global and local features of the images in the image dataset and obtaining image embedding vectors; a prompt encoder for encoding the point prompts and box prompts into high-dimensional vectors; and an image decoder for fusing the image embedding vectors and high-dimensional vectors, focusing on the target region of the fused image through the attention mechanism, and obtaining the material defect prediction image.
[0128] Further, the image encoder includes: a position encoding unit for dividing the images in the image dataset into small blocks of a target size, each small block undergoing a linear transformation, being mapped into an embedding vector of a fixed dimension, and adding position encoding to each small block;
[0129] A number of self-attention mechanisms and multi-layer perception units for using the embedding vectors and position encoding to focus on different positions in the image, capture different spatial information, and extract image embedding vectors.
[0130] Further, the prompt encoder includes: a prompt encoding unit for multiplying the spatial coordinates in the point prompts and box prompts by vectors of a Gaussian distribution to generate position encoding, and at the same time adding a one-dimensional description vector to obtain high-dimensional vectors; the one-dimensional description vector is: a one-dimensional vector describing the current point state.
[0131] Further, the image decoder includes: a number of Transformer units for fusing the image embedding vectors and high-dimensional vectors; a number of transposed convolution units for increasing the resolution of the fused image to the target value and processing it through the Sigmoid activation function to obtain the material defect prediction image.
[0132] Specifically:
[0133] Download and deploy the weights of the pre-trained vision large model SAM, and set the Dice threshold Dice of the labeled data of the vision large model T .
[0134] Input the CT slice image X containing defects in the material image , the point prompt coordinates (x, y) and the box prompt coordinates (a1, b1), (a2, b2) into the vision large model.
[0135] Such as Figure 2As shown in the figure, the vision large model consists of an image encoder, a prompt encoder, and an image decoder. The image encoder based on Vision Transformer (ViT) has been pre-trained on a large scale and can capture the global and local features of images; the prompt encoder encodes the input point prompts and box prompts into high-dimensional vectors and fuses them with the image features; the image decoder focuses on the target area through the attention mechanism and generates a high-resolution predicted image X out . Finally, the material defect image predicted by the vision large model is saved.
[0136] The image encoder consists of a basic ViT model, and the basic ViT model consists of 12 Transformer layers. Each Transformer layer has two main modules inside: the multi-head self-attention mechanism and the multi-layer perceptron block. First, the input image is first segmented into small patches of a fixed size, and each patch is linearly transformed and mapped into an embedding vector of a fixed dimension. Then, ViT adds position encoding to each patch to ensure that the model can perceive the spatial position of each image patch in the input image. These position encodings are added to the embedding vector of each small patch and used as the input of the Transformer. Then, the Transformer layer focuses on different parts of the image through the multi-head self-attention mechanism and the multi-layer perceptron block, captures different spatial information, and extracts the image embedding vector.
[0137] The prompt encoder multiplies the prompt space coordinates by a vector of Gaussian distribution to generate position encoding, and at the same time adds a one-dimensional description vector describing the current point state. Finally, the mapped feature channels are consistent with the channels of the image embedding.
[0138] The image decoder consists of two Transformer layers and two transposed convolutional layers. The Transformer layers are used to fuse the image embedding and the prompt encoding, and the transposed convolution is used to increase the embedding resolution to 256*256. Then, the embedding is processed by the Sigmoid activation function and the image size is adjusted by bilinear interpolation to match the input size.
[0139] As Figure 3 shown, Figure 3 (a) represents the CT slice image of the material with crack defects, Figure 3 (b) represents the material crack image predicted by the vision large model.
[0140] Furthermore, after obtaining the material defect prediction image, it includes: performing regional connectivity denoising on the material defect prediction image, removing the connected regions that do not conform to the target features, and generating the denoised image.
[0141] Further, obtaining the comparison result includes: quantitatively comparing the denoised image with the manually annotated image, calculating the Dice coefficient between each denoised image and the manually annotated image, and calculating the average value of the Dice coefficients, and taking the average value as the comparison result.
[0142] Specifically:
[0143] Perform region connectivity denoising on the material defect image predicted by the visual large model, remove the connected regions that do not conform to the target features, and generate the denoised image X'. out .
[0144] Quantitatively compare the processed predicted image X' out and the manually annotated image X label to calculate the Dice coefficient Dice of each slice image i , where X' out represents the predicted image of the post-processed visual large model, X label represents the manually annotated segmentation image, and i represents the serial number of the slice image.
[0145] Then calculate the average value of these data to obtain the average Dice coefficient of this part of the data. where i represents the serial number of the slice image, Dice i represents the Dice coefficient of the i-th slice, and m represents the total number of all slice images. Determine whether Dice average reaches the set Dice coefficient threshold Dice T . If it reaches the set value, proceed to the next step: input the divided second group of large amounts of data into the large model, add point hint information in the form of manual interaction, output and visualize the material defect image predicted by the visual large model. Manually perform a secondary check on the predicted image to finally obtain a large amount of high-quality data sets; if it does not reach the set value, repeat step 2, add a point hint information (x1, y1), and input it into the model until the optimal visual large model is obtained.
[0146] To more clearly show the technical solution of the present invention, the following takes a ceramic matrix composite material sample as an example for illustration, and the specific description is as follows:
[0147] This embodiment discloses a method for assisting in annotating a material defect data set based on a visual large model, including:
[0148] Step 1: Under an ultra-high temperature thermo-mechanical-oxygen environment, perform CT scanning and image reconstruction on a ceramic matrix composite sample. After image post-processing, a total of 200 1024*1024 slice images are selected as the original data of the crack dataset of the composite material CT images. Then, the CT images are cropped to images of size 256*256, obtaining 3200 CT images. On this basis, further screening is carried out to remove the images containing edge background parts. Finally, 1000 images are selected as the original data of the deep learning model dataset, and the CT images are saved in TIFF format. The selected 1000 CT images are randomly divided into two parts: the first part has 200 images, accounting for 20% of the total data volume, and the second part has 800 images, accounting for 80% of the total data volume.
[0149] Then, use the data annotation software labelMe to manually annotate the 200 CT images in the first part, annotating cracks in different directions. After annotation, set the pixel value of the background area to 0 and the pixel value of the crack area to 1, and save the annotated images in TIFF format.
[0150] Step 2: Generate point prompts and box prompts based on the annotated CT slice images of the composite material containing cracks. From the annotated images, randomly select a point coordinate (x, y) at the position where the pixel value is 1 as the point prompt information; calculate the position points of the crack annotation boundary: the left boundary a1, the right boundary a2, the upper boundary b1, and the lower boundary b2, and determine the diagonal coordinates of the box prompt through the boundary points: (a1, b1), (a2, b2).
[0151] Step 3: Implement the deployment and operation of the large vision model SAM through Pytorch, and load the pre-trained model weights of the large vision model. Set the Dice threshold Dice T = 0.8.
[0152] Step 4: Input the CT slice image X image of the composite material containing cracks, the point prompt coordinates (x, y), and the box prompt coordinates (a1, b1), (a2, b2) into the large vision model. The large vision model consists of an image encoder, a prompt encoder, and an image decoder.
[0153] Step 4.1: Input the CT slice image X imageRead it in torch format, scale it proportionally to a size of 1024*1024, and then input it into the image encoder. The image encoder consists of a basic ViT model, and the basic ViT model consists of 12 Transformer layers. Each Transformer layer has two main modules inside: the multi-head self-attention mechanism and the multi-layer perceptron block. First, the input image is first segmented into small patches of a fixed size. Each patch undergoes a linear transformation and is mapped into an embedding vector of a fixed dimension. Then, ViT adds position encoding to each patch to ensure that the model can perceive the spatial position of each image patch in the input image. These position encodings are added to the embedding vectors of each small patch and used as the input to the Transformer. Then, the Transformer layer focuses on different parts of the image through the multi-head self-attention mechanism and the multi-layer perceptron block, captures different spatial information, extracts the image embedding vector, and finally obtains an image embedding with a length of 256.
[0154] Step 4.2: Input the point prompt coordinates (x, y) and the box prompt coordinates (a1, b1), (a2, b2) into the prompt encoder. The prompt encoder multiplies the prompt spatial coordinates by a vector of the Gaussian distribution to generate the position encoding. At the same time, a one-dimensional description vector describing the current point state is added. The description vector of the point prompt describes whether the current point is a foreground or a background, and the prompt vector of the box prompt coordinates describes whether the point is the upper-left coordinate or the lower-right coordinate of the box prompt. Finally, the mapped feature channels are consistent with those of the image embedding.
[0155] Step 4.3: Input the image embedding and the prompt embedding into the image decoder to predict the segmentation mask X out . The image decoder consists of two Transformer layers and two transposed convolutional layers. The Transformer layer is used to fuse the image embedding and the prompt encoding, and the transposed convolution is used to increase the embedding resolution to 256*256. Then, the embedding is processed by the Sigmoid activation function and the image size is adjusted by bilinear interpolation to match the input size. Finally, the composite material crack image predicted by the vision large model is saved in TIFF format, with the background pixel value set to 0 and the pixel value of the marked crack area set to 1.
[0156] Step 5: Read the composite material CT image crack data predicted by the vision large model in Numpy format, use the breadth-first search algorithm to label the connected regions in the image, and generate a region label map. Calculate the area (number of pixels) of each connected component. If its area is less than 5 pixels, it is considered a noise point, and its pixel value is set to 0 and its annotation is removed. Finally, generate the denoised image X' out , and save it in TIFF format.
[0157] Step 6: Quantitatively compare the processed predicted image X′ out and the manually annotated image X label Based on Python, read X′ out and X′ out as Numpy format, and calculate the Dice coefficient Dice of each slice image i , where X′ out represents the predicted image of the post - processed visual large model, X label represents the segmented image manually annotated, and i represents the serial number of the slice image.
[0158] Then, take the average of these data to obtain the average dice coefficient of this part of the data. where i represents the serial number of the slice image, Dice i represents the Dice coefficient of the i - th slice, and m represents the total number of all slice images.
[0159] Judge whether Dice average reaches the set dice coefficient threshold Dice T . If it reaches the set value, go to Step 7. If it does not reach the set value, randomly select a point coordinate (x1, y1) from the positions where the pixel value of the annotated image is 1, add a point hint information, input the new point hint information into the hint encoder, and re - perform the model annotation and post - processing operations until the calculated Dice average reaches the preset 0.8;
[0160] Step 7: Input the second group of 800 CT images divided into the visual large model image encoder, add point hint information in the form of manual interaction, convert the position clicked by the annotator into point hint coordinates (m, n), input it into the hint encoder, obtain the predicted crack image and save it as TIFF format, and visualize it based on opencv.
[0161] On this basis, in order to further improve the accuracy and quality of data annotation, the annotator conducts a secondary inspection on the generated image dataset. The manual review at this stage focuses on refining the boundaries of the crack area, excluding model misjudgments, and calibrating the annotations in complex areas to ensure the high credibility and high precision of the dataset. Finally, the annotator constructs a high - quality material crack dataset only through the operations of clicking on the image and checking the annotation effect, laying a solid data foundation for the research on defect extraction in the material field.
[0162] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A defect dataset auxiliary annotation method based on a large visual model, characterized in that: include: Step 1: obtaining a number of defective CT slice images, cropping the CT slice images, and obtaining an image data set; Step 2: manually annotate a small amount of data in the image data set, obtain manually annotated images, and generate point prompts and box prompts; Step 3: input the image data set, the point prompt and the frame prompt into the visual big model to obtain a defect prediction image; Step 4: Set a target threshold, perform quantitative comparison between the defect prediction image and the manually annotated image, and determine whether the comparison result meets the target threshold. If the target threshold is met, obtain the optimal visual macro model; if the target threshold is not met, return to step 2 and add the point prompt until the optimal visual macro model is obtained. Step 5: Input the remaining data in the image data set into the optimal visual large model, and combine it with the point prompt information to obtain a large number of target defect annotation data sets.
2. The defect data set auxiliary annotation method based on a visual large model according to claim 1 is characterized in that: Generating the point prompt and the frame prompt includes: A point coordinate is randomly selected from the defect position marked in the manually marked image as the point prompt, the position point of the defect marking boundary is calculated, the diagonal coordinate is determined according to the position point, and the box prompt corresponding to the diagonal coordinate is obtained.
3. The defect data set auxiliary annotation method based on a visual large model according to claim 1 is characterized in that: The visual macro model includes: An image encoder, used to capture global and local features of an image in the image dataset and obtain an image embedding vector; A prompt encoder, for encoding the point prompt and the box prompt into a high-dimensional vector; An image decoder is used to fuse the image embedding vector and the high-dimensional vector, focus on the target area of the fused image through an attention mechanism, and obtain the defect prediction image.
4. The defect data set auxiliary annotation method based on a visual large model according to claim 3 is characterized in that: The image encoder comprises: A position encoding unit, used for dividing the images in the image data set into small blocks of a target size, each small block is linearly transformed, mapped into an embedding vector of a fixed dimension, and adding a position code to each small block; A plurality of self-attention mechanisms and multi-layer perception units are used to utilize the embedding vector and the position encoding, focus on different positions in the image, capture different spatial information, and extract the image embedding vector.
5. The defect data set auxiliary annotation method based on visual big model according to claim 3 is characterized in that: The prompt encoder comprises: A prompt encoding unit is used to multiply the spatial coordinates in the point prompt and the box prompt by a Gaussian distributed vector to generate a position code, and add a one-dimensional description vector to obtain the high-dimensional vector; the one-dimensional description vector is: a one-dimensional vector that describes the current point state.
6. The defect data set auxiliary annotation method based on a visual large model according to claim 3 is characterized in that: The image decoder comprises: A plurality of Transformer units, used for fusing the image embedding vector and the high-dimensional vector; A plurality of transposed convolution units are used to increase the resolution of the fused image to a target value, and process the fused image through a Sigmoid activation function to obtain the defect prediction image.
7. The defect data set auxiliary annotation method based on a visual large model according to claim 1 is characterized in that: After obtaining the defect prediction image, the method includes: The defect prediction image is subjected to regional connectivity denoising, connected regions that do not meet target features are eliminated, and a denoised image is generated.
8. The defect data set auxiliary annotation method based on a visual large model according to claim 7 is characterized in that: Obtaining the comparison result includes: The denoised image is quantitatively compared with the manually annotated image, the Dice coefficient between each denoised image and the manually annotated image is calculated, and the average value of the Dice coefficients is calculated, and the average value is used as the comparison result.
9. The defect data set auxiliary annotation method based on visual big model according to claim 8 is characterized in that: Calculating the Dice coefficient between each denoised image and the manually annotated image includes: Among them, X′ out represents the predicted image of the post-processed visual model, X label represents the manually annotated segmented image, and i represents the sequence number of the slice image.
10. The defect data set auxiliary annotation method based on a visual large model according to claim 8, characterized in that: Calculating the average value of the Dice coefficient includes: Among them, i represents the sequence number of the slice image, Dice i represents the Dice coefficient of the i-th slice, and m represents the total number of all slice images.
Citation Information
Patent Citations
Defect detection model training method and device, defect detection method and device and storage medium
CN111784673A
Defect labeling method based on computer vision large model
CN116912830A
Image processing method and device, equipment and medium
CN117711001A
Port water area traffic condition identification method based on satellite image and SAM
CN118135426A
Key target positioning and segmentation method and system based on human fuzzy intuition driving
CN118429422A