Data enhancement method and system for mobile phone screen surface defect image segmentation data set
By fine-tuning the Stable Diffusion model through LoRA to generate semantically controllable pseudo defect images, and combining the SegFormer and ViT-B/16 models to perform pseudo segmentation and annotation quality assessment, the problem of scarce mobile phone screen defect image data is solved, and efficient and automatic dataset generation and annotation are achieved.
Patent Information
- Application Number
- CN202510686356.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to generate large-scale, high-quality datasets of mobile phone screen defect images, resulting in a shortage of visual model training data and inefficient manual labeling.
The Stable Diffusion model is fine-tuned using LoRA technology to generate semantically controllable pseudo-defect images, and the SegFormer model is combined to automatically generate pseudo-segmentation annotations. The ViT-B/16 model is used for quality assessment to build a pseudo-segmentation annotation quality screening mechanism.
It has achieved automatic and efficient generation of large-scale mobile phone screen surface defect segmentation data, reduced annotation costs, improved the quality and effectiveness of the dataset, and provided a solid foundation for subsequent model training.
Smart Images

Figure CN120635628A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and artificial intelligence data generation, and in particular to a data enhancement method and system for a mobile phone screen surface defect image segmentation dataset. Background Art
[0002] With the rapid development of smartphones, the quality of mobile phone screens, as core display and interactive components, directly impacts product performance and user experience. During the actual manufacturing process, screen surface defects such as oil stains, scratches, and spots are easily generated due to factors such as contact contamination during assembly, physical wear during cutting and handling, and inadequate quality control of raw materials. These defects can negatively impact key functions such as touch sensitivity and display quality. Therefore, efficient and accurate detection has become a crucial step in ensuring product quality.
[0003] In recent years, automated inspection technologies based on computer vision have gradually replaced traditional manual visual inspection methods in the mobile phone manufacturing industry. While visual models offer significant advantages over manual inspection in terms of stability and consistency, their performance relies heavily on large-scale, high-quality, labeled sample data. Because mobile phone screen defects rarely occur in actual production, and some minor defects are difficult to visually observe and accurately label, available training data is relatively scarce. Furthermore, the collection and labeling of defective images often requires manual review and meticulous work, resulting in inefficiencies and further limiting the rate of data accumulation.
[0004] To alleviate the problem of insufficient training data, image data augmentation technology has become a common method for improving model performance. Traditional image data augmentation methods mainly include geometric or color space transformations such as rotation, cropping, flipping, and adding noise. Although these methods can expand the dataset size to a certain extent, they are difficult to introduce diverse defect characteristics at the semantic level.
[0005] Therefore, there is an urgent need for a data enhancement method that can achieve semantically controllable defect image generation based on a small number of defect samples, and combine it with an auxiliary annotation mechanism to improve the efficiency of obtaining defect images and their annotation information, thereby alleviating the constraints of sample scarcity on the performance of existing visual models. Summary of the Invention
[0006] The purpose of the present invention is to provide a data enhancement method and system for a mobile phone screen surface defect image segmentation dataset. By combining a semantically controllable pseudo-defect image generation method and a pseudo-segmentation annotation automatic generation method, as well as a pseudo-segmentation annotation quality screening mechanism, it can automatically and efficiently generate large-scale mobile phone screen surface defect segmentation data, effectively solving the problems of scarcity and high acquisition cost of mobile phone screen surface defect image segmentation data.
[0007] In a first aspect, the present invention provides a data enhancement method for a mobile phone screen surface defect image segmentation dataset, the method comprising the following steps:
[0008] S1: Data preprocessing: Unify the raw data into a unified format, analyze the visual features of mobile phone screen defect images, and summarize key image features as prompt words to construct an "image-text" pair dataset;
[0009] S2: Fine-tune the Stable Diffusion model based on LoRA technology, build a pseudo-defect image generation model, and use prompt words to guide the model to generate semantically controllable pseudo-defect images;
[0010] S3: Fine-tune the pre-trained SegFormer model to build a pseudo-segmentation and annotation generation model, segment the defect images in the original dataset and the pseudo-defect images generated in step S2, and automatically generate the corresponding pseudo-segmentation and annotation images;
[0011] S4: A pseudo segmentation annotation quality assessment model is constructed based on the ViT-B / 16 model to effectively screen pseudo segmentation annotations.
[0012] Furthermore, step S1, the data preprocessing stage, specifically includes the following steps:
[0013] S101: Perform edge copying and filling on the original defect image and its corresponding expert segmentation and annotation image and use the Lanczos4 interpolation method to unify the format to meet the input requirements of the Stable Diffusion model;
[0014] S102: Analyze and summarize the defect feature information visible to the naked eye in the original data set as prompt words;
[0015] S103: Manually annotate each image in the dataset with text, and construct a dataset in the form of “image-text” pairs.
[0016] Furthermore, step S2, the pseudo defect image generation stage, specifically includes the following steps:
[0017] S201: The StableDiffusion V1.5 (SD1.5) model is selected as the base model. The LoRA technology is introduced to insert a trainable low-rank matrix into the Cross-Attention module in its U-Net network structure. Only the newly added low-rank matrix weight parameters and the last two Transformer Blocks of the text encoder CLIP are trained and updated, and the remaining layers remain frozen.
[0018] S202: Setting training parameters and training the model using the “image-text” dataset constructed in step S1;
[0019] S203: Input different prompt words into the trained LoRA model to guide the model to generate pseudo-defect images with specified defect type, quantity, location, shape and other attributes, thereby achieving controllable generation of defect images at the semantic level.
[0020] Furthermore, step S3, the pseudo segmentation and annotation image generation stage, specifically includes the following steps:
[0021] S301: Select the pre-trained SegFormer as the base model and load its weights pre-trained on a general semantic segmentation dataset. During fine-tuning, freeze the parameters of the first two Transformer blocks and only train the parameters of the third and fourth Transformer blocks.
[0022] S302: training the model using the original defect image and its corresponding expert segmentation and annotation image;
[0023] S303: The original defect image and the pseudo defect image generated in step S2 are input into the trained segmentation model for segmentation, and the corresponding pseudo segmentation annotation image is automatically generated.
[0024] Furthermore, step S4 is the pseudo segmentation and annotation quality assessment and screening stage, which specifically includes the following steps:
[0025] S401: Define the pseudo segmentation labeling quality evaluation index Score, which is defined as follows:
[0026] Score=ω1×Dice+ω2×IoU+ω3×Pixel Accuracy
[0027] Among them, the Dice coefficient is an indicator to measure the similarity between the model prediction value and the true value, IoU is an indicator to measure the model segmentation accuracy, Pixel Accuracy is pixel accuracy, which indicates the proportion of the model prediction value to the true value, ω1, ω2, ω3 are the weights of each indicator, satisfying ω1+ω2+ω3=1;
[0028] S402: The ViT-B / 16 model is selected as the base model, and its pre-trained weights on ImageNet1K are loaded. To adapt it to the regression task, the model is modified: the input consists of two parts: the defect image and the segmented image, and the number of input channels of the model is increased from 3 to 4 to accommodate the input image. Secondly, for the regression task, the classification output head of the model is changed to a regression output head, which only outputs a single value as the predicted comprehensive score of the input image.
[0029] S403: The pseudo segmentation annotation image obtained in step S3 and its corresponding defect image are processed using the Lanczos4 interpolation method to meet the input requirements of the modified ViT-B / 16 model;
[0030] S404: Calculate the comprehensive score Score of the pseudo-segmentation and annotation image of the original defect image obtained in step S3, and use it as a supervision signal for the pseudo-segmentation and annotation quality assessment model to train the modified ViT-B / 16 model;
[0031] S405: The pseudo defect image generated in step S2 and the corresponding pseudo segmentation and annotation image obtained in step S3 are input into the trained pseudo segmentation and annotation quality assessment model, and the corresponding score prediction value is output;
[0032] S406: Set a threshold for the score, filter the pseudo segmentation annotations according to the set threshold, and save the "pseudo image-pseudo annotation" with a score prediction value higher than the set threshold in a separate folder for use in downstream tasks.
[0033] In a second aspect, the present invention provides a data enhancement system for a mobile phone screen surface defect segmentation dataset, the system comprising the following modules:
[0034] Data preprocessing module: used to standardize the original image data, unify the data format, and construct an "image-text" pair dataset for subsequent training and inference of the pseudo-defect image generation module;
[0035] Pseudo-defect image generation module: Based on the LoRA fine-tuning Stable Diffusion model, according to the input prompt information such as defect type, quantity, location, shape, etc., the model is guided to generate pseudo-defect images containing specified defect features;
[0036] Pseudo segmentation annotation generation module: used to automatically generate pseudo segmentation annotation images corresponding to pseudo defect images;
[0037] Pseudo segmentation and annotation quality assessment module: used to perform quality assessment on pseudo segmentation and annotation images;
[0038] Data screening module: used to screen and retain high-quality "pseudo-image-pseudo-annotation" data pairs based on the scoring results output by the quality assessment module, and eliminate low-quality data;
[0039] The modules work together to complete the complete data enhancement process from raw data standardization, pseudo defect image generation, pseudo segmentation and annotation image generation to quality screening.
[0040] In a third aspect, an embodiment of the present application further provides a computing device for implementing the data enhancement method for the above-mentioned mobile phone screen surface defect segmentation dataset, the computing device comprising at least one processor and a memory, wherein the memory stores an instruction set executable by the processor; when the instruction set is executed, the processor executes the method described in the first aspect to implement the data enhancement method for the mobile phone screen surface defect segmentation dataset.
[0041] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, it executes the method described in the first aspect, thereby realizing a data enhancement method for a mobile phone screen surface defect segmentation data set.
[0042] The beneficial effects of the present invention are:
[0043] The LoRA fine-tuning technology is introduced to conduct targeted training of the Stable Diffusion model. Guided by prompt words, it can generate pseudo mobile phone screen surface defect images with specified defect type, quantity, location, shape and other attributes, realizing the controllable generation of defect images at the semantic level, overcoming the problems of difficulty in obtaining real defect images and insufficient sample quantity.
[0044] Fine-tuning training is performed based on the pre-trained SegFormer model to segment the generated pseudo mobile phone screen surface defect images, and the corresponding pseudo-segmented annotated images are automatically generated to replace manual annotation operations, effectively reducing the annotation cost and improving the degree of automation of data construction.
[0045] Based on the ViT-B / 16 model, a pseudo-segmentation and annotation quality assessment model is constructed to quantitatively evaluate and screen the quality of pseudo-segmentation and annotation samples, thereby eliminating low-quality samples and retaining high-quality "pseudo-image-pseudo-annotation" pairs. This ensures that the constructed mobile phone screen surface defect segmentation dataset has high annotation accuracy and data validity, providing a solid data foundation for subsequent model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flow chart of a data enhancement method for a mobile phone screen surface defect image segmentation dataset provided by the present invention;
[0047] Figure 2 This is a schematic diagram of the structure of the semantically controllable defect image generation model based on LoRA fine-tuning of Stable Diffusion;
[0048] Figure 3 This is a schematic diagram of the structure of the defect image pseudo-segmentation and annotation generation model based on SegFormer;
[0049] Figure 4 This is a schematic diagram of the pseudo segmentation and annotation quality assessment model based on ViT-B / 16;
[0050] Figure 5 This is a structural diagram of a data enhancement system for a mobile phone screen surface defect image segmentation dataset provided by the present invention;
[0051] Figure 6 It is a structural schematic diagram of a computing device provided by the present invention. DETAILED DESCRIPTION
[0052] Some embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0053] Figure 1 A flowchart of a data enhancement method for a mobile phone screen surface defect image segmentation dataset provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the method includes:
[0054] S1: Data preprocessing stage: The original dataset is formatted uniformly, the visual features of mobile phone screen defect images are analyzed, and key image features are summarized as prompt words to construct an "image-text" dataset;
[0055] S2: Fine-tune the Stable Diffusion model based on LoRA technology, build a pseudo-defect image generation model, and use prompt words to guide the model to generate semantically controllable pseudo-defect images;
[0056] S3: Fine-tune the pre-trained SegFormer model to build a pseudo-segmentation and annotation generation model, segment the defect images in the original dataset and the pseudo-defect images generated in step S2, and automatically generate the corresponding pseudo-segmentation and annotation images;
[0057] S4: A pseudo segmentation annotation quality assessment model is constructed based on the ViT-B / 16 model to effectively screen pseudo segmentation annotations.
[0058] Furthermore, the data preprocessing stage in step S1 is specifically as follows:
[0059] This implementation uses the publicly available mobile phone screen surface defect segmentation dataset (PKU-Market-Phone) as the raw data. This dataset contains 1,200 images of three common surface defects: oil stains, scratches, and spots, with 400 images of each type. All images were captured with an industrial camera at a resolution of 1920×1080. The images are labeled with expert segmentation information, but do not include the corresponding text prompts.
[0060] Since the Stable Diffusion model requires an input image size of 512×512, in order to adapt to the model input format and preserve the image context as much as possible, this embodiment uses edge replication to vertically pad the initial image I to expand its size to 1920×1920 to maximize the continuity of the upper and lower background textures. The padded image I' is then scaled to 512×512 using the Lanczos4 interpolation method to meet the input requirements of the Stable Diffusion model.
[0061] To enhance the generative model's ability to understand and control defect images, this example analyzes visible defect feature information in the dataset as prompt words and summarizes the following image features:
[0062] Defect type: such as oil stains, scratches, spots;
[0063] Defect location: such as upper left, upper right, lower left, lower right, center of the screen, upper middle edge, lower middle edge, middle left edge, middle right edge;
[0064] Defect shape: elongated strip, circle, point, mesh, irregular shape;
[0065] Number of defects: single, two, a small number (3-5), multiple (more than 5);
[0066] Defect size: very small, medium, large area coverage;
[0067] Phone frame type: no frame, black narrow frame, black wide frame, white narrow frame, white wide frame;
[0068] Background color: gray background, blue-gray background, dark gray background;
[0069] In order to improve the efficiency of prompt word annotation, a custom Python script is used to automatically generate text description files through candidate prompt word boxes, saving annotation time and constructing a dataset in the form of "image-text" pairs.
[0070] Furthermore, the pseudo defect image generation stage in step S2 is specifically as follows:
[0071] This example uses Stable Diffusion V1.5 (SD1.5) as the base model, introduces LoRA technology, inserts a trainable low-rank matrix into the Cross-Attention module in its U-Net network structure, and only trains and updates the newly added low-rank matrix weight parameters and the last two Transformer Blocks of the text encoder CLIP, while keeping the remaining layers frozen. This allows for rapid adaptation while maintaining the original capabilities of the model. The fine-tuning model structure is as follows: Figure 2 As shown;
[0072] The specific parameter freezing and fine-tuning strategies are as follows:
[0073] Insert the LoRA module into the Cross-Attention layer in U-Net, and only train the newly added low-rank matrix weight parameters, while the remaining weights remain frozen.
[0074] Unfreeze the last two Transformer Blocks in the text encoder CLIP to participate in training updates, while the remaining layers remain frozen, which enhances the semantic guidance ability while reducing training overhead.
[0075] The VAE module always remains frozen and does not participate in the training process;
[0076] During training, the LoRA module is inserted as follows:
[0077]
[0078] Among them, W' is the fine-tuned weight, W is the original weight, A and B are newly added trainable low-rank matrices, and α is the scaling factor used to control the perturbation amplitude.
[0079] The training goal is to minimize the model's prediction error for the noise in the diffusion process:
[0080]
[0081] Among them, z t The intermediate latent variable in the t-th step of the diffusion process, ∈ is the real noise, ∈ θ is the model prediction noise, and T(y) is the conditional vector generated by the input prompt word y through the text encoder.
[0082] To further improve semantic consistency, CLIP semantic loss is introduced:
[0083]
[0084] in, and They represent the text and image encoders in the CLIP model respectively, and the goal is to maximize the cosine similarity of image-text matching.
[0085] The final training loss function is the weighted sum of the above two losses:
[0086]
[0087] Among them, λ is a hyperparameter that adjusts the influence of the two loss terms.
[0088] The key training parameter settings are as follows:
[0089] Training data volume: The PKU-Market-Phone dataset contains three distinct types of defect images. In this example, a LoRA model is trained for each of the three types of defect images, with 400 training images for each type.
[0090] Batch size: set to 4 to accommodate a 16GB GPU.
[0091] Learning rate: U-Net is set to 0.0001, Text Encoder is set to 0.00001;
[0092] Optimizer: Adopts AdamW8bit, which has the characteristics of low memory usage, fast convergence, and weight decay;
[0093] Number of training rounds (Epoch): set to 20;
[0094] LoRA dimension and scaling factor: dim and alpha are both set to 32;
[0095] After the parameters are set, three LoRA models are trained using three sets of mobile phone screen surface defect data with different defect types. The weight file of the LoRA module is obtained, which can be merged with the original Stable Diffusion model in the inference stage:
[0096] W final =W+α·W LoRA
[0097] Without affecting the original capabilities, the model can guide the model to generate pseudo-defect images with specified defect types, quantities, locations, shapes and other attributes by inputting different prompt words, thereby realizing the controllable generation of defect images at the semantic level.
[0098] Furthermore, the pseudo segmentation annotation generation stage in step S3 is specifically as follows:
[0099] This example uses the SegFormer model as the base model and loads its pre-trained weights on a general semantic segmentation dataset. During fine-tuning, the parameters of the first two Transformer Blocks are frozen, and only the parameters of the third and fourth Transformer Blocks are trained. The model structure is as follows: Figure 3 As shown;
[0100] The training dataset uses the PKU-Market-Phone dataset, which contains 1200 images with pixel-level annotations. The annotations are divided into "defect" and "background" categories. The defect category uniformly includes oil stains, scratches, and spots. Each image and its corresponding annotation can be represented as a sample pair (x i ,y i ).
[0101] The binary cross entropy loss function is used as the optimization target in training. The loss function is defined as follows:
[0102]
[0103] Among them, y i,j Represents the pixel value at the pixel position (i, j). If the pixel belongs to the defect area, then y i,j =1, otherwise y i,j =0; The probability value predicted by the model that the position is a defect is the output of the model; N is the number of training samples, and H×W is the total number of pixels in each image.
[0104] After the training is completed, the trained SegFormer model is used to segment the defect images of the dataset PKU-Market-Phone and the pseudo-defect images generated in step S2, and the corresponding pseudo-segmentation annotation images are automatically generated.
[0105] Furthermore, the pseudo segmentation and annotation quality assessment model construction stage in step S4 is specifically as follows:
[0106] Define the pseudo segmentation annotation quality evaluation indicators as follows:
[0107] Dice coefficient: The Dice coefficient is an indicator that measures the similarity between the model prediction and the true annotation. For each image, the overlap between the defect area predicted by the model and the true annotation defect area is calculated. The formula is:
[0108]
[0109] Among them, pred is the segmentation annotation map predicted by the segmentation model, and mask is the segmentation annotation map of the expert.
[0110] IoU: IoU is another indicator to measure the accuracy of model segmentation. It calculates the ratio of the intersection and union of the predicted area and the true area. Its formula is:
[0111]
[0112] Among them, pred is the segmentation annotation map predicted by the segmentation model, and mask is the segmentation annotation map of the expert.
[0113] Pixel accuracy: Pixel accuracy measures the proportion of pixels predicted to be defective that are actually defective. The formula is:
[0114]
[0115] Among them, y i,j is the true label, is the label predicted by the model, and H×W is the total number of pixels in the image.
[0116] In order to comprehensively consider the performance of the above evaluation indicators, a weighted combination method is used to calculate the comprehensive score of each pseudo segmentation annotation. The specific weighted combination formula is as follows:
[0117] Score=ω1×Dice+ω2×IoU+ω3×Pixel Accuracy
[0118] Among them, ω1, ω2, ω3 are the weights of each indicator, satisfying ω1+ω2+ω3=1;
[0119] Construct a pseudo segmentation annotation quality assessment model to predict the comprehensive score of the generated pseudo segmentation annotation. The specific steps are as follows:
[0120] This example uses the ViT-B / 16 model as the base model and loads the model's pre-trained weights on ImageNet1K. To adapt to regression tasks, this example makes some modifications to the model: First, the input includes two parts: the image and the segmentation map. The number of input channels of the model is changed from 3 to 4 to adapt to the input image; second, corresponding to the regression task, the classification output head of the model is changed to a regression output head. The output head only outputs a numerical value as the predicted value Score of the comprehensive score of the input image. The modified model structure is as follows: Figure 4 As shown;
[0121] The pseudo segmentation annotation image obtained in step S3 and its corresponding defect image are scaled to 224×224 using the Lanczos4 interpolation method to meet the input requirements of the modified ViT-B / 16 model;
[0122] Using the corresponding expert segmentation and annotation images in the PKU-Market-Phone dataset, calculate the comprehensive score of the pseudo-segmentation and annotation images of the original defect image obtained in step S3 according to the above comprehensive score calculation formula, and use it as the supervision signal for the segmentation and annotation quality assessment model;
[0123] Each training sample can be represented as a triple where x i is the defect image, is the pseudo segmentation annotation image predicted by the model, s i Score is the comprehensive score calculated based on expert annotations;
[0124] The loss function is used to minimize the difference between the predicted value and the actual value and update the model parameter weights. The loss function used in model training is mainly divided into two parts: mean square error loss (MSE) and optimal pair ranking loss (OPRL), which are defined as follows:
[0125]
[0126] y i Indicates the actual value, Represents the predicted value, N represents the number of samples, and △ is a minimum value representing the boundary tolerance.
[0127] The pseudo defect image generated in step S2 and the pseudo segmentation and annotation image obtained in step S3 are used as input and fed into the trained pseudo segmentation and annotation quality assessment model to obtain the predicted value of Score;
[0128] Set the Score confidence threshold to 0.95, and save the "pseudo image-pseudo annotation" pairs with Score prediction values higher than the confidence threshold in a separate folder for use in downstream tasks.
[0129] Figure 5 This is a structural diagram of a data enhancement system for mobile phone screen surface defect images provided by the present invention. Figure 5 As shown, the system includes:
[0130] Data preprocessing module: used to standardize the original image data, unify the data format, and construct an "image-text" pair dataset for subsequent training and inference of the pseudo-defect image generation module;
[0131] Pseudo-defect image generation module: Based on the LoRA fine-tuning Stable Diffusion model, according to the input prompt information such as defect type, quantity, location, shape, etc., the model is guided to generate pseudo-defect images containing specified defect features;
[0132] Pseudo segmentation annotation generation module: used to automatically generate pseudo segmentation annotation images corresponding to pseudo trapped images;
[0133] Quality assessment module: used to perform quality assessment on pseudo-segmentation labeled images
[0134] Data screening module: used to screen and retain high-quality "pseudo-image-pseudo-annotation" data pairs based on the scoring results output by the quality assessment module, and eliminate low-quality data;
[0135] The modules work together to complete the complete data enhancement process from raw data standardization, pseudo image synthesis, pseudo segmentation annotation generation to quality screening.
[0136] The system embodiment of the present invention can be deployed in a computing device including a processor and a memory. The system can be implemented by software, hardware, or a combination of software and hardware. Taking software implementation as an example, the system is an integrated architecture at a logical level. Its functions are implemented by the processor in the device calling the computer program instructions stored in the memory and loading them into the memory for execution, so as to execute the data enhancement method for the mobile phone screen surface defect image proposed by the present invention. From the hardware level, if Figure 6 The figure shows a schematic diagram of the structure of a computing device provided by the present invention. The device mainly includes a storage module, a processor module and an input and output module, which are used to support the basic computing capabilities and data interaction functions of the system operation. Specifically, the storage module is used to store the programs and data resources required for the operation of the system; the processor module includes one or more computing units such as CPU and GPU, which are used to perform tasks such as image processing, model reasoning, pseudo-defect image or pseudo-segmentation annotation generation and quality assessment; the input and output module is used to import and export image data and set parameters. Based on the above basic modules, the system integrates multiple functional sub-modules, including a data preprocessing module, a pseudo-defect image generation module, a pseudo-segmentation annotation generation module, a quality assessment module, and a data screening module. The modules work together to complete image standardization processing, pseudo-defect image generation, pseudo-segmentation annotation automatic generation, data quality assessment and screening, and support the construction of a mobile phone screen surface defect image segmentation dataset for subsequent model training and testing. On the other hand, the present invention also provides a computer-readable storage medium for storing a computer program running on the computing device, which is executed to implement a data enhancement method for mobile phone screen surface defect images.
[0137] The above embodiments are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineers and technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A data enhancement method for mobile phone screen surface defect image segmentation dataset, characterized in that: include: S1: Data preprocessing: Unify the raw data into a unified format, analyze the visual features of mobile phone screen defect images, and summarize key image features as prompt words to construct an "image-text" pair dataset; S2: Fine-tune the Stable Diffusion model based on LoRA technology, build a pseudo-defect image generation model, and use prompt words to guide the model to generate semantically controllable pseudo-defect images; S3: Fine-tune the pre-trained SegFormer model to build a pseudo-segmentation and annotation generation model, segment the defect images in the original dataset and the pseudo-defect images generated in step S2, and automatically generate the corresponding pseudo-segmentation and annotation images; S4: A pseudo-segmentation and annotation quality assessment model is constructed based on the ViT-B / 16 model to effectively screen pseudo-segmentation and annotation images.
2. The method according to claim 1, wherein The data preprocessing method comprises: S101: Perform edge copying and filling on the original defect image and its corresponding expert segmentation and annotation image and use the Lanczos4 interpolation method to unify the format to meet the input requirements of the Stable Diffusion model; S102: Analyze and summarize the defect feature information visible to the naked eye in the original data set as prompt words; S103: Manually annotate each image in the dataset with text, and construct a dataset in the form of "image-text" pairs.
3. The method according to claim 1 or 2, wherein: The pseudo-defect image generation method comprises: S201: The StableDiffusion V1.5 model is selected as the base model. LoRA technology is introduced to insert a trainable low-rank matrix into the Cross-Attention module in its U-Net network structure. Only the weight parameters of the newly added low-rank matrix and the last two Transformer Blocks of the text encoder CLIP are trained and updated, while the remaining layers remain frozen. S202: Set training parameters and use the "image-text" dataset constructed in step S1 to train the basic model; S203: Input different prompt words into the trained LoRA model to guide the LoRA model to generate pseudo-defect images with specified defect types, quantities, locations, and shapes, thereby achieving controllable generation of defect images at the semantic level.
4. The method according to claim 1, wherein The pseudo-segmentation and annotation image generation method comprises: S301: Select the pre-trained SegFormer as the base model and load its weights pre-trained on a general semantic segmentation dataset. During fine-tuning, freeze the parameters of the first two Transformer blocks and only train the parameters of the third and fourth Transformer blocks. S302: Using the original defect image and its corresponding expert segmentation and annotation image to train the basic model to obtain a segmentation model; S303: The original defect image and the pseudo defect image generated in step S2 are input into the trained segmentation model for segmentation, and the corresponding pseudo segmentation annotation image is automatically generated.
5. The method according to claim 1, wherein The pseudo segmentation annotation quality assessment and screening method includes: S401: Define the pseudo segmentation labeling quality evaluation index Score, which is defined as follows: Score=ω1×Dice+ω2×IoU+ω3×Pixel Accuracy Among them, the Dice coefficient is an indicator to measure the similarity between the model prediction value and the true value, IoU is an indicator to measure the model segmentation accuracy, Pixel Accuracy is pixel accuracy, which indicates the proportion of the model prediction value to the true value, ω1, ω2, ω3 are the weights of each indicator, satisfying ω1+ω2+ω3=1; S402: Select the ViT-B / 16 model as the base model, load its pre-trained weights on ImageNet1K, and modify the ViT-B / 16 model; S403: The pseudo segmentation and annotation image obtained in step S3 and its corresponding defect image are processed using the Lanczos4 interpolation method to meet the input requirements of the modified ViT-B / 16 model; S404: Calculate the comprehensive score Score of the pseudo-segmentation and annotation image of the original defect image obtained in step S3, and use it as a supervision signal for the pseudo-segmentation and annotation quality assessment model to train the modified ViT-B / 16 model to obtain the pseudo-segmentation and annotation quality assessment model; S405: The pseudo defect image generated in step S2 and the corresponding pseudo segmentation and annotation image obtained in step S3 are input into the trained pseudo segmentation and annotation quality assessment model, and the corresponding score prediction value is output; S406: Set a threshold for the score, filter the pseudo segmentation annotations according to the set threshold, and save the "pseudo images-pseudo annotations" with score prediction values higher than the set threshold in a separate folder for use in downstream tasks.
6. The method according to claim 1, wherein S402: The ViT-B / 16 model modification includes: the input includes two parts: the defect image and the segmentation image, and the number of input channels of the model is changed from 3 to 4 to adapt to the input image; secondly, corresponding to the regression task, the classification output head of the model is changed to a regression output head, which only outputs a numerical value as the predicted value Score of the comprehensive score of the input image.
7. A data enhancement system for mobile phone screen surface defect segmentation dataset, characterized in that: include: Data preprocessing module: used to standardize the raw image data, unify the data format, and construct an "image-text" pair dataset for subsequent training and inference of the pseudo-defect image generation module; Pseudo-defect image generation module: Based on the LoRA fine-tuning Stable Diffusion model, according to the input defect type, quantity, location, and shape prompt information, the model is guided to generate pseudo-defect images containing specified defect features; Pseudo segmentation annotation generation module: used to automatically generate pseudo segmentation annotation images corresponding to pseudo defect images; Pseudo segmentation and annotation quality assessment module: used to perform quality assessment on pseudo segmentation and annotation images; Data screening module: used to screen and retain high-quality "pseudo-image-pseudo-annotation" data pairs based on the scoring results output by the quality assessment module, and eliminate low-quality data; Each module works together to complete the complete data enhancement process from raw data standardization, pseudo defect image generation, pseudo segmentation and annotation image generation to quality screening.
8. A computing device, characterized in that The computing device includes at least one processor and a memory, wherein the memory stores an instruction set executable by the processor; when the instruction set is executed, the processor executes the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a computing device, performs the method according to any one of claims 1 to 6.
Citation Information
Cited By
VCSEL chip multi-modal defect pixel-level segmentation model and method based on global-local prior aggregation and dense link context fusion
CN122391279A
VCSEL chip multi-modal defect pixel-level segmentation model and method based on global-local prior aggregation and dense link context fusion
CN122391279B