Screen detection and screen detection model training method, device and equipment
By acquiring and classifying the feature vectors of each pixel point of the screen to be detected, the problem of low accuracy of micro defect detection in the prior art is solved, and high accuracy detection of micro defects in the screen is achieved.
Patent Information
- Application Number
- CN202010042468.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-01-15
AI Technical Summary
The prior art detects tiny defects in the screen with low accuracy, making it difficult to detect tiny defects including only a few pixel points.
By acquiring the feature vectors of each pixel point in the first image of the screen to be detected and classifying the pixel points based on these feature vectors, detection of tiny defects of the screen to be detected is realized.
Improves the accuracy of screen defect detection and can detect tiny screen defects at the pixel level.
Smart Images

Figure CN113205474B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device and equipment for screen detection and screen detection model training. Background Art
[0002] During the production process of each electronic device screen, the screen may be affected by the assembly process or other reasons, resulting in defects on the screen.
[0003] Currently, defects on the screen are mainly detected through manual detection or various screen detection algorithms. However, when detecting defects on the screen through manual detection or screen detection algorithms, the accuracy of detecting smaller screen defects is low, and it is difficult to detect tiny defects that only include a few pixels. Summary of the invention
[0004] The embodiments of the present application provide a method, device and equipment for screen detection and screen detection model training to improve the accuracy of detecting tiny defects in the screen.
[0005] In the first aspect, an embodiment of the present application provides a screen detection method. When the screen needs to be detected, the feature vector of each pixel in the first image can be first obtained, and the pixels in the first image can be classified according to the feature vector of each pixel in the first image. The detection result of the screen to be detected is obtained according to the classification result, wherein the first image is an image obtained by photographing the screen to be detected.
[0006] In the above process, after the screen to be inspected is photographed to obtain a first image, a feature vector of each pixel in the first image is obtained, and then the pixels in the first image are classified according to the feature vectors of the pixels to obtain a detection result of the screen to be inspected. The detection process is performed on the pixels in the first image, and independent judgment is performed based on the feature vector of each pixel, which can achieve pixel-level segmentation of the first image. Therefore, tiny screen defects at the pixel level in the screen to be inspected can be detected, thereby improving the detection accuracy of tiny defects in screen defect detection.
[0007] In a possible implementation, the feature vector of each pixel in the first image can be obtained in the following manner: based on the first image, multiple three-dimensional feature images are determined, the first image includes M pixels in the horizontal direction, the first image includes N pixels in the vertical direction, each three-dimensional feature image includes M pixels in the horizontal direction, each three-dimensional feature image includes N pixels in the vertical direction, the number of channels of each three-dimensional feature image is C, and C is a preset number of categories; based on the multiple three-dimensional feature images, the feature vector of each pixel in the first image is obtained.
[0008] In a possible implementation, multiple three-dimensional feature images can be determined in the following manner: multiple feature extraction processes are performed on the first image to obtain multiple three-dimensional feature images; wherein the number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0009] In the above process, multiple three-dimensional feature images are obtained by performing feature extraction processing on the first image, wherein the number of feature extraction operations in each feature extraction process is different, and the multiple three-dimensional feature images obtained reflect the features of different levels in the first image. The feature vectors of the pixels in the first image obtained based on the multiple three-dimensional feature images are more conducive to the subsequent classification of the pixels.
[0010] In a possible implementation, for any feature extraction process, a feature extraction process is performed on the first image to obtain a three-dimensional feature image, including: performing a convolution operation and a downsampling operation on the first image to obtain K downsampled feature images, and the size of the i-th downsampled feature image is i is 1, 2, ..., K in sequence; convolution operation and upsampling operation are performed on the K down-sampled feature images to obtain a three-dimensional feature image.
[0011] In a possible implementation, K downsampled feature images can be obtained in the following manner: a convolution operation is performed on the first image to obtain a first downsampled feature image; the first operation is performed i times on the first downsampled feature image in sequence to obtain an i+1th downsampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a downsampling operation and a convolution operation.
[0012] In the above process, for any feature extraction process, after performing a convolution operation and a downsampling operation on the first image, a downsampled feature image of different scales will be obtained. After performing a convolution operation on the first image to obtain the first downsampled feature image, the first operation is performed on the first downsampled feature image for different times to obtain downsampled feature images at different scales, thereby extracting features of the first image at different scales.
[0013] In a possible implementation, a three-dimensional feature image can be obtained in the following manner: downsampling, convolution and upsampling operations are performed on the Kth downsampled feature image to obtain the Kth upsampled feature image; the i-th upsampled feature image and the i-th downsampled feature image are merged, convolved and upsampled in sequence to obtain the i-1th upsampled feature image, where i is K, K-1, ..., 2 in sequence; and the first upsampled image is convolved to obtain a three-dimensional feature image.
[0014] In the above process, by performing downsampling operations, convolution operations and upsampling operations on the upsampled feature image, the i-th downsampled feature image and the i-th upsampled feature image can be merged before the upsampling operation, so that the features in the i-th downsampled feature image are retained as much as possible to the next layer.
[0015] In a possible implementation, the feature vector of each pixel in the first image can be obtained in the following manner: according to the pixel value of the pixel in each three-dimensional feature image, a target three-dimensional feature image is determined, the target three-dimensional feature image includes M pixels in the horizontal direction, the target three-dimensional image includes N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C; according to the pixel value of each pixel in the target three-dimensional feature image, a feature vector of each pixel in the first image is determined.
[0016] In a possible implementation, the target three-dimensional feature image can be determined in the following manner: according to the pixel values of the M*N pixels in the x-th channel in each three-dimensional feature image, the pixel values of the M*N pixels in the x-th channel of the target three-dimensional feature image are determined, and x is 1, 2, ..., C in sequence; wherein the pixel value of the (a, b)-th pixel point in the x-th channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)-th pixel point in the x-th channel of multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
[0017] In one possible implementation, for the (a, b)th pixel in the first image, the feature vector of the (a, b)th pixel in the first image can be determined in the following manner: based on the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image, the feature vector of the (a, b)th pixel in the first image is determined.
[0018] In the above process, the feature vector of each pixel in the first image is obtained by the pixel value of each pixel in each three-dimensional feature image, wherein the target three-dimensional feature image is obtained according to each three-dimensional feature image, the size of each three-dimensional feature image and the target three-dimensional feature image is M*N, the number of channels of each three-dimensional feature image and the target three-dimensional feature image is C, and the pixel value of the pixel on any channel in the target three-dimensional feature image is the maximum value among the pixels at the same position on the corresponding channel of each target three-dimensional feature image. The feature vector of each pixel in the first image can be obtained by the pixel value of each pixel on the target three-dimensional feature image on C channels.
[0019] In a possible implementation, the feature vector of each pixel in the first image can be obtained in another way as follows: the first image is input into a feature extraction model to obtain a feature vector of each pixel in the first image, wherein the feature extraction model is learned from multiple groups of first samples, each group of first samples includes a first sample image and feature vectors of pixels on the first sample image, and the first sample image is an image obtained by photographing the first sample screen.
[0020] Through training of multiple groups of first samples, the feature extraction model can have the function of extracting feature vectors of pixels in the image. At this time, the first image is input into the feature extraction model, and the feature vector of each pixel in the first image output by the feature extraction model can be obtained.
[0021] In a possible implementation, the pixels in the first image can be classified in the following manner, and the detection result of the screen to be detected can be obtained based on the classification result: the feature vector of the pixel in the first image is input into a preset model to obtain the category of each pixel in the first image, wherein the preset model is learned from multiple groups of second samples, each group of second samples includes the feature vector and annotation information of the pixel on the second sample image, the second sample image is an image obtained by shooting the second sample screen, the annotation information is information on the category of the pixel on the second sample image, the category of any pixel in the first image is one of the preset categories, and the number of preset categories is C; according to the category of each pixel in the first image, the detection result of the screen to be detected is obtained.
[0022] In the above process, the preset model is a model obtained by pre-training, wherein the preset model is obtained by learning multiple groups of second samples, each group of second samples includes feature vectors and annotation information of pixels on the second sample image, the second sample image is an image obtained by shooting the second sample screen, the annotation information is information on the category of the pixels on the second sample image, and the category of the pixels is one of C preset categories. According to the learning of the second training sample, the preset model can have the function of classifying the pixels in the image. At this time, the feature vector of the pixel in the first image is input into the preset model, and the category of each pixel in the first image output by the preset model can be obtained. Further, if the feature vector of each pixel in the first image is obtained according to the feature extraction model, the feature extraction model and the preset model can also be used as two parts of a model to train the overall model. After the training is completed, after the first image is input into the overall model, the overall model first extracts the feature vector of each pixel in the first image, and then classifies the pixels in the first image according to the feature vector of each pixel, so as to realize the defect detection in the screen to be detected.
[0023] In a possible implementation, the category of each pixel in the first image can be obtained by the following method: inputting the feature vector of the pixel in the first image into a preset model to obtain C-1 first output images and one second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel in the second output image is used to indicate the probability that the category of the (c, d)th pixel on the first image is normal, c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N; according to the grayscale value of each pixel on the C-1 first output images and the grayscale value of each pixel on the second output image, the category of each pixel in the first image is obtained.
[0024] In a possible implementation, the category of each pixel in the first image can be obtained by the following method: for the (c, d)th pixel in the first image, obtain the grayscale value of the (c, d)th pixel on each first output image in C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image; determine the pixel with the largest grayscale value according to the grayscale value of the (c, d)th pixel on each first output image and the grayscale value of the (c, d)th pixel on the second output image; if the image where the pixel with the largest grayscale value is located is the second output image, determine the category of the (c, d)th pixel in the first image to be normal; otherwise, determine the defect category indicated by the first output image where the pixel with the largest grayscale value is located as the category of the (c, d)th pixel in the first image.
[0025] In the above process, a method for determining the category of pixels in the first image is shown. The preset model outputs C images, each of which has a size of M*N. Each first output image indicates a defect category, and the grayscale value of each pixel in the first output image indicates the probability that the pixel at the same position in the first image is the defect category indicated by the first output image. The second output image indicates a normal pixel, and the grayscale value of each pixel in the second output image indicates the probability that the pixel at the same position in the first image is a normal pixel. After obtaining C-1 first output images and one second output image, the category of each pixel in the first image can be determined.
[0026] In a second aspect, an embodiment of the present application provides a screen detection model training method, comprising: obtaining a training sample, the training sample comprising a sample image and annotation information of pixels in the sample image, the annotation information of the pixels in the sample image being information that annotates the categories of the pixels in the sample image; inputting the sample image into a screen detection model to obtain a training output category of the pixels in the sample image; adjusting parameters of the screen detection model according to the training output category of the pixels in the sample image and the annotation information of the pixels in the sample image, until the error between the training output category of the pixels in the sample image and the annotation information of the pixels in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
[0027] In the above process, the screen detection model is trained through training samples. After the screen detection model processes the sample image to obtain the training output category of the pixels in the sample image, the parameters in the screen detection model are adjusted according to the training output category and the actual category of the pixels in the sample image until the error between the two is small. The model training is completed. At this time, the trained screen detection model can classify the pixels of the input image.
[0028] In a possible implementation, the screen detection model includes a feature extraction network and a classification network; the training output category of the pixel points in the sample image can be obtained by the following method: feature extraction processing is performed on the sample image according to the feature extraction network to obtain multiple sample three-dimensional feature images, the sample image includes M pixels in the horizontal direction, and the sample image includes N pixels in the vertical direction. Each sample three-dimensional feature image includes M pixels in the horizontal direction, and each sample three-dimensional feature image includes N pixels in the vertical direction. The number of channels of each sample three-dimensional feature image is C, and C is the number of categories for classifying the pixels in the sample image; multiple sample three-dimensional feature images are processed according to the classification network to obtain the training output category of the pixels in the sample image.
[0029] In a possible implementation, multiple sample three-dimensional feature images can be obtained by the following method: multiple feature extraction processes are performed on the sample images according to the feature extraction network to obtain multiple sample three-dimensional feature images; wherein the number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0030] In the above process, a plurality of sample three-dimensional feature images are obtained by performing feature extraction processing on the sample image, wherein the number of feature extraction operations in each feature extraction processing is different, thereby extracting features of different levels in the sample image.
[0031] In a possible implementation, the feature extraction network includes a convolution layer, a pooling layer, and an upsampling layer; for any feature extraction process, a sample three-dimensional feature image can be obtained by the following method: performing convolution operations and downsampling operations on the sample image according to the convolution layer and the pooling layer to obtain K sample downsampled feature images, and the size of the i-th sample downsampled feature image is i is 1, 2, ..., K in sequence; convolution operation and upsampling operation are performed on the K sample downsampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image.
[0032] In a possible implementation, K sample down-sampled feature images can be obtained by the following method: performing a convolution operation on the sample image according to the convolution layer to obtain the first sample down-sampled feature image; performing the first operation i times in sequence on the first sample down-sampled feature image according to the convolution layer and the pooling layer to obtain the i+1th sample down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0033] In the above process, for any feature extraction process, after performing a convolution operation and a downsampling operation on the sample image, a sample downsampling feature image of different scales will be obtained. After performing a convolution operation on the sample image to obtain the first sample downsampling feature image, the first operation is performed on the first sample downsampling feature image for different times to obtain sample downsampling feature images at different scales, thereby extracting features of the sample image at different scales.
[0034] In a possible implementation, the sample three-dimensional feature image can be obtained by the following method: according to the pooling layer, the convolution layer and the upsampling layer, the K-th sample down-sampled feature image is down-sampled, convolved and up-sampled to obtain the K-th sample up-sampled feature image; according to the convolution layer and the upsampling layer, the i-th sample up-sampled feature image and the i-th sample down-sampled feature image are merged, convolved and up-sampled in sequence to obtain the i-1-th sample up-sampled feature image, where i is K, K-1, ..., 2 in sequence; according to the convolution layer, the first sample up-sampled image is convolved to obtain the sample three-dimensional feature image.
[0035] In the above process, by performing downsampling operations, convolution operations and upsampling operations on the sample upsampled feature image, the i-th sample downsampled feature image and the i-th sample upsampled feature image can be merged before the upsampling operation, so that the features in the i-th sample downsampled feature image are retained to the next layer as much as possible.
[0036] In a third aspect, an embodiment of the present application provides a screen detection device, including:
[0037] An acquisition module, used to acquire a feature vector of each pixel in a first image, where the first image is an image obtained by photographing a screen to be detected;
[0038] A classification module is used to classify the pixels in the first image according to the feature vector of each pixel in the first image, and obtain the detection result of the screen to be detected according to the classification result.
[0039] In a possible implementation, the acquisition module is specifically used to:
[0040] Determine a plurality of three-dimensional feature images according to the first image, wherein the first image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the three-dimensional feature images is C, where C is a preset number of categories;
[0041] A feature vector of each pixel in the first image is obtained according to the multiple three-dimensional feature images.
[0042] In a possible implementation, the acquisition module is specifically used to:
[0043] Performing feature extraction processing on the first image multiple times to obtain the multiple three-dimensional feature images;
[0044] The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0045] In a possible implementation, for any feature extraction process, the acquisition module is specifically used to:
[0046] Perform convolution and downsampling operations on the first image to obtain K downsampled feature images. The size of the i-th downsampled feature image is The i is 1, 2, ..., K in sequence;
[0047] A convolution operation and an upsampling operation are performed according to the K down-sampled feature images to obtain the three-dimensional feature image.
[0048] In a possible implementation, the acquisition module is specifically used to:
[0049] Performing a convolution operation on the first image to obtain a first downsampled feature image;
[0050] The first operation is performed i times in sequence on the first down-sampled feature image to obtain an (i+1)th down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0051] In a possible implementation, the acquisition module is specifically used to:
[0052] Performing a downsampling operation, a convolution operation, and an upsampling operation on the K-th down-sampled feature image to obtain a K-th up-sampled feature image;
[0053] Performing a merging operation, a convolution operation, and an upsampling operation on the i-th upsampled feature image and the i-th downsampled feature image in sequence to obtain an i-1-th upsampled feature image, where i is K, K-1, ..., 2 in sequence;
[0054] A convolution operation is performed on the first up-sampled image to obtain the three-dimensional feature image.
[0055] In a possible implementation, the acquisition module is specifically used to:
[0056] Determine a target three-dimensional feature image according to the pixel values of the pixels in each three-dimensional feature image, wherein the target three-dimensional feature image includes M pixels in the horizontal direction, and N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C;
[0057] A feature vector of each pixel in the first image is determined according to the pixel value of each pixel in the target three-dimensional feature image.
[0058] In a possible implementation, the acquisition module is specifically used to:
[0059] Determine the pixel values of the M*N pixels of the xth channel of the target three-dimensional feature image according to the pixel values of the M*N pixels of the xth channel in each three-dimensional feature image, where x is 1, 2, ..., C in sequence;
[0060] Among them, the pixel value of the (a, b)th pixel point in the xth channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)th pixel point in the xth channel of the multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
[0061] In a possible implementation manner, for the (a, b)th pixel point in the first image, the acquisition module is specifically configured to:
[0062] A feature vector of the (a, b)th pixel in the first image is determined according to the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image.
[0063] In a possible implementation, the acquisition module is specifically used to:
[0064] The first image is input into a feature extraction model to obtain a feature vector for each pixel in the first image, wherein the feature extraction model is learned from multiple groups of first samples, each group of first samples includes a first sample image and a feature vector of a pixel on the first sample image, and the first sample image is an image obtained by photographing the first sample screen.
[0065] In a possible implementation, the classification module is specifically used to:
[0066] Inputting feature vectors of pixels in the first image into a preset model to obtain a category of each pixel in the first image, wherein the preset model is learned from multiple groups of second samples, each group of second samples includes feature vectors and annotation information of pixels on a second sample image, the second sample image is an image obtained by shooting a second sample screen, the annotation information is information annotating the category of pixels on the second sample image, and the category of any pixel in the first image is one of the preset categories, and the number of preset categories is C;
[0067] A detection result of the screen to be detected is obtained according to the category of each pixel in the first image.
[0068] In a possible implementation, the classification module is specifically used to:
[0069] Inputting the feature vector of the pixel points in the first image into the preset model, obtaining C-1 first output images and one second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel point in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel point on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel point in the second output image is used to indicate the probability that the category of the (c, d)th pixel point on the first image is normal, wherein c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N;
[0070] According to the grayscale value of each pixel on the C-1 first output images and the grayscale value of each pixel on the second output image, the category of each pixel in the first image is obtained.
[0071] In a possible implementation, the classification module is specifically used to:
[0072] For the (c, d)th pixel in the first image, obtain the grayscale value of the (c, d)th pixel on each of the C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image;
[0073] Determine a pixel with a maximum grayscale value according to the grayscale value of the (c, d)th pixel on each first output image and the grayscale value of the (c, d)th pixel on the second output image;
[0074] If the image where the pixel with the maximum gray value is located is the second output image, determining that the category of the (c, d)th pixel in the first image is normal;
[0075] Otherwise, the defect category indicated by the first output image where the pixel with the maximum gray value is located is determined as the category of the (c, d)th pixel in the first image.
[0076] In a fourth aspect, an embodiment of the present application provides a screen detection model training device, comprising:
[0077] A training module, used to obtain training samples, wherein the training samples include a sample image and annotation information of pixels in the sample image, wherein the annotation information of pixels in the sample image is information that annotates the categories of the pixels in the sample image;
[0078] A processing module, used for inputting the sample image into a screen detection model to obtain a training output category of pixels in the sample image;
[0079] An adjustment module is used to adjust the parameters of the screen detection model according to the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image, until the error between the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
[0080] In a possible implementation, the screen detection model includes a feature extraction network and a classification network; the processing module is specifically used for:
[0081] Performing feature extraction processing on the sample image according to the feature extraction network to obtain a plurality of sample three-dimensional feature images, wherein the sample image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the sample three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the sample three-dimensional feature images is C, and C is the number of categories for classifying the pixels in the sample image;
[0082] The plurality of sample three-dimensional feature images are processed according to the classification network to obtain training output categories of pixel points in the sample images.
[0083] In a possible implementation manner, the processing module is specifically configured to:
[0084] Performing multiple feature extraction processes on the sample image according to the feature extraction network to obtain the multiple sample three-dimensional feature images;
[0085] The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0086] In a possible implementation, the feature extraction network includes a convolution layer, a pooling layer, and an upsampling layer; for any feature extraction process, the processing module is specifically used to:
[0087] The convolution operation and the downsampling operation are performed on the sample image according to the convolution layer and the pooling layer to obtain K sample downsampling feature images, and the size of the i-th sample downsampling feature image is The i is 1, 2, ..., K in sequence;
[0088] The convolution operation and the upsampling operation are performed on the K sample down-sampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image.
[0089] In a possible implementation manner, the processing module is specifically configured to:
[0090] Performing a convolution operation on the sample image according to the convolution layer to obtain a first sample down-sampled feature image;
[0091] According to the convolution layer and the pooling layer, the first operation is performed i times in sequence on the first sample down-sampled feature image to obtain the i+1th sample down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0092] In a possible implementation manner, the processing module is specifically configured to:
[0093] Performing downsampling operations, convolution operations, and upsampling operations on the K-th sample downsampling feature image according to the pooling layer, the convolution layer, and the upsampling layer to obtain a K-th sample upsampling feature image;
[0094] According to the convolution layer and the upsampling layer, a merging operation, a convolution operation and an upsampling operation are sequentially performed on the i-th sample upsampled feature image and the i-th sample downsampled feature image to obtain the i-1-th sample upsampled feature image, where i is K, K-1, ..., 2 in sequence;
[0095] A convolution operation is performed on the first sample upsampled image according to the convolution layer to obtain the sample three-dimensional feature image.
[0096] In a fifth aspect, an embodiment of the present application provides a screen detection device, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor runs the computer program to execute the screen detection method as described in any one of the first aspects.
[0097] In a sixth aspect, an embodiment of the present application provides a screen detection model training device, comprising: a memory and a processor, the memory storing a computer program, the processor running the computer program to execute the screen detection model training method as described in any one of the second aspects.
[0098] In a seventh aspect, an embodiment of the present application provides a screen detection system, including an image acquisition device and a screen detection device, wherein:
[0099] The image acquisition device is used to photograph the screen to be detected, obtain a first image, and send the first image to the screen detection device;
[0100] The screen detection device is used to process the first image according to the method described in any one of the first aspects to obtain a detection result of the screen to be detected.
[0101] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a computer program, and when the computer program is executed by one or more processors, the computer program implements the screen detection method described in any one of the first aspects, or, when the computer program is executed by one or more processors, the screen detection model training method described in any one of the second aspects.
[0102] The screen detection and screen detection model training method, device and equipment provided in the embodiments of the present application first obtain the feature vector of each pixel in the first image, and then classify the pixels in the first image according to the feature vector of each pixel in the first image to obtain the classification result. Since the first image is an image obtained by photographing the screen to be detected, the detection result of the screen to be detected can be obtained according to the classification result. The detection process is performed on the pixels in the first image, and independent judgment is made based on the feature vector of each pixel, which can realize pixel-level segmentation of the first image, so that tiny screen defects at the pixel level in the screen to be detected can be detected, thereby improving the accuracy of screen defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 A schematic diagram of a hybrid pixel provided in an embodiment of the present application;
[0104] Figure 2A screen detection schematic diagram provided by the prior art Figure 1 ;
[0105] Figure 3 A screen detection diagram provided by the prior art Figure 2 ;
[0106] Figure 4 A schematic diagram of an application scenario of a screen detection method provided in an embodiment of the present application;
[0107] Figure 5 A schematic diagram of a process for obtaining a feature vector of a pixel point provided in an embodiment of the present application;
[0108] Figure 6 A schematic diagram of feature extraction processing provided in an embodiment of the present application;
[0109] Figure 7 A schematic diagram of a convolution operation provided in an embodiment of the present application;
[0110] Figure 8 A schematic diagram of downsampling operation provided in an embodiment of the present application;
[0111] Fig. 9 A schematic diagram of an upsampling operation provided in an embodiment of the present application;
[0112] Fig.10 A schematic diagram of determining a target three-dimensional feature image provided in an embodiment of the present application;
[0113] Fig.11 A schematic diagram of a flow chart of another method for obtaining a feature vector of a pixel point provided in an embodiment of the present application;
[0114] Fig.12 A schematic diagram of a process for classifying pixels in a first image provided in an embodiment of the present application;
[0115] Fig.13 A schematic diagram of an output image provided in an embodiment of the present application;
[0116] Fig.14 Schematic diagram of pixel defect category detection provided in the embodiment of the present application Figure 1 ;
[0117] Fig.15 Schematic diagram of pixel defect category detection provided in the embodiment of the present application Figure 2 ;
[0118] Fig.16 A schematic diagram of a flow chart of a screen detection method provided in an embodiment of the present application;
[0119] Fig.17 A schematic diagram of a screen detection module provided in an embodiment of the present application;
[0120] Fig.18 A flowchart of a screen detection model training method provided in an embodiment of the present application;
[0121] Fig.19 A schematic diagram of training samples provided in an embodiment of the present application;
[0122] Fig. 20 A schematic diagram of the structure of a screen detection model provided in an embodiment of the present application;
[0123] Fig.21 A schematic diagram of a feature extraction provided in an embodiment of the present application;
[0124] Fig. 22 A first image provided by an embodiment of the present application;
[0125] Fig.23A A schematic diagram of line defect detection provided in an embodiment of the present application;
[0126] Fig. 23B Schematic diagram of point defect detection provided in the embodiment of the present application Figure 1 ;
[0127] Fig.23C Schematic diagram of point defect detection provided in the embodiment of the present application Figure 2 ;
[0128] Fig.24 A schematic diagram of the structure of a screen detection device provided in an embodiment of the present application;
[0129] Fig.25 A schematic diagram of the structure of a screen detection model training device provided in an embodiment of the present application;
[0130] Fig.26 A schematic diagram of the structure of a screen detection system provided in an embodiment of the present application;
[0131] Fig. 27 A schematic diagram of the hardware structure of a screen detection device provided in an embodiment of the present application;
[0132] Fig.28 A schematic diagram of the hardware structure of the screen detection model training device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0133] First, the concepts involved in this application are explained.
[0134] TFT: Thin Film Transistor.
[0135] LCD: Liquid Crystal Display, liquid crystal display, is an active matrix liquid crystal display driven by TFT.
[0136] OLED: Organic Light-Emitting Diode, organic light-emitting diode.
[0137] Screen Defect: The display screen is usually composed of multiple materials and substrate layers bonded together. It is almost impossible to bond all these layers with absolute precision every time. Various seams, migrations, contaminants, bubbles or other defects may be introduced. For example, defects on LCDs may include: impurities or foreign particles in the liquid crystal matrix, uneven distribution of the LCD matrix during the manufacturing process, uneven TFT thickness, uneven spacing between substrates, uneven brightness distribution of the backlight source, LCD panel defects, etc. All of the above factors may cause inconsistencies in light passing through the display, causing inconvenience to the user. This phenomenon is called a screen defect.
[0138] Image Segmentation: Image segmentation refers to the process of subdividing a digital image into multiple image sub-regions (a collection of pixel points).
[0139] Defect Inspection (Defect Detection): usually refers to the location and classification of object defects. The defect detection in this application is aimed at defect detection on the screen.
[0140] Mura defect: Mura defect is a common visual defect in TFT-LCD. It refers to the phenomenon of various traces caused by uneven brightness of the display screen. It manifests as low contrast, non-uniform brightness area, blurred edges, and the area of mura defect is usually larger than 1 pixel.
[0141] Pixel: also known as pixel point or pixel point, is the smallest unit that makes up a digital image.
[0142] Mixed pixels: The image signal obtained by the sensor is recorded in pixels. If a pixel contains only one type of substance, it is called a pure pixel. However, in most cases, a pixel often contains multiple types of substances. This type of pixel is a mixed pixel. For example, for a liquid crystal display, the imaging principle is due to the activation of liquid crystals of different brightness. Each liquid crystal corresponds to a pixel in the liquid crystal display. When a camera is used to shoot a liquid crystal display, each pixel in the obtained image includes several liquid crystals in the liquid crystal display, that is, it corresponds to multiple pixels in the liquid crystal display. These multiple pixels are the multiple types of substances contained in one pixel in the image taken by the camera. Next, we will combine Figure 1 An introduction to mixed pixels.
[0143] Figure 1 A schematic diagram of a mixed pixel provided in an embodiment of the present application, such as Figure 1 As shown, it includes a first image 11 of the screen to be detected and a second image 12 of the camera. In the first image 11, each box is a pixel of the screen to be detected, that is, a pixel point. For the convenience of explanation, Figure 1 Each pixel in the first image 11 is filled differently.
[0144] In the second imaging 12, each box is a picture element of the camera, that is, a pixel. Since the screen to be detected has its own resolution and the camera also has its own resolution, when the resolutions of the screen to be detected and the camera are different, at the same imaging size, the number of pixels included in the screen to be detected and the number of pixels included in the camera imaging are different.
[0145] The screen to be inspected is photographed with a camera to obtain image 13. It can be seen that when the resolution of the screen to be inspected and the camera are different, one pixel in image 13 corresponds to multiple pixels in the first imaging 11. Taking pixel 14 at the center of image 13 as an example, pixel 14 corresponds to the imaging of four pixels in the first imaging 11, that is, one pixel in image 13 contains multiple material types, which is a mixed pixel problem.
[0146] Accuracy: Accuracy is used to measure the proportion of detected screen defects that are true defects.
[0147] Recall: Recall measures the percentage of true screen defects that are detected.
[0148] Halcon: An algorithm library that provides various screen detection algorithms.
[0149] Due to the complex process and huge output in the screen production process, it is difficult to avoid defective products flowing into mobile phones, TVs and other devices. At the same time, during the screen assembly process of mobile phones, TVs and other screen-equipped devices, there is a high possibility of accidental damage to the screen due to the assembly process. Therefore, in the manufacturing process of mobile phones, TVs and other devices, the screen needs to be inspected for defects before and after the screen assembly is completed. The finished products that pass the test flow into the next process to prevent defective products from flowing into the next process. If these problem screens are not detected in time, the inferior screens will flow into the market along with the finished products of mobile phones, TVs and other devices, which will have a great impact on the use of various devices.
[0150] Figure 2 A screen detection diagram provided by the prior art Figure 1 ,like Figure 2As shown in the figure, since the grayscale of the defective pixel is not uniform with that of the surrounding pixels, the defective pixel is segmented by calculating the gradient. Specifically, the image is input first, the input image is preprocessed, and then the gradient image is calculated to find the connected domain and the connected area, and finally compared with the preset threshold to determine whether each pixel is a defective pixel.
[0151] Figure 2 The main disadvantages of the example solution include: first, the difference between the defective pixel and the surrounding pixels may not be large sometimes, and it is difficult to find a suitable threshold to ensure high accuracy and recall of detection; second, the connected areas of different defect categories vary greatly, ranging from defects as small as 1-2 pixels to scratches or penetration lines of thousands of pixels. Even for the same type of defects (such as scratches, line defects, etc.), the area occupied cannot be determined, and it is difficult to find a suitable threshold to ensure high accuracy and recall.
[0152] Figure 3 A screen detection diagram provided by the prior art Figure 2 ,like Figure 3 As shown in the figure, the scheme adopted is to detect the screen by comparing the difference between the normal image and the defective image. Specifically, the picture is first input, and then Fourier transform is performed to convert the input image into the frequency domain. Since defects are often high-frequency noise, the noise can be filtered out by frequency domain filtering, and then the non-defective image is reconstructed by inverse Fourier transform. By comparing the difference between the input image and the non-defective image, defects are found, including point mura defects, regional mura defects and linear mura defects.
[0153] As screen resolutions have become higher and higher in recent years, the liquid crystal layout of LCD screens has become more and more complex. The resolution of the camera used for detection is generally larger than the size of the liquid crystal, which easily causes mixed pixel problems during imaging, making the same screen show obvious texture characteristics. Defective pixels only appear as local non-uniformity, and the grayscale difference from other normal areas may not be large. It is also difficult to separate defects from the frequency domain. Therefore, the screen detection accuracy of this solution is also low and cannot meet production requirements.
[0154] Due to the production process and the pursuit of ultra-high resolution and ultra-high image quality, the quality of the screen needs to be strictly controlled. The current screen detection solution still has the problem of low accuracy, high false detection and missed detection rates. In order not to affect product quality, manual inspection has to be added. Manual inspection not only affects production efficiency, but also has the disadvantages of strong subjectivity, inconsistent standards, and increased costs.
[0155] In order to overcome the above shortcomings and improve the accuracy and recall rate of automatic inspection, the embodiment of the present application proposes a screen detection method to realize automatic detection of screen defects and locate and classify the defect positions.
[0156] Combine the following Figure 4 An application scenario of an embodiment of the present application is introduced.
[0157] Figure 4 A schematic diagram of an application scenario of a screen detection method provided in an embodiment of the present application, such as Figure 4 As shown, it includes a conveyor belt 41, a screen 42, a visual camera 43, a robot arm 44 and a client 45, wherein the screen 42 is placed on the conveyor belt 41 and moves with the movement of the conveyor belt 41.
[0158] The visual camera 43 is used to take a picture of the screen 42. When the screen 42 moves to a predetermined position along the conveyor belt 41, the visual camera 43 takes an image of the screen 42, and then sends the captured image to the client 45. The client 45 detects defects on the screen 42 based on the image sent by the visual camera 43. When defects are detected on the screen 42, the screen is identified as a defective product, and the robot arm 44 is controlled to intercept the defective screen to prevent the defective screen from entering the next process. If the client detects that there are no defects on the screen 42, no control instructions are sent to the robot arm 44 to intercept the screen 42.
[0159] exist Figure 4 In the example scenario, the visual camera 43 and the client 45 are two independent devices. In some scenarios, the visual camera 43 and the client 45 can be set in one device, which is a device with a camera function and sufficient processing and computing capabilities.
[0160] Since the screen 42 needs to be inspected for defects, the visual camera 43 needs to capture the screen 42 to obtain a corresponding image. When the image captured by the visual camera 43 includes other areas in addition to the area corresponding to the screen 42, the image needs to be preprocessed to remove other areas and only retain the area corresponding to the screen 42, and then the preprocessed image is analyzed and processed to determine whether there is a screen defect on the screen 42.
[0161] Figure 4The example application scenario can be applied to the production and manufacturing process of the screen, or before and after the screen assembly of the screen-equipped device, so as to control the quality of the screen. Among them, the defects of the screen are detected before the screen is assembled, and the defective screen caused by process problems can be detected. After the defective screen is intercepted before the screen is assembled, the screen without defects is assembled, and the screen can be inspected for defects again after the assembly is completed. For the screen that has no defects detected before the screen assembly but has defects detected after the screen assembly, it can be determined that it is a defect that occurred during the screen assembly process.
[0162] When the screen 42 is inspected for defects, the screen 42 may be in an off state or in an on state. When the screen 42 is in an off state, the liquid crystal molecules in the screen 42 are not activated, and at this time, the inspection is mainly performed to see if there are scratches on the screen 42. When the screen 42 is in an on state, the liquid crystal molecules in the screen 42 are activated, and at this time, the point defects, line defects, light leakage defects, etc. in the screen 42 can be inspected one by one according to the image of the screen 42 captured by the visual camera 43. In the subsequent embodiments of the present application, the example of when the visual camera 43 captures the screen 42 and the screen 42 is in an on state is used for explanation.
[0163] The technical solution shown in the present application is described in detail below through specific embodiments. It should be noted that the following embodiments can exist independently or in combination with each other, and the same or similar contents will not be described repeatedly in different embodiments.
[0164] For ease of understanding, two methods of obtaining a feature vector for each pixel in the first image are first introduced. Figure 5-Figure 10 The embodiment shown is a method of obtaining a feature vector of a pixel point. Fig.11 The illustrated embodiment is another way of obtaining the feature vector of a pixel point.
[0165] Figure 5 A flowchart of a method for obtaining a feature vector of a pixel point provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the method may include:
[0166] S51, determining a plurality of three-dimensional feature images according to the first image.
[0167] The first image is an image obtained by photographing the screen to be detected. Optionally, when other areas except the screen to be detected are photographed in the first image, the first image can be preprocessed to retain only the image area related to the screen to be detected.
[0168] The size of the first image is M*N, that is, the first image includes M pixels in the horizontal direction and N pixels in the vertical direction, and both M and N are positive integers greater than 0. Among the multiple three-dimensional feature images, each three-dimensional feature image also includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each three-dimensional feature image is C, where C is the preset number of categories. The preset number of categories C is a preset value, and after the pixels in the first image are subsequently classified, the category of any pixel in the first image is one of the C preset categories.
[0169] Optionally, multiple three-dimensional feature images can be obtained by performing feature extraction processing on the first image multiple times, wherein the number of feature extraction operations included in each two feature extraction processing is different, and the feature extraction operations include convolution operations and sampling operations.
[0170] By performing feature extraction processing on the first image at different times, features of the first image at different scales can be extracted. The number of times the feature extraction processing is performed on the first image can be set according to actual needs, and this application does not specifically limit this. A feature extraction processing is described below.
[0171] The number of times the feature extraction process is performed on the first image can be determined according to actual needs. For example, the feature extraction can be performed once, twice, or three times on the first image. For any feature extraction process, the specific operation is to perform a convolution operation and a downsampling operation on the first image to obtain K downsampled feature images. The size of the i-th downsampled feature image is i is 1, 2, ..., K. Then, a convolution operation and an upsampling operation are performed according to the K downsampled feature images to obtain a three-dimensional feature image, wherein K is a positive integer and the value of K can be preset.
[0172] The following will be combined Figure 6 The feature extraction process is described with an example.
[0173] Figure 6 The feature extraction process diagram provided in the embodiment of the present application is as follows: Figure 6 As shown, first, a convolution operation is performed on the first image to obtain a first downsampled feature image.
[0174] Then, the first operation is performed i times on the first downsampled feature image to obtain the i+1th downsampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a downsampling operation and a convolution operation. Figure 6In the above, firstly, a downsampling operation is performed on the first downsampled feature image to obtain a first scale image, and then a convolution operation is performed on the first scale image to obtain a second downsampled feature image; then a downsampling operation is performed on the second downsampled feature image to obtain a second scale image, and a convolution operation is performed on the second scale image to obtain a third downsampled feature image. Figure 6 In the example, K=3, the size of the first downsampled feature image is the same as that of the first image, which is M*N, and the size of the second downsampled feature image is The size of the third downsampled feature image is
[0175] After obtaining K down-sampled feature images, perform down-sampling operations, convolution operations, and up-sampling operations on the K-th down-sampled feature image to obtain the K-th up-sampled feature image. Figure 6 As shown in , a downsampling operation is performed on the third downsampled feature image to obtain a third scale image, a convolution operation is performed on the third scale image to obtain a fourth downsampled feature image, and an upsampling operation is performed on the fourth downsampled feature image to obtain a third upsampled feature image.
[0176] The i-th up-sampled feature image and the i-th down-sampled feature image are sequentially merged, convolved, and up-sampled to obtain the i-1-th up-sampled feature image, where i is K, K-1, ..., 2. Figure 6 As shown in , the third up-sampled feature image and the third down-sampled feature image are merged and convolved to obtain a second convolution merged image, and then the second convolution merged image is up-sampled to obtain a second up-sampled feature image. The merging operation on the third down-sampled feature image and the third up-sampled feature image is to splice the third down-sampled feature image and the third up-sampled feature image together. For example, if the size of the third down-sampled feature image is X*Y*C1 and the size of the third up-sampled feature image is X*Y*C2, the size of the image after the merging operation is X*Y*(C1+C2).
[0177] Then, the second down-sampled feature image and the second up-sampled feature image are merged and convolved to obtain a first convolution merged image, and then the first convolution merged image is up-sampled to obtain a first up-sampled feature image.
[0178] Perform convolution operation on the first upsampled image to obtain a three-dimensional feature image. Figure 6 As shown in , finally, the first down-sampled feature image and the first up-sampled feature image are merged and convolved to obtain a three-dimensional feature image.
[0179] exist Figure 6 In this paper, the convolution operation and sampling operation of the image are involved. Figure 7-Figure 9 The convolution operation and sampling operation of the image are explained respectively.
[0180] Figure 7 A schematic diagram of the convolution operation provided in the embodiment of the present application is shown in FIG. Figure 7 As shown, it includes the first image on the left and the convolution kernel on the right, where Figure 7 In the example, the first image is a three-channel image, including RGB channels. The size of the first image is 8*8*3, where 8*8 is the length and width of the first image, indicating that the first image includes 8 pixels in both the horizontal and vertical directions. 3 represents the three channels of the first image, and each pixel in the first image has a corresponding pixel value in the three channels. Figure 7 Only the 9 pixel values in the upper left corner of one channel of the first image are shown.
[0181] Figure 7 The size of the convolution kernel in the example is 3*3*3, which means that the length and width of the convolution kernel are both 3, and the depth is also 3. When the convolution operation is performed on the first image, the depth of the convolution kernel needs to be equal to the number of channels of the first image.
[0182] The first image is convolved according to the convolution kernel, such as Figure 7 As shown in , the convolution kernel includes three 3*3 weight matrices, one of which is
[0183] Since there are no pixels around the pixel in the upper left corner of the first image, pixels with a pixel value of 0 can be filled around the pixel. Then, according to the center of the weight matrix in the convolution kernel and the pixel in the upper left corner of the first image, the elements at the corresponding positions are multiplied and added to obtain the value of the first image after convolution processing on this channel.
[0184] For example, in Figure 7 In the weight matrix Perform the above processing on the pixels in the black frame in the first image:
[0185] 0*1+0*0+0*(-1)+0*1+2*0+3*(-1)+0*1+3*0+1*(-1)=-4.
[0186] Figure 7 The example of processing a pixel point in the first image according to a weight matrix in the convolution kernel is shown in FIG. When the first image includes multiple channels, the pixel points of each channel in the first image can be processed similarly according to the weight matrix in the convolution kernel to obtain the first down-sampled feature image. It should be noted that in Figure 7In the figure, the pixel values of the pixels in the first image and the weight matrix in the convolution kernel are examples and do not constitute a limitation on the pixel values of the pixels in the first image and the convolution kernel.
[0187] Optionally, the convolution kernel for performing convolution processing on the first image in the convolution operation may include one or more convolution kernels. If there is one convolution kernel, the number of channels of the first down-sampled feature image obtained is 1; if there are multiple convolution kernels, the number of channels of the first down-sampled feature image obtained is multiple, and the number of channels of the first down-sampled feature image is equal to the number of convolution kernels.
[0188] Figure 7 The process of performing convolution operation on the first image to obtain the first downsampled feature image is illustrated in FIG. When performing convolution operation on other images, the process is similar to Figure 7 When performing convolution operations on different images, the selected convolution kernels can be different, and the number of convolution kernels can also be different.
[0189] Figure 8 A schematic diagram of the downsampling operation provided in the embodiment of the present application is shown in FIG. Figure 8 As shown, it includes a first down-sampled feature image 81, wherein the first down-sampled feature image 81 is obtained by performing a convolution operation on the first image. Figure 8 The downsampling operation of the first down-sampled feature image 81 is taken as an example for explanation.
[0190] The first downsampled feature image 81 is an 8*8 image, including 64 pixels in total. Figure 8 In the Figure 8 In FIG. 8 , only the pixel value of one channel of each pixel point in the first down-sampled feature image 81 is indicated. If the first down-sampled feature image 81 includes multiple channels, the pixel value of each channel can be processed in the same way.
[0191] A downsampling operation is performed on the first downsampled feature image 81. Figure 8 In the example, every four pixels of the first down-sampled feature image 81 are converted into one pixel in the first scale image 82. When performing the down-sampling operation, the pixel value of the pixel in the first scale image 82 can be obtained by the pixel value of every four pixels in the first down-sampled feature image 81. For example, the average of every four pixels in the first down-sampled feature image 81 can be calculated to obtain the pixel value of one pixel in the first scale image 82; or the maximum value of every four pixels in the first down-sampled feature image 81 can be used as the pixel value of the corresponding pixel in the first scale image 82. Figure 8The example in FIG. 8 is to use the maximum value of every four pixels in the first down-sampled feature image 81 as the pixel value of the corresponding pixel in the first scale image 82 to obtain the first scale image 82 .
[0192] For example, in Figure 8 , the pixel values of the four pixels in the upper left corner of the first downsampled feature image 81 are 100, 120, 210 and 110 respectively, and the maximum value of the pixel values of the four pixels is 210. At this time, the pixel value of the corresponding pixel in the upper left corner of the first scale image 82 is 210.
[0193] According to the above method, the pixel value of each pixel in the first scale image 82 is obtained. Figure 8 It can be seen that the first down-sampled feature image 81 is an 8*8 image, and the first scale image 82 obtained after the above conversion is a 4*4 image.
[0194] Figure 8 The process of downsampling the first downsampled feature image to obtain the first scale image is illustrated in FIG. After obtaining the first scale image, the first scale image is convolved to obtain the second downsampled feature image. The convolution process is similar to Figure 7 After obtaining the second down-sampled feature image, the process of obtaining the third down-sampled feature image according to the second down-sampled feature image until K down-sampled feature images are obtained is similar to the process of obtaining the second down-sampled feature image according to the first down-sampled feature image, and will not be repeated here.
[0195] Fig. 9 The upsampling operation diagram provided in the embodiment of the present application is as follows: Fig. 9 As shown, it includes a first convolution merged image 91 and a first upsampled feature image 92. The first upsampled feature image 92 is obtained by upsampling the first convolution merged image 91. The first convolution merged image 91 is a 2*2 image, each small box represents a pixel point, and the pixel value of each pixel point is as follows Fig. 9 As shown, taking the pixel at the upper left corner of the first convolution merged image 91 as an example, the pixel value of the pixel is 5.
[0196] After the first convolution merged image 91 is upsampled, the pixel point in the upper left corner of the first convolution merged image 91 corresponds to the four pixel points in the upper left corner of the first upsampled feature image 92, and the pixel values of the four pixel points in the upper left corner of the first upsampled feature image 92 are related to the pixel value of the pixel point in the upper left corner of the first convolution merged image 91.
[0197] For example, one possible implementation is that the pixel values of the four pixels in the upper left corner of the first up-sampled feature image 92 are all equal to the pixel value of the pixel in the upper left corner of the first convolution merged image 91, which is 5. Fig. 9 As shown, according to this method, the pixel value of each pixel point on the first up-sampled feature image 92 can be obtained.
[0198] It should be noted that Fig. 9 The method for obtaining the pixel value of the pixel point after the up-sampling operation in the example is only an example, and the actual method for obtaining the pixel value can be determined according to needs. Fig. 9 The process of upsampling the first convolution merged image to obtain the first upsampled feature image is illustrated in FIG. 1 . The process of upsampling the i-th convolution merged image to obtain the i-th upsampled feature image is similar to the process of upsampling the i-th convolution merged image to obtain the i-th upsampled feature image. Fig. 9 The example is similar to that of , so I will not repeat it here.
[0199] S52: Obtain a feature vector of each pixel in the first image according to the multiple three-dimensional feature images.
[0200] Specifically, first, according to the pixel values of the pixels in each three-dimensional feature image, a target three-dimensional feature image is determined, the target three-dimensional feature image includes M pixels in the horizontal direction, the target three-dimensional image includes N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C. Then, according to the pixel values of each pixel in the target three-dimensional feature image, a feature vector of each pixel in the first image is determined.
[0201] Fig.10 A schematic diagram of determining a target three-dimensional feature image provided in an embodiment of the present application, such as Fig.10 As shown, there are three three-dimensional feature images, namely three-dimensional feature image 101, three-dimensional feature image 102 and three-dimensional feature image 103. The number of channels of each three-dimensional feature image is 3, and each three-dimensional feature image includes 5 pixels in the horizontal direction and 4 pixels in the vertical direction.
[0202] According to the pixel values of the M*N pixels of the xth channel in each three-dimensional feature image, determine the pixel values of the M*N pixels of the xth channel of the target three-dimensional feature image, where x is 1, 2, ..., C in sequence;
[0203] Among them, the pixel value of the (a, b)th pixel point in the xth channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)th pixel point in the xth channel of multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
[0204] like Fig.10As shown, taking the (1,1)th pixel on the first channel as an example, the pixel value of the (1,1)th pixel on the first channel of the three-dimensional feature image 101 is 2, the pixel value of the (1,1)th pixel on the first channel of the three-dimensional feature image 102 is 73, and the pixel value of the (1,1)th pixel on the first channel of the three-dimensional feature image 103 is 33. The maximum value of the three pixel values is 73, so the pixel value of the (1,1)th pixel on the first channel of the target three-dimensional feature image 104 is 73. According to Fig.10 In an exemplary manner, the pixel value of each pixel point on each channel of the target three-dimensional feature image 104 is obtained, thereby determining the target three-dimensional feature image.
[0205] After obtaining the target three-dimensional feature image, the feature vector of each pixel in the first image can be determined according to the pixel value of each pixel in the target three-dimensional feature image. Specifically, for the (a, b)th pixel in the first image, the feature vector of the (a, b)th pixel in the first image is determined according to the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image. For example, to obtain the feature vector of the (1, 1)th pixel in the first image, the pixel value of the (1, 1)th pixel in each channel of the target three-dimensional feature image is obtained. For example, the target three-dimensional feature image includes 3 channels, and the pixel values of the (1, 1)th pixel in the 3 channels are 73, 86, and 46 respectively, then the feature vector of the (1, 1)th pixel in the first image is (73, 86, 46).
[0206] Figure 5-Figure 10 A method for obtaining the feature vector of a pixel point is illustrated. Another method is described below.
[0207] Fig.11 A schematic diagram of a flow chart of another method for obtaining a feature vector of a pixel provided in an embodiment of the present application, such as Fig.11 As shown, the method may include:
[0208] S111, obtaining a feature extraction model.
[0209] The feature extraction model in the embodiment of the present application is obtained by learning multiple groups of first samples, and each group of first samples includes a first sample image and a feature vector of a pixel in the first sample image. The first sample image is an image obtained by photographing the first sample screen. If the photographed image includes other areas in addition to the area corresponding to the first sample screen, the first sample image needs to be preprocessed so that the first sample image only includes the area corresponding to the first sample screen.
[0210] S112, inputting the first image into a feature extraction model to obtain a feature vector for each pixel in the first image.
[0211] After the feature extraction model is trained, the feature extraction model has the function of obtaining the feature vector of the pixel points of the image. At this time, the first image is input into the feature extraction model to obtain the feature vector of each pixel point in the first image output by the feature extraction model.
[0212] After obtaining the feature vector of each pixel in the first image, it is necessary to classify the pixels in the first image according to the feature vector of the pixel in the first image. Fig.12 This process is described.
[0213] Fig.12 A schematic diagram of a process for classifying pixels in a first image provided in an embodiment of the present application, such as Fig.12 As shown, the method may include:
[0214] S121, inputting the feature vector of the pixel points in the first image into a preset model to obtain the category of each pixel point in the first image.
[0215] The preset model is learned from multiple groups of second samples, each group of second samples includes a second sample image and annotation information, the second sample image is an image obtained by shooting the second sample screen, and the annotation information is information annotating the categories of the pixels on the second sample image. The second sample image and the first sample image can be the same image or different images.
[0216] In the embodiment of the present application, there are multiple categories of pixels, for example, normal category, point defect category, line defect category, light leakage defect category, etc. Point defect refers to a defect with a small area on the screen involving several pixels; line defect refers to a narrow and long defect on the screen, etc.
[0217] In the process of training the preset model, in order to enable the preset model to have the function of identifying various types of defects on the screen to be detected, it is necessary to use samples containing different types of defects for training. Among them, the second sample image needs to include pixels of different categories such as point defects and line defects, and also needs to include pixels of normal categories. The number of categories of pixels in the preset model is C, and C is the number of preset categories.
[0218] After determining that multiple groups of second samples are obtained, the multiple groups of second samples can be input into a preset model, and the preset model can learn the multiple groups of second samples. Since the multiple groups of second samples include multiple second sample images, and the multiple second sample images include normal pixels and abnormal pixels of different defect categories, the preset model can learn the features of normal pixels and the features of pixels of various defect categories. At the same time, since the samples include annotation information for different defect categories, the pixels can be classified according to the learned features of normal pixels and the features of pixels of various defect categories.
[0219] After the preset model learns the multiple groups of second samples, the preset model has the function of classifying pixels. The classification of pixels is achieved according to the characteristics of different pixels. After the preset model is trained, the feature vector of the pixel in the first image is input into the preset model, and the category of each pixel in the first image output by the preset model can be obtained, wherein the category of each pixel is one of C preset categories.
[0220] Specifically, the feature vector of the pixel points in the first image is input into the preset model to obtain C-1 first output images and one second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel point in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel point on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel point in the second output image is used to indicate the probability that the category of the (c, d)th pixel point on the first image is normal, c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N; then, according to the grayscale value of each pixel point on the C-1 first output images and the grayscale value of each pixel point on the second output image, the category of each pixel point in the first image is obtained.
[0221] In the embodiment of the present application, after the feature vector of the pixel point in the first image is input into the preset model, C-1 first output images and one second output image are obtained, wherein C is the preset number of categories, each of the C-1 first output images indicates a defect category, the grayscale value of the pixel point in each first output image is used to indicate the probability that the defect category of the pixel point at the corresponding position on the first image is the defect category indicated by the first output image, and the grayscale value of the pixel point in the second output image is used to indicate the probability that the category of the pixel point at the corresponding position on the first image is normal. Then, the category of each pixel point in the first image is obtained according to the grayscale value of each pixel point on the C-1 first output images and the grayscale value of each pixel point on the second output image.
[0222] The defect types of pixels in the screen include, for example, point defects, line defects, and light leakage defects, and point defects also include black point defects and white point defects, etc. For the defect type to be detected, when training the preset model, different defect types on the sample images in the sample will first be marked differently. After the preset model training is completed, the feature vector of each pixel in the first image is input into the preset model to obtain C-1 first output images and one second output image, where C is the number of preset categories. For example, according to the preset model, only line defects in the first image are detected, and the pixel categories include normal and line defect categories, then C is equal to 2; for example, according to the preset model, white point defects, black point defects, line defects and light leakage defects in the first image are detected, and the pixel categories include normal, white point defects, black point defects, line defects and light leakage defects, then C is equal to 5. After the preset model training is completed, C is a fixed value.
[0223] Each pixel in each first output image corresponds to each pixel in the first image one-to-one, and each pixel in the second output image corresponds to each pixel in the first image one-to-one. In the embodiment of the present application, the first output image and the second output image are processed to a certain extent, and the grayscale value of each pixel in each first output image reflects the probability that the defect category of the pixel is the defect category indicated by the first output image, and the grayscale value of each pixel in the second output image reflects the probability that the pixel is a normal pixel.
[0224] The following will be combined Fig.13 This process is described.
[0225] Fig.13 The output image schematic diagram provided in the embodiment of the present application is as follows: Fig.13 As shown, it includes three first output images and one second output image, namely, a first output image 131, a first output image 132, a first output image 133, and a second output image 134. Fig.13 In the example, line defects, point defects and light leakage defects are illustrated. Line defects are defects in narrow and long areas on the screen to be detected, which may involve abnormalities in multiple pixels. For example, if there are abnormalities in the pixels of a 5*200 area on the screen, this area is a line defect. Point defects are defects in small areas on the screen to be detected, which may involve abnormalities in several or dozens of pixels. Light leakage defects usually occur at the edge of the screen. One possible cause of light leakage defects is that the edge of the screen is not pressed tightly during assembly, resulting in light leakage defects. If other types of defects are actually included, corresponding processing can also be performed.
[0226] exist Fig.13, the first output image 131 is an output image obtained by detecting line defects, and the grayscale value of each pixel on the first output image 131 is used to indicate the probability that the defect category of the pixel is a line defect; the first output image 132 is an output image obtained by detecting point defects, and the grayscale value of each pixel on the first output image 132 is used to indicate the probability that the defect category of the pixel is a point defect; the first output image 133 is an output image obtained by detecting light leakage defects, and the grayscale value of each pixel on the first output image 133 is used to indicate the probability that the defect category of the pixel is a light leakage defect; the second output image 134 is an output image obtained by detecting normal pixels, and the grayscale value of each pixel on the second output image 134 is used to indicate the probability that the pixel is a normal pixel.
[0227] Since the grayscale value of a pixel point ranges from 0 to 255, a possible implementation is to determine the grayscale value of the pixel point on the first output image or the second output image according to the probability of the category to which the pixel point belongs on the first output image or the second output image. For example, if the probability of a certain pixel point being a line defect is 0.8, the probability of being a point defect is 0.06, the probability of being a light leakage defect is 0.06, and the probability of being a normal pixel point is 0.08, then the grayscale value of the pixel point on the first output image 131 is:
[0228] 0.8*255=204;
[0229] The gray value of the pixel point on the first output image 132 is:
[0230] 0.06*255=15.3;
[0231] The gray value of the pixel point on the first output image 133 is:
[0232] 0.06*255=15.3;
[0233] The gray value of the pixel point on the second output image 134 is:
[0234] 0.08*255=20.4.
[0235] Optionally, the grayscale value obtained by the above calculation may be rounded to obtain the grayscale value of the actual output image.
[0236] Fig.14 Schematic diagram of pixel defect category detection provided in the embodiment of the present application Figure 1 ,like Fig.14 As shown, it includes a first image 140 and a first output image 131, a first output image 132, a first output image 133 and a second output image 134, wherein for the convenience of description, Fig.14Each box in represents a pixel, the color of the pixel in the first image 140 does not represent the grayscale of the pixel in the first image, and the color of the pixel in the first output image 131, the first output image 132, the first output image 133 and the second output image 134 represents the grayscale of the pixel.
[0237] Specifically, for the (c, d)th pixel in the first image, the grayscale value of the (c, d)th pixel on each of the C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image are obtained. Then, according to the grayscale value of the (c, d)th pixel on each of the first output images and the grayscale value of the (c, d)th pixel on the second output image, the pixel with the maximum grayscale value is determined.
[0238] Taking four pixels in the first image 140 as an example, to determine the category of each pixel in the first image 140, first find the corresponding pixels in three first output images and one second output image. Fig.14 As shown, it is now necessary to detect the defect categories of the pixels A, B, C and D in the first image 140. First, according to the positions of A, B, C and D in the first image 140, the corresponding pixels A1, B1, C1 and D1 in the first output image 131, the corresponding pixels A2, B2, C2 and D2 in the first output image 132, the corresponding pixels A3, B3, C3 and D3 in the first output image 133, and the corresponding pixels A4, B4, C4 and D4 in the second output image 134 are determined.
[0239] If the image where the pixel with the largest grayscale value is located is the second output image, the category of the (c, d)th pixel in the first image is determined to be normal; otherwise, the defect category indicated by the first output image where the pixel with the largest grayscale value is located is determined to be the category of the (c, d)th pixel in the first image.
[0240] Fig.15 Schematic diagram of pixel defect category detection provided in the embodiment of the present application Figure 2 ,like Fig.15 As shown, for pixel A, its corresponding four pixels in the output image are A1, A2, A3 and A4, and the grayscale values of these four pixels are 200, 12, 15 and 28 respectively. Among the grayscale values of the four pixels, A1 has the highest grayscale value. A1 is a pixel on the first output image 131, and the first output image 131 is an output image obtained by detecting line defects. Therefore, the probability that pixel A is a line defect is the highest, and the category of pixel A is determined to be a line defect.
[0241] Similarly, for pixel B, its corresponding four pixels in the output image are B1, B2, B3 and B4, and the grayscale values of these four pixels are 20, 196, 10 and 24 respectively. Among the grayscale values of the four pixels, B2 has the highest grayscale value. B2 is a pixel on the first output image 132, and the first output image 132 is an output image obtained by detecting point defects. Therefore, the probability that pixel B is a point defect is the highest, and the category of pixel B is determined to be a point defect.
[0242] For pixel C, its corresponding four pixels in the output image are C1, C2, C3 and C4, and the grayscale values of these four pixels are 5, 10, 216 and 24 respectively. Among the grayscale values of the four pixels, C3 has the highest grayscale value. C3 is a pixel on the first output image 133, and the first output image 133 is an output image obtained by detecting light leakage defects. Therefore, the probability that pixel C is a light leakage defect is the highest, and the category of pixel C is determined to be a light leakage defect.
[0243] For pixel D, its corresponding four pixel points in the output image are D1, D2, D3 and D4, and the grayscale values of these four pixel points are 5, 10, 20 and 220 respectively. Among the grayscale values of the four pixel points, the grayscale value of D4 is the highest. D4 is a pixel point on the second output image 134, and the second output image 134 is an output image obtained by detecting normal type pixels. Therefore, the probability that pixel D is of normal type is the highest, and the category of pixel D is determined to be normal.
[0244] For the defect category of the pixel, Figure 13-Figure 15 In the example, line defects, point defects and light leakage defects are used as examples for explanation. In actual situations, if other types of defects are included, such as black dot defects and white dot defects among point defects, the processing methods are similar and will not be repeated here.
[0245] The feature vector of each pixel in the first image referred to in S121 can be expressed as Figure 5-Figure 10 The method in the exemplary embodiment can also be obtained by Fig.11 The method in the exemplary embodiment is obtained.
[0246] It should be noted that if you choose Fig.11 The method in the exemplary embodiment is to obtain the feature vector of each pixel in the first image, firstly, the feature vector of each pixel in the first image is obtained by the feature extraction model, and then the category of each pixel in the first image is obtained by the preset model. Optionally, in the embodiment of the present application, the feature extraction model and the preset model can be trained as one model.
[0247] S122, obtaining a detection result of the screen to be detected according to the category of each pixel in the first image.
[0248] After obtaining the category of each pixel in the first image according to the above method, the specific defects in the first image and the corresponding defect categories can be obtained, and then the corresponding detection results of the screen to be detected can be obtained, including the number of defects on the screen to be detected, the defect category of each defect and the specific position of each defect on the screen to be detected, etc.
[0249] The screen detection method provided in the embodiment of the present application first obtains the feature vector of each pixel in the first image, and then classifies the pixels in the first image according to the feature vector of each pixel in the first image to obtain the classification result. Since the first image is an image obtained by photographing the screen to be detected, the detection result of the screen to be detected can be obtained according to the classification result. The detection process is performed on the pixels in the first image, and independent judgment is performed based on the feature vector of each pixel, which can realize pixel-level segmentation of the first image, so that tiny screen defects at the pixel level in the screen to be detected can be detected, thereby improving the accuracy of screen defect detection. This solution can perform automatic defect detection on the screen, does not require manual preset thresholds, does not require manual intervention, does not require image resampling, is not limited by image size, and is not limited by image acquisition equipment.
[0250] Next, combine Fig.16 , a screen detection method is introduced.
[0251] Fig.16 A schematic diagram of the flow of the screen detection method provided in the embodiment of the present application is shown in FIG. Fig.16 As shown, the method may include:
[0252] S161, obtaining a feature vector of each pixel in a first image, where the first image is an image obtained by photographing the screen to be inspected.
[0253] It should be noted that this can be done through Figure 5-Figure 10 The method shown in the embodiment obtains the feature vector of each pixel in the first image, or, by Fig.11 The method shown obtains the feature vector of each pixel in the first image, which will not be described in detail here.
[0254] S162, classifying the pixels in the first image according to the feature vector of each pixel in the first image, and obtaining a detection result of the screen to be detected according to the classification result.
[0255] It should be noted that this can be done through Figure 12-Figure 15 The method shown in the embodiment classifies the pixels in the first image to obtain the detection result of the screen to be detected, which will not be described in detail here.
[0256] Optionally, the solution provided in the embodiment of the present application mainly includes four modules: Fig.17 A schematic diagram of a screen detection module provided in an embodiment of the present application, such as Fig.17 As shown, it includes a data module, a feature extraction module, a training module and a prediction module, wherein:
[0257] The data module is used to collect data. Specifically, the data module collects an image of the sample screen and provides image annotation information. The annotation information may be, for example, a mask image of the same size as the image, marking the corresponding defect category at the corresponding defect position with a specified color.
[0258] The feature extraction module is used to extract the feature vector of each pixel on the first image, wherein the method of extracting the feature vector of the pixel is as described in the above embodiment, and each pixel corresponds to a feature vector.
[0259] The training module is used to train the preset model based on the feature vector of each pixel on the image and the mask image marked with the defect category, so that the multi-scale feature vector of each pixel can be output as the correct defect type after passing through the preset model.
[0260] The prediction module is used to make predictions based on the feature vector of the pixel point and the trained preset model.
[0261] Using the preset model trained by the training module, the feature vector of the pixel in the first image is input into the preset model to obtain the category of each pixel in the first image, and the defect category of the screen to be detected is obtained according to the category of each pixel in the first image. If the training model adopts the method of obtaining the feature vector of the pixel by the feature extraction model, the trained model can perform the function of extracting the feature vector of the pixel and classifying the pixel according to the feature vector. At this time, the first image is input into the feature extraction model to obtain the feature vector of each pixel in the first image, and the preset model is used to classify the feature vector, and the category information is output to predict the defect category of the first image.
[0262] Fig.18 A flowchart of the screen detection model training method provided in the embodiment of the present application is shown in FIG. Fig.18 As shown, including:
[0263] S181, obtaining a training sample, wherein the training sample includes a sample image and annotation information of pixels in the sample image, wherein the annotation information of pixels in the sample image is information that annotates categories of pixels in the sample image.
[0264] Before training the screen detection model, training samples must be obtained. In the embodiment of the present application, the training samples include one or more groups, each group of training samples includes a sample image and annotation information of pixels in the sample image, and the annotation information annotates the categories of pixels in the sample image.
[0265] Fig.19 A schematic diagram of a training sample provided in an embodiment of the present application, such as Fig.19 As shown, the left side is a sample image 191, which is an image obtained by photographing a sample screen, and the right side is a defect mask image 192 of the sample image 191. The defect information on the sample image 191 is marked on the defect mask image 192, and different defects are marked in different ways. The sample image 191 and the defect mask image 192 constitute a set of training samples.
[0266] S182, inputting the sample image into a screen detection model to obtain a training output category of pixels in the sample image.
[0267] After obtaining the training samples, for a group of training samples, the sample images are input into the screen detection model. The screen detection model processes the sample images, such as convolution processing, sampling processing, etc., and then obtains the training output categories of the pixels in the sample images.
[0268] S183, adjusting the parameters of the screen detection model according to the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image, until the error between the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
[0269] The training output category is the category obtained by the screen detection model by classifying the pixels in the sample image. The training output category may have errors with the categories of the pixels in the sample image annotated in the annotation information. Therefore, the parameters of the screen detection model are adjusted according to the training output category of the pixels in the sample image and the annotation information of the pixels in the sample image. After multiple adjustments to the parameters in the screen detection model, when the error between the training output category of the pixels in the sample image and the annotation information of the pixels in the sample image is less than or equal to the preset error, a trained screen detection model is obtained.
[0270] Optionally, the screen detection model in the embodiment of the present application includes a feature extraction network and a classification network, and the feature extraction network is a feature extraction model of a multi-scale U-Net structure. Fig. 20 A schematic diagram of the structure of the screen detection model provided in the embodiment of the present application, such as Fig. 20As shown, the feature extraction network is composed of a multi-scale U-Net network structure, wherein the sample image is subjected to multiple feature extraction processes to obtain multiple sample three-dimensional feature images, wherein the sample image includes M pixels in the horizontal direction and N pixels in the vertical direction, each sample three-dimensional feature image includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each sample three-dimensional feature image is C, and C is the number of categories for classifying the pixels in the sample image. Fig. 20 Each branch in is a feature extraction process. After obtaining multiple sample three-dimensional feature images, the multiple sample three-dimensional feature images are processed according to the classification network to obtain the training output categories of the pixel points in the sample images.
[0271] Optionally, the sample image can be subjected to multiple feature extraction processes according to the feature extraction network to obtain multiple sample three-dimensional feature images, wherein the number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0272] The feature extraction network includes a convolution layer, a pooling layer, and an upsampling layer. For any feature extraction process, the convolution operation and downsampling operation can be performed on the sample image according to the convolution layer and the pooling layer to obtain K sample downsampled feature images. The size of the i-th sample downsampled feature image is i is 1, 2, ..., K in sequence; then, convolution operations and upsampling operations are performed on the K sample downsampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature images.
[0273] Among them, an optional implementation method for obtaining K sample down-sampled feature images is to perform a convolution operation on the sample image according to the convolution layer to obtain the first sample down-sampled feature image; perform the first operation on the first sample down-sampled feature image i times in sequence according to the convolution layer and the pooling layer to obtain the i+1th sample down-sampled feature image, i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0274] Fig. 20 The first column of boxes in the figure represents a 1x1 convolution, which is used to increase the dimension and increase nonlinearity. After multiple feature extractions, global pooling is performed to combine the results of each feature extraction and output them as the final feature map. Fig. 20 The f(z) in the equation is the classification of pixels and the output result.
[0275] After obtaining K down-sampled feature images, optionally, a down-sampling operation, a convolution operation and an up-sampling operation can be performed on the K-th sample down-sampled feature image according to the pooling layer, the convolution layer and the up-sampling layer to obtain the K-th sample up-sampled feature image; then, a merging operation, a convolution operation and an up-sampling operation are performed on the i-th sample up-sampled feature image and the i-th sample down-sampled feature image in sequence according to the convolution layer and the up-sampling layer to obtain the i-1-th sample up-sampled feature image, where i is K, K-1, ..., 2 in sequence; and a convolution operation is performed on the first sample up-sampled image according to the convolution layer to obtain a sample three-dimensional feature image.
[0276] Although U-Net itself has the ability to extract some multi-scale features, it is difficult for a single U-Net structure to achieve good results for problems with huge scale differences such as screen defects. Therefore, the embodiment of the present application designs a feature extraction network of a multi-scale U-Net combination. Using the feature extraction network, the sample image is subjected to multiple feature extraction processes to extract the feature vectors of the pixel points of the sample image under different receptive fields to distinguish the defect categories, which can solve the problem of feature vector extraction for multiple defect types. Finally, multiple sample three-dimensional feature images are globally pooled, and the pixels in the sample image are classified using a classification network.
[0277] Fig.21 A schematic diagram of a feature extraction provided in an embodiment of the present application, such as Fig.21 As shown, it consists of two parts. The first part (left side) is the feature extraction part of U-Net, which is composed of a basic convolutional network. This part will have a new scale after each pooling layer. The second part (right side) is the upsampling part of U-Net. Each time upsampling is performed, the scale of the corresponding channels of the feature extraction part of U-Net will be fused and spliced, so that the features of this scale can be retained as much as possible to the next layer. Through multi-level feature extraction, as much information of the original sample image as possible can be retained, making the subsequent classification more accurate.
[0278] In order to verify the effect of the solution of the present application, the following test was conducted.
[0279] Table 1 is a statistical table of defects detected by the screen detection solution provided by the embodiment of the present application. As shown in Table 1, after data testing on the production line, the statistical results obtained are as follows:
[0280] Table 1
[0281]
[0282] In Table 1, the traditional visual inspection algorithm package Halcon has weak detection capabilities for the four defects of black dots, white dots, lines, and light leakage. Halcon has no detection capabilities for black dots and light leakage, and Halcon detected 3 and 5 white dots and line defects respectively, which are far less than the solution of this application. Since light leakage is a subjective defect and is affected by the labeled samples, accuracy information is more important.
[0283] Fig. 22 The first image provided in the embodiment of the present application is obtained by photographing the screen to be inspected with a camera. The first image is processed to obtain a defect detection result on the screen to be inspected. Fig.23A A schematic diagram of line defect detection provided in an embodiment of the present application, Fig. 23B Schematic diagram of point defect detection provided in the embodiment of the present application Figure 1 , Fig.23C Schematic diagram of point defect detection provided in the embodiment of the present application Figure 2 ,It can be seen that for line defects, the ,detection results are complete, and for point defects, the scheme of ,this application also has high detection accuracy and recall rate.
[0284] Fig.24 A schematic diagram of the structure of the screen detection device provided in the embodiment of the present application is shown in FIG. Fig.24 As shown, the screen detection device 24 may include an acquisition module 241 and a classification module 242, wherein:
[0285] The acquisition module 241 is used to acquire a feature vector of each pixel in a first image, where the first image is an image obtained by photographing the screen to be detected;
[0286] The classification module 242 is used to classify the pixels in the first image according to the feature vector of each pixel in the first image, and obtain the detection result of the screen to be detected according to the classification result.
[0287] In a possible implementation, the acquisition module 241 is specifically used to:
[0288] Determine a plurality of three-dimensional feature images according to the first image, wherein the first image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the three-dimensional feature images is C, where C is a preset number of categories;
[0289] A feature vector of each pixel in the first image is obtained according to the multiple three-dimensional feature images.
[0290] In a possible implementation, the acquisition module 241 is specifically used to:
[0291] Performing feature extraction processing on the first image multiple times to obtain the multiple three-dimensional feature images;
[0292] The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0293] In a possible implementation, for any feature extraction process, the acquisition module 241 is specifically used to:
[0294] Perform convolution and downsampling operations on the first image to obtain K downsampled feature images. The size of the i-th downsampled feature image is The i is 1, 2, ..., K in sequence;
[0295] A convolution operation and an upsampling operation are performed according to the K down-sampled feature images to obtain the three-dimensional feature image.
[0296] In a possible implementation, the acquisition module 241 is specifically used to:
[0297] Performing a convolution operation on the first image to obtain a first downsampled feature image;
[0298] The first operation is performed i times in sequence on the first down-sampled feature image to obtain an (i+1)th down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0299] In a possible implementation, the acquisition module 241 is specifically used to:
[0300] Performing a downsampling operation, a convolution operation, and an upsampling operation on the K-th down-sampled feature image to obtain a K-th up-sampled feature image;
[0301] Performing a merging operation, a convolution operation, and an upsampling operation on the i-th upsampled feature image and the i-th downsampled feature image in sequence to obtain an i-1-th upsampled feature image, where i is K, K-1, ..., 2 in sequence;
[0302] A convolution operation is performed on the first up-sampled image to obtain the three-dimensional feature image.
[0303] In a possible implementation, the acquisition module 241 is specifically used to:
[0304] Determine a target three-dimensional feature image according to the pixel values of the pixels in each three-dimensional feature image, wherein the target three-dimensional feature image includes M pixels in the horizontal direction, and N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C;
[0305] A feature vector of each pixel in the first image is determined according to the pixel value of each pixel in the target three-dimensional feature image.
[0306] In a possible implementation, the acquisition module 241 is specifically used to:
[0307] Determine the pixel values of the M*N pixels of the xth channel of the target three-dimensional feature image according to the pixel values of the M*N pixels of the xth channel in each three-dimensional feature image, where x is 1, 2, ..., C in sequence;
[0308] Among them, the pixel value of the (a, b)th pixel point in the xth channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)th pixel point in the xth channel of the multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
[0309] In a possible implementation, for the (a, b)th pixel in the first image, the acquisition module 241 is specifically configured to:
[0310] A feature vector of the (a, b)th pixel in the first image is determined according to the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image.
[0311] In a possible implementation, the acquisition module 241 is specifically used to:
[0312] The first image is input into a feature extraction model to obtain a feature vector for each pixel in the first image, wherein the feature extraction model is learned from multiple groups of first samples, each group of first samples includes a first sample image and a feature vector of a pixel on the first sample image, and the first sample image is an image obtained by photographing the first sample screen.
[0313] In a possible implementation, the classification module 242 is specifically used for:
[0314] Inputting feature vectors of pixels in the first image into a preset model to obtain a category of each pixel in the first image, wherein the preset model is learned from multiple groups of second samples, each group of second samples includes feature vectors and annotation information of pixels on a second sample image, the second sample image is an image obtained by shooting a second sample screen, the annotation information is information annotating the category of pixels on the second sample image, and the category of any pixel in the first image is one of the preset categories, and the number of preset categories is C;
[0315] A detection result of the screen to be detected is obtained according to the category of each pixel in the first image.
[0316] In a possible implementation, the classification module 242 is specifically used for:
[0317] Inputting the feature vector of the pixel points in the first image into the preset model, obtaining C-1 first output images and one second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel point in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel point on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel point in the second output image is used to indicate the probability that the category of the (c, d)th pixel point on the first image is normal, wherein c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N;
[0318] According to the grayscale value of each pixel on the C-1 first output images and the grayscale value of each pixel on the second output image, the category of each pixel in the first image is obtained.
[0319] In a possible implementation, the classification module 242 is specifically used for:
[0320] For the (c, d)th pixel in the first image, obtain the grayscale value of the (c, d)th pixel on each of the C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image;
[0321] Determine a pixel with a maximum grayscale value according to the grayscale value of the (c, d)th pixel on each first output image and the grayscale value of the (c, d)th pixel on the second output image;
[0322] If the image where the pixel with the maximum gray value is located is the second output image, determining that the category of the (c, d)th pixel in the first image is normal;
[0323] Otherwise, the defect category indicated by the first output image where the pixel with the maximum gray value is located is determined as the category of the (c, d)th pixel in the first image.
[0324] It should be noted that the screen detection device shown in the embodiment of the present application can execute the screen detection method shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be repeated here.
[0325] Fig.25A schematic diagram of the structure of the screen detection model training device provided in the embodiment of the present application, such as Fig.25 As shown, the screen detection model training device 25 includes a training module 251, a processing module 252 and an adjustment module 253, wherein:
[0326] The training module 251 is used to obtain training samples, wherein the training samples include a sample image and annotation information of pixels in the sample image, wherein the annotation information of pixels in the sample image is information that annotates the categories of the pixels in the sample image;
[0327] The processing module 252 is used to input the sample image into the screen detection model to obtain the training output category of the pixel points in the sample image;
[0328] The adjustment module 253 is used to adjust the parameters of the screen detection model according to the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image, until the error between the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
[0329] In a possible implementation, the screen detection model includes a feature extraction network and a classification network; the processing module 252 is specifically used for:
[0330] Performing feature extraction processing on the sample image according to the feature extraction network to obtain a plurality of sample three-dimensional feature images, wherein the sample image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the sample three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the sample three-dimensional feature images is C, and C is the number of categories for classifying the pixels in the sample image;
[0331] The plurality of sample three-dimensional feature images are processed according to the classification network to obtain training output categories of pixel points in the sample images.
[0332] In a possible implementation, the processing module 252 is specifically configured to:
[0333] Performing multiple feature extraction processes on the sample image according to the feature extraction network to obtain the multiple sample three-dimensional feature images;
[0334] The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
[0335] In a possible implementation, the feature extraction network includes a convolution layer, a pooling layer, and an upsampling layer; for any feature extraction process, the processing module 252 is specifically used to:
[0336] The convolution operation and the downsampling operation are performed on the sample image according to the convolution layer and the pooling layer to obtain K sample downsampling feature images, and the size of the i-th sample downsampling feature image is The i is 1, 2, ..., K in sequence;
[0337] The convolution operation and the upsampling operation are performed on the K sample down-sampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image.
[0338] In a possible implementation, the processing module 252 is specifically configured to:
[0339] Performing a convolution operation on the sample image according to the convolution layer to obtain a first sample down-sampled feature image;
[0340] According to the convolution layer and the pooling layer, the first operation is performed i times in sequence on the first sample down-sampled feature image to obtain the i+1th sample down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
[0341] In a possible implementation, the processing module 252 is specifically configured to:
[0342] Performing downsampling operations, convolution operations, and upsampling operations on the K-th sample downsampling feature image according to the pooling layer, the convolution layer, and the upsampling layer to obtain a K-th sample upsampling feature image;
[0343] According to the convolution layer and the upsampling layer, a merging operation, a convolution operation and an upsampling operation are sequentially performed on the i-th sample upsampled feature image and the i-th sample downsampled feature image to obtain the i-1-th sample upsampled feature image, where i is K, K-1, ..., 2 in sequence;
[0344] A convolution operation is performed on the first sample upsampled image according to the convolution layer to obtain the sample three-dimensional feature image.
[0345] The screen detection model training device shown in the embodiment of the present application can execute the screen detection model training method shown in the above method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.
[0346] Fig.26 A schematic diagram of the structure of the screen detection system provided in the embodiment of the present application is shown in FIG. Fig.26As shown, it includes an image acquisition device 261 and a screen detection device 262, wherein:
[0347] The image acquisition device 261 is used to photograph the screen to be detected, obtain a first image, and send the first image to the screen detection device 262;
[0348] The screen detection device 262 is used to process the first image according to the screen detection method in the above embodiment to obtain a detection result of the screen to be detected.
[0349] The image acquisition device 261 is a device for photographing the screen to be detected and obtaining a first image. The image acquisition device 261 can be a camera or other device with image acquisition capabilities. The screen detection device 262 is used to perform Fig.24 The actions performed by the acquisition module and the classification model in the example. In some embodiments, the image acquisition device 261 and the screen detection device 262 are two independent devices. In some embodiments, the image acquisition device 261 and the screen detection device 262 can also be integrated into one device, which is not limited in the embodiments of the present application.
[0350] Fig. 27 This is a hardware structure diagram of the screen detection device provided in the embodiment of the present application. Fig. 27 The screen detection device 270 includes: a memory 271 and a processor 272, wherein the memory 271 and the processor 272 communicate with each other; illustratively, the memory 271 and the processor 272 can communicate with each other through a communication bus 273, the memory 271 is used to store a computer program, and the processor 272 executes the computer program to implement the various steps of the screen detection method in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0351] Optionally, the memory 271 may be independent or integrated with the processor 272 .
[0352] The screen detection device provided in this embodiment can be used to execute the screen detection method in the above embodiments. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0353] Optionally, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0354] Fig.28 A schematic diagram of the hardware structure of the screen detection model training device provided in the embodiment of the present application. Fig.28 The screen detection model training device 280 includes: a memory 281 and a processor 282, wherein the memory 281 and the processor 282 communicate; illustratively, the memory 281 and the processor 282 can communicate via a communication bus 283, the memory 281 is used to store a computer program, and the processor 282 executes the computer program to implement the various steps of the screen detection model training method in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0355] Optionally, the memory 281 may be independent or integrated with the processor 282 .
[0356] The screen detection model training device provided in this embodiment can be used to execute the screen detection model training method in the above embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0357] Optionally, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the application may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0358] The present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to implement the screen detection method described in the above embodiment, or the computer program is used to implement the screen detection model training method described in the above embodiment.
[0359] The embodiment of the present application also provides a chip or integrated circuit, including: a memory and a processor;
[0360] The memory is used to store program instructions and sometimes also to store intermediate data;
[0361] The processor is used to call the program instructions stored in the memory to implement the screen detection method or screen detection model training method as described above.
[0362] Optionally, the memory may be independent or integrated with the processor. In some implementations, the memory may be located outside the chip or integrated circuit.
[0363] An embodiment of the present application also provides a program product, which includes a computer program stored in a storage medium, and the computer program is used to implement the above-mentioned screen detection method or screen detection model training method.
[0364] All or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned memory (storage medium) includes: read-only memory (English: read-only memory, abbreviated: ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape (English: magnetic tape), floppy disk (English: floppy disk), optical disc (English: optical disc) and any combination thereof.
[0365] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0366] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0367] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0368] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
[0369] In the present application, the term "include" and its variations may refer to non-restrictive inclusion; the term "or" and its variations may refer to "and / or". The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. In the present application, "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previously associated objects are in an "or" relationship.
Claims
1. A screen detection method, characterized in that: include: Obtaining a feature vector of each pixel in a first image, where the first image is an image obtained by photographing a screen to be detected; The category of any pixel point in the first image is one of the preset categories, and the number of preset categories is C; Input the feature vector of each pixel in the first image into a preset model to obtain C-1 first output images and one second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel in the second output image is used to indicate the probability that the category of the (c, d)th pixel on the first image is normal, wherein c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N; Obtaining a category of each pixel in the first image according to the grayscale value of each pixel on the C-1 first output image and the grayscale value of each pixel on the second output image; A detection result of the screen to be detected is obtained according to the category of each pixel in the first image.
2. The method according to claim 1, characterized in that Obtain the feature vector of each pixel in the first image, including: Determine a plurality of three-dimensional feature images according to the first image, wherein the first image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the three-dimensional feature images is C, where C is a preset number of categories; According to the multiple three-dimensional feature images, a feature vector of each pixel in the first image is obtained.
3. The method according to claim 2, characterized in that Determining a plurality of three-dimensional feature images according to the first image includes: Performing feature extraction processing on the first image multiple times to obtain the multiple three-dimensional feature images; The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
4. The method according to claim 3, characterized in that For any feature extraction process, performing feature extraction process on the first image to obtain the three-dimensional feature image includes: Perform convolution and downsampling operations on the first image to obtain K downsampled feature images. The size of the i-th downsampled feature image is , i is 1, 2, ..., K in sequence; A convolution operation and an upsampling operation are performed according to the K down-sampled feature images to obtain the three-dimensional feature image.
5. The method according to claim 4, characterized in that Performing a convolution operation and a downsampling operation on the first image to obtain K downsampled feature images, including: Performing a convolution operation on the first image to obtain a first downsampled feature image; The first operation is performed i times in sequence on the first down-sampled feature image to obtain an (i+1)th down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
6. The method according to claim 4 or 5, characterized in that: Performing a convolution operation and an upsampling operation according to the K down-sampled feature images to obtain the three-dimensional feature image includes: Performing a downsampling operation, a convolution operation, and an upsampling operation on the K-th down-sampled feature image to obtain a K-th up-sampled feature image; Performing a merging operation, a convolution operation, and an upsampling operation on the i-th upsampled feature image and the i-th downsampled feature image in sequence to obtain an i-1-th upsampled feature image, where i is K, K-1, ..., 2 in sequence; A convolution operation is performed on the first up-sampled image to obtain the three-dimensional feature image.
7. The method according to any one of claims 2 to 5, characterized in that: Acquiring a feature vector of each pixel in the first image according to the multiple three-dimensional feature images includes: Determine a target three-dimensional feature image according to the pixel values of the pixels in each three-dimensional feature image, wherein the target three-dimensional feature image includes M pixels in the horizontal direction, and N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C; A feature vector of each pixel in the first image is determined according to the pixel value of each pixel in the target three-dimensional feature image.
8. The method according to claim 7, characterized in that Determine the target three-dimensional feature image according to the pixel value of each pixel point in each three-dimensional feature image, including: According to the xth channel in each 3D feature image The pixel value of pixels determines the xth channel of the target three-dimensional feature image. The pixel value of pixels, wherein x is 1, 2, ..., C in sequence; Among them, the pixel value of the (a, b)th pixel point in the xth channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)th pixel point in the xth channel of the multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
9. The method according to claim 8, characterized in that For the (a, b)th pixel point in the first image, determining a feature vector of the (a, b)th pixel point in the first image according to the pixel values of each pixel point in the target three-dimensional feature image, including: A feature vector of the (a, b)th pixel in the first image is determined according to the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image.
10. The method according to claim 1, characterized in that Obtain the feature vector of each pixel in the first image, including: The first image is input into a feature extraction model to obtain a feature vector for each pixel in the first image, wherein the feature extraction model is learned from multiple groups of first samples, each group of first samples includes a first sample image and feature vectors of pixels in the first sample image, and the first sample image is an image obtained by photographing the first sample screen.
11. The method according to any one of claims 2-5, 8-10, characterized in that: The preset model is learned from multiple groups of second samples, each group of second samples includes feature vectors and annotation information of pixels on a second sample image, the second sample image is an image obtained by shooting the second sample screen, and the annotation information is information that annotates the categories of pixels on the second sample image.
12. The method according to claim 1, characterized in that Obtaining a category of each pixel in the first image according to the grayscale value of each pixel on the C-1 first output image and the grayscale value of each pixel on the second output image, including: For the (c, d)th pixel in the first image, obtain the grayscale value of the (c, d)th pixel on each of the C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image; Determine a pixel with a maximum grayscale value according to the grayscale value of the (c, d)th pixel on each first output image and the grayscale value of the (c, d)th pixel on the second output image; If the image where the pixel with the maximum gray value is located is the second output image, determining that the category of the (c, d)th pixel in the first image is normal; Otherwise, the defect category indicated by the first output image where the pixel with the maximum grayscale value is located is determined as the category of the (c, d)th pixel in the first image.
13. The method according to claim 1, characterized in that The method is applied to a screen detection model, which includes a feature extraction network and the preset model; the screen detection model is trained using the following method: Acquire a training sample, wherein the training sample includes a sample image and annotation information of pixels in the sample image, wherein the annotation information of the pixels in the sample image is information that annotates the categories of the pixels in the sample image; Performing feature extraction processing on the sample image according to the feature extraction network to obtain a plurality of sample three-dimensional feature images, wherein the sample image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the sample three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the sample three-dimensional feature images is C, and C is the number of categories for classifying the pixels in the sample image; Processing the plurality of sample three-dimensional feature images according to the preset model to obtain training output categories of pixel points in the sample images; According to the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image, the parameters of the screen detection model are adjusted until the error between the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
14. The method according to claim 13, characterized in that Performing feature extraction processing on the sample image according to the feature extraction network to obtain a plurality of sample three-dimensional feature images, including: Performing multiple feature extraction processes on the sample image according to the feature extraction network to obtain the multiple sample three-dimensional feature images; The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
15. The method according to claim 14, characterized in that The feature extraction network includes a convolution layer, a pooling layer and an upsampling layer; for any feature extraction process, the sample image is subjected to feature extraction process according to the feature extraction network to obtain the sample three-dimensional feature image, including: According to the convolution layer and the pooling layer, a convolution operation and a downsampling operation are performed on the sample image to obtain K sample downsampling feature images, and the size of the i-th sample downsampling feature image is , i is 1, 2, ..., K in sequence; The convolution operation and the upsampling operation are performed on the K sample down-sampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image.
16. The method according to claim 15, characterized in that Performing a convolution operation and a downsampling operation on the sample image according to the convolution layer and the pooling layer to obtain K sample downsampling feature images, including: Performing a convolution operation on the sample image according to the convolution layer to obtain a first sample down-sampled feature image; According to the convolution layer and the pooling layer, the first operation is performed i times in sequence on the first sample down-sampled feature image to obtain the i+1th sample down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
17. The method according to claim 15 or 16, characterized in that The method further comprises: performing a convolution operation and an upsampling operation on the K sample down-sampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image, including: Performing downsampling operations, convolution operations, and upsampling operations on the K-th sample downsampling feature image according to the pooling layer, the convolution layer, and the upsampling layer to obtain a K-th sample upsampling feature image; According to the convolution layer and the upsampling layer, a merging operation, a convolution operation and an upsampling operation are sequentially performed on the i-th sample upsampled feature image and the i-th sample downsampled feature image to obtain the i-1-th sample upsampled feature image, where i is K, K-1, ..., 2 in sequence; A convolution operation is performed on the first sample upsampled image according to the convolution layer to obtain the sample three-dimensional feature image.
18. A screen detection device, characterized in that: include: An acquisition module, used to acquire a feature vector of each pixel in a first image, where the first image is an image obtained by photographing a screen to be detected; The category of any pixel point in the first image is one of the preset categories, and the number of preset categories is C; a classification module, configured to input a feature vector of each pixel in the first image into a preset model to obtain C-1 first output images and a second output image, wherein each first output image indicates a defect category, the grayscale value of the (c, d)th pixel in each first output image i is used to indicate the probability that the defect category of the (c, d)th pixel on the first image is the defect category indicated by the first output image i, and the grayscale value of the (c, d)th pixel in the second output image is used to indicate the probability that the category of the (c, d)th pixel on the first image is normal, wherein c is a positive integer less than or equal to M, and d is a positive integer less than or equal to N; Obtaining a category of each pixel in the first image according to the grayscale value of each pixel on the C-1 first output image and the grayscale value of each pixel on the second output image; A detection result of the screen to be detected is obtained according to the category of each pixel in the first image.
19. The device according to claim 18, characterized in that The acquisition module is specifically used for: Determine a plurality of three-dimensional feature images according to the first image, wherein the first image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and the number of channels of each of the three-dimensional feature images is C, where C is a preset number of categories; According to the multiple three-dimensional feature images, a feature vector of each pixel in the first image is obtained.
20. The device according to claim 19, characterized in that The acquisition module is specifically used for: Performing feature extraction processing on the first image multiple times to obtain the multiple three-dimensional feature images; The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
21. The device according to claim 20, characterized in that For any feature extraction process, the acquisition module is specifically used to: Perform convolution and downsampling operations on the first image to obtain K downsampled feature images. The size of the i-th downsampled feature image is , i is 1, 2, ..., K in sequence; A convolution operation and an upsampling operation are performed according to the K down-sampled feature images to obtain the three-dimensional feature image.
22. The device according to claim 21, characterized in that The acquisition module is specifically used for: Performing a convolution operation on the first image to obtain a first downsampled feature image; The first operation is performed i times in sequence on the first down-sampled feature image to obtain an (i+1)th down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
23. The device according to claim 21 or 22, characterized in that The acquisition module is specifically used for: Performing a downsampling operation, a convolution operation, and an upsampling operation on the K-th down-sampled feature image to obtain a K-th up-sampled feature image; Performing a merging operation, a convolution operation, and an upsampling operation on the i-th upsampled feature image and the i-th downsampled feature image in sequence to obtain an i-1-th upsampled feature image, where i is K, K-1, ..., 2 in sequence; A convolution operation is performed on the first up-sampled image to obtain the three-dimensional feature image.
24. The device according to any one of claims 19 to 22, characterized in that The acquisition module is specifically used for: Determine a target three-dimensional feature image according to the pixel values of the pixels in each three-dimensional feature image, wherein the target three-dimensional feature image includes M pixels in the horizontal direction, and N pixels in the vertical direction, and the number of channels of the target three-dimensional feature image is C; A feature vector of each pixel in the first image is determined according to the pixel value of each pixel in the target three-dimensional feature image.
25. The device according to claim 24, characterized in that The acquisition module is specifically used for: According to the xth channel in each 3D feature image The pixel value of pixels determines the xth channel of the target three-dimensional feature image. The pixel value of pixels, wherein x is 1, 2, ..., C in sequence; Among them, the pixel value of the (a, b)th pixel point in the xth channel of the target three-dimensional feature image is: the maximum value of the pixel values of the (a, b)th pixel point in the xth channel of the multiple three-dimensional feature images, a is a positive integer less than or equal to M, and b is a positive integer less than or equal to N.
26. The device according to claim 25, characterized in that For the (a, b)th pixel in the first image, the acquisition module is specifically used to: A feature vector of the (a, b)th pixel in the first image is determined according to the value of the (a, b)th pixel in the C channels of the target three-dimensional feature image.
27. The device according to claim 18, characterized in that The acquisition module is specifically used for: The first image is input into a feature extraction model to obtain a feature vector for each pixel in the first image, wherein the feature extraction model is learned from multiple groups of first samples, each group of first samples includes a first sample image and a feature vector of a pixel on the first sample image, and the first sample image is an image obtained by photographing the first sample screen.
28. The device according to any one of claims 19-22, 25-27, characterized in that The classification module is specifically used for: The preset model is learned from multiple groups of second samples, each group of second samples includes feature vectors and annotation information of pixels on a second sample image, the second sample image is an image obtained by shooting the second sample screen, and the annotation information is information that annotates the categories of pixels on the second sample image.
29. The device according to claim 18, characterized in that The classification module is specifically used for: For the (c, d)th pixel in the first image, obtain the grayscale value of the (c, d)th pixel on each of the C-1 first output images and the grayscale value of the (c, d)th pixel on the second output image; Determine a pixel with a maximum grayscale value according to the grayscale value of the (c, d)th pixel on each first output image and the grayscale value of the (c, d)th pixel on the second output image; If the image where the pixel with the maximum gray value is located is the second output image, determining that the category of the (c, d)th pixel in the first image is normal; Otherwise, the defect category indicated by the first output image where the pixel with the maximum grayscale value is located is determined as the category of the (c, d)th pixel in the first image.
30. The device according to claim 18, characterized in that The device is applied to a screen detection model, which includes a feature extraction network and the preset model; The training device of the screen detection model comprises: A training module, used to obtain training samples, wherein the training samples include a sample image and annotation information of pixels in the sample image, wherein the annotation information of pixels in the sample image is information that annotates the categories of the pixels in the sample image; a processing module, configured to perform feature extraction processing on the sample image according to the feature extraction network to obtain a plurality of sample three-dimensional feature images, wherein the sample image includes M pixels in the horizontal direction and N pixels in the vertical direction, each of the sample three-dimensional feature images includes M pixels in the horizontal direction and N pixels in the vertical direction, and each of the sample three-dimensional feature images has a channel number C, where C is the number of categories for classifying the pixels in the sample image; Processing the plurality of sample three-dimensional feature images according to the preset model to obtain training output categories of pixel points in the sample images; An adjustment module is used to adjust the parameters of the screen detection model according to the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image, until the error between the training output category of the pixel points in the sample image and the annotation information of the pixel points in the sample image is less than or equal to a preset error, thereby obtaining a trained screen detection model.
31. The device according to claim 30, characterized in that The processing module is specifically used for: Performing multiple feature extraction processes on the sample image according to the feature extraction network to obtain the multiple sample three-dimensional feature images; The number of feature extraction operations included in each two feature extraction processes is different, and the feature extraction operations include convolution operations and sampling operations.
32. The device according to claim 31, characterized in that The feature extraction network includes a convolution layer, a pooling layer and an upsampling layer; for any feature extraction process, the processing module is specifically used to: According to the convolution layer and the pooling layer, a convolution operation and a downsampling operation are performed on the sample image to obtain K sample downsampling feature images, and the size of the i-th sample downsampling feature image is , i is 1, 2, ..., K in sequence; The convolution operation and the upsampling operation are performed on the K sample down-sampled feature images according to the convolution layer and the upsampling layer to obtain the sample three-dimensional feature image.
33. The device according to claim 32, characterized in that The processing module is specifically used for: Performing a convolution operation on the sample image according to the convolution layer to obtain a first sample down-sampled feature image; According to the convolution layer and the pooling layer, the first operation is performed i times in sequence on the first sample down-sampled feature image to obtain the i+1th sample down-sampled feature image, where i is 1, 2, ..., K-1 in sequence, and the first operation includes a down-sampling operation and a convolution operation.
34. The device according to claim 32 or 33, characterized in that The processing module is specifically used for: Performing downsampling operations, convolution operations, and upsampling operations on the K-th sample downsampling feature image according to the pooling layer, the convolution layer, and the upsampling layer to obtain a K-th sample upsampling feature image; According to the convolution layer and the upsampling layer, a merging operation, a convolution operation and an upsampling operation are sequentially performed on the i-th sample upsampled feature image and the i-th sample downsampled feature image to obtain the i-1-th sample upsampled feature image, where i is K, K-1, ..., 2 in sequence; A convolution operation is performed on the first sample upsampled image according to the convolution layer to obtain the sample three-dimensional feature image.
35. A screen detection device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor runs the computer program to execute the screen detection method according to any one of claims 1 to 17.
36. A screen detection system, characterized in that: It includes image acquisition equipment and screen detection equipment, including: The image acquisition device is used to photograph the screen to be detected, obtain a first image, and send the first image to the screen detection device; The screen detection device is used to process the first image according to the method according to any one of claims 1 to 17 to obtain a detection result of the screen to be detected.
37. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the screen detection method according to any one of claims 1 to 17 is implemented.
38. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by one or more processors, the screen detection method according to any one of claims 1 to 17 is implemented.