Display screen single classification anomaly detection method and device and storage medium

By constructing a single-class detection model and loss function, combining shallow feature enhancement and multi-classification head training, the feature representation of the new display screen is optimized, which solves the problem that the new display single-class detection model lacks special optimization for single-class abnormal data during feature extraction, and improves detection accuracy and robustness.

CN120388020AActive Publication Date: 2025-07-29SHENZHEN SEICHITECH TECHN CO LTD

Patent Information

Application Number
CN202510884048.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing new display single-class anomaly detection model lacks special optimization of single-class anomaly data during feature extraction, resulting in a dispersed distribution of feature within the class, which is unable to effectively distinguish abnormal samples from normal samples, reducing detection accuracy.

Method used

A single-class detection model and a single-class loss function are constructed, feature maps are generated through shallow feature enhancement processing, shallow features are extracted and fused, deep features are generated, and feature representation is optimized through joint training of multi-classification heads and sub-category pseudo-labels, and anomaly detection capabilities are enhanced.

Benefits of technology

The accuracy of single-class detection and abnormal detection of the new display defect detection model is improved, and the classification of abnormal images can be refined and more fine-grained detection capabilities are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388020A_ABST
    Figure CN120388020A_ABST
Patent Text Reader

Abstract

The invention discloses a display screen single-classification anomaly detection method and device and a storage medium, which are used for improving the accuracy of single-classification detection and anomaly detection of a novel display screen defect detection model. Constructing a single classification detection model and a single classification loss function; inputting the training image into a single-classification detection model; performing shallow feature enhancement processing on the training image to generate a feature map; extracting shallow layer features of the training image and the feature map, performing feature fusion, extracting depth features, and generating a dichotomy prediction result for the depth features; calculating a first loss value, and updating the single-classification detection model according to the first loss value; generating a sub-category pseudo label for each training image; expanding a multi-classification head on the single-classification detection model, and constructing a multi-classification loss function and a sub-classification target; and jointly training a single-classification detection model through the depth features, the multi-classification head, the multi-classification loss function and the sub-classification target until the single-classification head reaches a training target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of display screen detection, and in particular, to a method, device, and storage medium for single-class abnormal detection of a display screen. Background Technique

[0002] With the development of industry, the surface defect detection and abnormal detection of display screens have become a hot research direction in the industrial field. Traditional surface abnormal detection methods rely on qualified experts for manual operations. However, this method is not only inefficient but also highly dependent on the subjective judgment of operators, making it difficult to ensure the accuracy of detection. Therefore, in the prior art, automated abnormal detection devices for industrial production lines are combined to ensure product quality, reduce manual testing costs, and significantly improve production efficiency. With the rapid development of computer image processing technology, many advanced algorithms have been widely applied to the field of abnormal detection of display screen materials, greatly improving the accuracy of detection. In recent years, deep learning has also been widely applied to the abnormal detection of display screen materials and achieved remarkable results.

[0003] With the update and iteration of display screens, the functionality of new display screens is becoming more and more powerful, and their structures are gradually becoming complex and precise. For example, folding screens, flexible screens, splicing screens, etc., as well as display screens with multiple internal microcircuit layers, and currently, in the R & D process of display screens, it is inclined to combine multiple precise structures precisely, which increases the complexity of the defects of new display screens, and the characteristics of new defects are more difficult to detect with the naked eye, making it even more difficult to detect. Nowadays, algorithms based on deep learning have shown better performance than traditional methods in the single-class visual detection tasks of many new display screen defects. However, existing single-class visual detection algorithms for new display screen defects usually rely on supervised learning, which requires a large amount of labeled data. Because the update and iteration speed of new display screens is fast, and the specificities of different types of display screen defects are gradually increasing, the core training samples (abnormal samples) of the original training dataset are difficult to continue to be used, and because of the difficulty of annotation, it is usually difficult to obtain the captured images of display screens with new abnormal labels, resulting in less training data with new abnormal labels or defect labels, which leads to the inefficiency of training and the limitations of application of the single-class abnormal detection model for new display screens.

[0004] Secondly, in the field of single-class anomaly detection for new display screens, supervised anomaly detection methods are used. Such supervised anomaly detection methods usually rely on a training set containing labeled data and are trained using classical machine learning models or deep neural networks (such as support vector machines, random forests, XGBoost, etc.). By learning the feature differences between normal data and abnormal data, the single-class anomaly detection model can achieve anomaly detection. The advantage of the supervised method is that when the labeled data is sufficient, the model can achieve high-precision anomaly detection. However, the defect of the single-class anomaly detection model for new display screens trained based on the supervised method is that most current single-class detection methods based on deep learning directly adopt the high-dimensional features of pre-trained networks (such as the features extracted by ResNet). These high-dimensional features of new display screens are usually trained in multi-class scenarios and lack adaptability to the target single-class data. The lack of special optimization for single-class abnormal data in the feature extraction process leads to a relatively scattered intra-class feature distribution, which in turn makes it impossible for the single-class detection model of new display screens to effectively distinguish abnormal samples from normal samples in the feature space, reducing the precision of single-class detection and anomaly detection of the new display screen defect detection model. Summary of the Invention

[0005] This application discloses a method, device, and storage medium for single-class anomaly detection of display screens, which are used to improve the precision of single-class detection and anomaly detection of a new display screen defect detection model.

[0006] In a first aspect, an embodiment of this application provides a method for single-class anomaly detection of display screens, including: constructing a single-class detection model and a single-class loss function; obtaining training images of new display screens labeled with single-class labels, and inputting the training images into the single-class detection model; performing shallow feature enhancement processing on the training images according to the single-class labels to generate feature maps; respectively extracting the shallow features of the training images and the feature maps, performing feature fusion on the two sets of shallow features, then extracting deep features, and generating a binary classification prediction result for the deep features according to the single-class head; calculating a first loss value according to the single-class loss function and the binary classification prediction result, and updating the single-class detection model according to the first loss value until the initial training target is reached; performing clustering processing on the deep features to generate sub-class pseudo-labels for each training image; expanding a multi-class head on the single-class detection model, and constructing a multi-class loss function and a sub-class classification target; jointly training the feature extractor, the two classification heads, and the sub-class pseudo-labels in the single-class detection model through the deep features, the multi-class head, the multi-class loss function, and the sub-class classification target until the single-class head reaches the training target.

[0007] Optionally, after the step of jointly training the feature extractor, two classification heads, and sub-category pseudo-labels in the single-class detection model through depth features, a multi-classification head, a multi-classification loss function, and sub-category classification targets, the method further includes: removing the extended multi-classification head to obtain unlabeled images in the new scenario; copying the unlabeled images, and performing perturbation processing on the copied unlabeled images; performing shallow feature enhancement processing on the unlabeled images to generate scene feature maps; extracting shallow features from the unlabeled images, the perturbed unlabeled images, and the scene feature maps to generate first features, second features, and third features respectively; fusing the first features and the third features to generate a first fused feature, and then performing depth feature extraction to generate first depth features; fusing the second features and the third features to generate a second fused feature, and then performing depth feature extraction to generate second depth features; inputting the first depth features and the second depth features into the single-classification head to generate conventional prediction results and perturbed prediction results; constructing a consistency loss function, and generating a consistency loss according to the conventional prediction results, the perturbed prediction results, and the consistency loss function; updating the parameters applicable to the new scenario of the single-class detection model according to the consistency loss.

[0008] Optionally, after the step of updating the parameters applicable to the new scenario of the single-class detection model according to the consistency loss, the method further includes: when the conventional prediction results and the perturbed prediction results show that there are unlabeled images at the classification boundary, precisely annotating these unlabeled images; determining the precisely annotated unlabeled images as untrained images, and using the untrained images to optimize the single-class detection model during the production process.

[0009] Optionally, the new display screen is a folding screen with a microcircuit layer provided in the folding area, the single-class detection model is a folding screen single-class detection model, and the unlabeled images are real-time captured images of the folding area of the folding screen on the production line; the step of copying the unlabeled images and performing perturbation processing on the copied unlabeled images includes: copying the unlabeled images; introducing light source perturbation into the copied unlabeled images according to the curvature parameters of the folding screen in the folding area; introducing noise perturbation into the unlabeled images with introduced perturbation by means of cropping and rotation.

[0010] Optionally, obtain training images of the new display screen labeled with single-category labels, and input the training images into the single-classification detection model, including: obtaining the captured images of the new display screen, and dividing the captured images into abnormal images and non-abnormal images; combining the abnormal images and non-abnormal images, and expanding the data volume through image enhancement technology to generate the first abnormal augmented image; identifying the reference regions in the non-abnormal images, replacing the reference regions with the corresponding abnormal regions in the abnormal images, and performing assimilation processing on the replaced abnormal regions to generate the second abnormal augmented image, where the reference region is the region where the permission influence weight reaches the preset value; performing perturbation processing on the non-abnormal images to generate the third abnormal augmented image; labeling the abnormal images, non-abnormal images, the first abnormal augmented image, the second abnormal augmented image, and the third abnormal augmented image with single-category labels, and determining them as training images; inputting the training images into the single-classification detection model.

[0011] Optionally, the new display screen is a folding screen with a microcircuit layer provided in the folding area, the single-classification detection model is a folding screen single-classification detection model, the training image is the captured image of the folding area of the folding screen, and the single-category label is the folding area single-category label; performing shallow feature enhancement processing on the training image according to the single-category label to generate a feature map, including: determining that the gain data of the training image is the folding screen curvature and the folding screen microcircuit reflection according to the folding area single-category label; determining the curvature parameters of the folding area of the training image; setting the first folding screen and the second folding screen according to the curvature parameters, and projecting light with a preset brightness value onto the first folding screen and the second folding screen through an external light source, where the first folding screen is a folding screen without a microcircuit layer provided in the folding area, and the second folding screen is a folding screen with a microcircuit layer provided in the folding area; respectively collecting the images of the first folding screen and the second folding screen, and calculating the microcircuit reflection parameters according to the folding areas of the two images; constructing a shallow feature gain convolution kernel according to the curvature parameters and the circuit reflection parameters; performing shallow feature enhancement processing on the training image through the shallow feature gain convolution kernel to generate a feature map.

[0012] Optionally, the steps of expanding a multi-classification head on the single-classification detection model and constructing a multi-classification loss function and sub-category classification targets include: expanding a multi-classification head on the single-classification detection model for detecting the probability distribution of sub-category pseudo-labels; determining the gray discrimination weight values of each sub-category pseudo-label on the new display screen; generating gray error weight values according to the curvature parameters and microcircuit reflection parameters of the new display screen; constructing a multi-classification cross-entropy function according to the gray discrimination weight values and the gray error weight values; constructing sub-category classification targets.

[0013] In a second aspect, an embodiment of the present application provides an apparatus for detecting anomalies in single-classification of a display screen, including: a first construction unit configured to construct a single-classification detection model and a single-classification loss function; a first acquisition unit configured to acquire training images of a new type of display screen labeled with single-class labels and input the training images into the single-classification detection model; a first generation unit configured to perform shallow feature enhancement processing on the training images according to the single-class labels to generate feature maps; a second generation unit configured to respectively extract shallow features of the training images and the feature maps, perform feature fusion on the two sets of shallow features, then extract deep features, and generate a binary classification prediction result for the deep features according to a single-classification head; a first training unit configured to calculate a first loss value according to the single-classification loss function and the binary classification prediction result, and update the single-classification detection model according to the first loss value until an initial training target is reached; a third generation unit configured to perform clustering processing on the deep features to generate sub-class pseudo-labels for each training image; a second construction unit configured to expand a multi-classification head on the single-classification detection model and construct a multi-classification loss function and a sub-class classification target; a second training unit configured to jointly train a feature extractor, two classification heads, and sub-class pseudo-labels in the single-classification detection model through the deep features, the multi-classification head, the multi-classification loss function, and the sub-class classification target until the single-classification head reaches the training target.

[0014] Optionally, after the second training unit, the apparatus further includes: a second acquisition unit configured to remove the expanded multi-classification head and acquire unlabeled images in a new scenario; a perturbation unit configured to copy the unlabeled images and perform perturbation processing on the copied unlabeled images; a fourth generation unit configured to perform shallow feature enhancement processing on the unlabeled images to generate scene feature maps; a fifth generation unit configured to extract shallow features of the unlabeled images, the perturbed unlabeled images, and the scene feature maps to respectively generate a first feature, a second feature, and a third feature; a sixth generation unit configured to fuse the first feature and the third feature to generate a first fused feature, then perform deep feature extraction to generate a first deep feature; a seventh generation unit configured to fuse the second feature and the third feature to generate a second fused feature, then perform deep feature extraction to generate a second deep feature; an eighth generation unit configured to input the first deep feature and the second deep feature into the single-classification head to generate a normal prediction result and a perturbed prediction result; a ninth generation unit configured to construct a consistency loss function and generate a consistency loss according to the normal prediction result, the perturbed prediction result, and the consistency loss function; an update unit configured to update the parameters of the single-classification detection model applicable to the new scenario according to the consistency loss.

[0015] Optionally, after the updating unit, the device further includes: a labeling unit, configured to precisely label the unlabeled images when the conventional prediction result and the perturbation prediction result indicate that there are unlabeled images at the classification boundary; an optimization unit, configured to determine the precisely labeled unlabeled images as untrained images, and use the untrained images to optimize the single-class detection model during the production process.

[0016] Optionally, the new display screen is a folding screen with a microcircuit layer in the folding area, the single-class detection model is a folding screen single-class detection model, and the unlabeled images are real-time captured images of the folding area of the folding screen on the production line; the perturbation unit includes: copying the unlabeled images; introducing light source perturbation into the copied unlabeled images according to the curvature parameter of the folding screen in the folding area; introducing noise perturbation into the unlabeled images with perturbation introduced by means of cropping and rotation.

[0017] Optionally, the first acquisition unit includes: acquiring the captured image of the new display screen, and dividing the captured image into abnormal images and non-abnormal images; combining the abnormal images and the non-abnormal images, and expanding the data volume through image enhancement technology to generate a first abnormally augmented image; identifying the reference area in the non-abnormal images, replacing the reference area with the corresponding abnormal area in the abnormal images, and performing assimilation processing on the replaced abnormal area to generate a second abnormally augmented image, where the reference area is the area where the permission influence weight reaches a preset value; disturbing the non-abnormal images to generate a third abnormally augmented image; labeling the abnormal images, non-abnormal images, first abnormally augmented image, second abnormally augmented image, and third abnormally augmented image with single-class labels, and determining them as training images; inputting the training images into the single-class detection model.

[0018] Optionally, the new display screen is a folding screen with a microcircuit layer in the folding area, the single-class detection model is a folding screen single-class detection model, the training images are the captured images of the folding area of the folding screen, and the single-class label is the single-class label of the folding area; the first generation unit includes: determining the gain data of the training images as the folding screen curvature and the microcircuit reflection of the folding screen according to the single-class label of the folding area; determining the curvature parameter of the folding area of the training images; setting a first folding screen and a second folding screen according to the curvature parameter, and projecting light rays with a preset brightness value onto the first folding screen and the second folding screen through an external light source, where the first folding screen is a folding screen without a microcircuit layer in the folding area, and the second folding screen is a folding screen with a microcircuit layer in the folding area; respectively collecting the images of the first folding screen and the second folding screen, and calculating the microcircuit reflection parameter according to the folding areas of the two images; constructing a shallow feature gain convolution kernel according to the curvature parameter and the circuit reflection parameter; performing shallow feature enhancement processing on the training images through the shallow feature gain convolution kernel to generate a feature map.

[0019] Optionally, the second construction unit includes: expanding a multi-classification head for detecting the probability distribution of sub-category pseudo-labels on a single-classification detection model; Determine the gray-scale discrimination weight value of each sub-category pseudo-label on the new display screen; generate a gray-scale error weight value according to the curvature parameter and the microcircuit reflection parameter of the new display screen; construct a multi-classification cross-entropy function according to the gray-scale discrimination weight value and the gray-scale error weight value; construct a sub-category classification target.

[0020] In a third aspect, an embodiment of the present application provides a device for single-classification anomaly detection of a display screen, including: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the methods in the first aspect and any optional methods of the first aspect.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the methods in the first aspect and any optional methods of the first aspect.

[0022] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages: The present application first constructs a single-classification detection model and a single-classification loss function. Next, obtain the training images labeled with single-class labels of the new display screen, and input the training images into the single-classification detection model. Then, perform shallow feature enhancement processing on the training images according to the single-class labels to generate feature maps. Then, extract the shallow features of the training images and the feature maps respectively, then perform feature fusion on the two sets of shallow features, and then extract deep features. According to the single-classification head, generate a binary classification prediction result for the deep features. Calculate the first loss value according to the single-classification loss function and the binary classification prediction result, and update the single-classification detection model according to the first loss value until the initial training target is reached. Cluster the deep features to generate sub-category pseudo-labels for each training image. Expand a multi-classification head on the single-classification detection model, and construct a multi-classification loss function and a sub-category classification target. Jointly train the feature extractor, the two classification heads, and the sub-category pseudo-labels in the single-classification detection model through the deep features, the multi-classification head, the multi-classification loss function, and the sub-category classification target until the single-classification head reaches the training target.

[0023] By performing specific shallow feature enhancement processing on the training images to generate feature maps, then extracting the shallow features of the training images and the feature maps respectively, and then fusing the two sets of shallow features, and then extracting deep features. At this time, the deep features extracted have completed the special optimization for single-class abnormal data, and the situation where the within-class feature distribution is relatively scattered has been initially solved. Next, a multi-classification head is extended on the single-classification detection model, and a multi-classification loss function and sub-category classification targets are constructed. Then, the feature extractor, two classification heads, and sub-category pseudo-labels in the single-classification detection model are jointly trained through the detection results of the multi-classification head until the single-classification head fully reaches the training target. Through the multi-classification head, the feature representation is further optimized, and the optimized enhanced feature representation can help the single-classification detection model further refine the classification of abnormal images in the training images, thereby providing a more fine-grained abnormal detection ability for this branch of the single-classification head and improving the accuracy of single-classification detection and abnormal detection of the new display defect detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 FIG. is a schematic diagram of an embodiment of the method for single-classification abnormal detection of the display screen in the present application; Figure 2 FIG. is a schematic diagram of an embodiment of the model update method based on a new scenario in the present application; Figure 3 FIG. is a schematic diagram of an embodiment of the method for model optimization in the present application; Figure 4 FIG. is a schematic diagram of an embodiment of the perturbation processing method based on unlabeled images in the present application; Figure 5 FIG. is a schematic diagram of an embodiment of the processing method for training images in the present application; Figure 6 FIG. is a schematic diagram of an embodiment of the method for shallow feature enhancement processing in the present application; Figure 7 FIG. is a schematic diagram of an embodiment of the extension method for the single-classification detection model in the present application; Figure 8 FIG. is a schematic diagram of an embodiment of the device for single-classification abnormal detection of the display screen in the present application; Figure 9 FIG. is a schematic diagram of another embodiment of the device for single-classification abnormal detection of the display screen in the present application; Figure 10 This is a schematic diagram of different stages of the single-classification detection model of this application. Detailed implementation manners

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0027] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0030] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0031] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0032] In the prior art, in the field of single-class anomaly detection for new display screens, supervised anomaly detection methods are used. Such supervised anomaly detection methods usually rely on a training set containing labeled data and are trained using classical machine learning models or deep neural networks (such as support vector machines, random forests, XGBoost, etc.). By learning the feature differences between normal data and abnormal data, the single-class anomaly detection model can achieve anomaly detection. The advantage of the supervised method is that when the labeled data is sufficient, the model can achieve high-precision anomaly detection. However, the defect of the single-class anomaly detection model for new display screens trained based on the supervised method is that: most of the current single-class detection methods based on deep learning directly adopt the high-dimensional features of the pre-trained network (such as the features extracted by ResNet). These high-dimensional features of new display screens are usually trained in multi-class scenarios and lack adaptability to the target single-class data. There is a lack of special optimization for single-class abnormal data in the feature extraction process, resulting in a relatively scattered intra-class feature distribution. Furthermore, the single-class detection model for new display screens cannot effectively distinguish abnormal samples from normal samples in the feature space, reducing the accuracy of single-class detection and anomaly detection of the new display screen defect detection model.

[0033] Based on this, the present application discloses a method, device, and storage medium for single-class anomaly detection of a display screen, which are used to improve the accuracy of single-class detection and anomaly detection of a new display screen defect detection model.

[0034] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0035] The method of the present application can be applied to a server, device, terminal, or other device with logical processing capabilities. In this regard, the present application makes no limitation. For the convenience of description, the following describes the execution subject as a terminal.

[0036] Please refer to Figure 1 , an embodiment of a method for single-class anomaly detection of a display screen provided by the present application includes: 101. Construct a single-class detection model and a single-class loss function.

[0037] Please refer to Figure 10, the single-class detection model has multiple stages, one is the training stage (upper half), and the other is the production detection stage (lower half). When using a deep learning model (single-class detection model) for image classification tasks, the goal is to classify training images (labeled images, where the label is a single-class label) into normal images and abnormal images (where abnormal images include abnormal data and disrupted data). In terms of model design, in this embodiment, a traditional convolutional neural network (CNN) framework (such as ResNet50) is adopted as the backbone network and improved on this basis. In this embodiment, it is mainly the training stage of the upper half, including a shallow feature map extraction network, above the shallow feature map is a shallow training image extraction network, a feature fusion module, a deep feature extraction module (for feature extraction), a binary classification head module (single-class head), a multi-classification head module, and a clustering analysis module ( Figure 10 The K-means method is used in Figure 10 . The K-means method (K-means clustering algorithm) in Figure 10 is a classic unsupervised learning algorithm used to divide a dataset into K different clusters such that data points within the same cluster are highly similar while those between different clusters are less similar.

[0038] To alleviate the data imbalance problem, during the supervised training process, in this embodiment, the single-class loss function adopts a weighted binary cross-entropy loss function. For a small amount of abnormal data, a higher training weight needs to be given to ensure that the model can better learn the features of abnormal data, thereby improving the recognition ability of the single-class detection model for abnormal images. The single-class loss function L BCE constructed in this embodiment is as follows:

[0039] where N is the number of labeled images (training images), is the single-class abnormal class weight, is the single-class normal class weight, is the single-class label of the i-th training image, is the single-class prediction result (binary classification prediction result) of the model for the i-th training image.

[0040] 102. Obtain training images with single-category labels of a new type of display screen, and input the training images into a single-classification detection model.

[0041] In this embodiment, the terminal collects training images from the new type of display screen, determines the defects on the training images, and labels the training images. Each training image has a corresponding single-category label at this time. Next, the terminal inputs the training images into the single-classification detection model.

[0042] 103. Perform shallow feature enhancement processing on the training images according to the single-category labels to generate feature maps.

[0043] The terminal determines the enhancement parameters and enhancement methods for the shallow feature enhancement processing according to the types existing in the single-category labels, and then performs shallow feature enhancement processing on the training images to generate feature maps of shallow feature enhancement. The specific determination and implementation details of the shallow feature enhancement processing will be described in subsequent embodiments.

[0044] 104. Extract the shallow features of the training images and the feature maps respectively, perform feature fusion on the two sets of shallow features, then extract deep features, and generate a binary classification prediction result for the deep features according to the single-classification head.

[0045] In this embodiment, the shallow features of the feature maps and the shallow features of the training images (labeled images) are extracted through a shallow feature map extraction network and a shallow series of image extraction networks respectively, and feature fusion processing is performed through a feature fusion module. Then, a deep feature extraction module is used to extract deep features. The obtained deep features have specifically optimized the single-class abnormal data and solved the problem of relatively scattered intra-class feature distributions. Then, the deep features are input into the single-classification head (binary classification head) for single-category defect abnormality analysis to obtain a binary classification prediction result.

[0046] In this embodiment, by combining the features of traditional image processing with deep learning methods and adopting an attention mechanism and multiple shallow feature extraction modules, the recognition ability of the single-classification detection model for abnormal images is enhanced, and the accuracy of abnormal detection is further improved.

[0047] 105. Calculate a first loss value according to the single-classification loss function and the binary classification prediction result, and update the single-classification detection model according to the first loss value until the initial training goal is reached.

[0048] In this embodiment, after the terminal obtains the binary classification prediction result, it inputs the binary classification prediction result into the single-classification loss function L BCECalculate the loss value to obtain the first loss value, and update the single-class detection model according to the first loss value until the initial training objective is achieved. For a small amount of abnormal data, this embodiment gives a higher training weight to ensure that the model can better learn the characteristics of the abnormal data, thereby improving the recognition ability of the single-class detection model for abnormal images. Specifically, in this embodiment, the weight of the single-class abnormal class is set to 3, and the weight of the single-class normal class is set to 1.

[0049] 106. Cluster the depth features to generate sub-category pseudo-labels for each training image.

[0050] When the initial training of the single-class head is completed, the single-class detection ability of the single-class detection model cannot fully adapt to the defect abnormalities of the new display screen, and it is necessary to further refine the classification ability of the single-class detection model for abnormal images in the training images. Specifically, the terminal first clusters the extracted depth features. In this embodiment, K-mean clustering is used to generate sub-category pseudo-labels for each training image. Specifically, after the initial training is completed, the single-class detection model will further establish the sub-category of each image by distinguishing the image features. The core idea of this stage is: cluster the feature vectors learned by the network to identify different types of abnormal images and assign a specific sub-category to each image. For example, the K-means clustering algorithm can be used to cluster the features of normal images and abnormal images respectively to determine K sub-categories. According to the clustering results, we assign a pseudo-label (the pseudo-label corresponding to the feature map) to each training image, and the pseudo-label represents the sub-category index to which the image belongs.

[0051] 107. Expand a multi-class head on the single-class detection model, and construct a multi-class loss function and a sub-category classification objective.

[0052] 108. Jointly train the feature extractor, the two classification heads, and the sub-category pseudo-labels in the single-class detection model through the depth features, the multi-class head, the multi-class loss function, and the sub-category classification objective until the single-class head reaches the training objective.

[0053] Next, the terminal expands a multi-class head in the entire single-class detection model network, constructs a sub-category classification objective, and constructs a multi-class loss function. By jointly training this multi-class network (multi-class head) and the binary-class network (binary-class head). It should be noted that the sub-category classification network and the binary-class network share the same feature extractor, that is, they jointly use a set of depth features to ensure the unity of feature extraction. By iteratively updating the feature extractor, the two classifiers, and the sub-category pseudo-labels, the terminal can continuously optimize the feature representation and thus improve the classification performance.

[0054] Use the multi-class cross-entropy loss L subPerform constraints, and the formula is:

[0055] where C is the number of sub - categories, is the label of the image corresponding to the i - th sub - category, is the model prediction result of the image corresponding to the i - th sub - category.

[0056] The advantage of this method is that the enhanced feature representation can help the single - class detection model further refine the classification of abnormal images, providing a more fine - grained abnormal detection ability. The single - class detection model can not only distinguish normal and abnormal images, but also identify and distinguish different types of abnormalities, improving the accuracy and robustness of single - class detection for abnormal detection.

[0057] First, construct a single - class detection model and a single - class loss function. Next, obtain the training images with single - class labels annotated for the new type of display screen, and input the training images into the single - class detection model. Then, perform shallow - feature enhancement processing on the training images according to the single - class labels to generate feature maps. Then, extract the shallow features of the training images and the feature maps respectively, then perform feature fusion on the two sets of shallow features, and then extract deep features. According to the single - class head, generate binary classification prediction results for the deep features. Calculate the first loss value according to the single - class loss function and the binary classification prediction results, and update the single - class detection model according to the first loss value until the initial training goal is reached. Cluster the deep features to generate sub - category pseudo - labels for each training image. Expand a multi - class head on the single - class detection model, and construct a multi - class loss function and a sub - category classification target. Jointly train the feature extractor, the two classification heads, and the sub - category pseudo - labels in the single - class detection model through the deep features, the multi - class head, the multi - class loss function, and the sub - category classification target until the single - class head reaches the training goal.

[0058] First, the feature parameters of traditional image detection are usually represented by an image with the same size as the original image, so as to obtain the parameter features of the image and retain the feature values at the corresponding positions of the image. In this embodiment, an attention mechanism is introduced on this basis, and a shallow - feature extraction module is designed to extract the shallow features of the image. This shallow - feature extraction module can be fused with the global features of the image, guiding the single - class detection model to focus on the most representative feature parts in the image, thereby improving the accuracy of image single - class classification.

[0059] Specifically, through specific shallow - feature enhancement processing on the training images, generate feature maps, then extract the shallow features of the training images and the feature maps respectively, and then perform feature fusion on the two sets of shallow features, and then extract deep features. At this time, the deep features extracted have completed the special optimization of single - class abnormal data, and the situation where the within - class feature distribution is relatively scattered has been initially solved.

[0060] Next, expand a multi-classification head on the single-classification detection model, construct a multi-classification loss function and sub-category classification targets, and then jointly train the feature extractor, two classification heads, and sub-category pseudo-labels in the single-classification detection model through the detection results of the multi-classification head until the single-classification head fully reaches the training target. Further optimize the feature representation through the multi-classification head. The optimized and enhanced feature representation can help the single-classification detection model further refine the classification of abnormal images in the training images, and then provide a more fine-grained abnormal detection ability for this branch of the single-classification head, improving the accuracy of single-classification detection and abnormal detection of the new display panel defect detection model.

[0061] For this reason, in this embodiment, an active learning strategy is introduced, and feature clustering is added to divide sub-labels, optimizing the feature representation of the target category to make it have stronger generalization ability while ensuring compactness. This method makes full use of data resources and greatly improves the adaptability of the model to unknown scenarios.

[0062] Please refer to Figure 2 , this application provides an embodiment of a model update method based on a new scenario, including: 201. Remove the expanded multi-classification head to obtain unlabeled images in the new scenario.

[0063] In this embodiment, although traditional image processing methods can extract image features and perform threshold control, they cannot effectively cope with the variable and complex display panel defects, and it is difficult to handle continuously changing abnormal types and new scenarios. And the existing technologies lack the ability to handle new scenarios and unlabeled data, resulting in the model being unable to quickly adapt when facing new scenarios, and still requiring a large amount of manual intervention and data annotation. For this reason, in this embodiment, an active learning strategy is introduced, and feature clustering is added to divide sub-labels, optimizing the feature representation of the target category to make it have stronger generalization ability while ensuring compactness. This method makes full use of data resources and greatly improves the adaptability of the model to unknown scenarios.

[0064] Specifically, the terminal removes the expanded multi-classification head to obtain unlabeled images in the new scenario because the auxiliary training of the multi-classification head is no longer required in the actual production process.

[0065] 202. Copy the unlabeled images and perform perturbation processing on the copied unlabeled images.

[0066] The terminal copies the unlabeled images and performs perturbation processing on the copied unlabeled images. Specifically, the terminal performs perturbation processing on the copied unlabeled images according to the defect types targeted by the single-classification detection model. The specific steps will be described in detail later.

[0067] 203. Perform shallow feature enhancement processing on the unlabeled images to generate scene feature maps.

[0068] In this embodiment, the same method of shallow feature enhancement processing as in Embodiment 1 is adopted. If there are multiple processing methods, multiple unlabeled images need to be copied, and then multiple shallow feature enhancement processing methods are used for separate processing, and then several scene feature maps are generated (also called feature maps in the steps in the Figure 10 lower half).

[0069] 204. Extract shallow features from the unlabeled images, the perturbed unlabeled images, and the scene feature maps to generate the first feature, the second feature, and the third feature respectively.

[0070] 205. Fuse the first feature and the third feature to generate the first fused feature, and then perform deep feature extraction to generate the first deep feature.

[0071] 206. Fuse the second feature and the third feature to generate the second fused feature, and then perform deep feature extraction to generate the second deep feature.

[0072] After the terminal generates the perturbed unlabeled images and several scene feature maps, similar to Embodiment 1, it is necessary to extract shallow features from the perturbed unlabeled images, the unperturbed unlabeled images, and the scene feature maps to generate the first feature, the second feature, and the third feature.

[0073] Then the terminal fuses the first feature and the third feature to generate the first fused feature, and then performs deep feature extraction to generate the first deep feature. Then it fuses the second feature and the third feature to generate the second fused feature, and then performs deep feature extraction to generate the second deep feature. The specific extraction method is similar to that in Embodiment 1. One of the first deep feature and the second deep feature has perturbations and the other does not. It should be noted that because of the length of the scene feature maps, it is necessary to analyze each scene feature map separately, so that there are actually multiple groups of the first deep feature and the second deep feature, and the quantity is similar to the number of scene feature maps.

[0074] 207. Input the first deep feature and the second deep feature into the single classification head to generate the normal prediction result and the perturbation prediction result.

[0075] The terminal inputs the first deep feature and the second deep feature into the single classification head to generate the normal prediction result and the perturbation prediction result, and will also generate multiple groups of normal prediction results and perturbation prediction results according to the number of scene feature maps.

[0076] 208. Construct a consistency loss function, and generate a consistency loss according to the normal prediction result, the perturbation prediction result, and the consistency loss function.

[0077] At this time, the terminal constructs a consistency loss function and generates a consistency loss based on the conventional prediction result, the perturbed prediction result, and the consistency loss function. To enable the model to respond to new scenarios more quickly and overcome the dependence on human subjective judgment in abnormal states, the present invention adopts an active learning training method. This method can help the model more accurately identify abnormal states and continuously learn and optimize its performance during the test production process. During the test production process, most of the data are unlabeled images without tags, so the following steps are required for processing: First, the unlabeled images and their feature parameters after image processing are input into the model to obtain features ; then, the data is perturbed, and after the perturbation process, it is input into the model again to obtain new features . The perturbation methods can include: when the image is input, the same noise is added to the labeled image and its feature parameters after image processing, or data augmentation methods such as cropping and rotation are used for processing; in addition, perturbation processing can also be directly performed at the feature level by adding noise or other forms of perturbation.

[0078] After obtaining the processed features, the data is input into the binary classification head for prediction and compared with the original prediction result. For samples with large differences in results, the consistency of the model output is required through the consistency loss to ensure consistent classification results between the perturbed input and the original input.

[0079] Among them, the consistency loss L U is:

[0080] M is the number of unlabeled images, is the prediction result of the i-th unlabeled image, is the prediction result after perturbation of the i-th unlabeled image.

[0081] It should be noted that because there are multiple groups of conventional prediction results and perturbed prediction results, there are multiple consistency losses. The weights can be superimposed according to different processing means of the scene feature map and then normalized to generate a single consistency loss.

[0082] 209. Update the parameters suitable for the new scenario of the single-class detection model according to the consistency loss.

[0083] Finally, the terminal updates the parameters suitable for the new scenario of the single-class detection model according to the consistency loss.

[0084] Training the single-class detection model based on the new scenario in the above manner has the following beneficial effects: 1. Adopt an active learning strategy: Through the active learning mechanism, the model can automatically select the most challenging unlabeled samples for annotation, which not only reduces manual intervention but also enables continuous learning to handle new scenarios and changes, improving the model's adaptability.

[0085] 2. Data augmentation methods alleviate the data imbalance problem: By means of image enhancement, abnormal part transplantation, and introduction of perturbed data, etc., the dataset is expanded and the model's ability to recognize abnormal data is enhanced, overcoming the challenge of data imbalance.

[0086] 3. Model self-iteration and optimization: Through iterative training and continuous learning, the model can continuously optimize and adjust during the test production process, quickly adapt to new production scenarios and abnormal types, and improve the reliability and accuracy of anomaly detection.

[0087] Please refer to Figure 3 , this application provides an embodiment of a method for model optimization, including: 301. When the conventional prediction result and the perturbed prediction result show that there are unlabeled images at the classification boundary, these unlabeled images are precisely annotated.

[0088] 302. Determine the precisely annotated unlabeled images as untrained images, and use the untrained images to optimize the single-class detection model during the production process.

[0089] In this embodiment, if some unlabeled images are near the classification boundary and the prediction results are uncertain, these unlabeled images are output as uncertain samples and handed over to experts for manual annotation, and the single-class detection model is further optimized with these labeled samples. In the active learning process, the training batch can be set. Usually, after training 10 batches, control is carried out according to a preset threshold, samples located at the classification boundary are collected and sent to experts for annotation. This process can effectively improve the generalization ability of the single-class detection model and ensure that the single-class detection model can continuously self-improve in practical applications and gradually improve the accuracy of anomaly detection.

[0090] Through active learning, the single-class detection model can not only quickly adapt to the new environment without a large amount of labeled data but also enhance the robustness and accuracy of anomaly detection through continuous feedback optimization during the actual production process.

[0091] Please refer to Figure 4 , this application provides an embodiment of a perturbation processing method based on unlabeled images, including: 401. Copy the unlabeled images.

[0092] 402. Introduce light source perturbation into the copied label-free image according to the curvature parameter of the folding screen in the folding area.

[0093] 403. Introduce noise perturbation to the label-free image with perturbation introduced by means of cropping and rotation.

[0094] In this embodiment, the new display screen is a folding screen with a microcircuit layer provided in the folding area (the folding area refers to an area on the new display screen, which is composed of multiple display screen levels, including the microcircuit layer, the pixel display layer, etc.). The single-class detection model is a folding screen single-class detection model, and the label-free image is a real-time captured image of the folding area of the folding screen on the production line. The new display screen combines a flexible foldable and bendable structure, and the pixel points can be stretched during the folding process and have the ability to recover the structure. The microcircuit layer is an improvement in the level structure of the new display screen. In order to increase the functionality of the pixel point layer (display layer) and other structure layers of the new display screen, it can be controlled by setting microcircuits. Then, multiple microcircuits can be integrated into a structure layer (microcircuit layer) of the display screen and attached to the hierarchical structure of the new display screen, that is, a new microcircuit layer is added to the original level. Since most of the microcircuits in the microcircuit layer have metal lines that can reflect light sources, there is a phenomenon of light source reflection in the circuit area of the microcircuit layer. Whether the pixel points of the new display layer emit light or an external light source irradiates the new display screen, a certain degree of reflection can occur in the circuit area, resulting in visual non-uniformity of the display screen. To reduce this situation, usually the microcircuits in the microcircuit layer are placed on the curved edges on both sides of the flexible screen. However, for the folding screen, which also belongs to the flexible screen, the circuit area is usually set in the folding area because the folding area is not likely to affect the user's perception during use. But this increases the difficulty of defect detection for this type of new display screen. Because the pixel integration of the new display screen is getting higher and higher, the pixel points are arranged closely and orderly. In the folding screen, the pixel points in the folding area will change the initial arrangement of the pixel points as the folding curvature increases. The light emitted by the pixel points in the folding area will be emitted in different directions, but the circuit area will affect the displayed light source according to its own reflection degree, and the reflection degree is also affected by the curvature and is different. One makes the light source in this area weaken and diverge through bending and scattering (the more severe the folding, the more serious the light source scattering), and the other enhances the displayed light source through the reflection degree. This makes the defect anomalies in the folding area more easily affected by the folding area and the circuit area. For the single-class detection model of the new display screen, the detection quality decreases, especially the folding area is an important defect detection area.

[0095] In the defect detection of the new folding screen, it is necessary to collect images of the folding area with different folding degrees (controlling the curvature) and perform defect detection on the collected images.

[0096] In this embodiment, after the terminal copies the tagless image, it is necessary to introduce light source perturbation into the copied tagless image according to the curvature parameter of the folding screen in the folding area. Because when the folding screen is folded, the external light source can affect both the circuit area on the display screen and the hierarchical reflection in the folding area, so the light source characteristics of the external light source are extracted and introduced into the training image for perturbation processing. Then, noise perturbation is introduced into the tagless image with perturbation through cropping and rotation. After this perturbation processing, it is possible to better obtain an accurate consistency loss. After updating and iterating the parameters of the single-class detection model through the consistency loss, a better model detection effect can be obtained.

[0097] Please refer to Figure 5 , an embodiment of a method for processing training images provided by this application includes: 501. Obtain the captured image of the new display screen, and divide the captured image into abnormal images and non-abnormal images.

[0098] 502. Combine the abnormal images and non-abnormal images, and expand the data volume through image enhancement technology to generate the first abnormally augmented image.

[0099] 503. Identify the reference area in the non-abnormal image, replace the reference area with the corresponding abnormal area in the abnormal image, and perform assimilation processing on the replaced abnormal area to generate the second abnormally augmented image. The reference area is the area where the permission influence weight reaches the preset value.

[0100] 504. Perform perturbation processing on the non-abnormal image to generate the third abnormally augmented image.

[0101] 505. Perform single-class label annotation on the abnormal images, non-abnormal images, first abnormally augmented images, second abnormally augmented images, and third abnormally augmented images, and determine them as training images.

[0102] 506. Input the training images into the single-class detection model.

[0103] In the prior art, the cost of data annotation is high, and it is very difficult to obtain annotated data containing accurate abnormal samples. Especially when dealing with complex and dynamically changing data sets, the annotation cost and time consumption are very high. And there is also the problem of sample imbalance. In actual production scenarios, abnormal data is usually extremely rare, far less than normal data. This class imbalance makes the supervised model tend to normal data during training, affecting the detection effect.

[0104] Therefore, in this embodiment, for supervised learning data, this embodiment uses the image and its feature parameters and labels after image processing for training. For unsupervised learning data, only the image and its feature parameters after image processing are used. In the application of supervised data, to solve the problem of lack of abnormal data, this embodiment adopts a series of image enhancement techniques to increase the diversity and richness of the data set, specifically including the following methods: 1. Combination of normal data and abnormal data: Combine a small amount of abnormal images with normal images and expand the data volume through image enhancement techniques (such as image shearing, rotation, color transformation, etc.). This method effectively improves the generalization ability of the model by artificially creating more variant samples.

[0105] 2. Transplantation of abnormal parts: Embed a small amount of abnormal parts into normal images to artificially synthesize abnormal data. The specific method is as follows: First, identify the typical regions in the normal images, and then replace these regions with real abnormal regions, thereby enhancing the diversity of abnormal samples. In this way, various types of abnormal situations can be effectively simulated.

[0106] 3. Introduction of disrupted data: Introduce disrupted data, which does not represent any specific normal or abnormal state, but creates "abnormal" image data by adding noise or other forms of image transformation (such as image blurring, random cropping, etc.). The disrupted data is regarded as "abnormal data", and its role is to help the model better identify and distinguish normal and abnormal images during the training process.

[0107] Through these image enhancement techniques, this embodiment can effectively overcome the problem of insufficient data, improve the training effect of the model and the accuracy of anomaly detection, so as to achieve more reliable anomaly detection.

[0108] Only using normal image data for training, in order to improve the generalization ability of the model, and combining image enhancement methods with cross-domain data to enrich the data set.

[0109] Please refer to Figure 6 , this application provides an embodiment of a method for shallow feature enhancement processing, including: 601. Determine that the gain data of the training image is the folding screen curvature and the folding screen microcircuit reflection according to the single-category label of the folding area.

[0110] In this embodiment, the new display screen is a folding screen with a microcircuit layer provided in the folding area. At this time, the training image is the captured image of the folding area of the folding screen during folding. The single-classification detection model is a folding screen single-classification detection model, and the single-class label is a folding area single-class label. The terminal determines that the gain data of the training image is the folding screen curvature and the folding screen microcircuit reflection according to the folding area single-class label, because the folding screen curvature and the folding screen microcircuit reflection are the parameters that have the greatest impact on defects in this type of new folding screen.

[0111] 602. Determine the curvature parameter of the folding area of the training image.

[0112] After the terminal determines that the gain data of the training image is the folding screen curvature and the folding screen microcircuit reflection according to the folding area single-class label, it then determines the curvature parameter of the folding area of the training image.

[0113] 603. Set the first folding screen and the second folding screen according to the curvature parameter, and project light with a preset brightness value onto the first folding screen and the second folding screen through an external light source. The first folding screen is a folding screen without a microcircuit layer provided in the folding area, and the second folding screen is a folding screen with a microcircuit layer provided in the folding area.

[0114] 604. Collect images of the first folding screen and the second folding screen respectively, and calculate the microcircuit reflection parameter according to the folding areas of the two images.

[0115] Next, the terminal needs to fold the two folding screens to a corresponding degree according to the curvature parameter. The first folding screen is a folding screen without a microcircuit layer provided in the folding area (an ordinary folding screen of the same model), and the second folding screen is a folding screen with a microcircuit layer provided in the folding area (a new folding screen of the same model). Then, project light with a preset brightness value onto the first folding screen and the second folding screen through an external light source.

[0116] Next, the terminal collects images of the first folding screen and the second folding screen respectively, calculates the brightness value according to the folding areas of the two images, and then calculates the microcircuit reflection parameter according to the brightness difference.

[0117] 605. Construct a shallow feature gain convolution kernel according to the curvature parameter and the circuit reflection parameter.

[0118] 606. Perform shallow feature enhancement processing on the training image through the shallow feature gain convolution kernel to generate a feature map.

[0119] Next, the terminal constructs a shallow feature gain convolution kernel according to the curvature parameter and the circuit reflection parameter: When constructing a shallow feature gain convolutional kernel for a foldable screen, in addition to considering spatial proximity and pixel similarity, the curvature parameter (curvature information) and circuit reflection parameter (reflectivity information) of the new foldable screen are incorporated into the convolutional kernel. The curvature information can help us retain edge information in areas with large curvature changes such as fold edges, while the reflectivity information can be used to adjust the filtering intensity in areas with different circuit reflectivities, thereby reducing the filtering intensity in areas with large reflectivity differences (such as high-light or shadow areas) and retaining details.

[0120] The curvature and reflectivity information of the foldable screen include the curvature parameter κ(p) and the circuit reflection parameter R(p). The curvature parameter κ(p) represents the curvature value at pixel point p. The greater the curvature, the higher the degree of bending in that area. The circuit reflection parameter R(p) represents the reflectivity value at pixel point p. The reflectivity can reflect the gloss or light intensity of the surface, and areas with large reflectivity differences may contain important details or edges.

[0121] To construct a shallow feature gain convolutional kernel sensitive to the curvature parameter and circuit reflection parameter, in order to incorporate the curvature and reflectivity information into the traditional convolutional kernel, the weights of the traditional convolutional kernel depend not only on spatial proximity and pixel similarity but also on curvature differences and reflectivity differences. Specifically, the shallow feature gain convolutional kernel G(p,q) can be expressed as:

[0122] where: is the spatial proximity Gaussian kernel, which measures the spatial distance between p and q.

[0123] : the pixel similarity Gaussian kernel, which measures the pixel value difference between p and q, and are the pixel values at points p and q respectively.

[0124] is the curvature similarity Gaussian kernel, which measures the curvature difference between p and q, and are the curvatures at points p and q respectively.

[0125] is the reflectivity similarity Gaussian kernel, which measures the reflectivity difference between p and q, and are the reflectivities at points p and q respectively.

[0126] The specific form is as follows: Spatial proximity Gaussian kernel :

[0127] Gaussian kernel for pixel similarity :[[]]

[0128] Gaussian kernel for curvature similarity :[[]]

[0129] Gaussian kernel for reflectance similarity :[[]]

[0130] Among them,[[ID=2③]] is the standard deviation of the spatial distance, used to control the influence of spatial proximity. A larger will result in a wider smoothing. is the standard deviation of the pixel value difference, used to control the influence of pixel similarity. A smaller will enhance the sensitivity to pixel value differences, thus better retaining edges. While is the standard deviation on the curvature, used to control the influence of curvature similarity. A smaller will enhance the sensitivity to curvature differences, thus reducing the filtering intensity in areas with large curvature changes (such as folded edges). is the standard deviation of the reflection parameter of the circuit area, used to control the influence of reflectance similarity. A smaller will enhance the sensitivity to reflectance differences, thus reducing the filtering intensity in areas with large reflectance changes (such as highlight or shadow areas).

[0131] By incorporating curvature and reflectance information into the traditional convolution kernel, a convolution kernel with curvature and reflectance-sensitive shallow feature gain is constructed. This convolution kernel can adaptively adjust the filtering intensity according to the curvature and reflectance changes of the folding screen, thus retaining shallow feature information and details in areas with large curvature or reflectance changes (such as the folding area plus the circuit area), while enhancing the smoothing effect in areas with small curvature and reflectance changes. This method is particularly suitable for processing new folding screen images (where the folding area and the circuit area coincide).

[0132] Finally, the terminal performs shallow feature enhancement processing on the training image through the shallow feature gain convolution kernel to generate a feature map, and the shallow features in this folding area of the processed feature map can be strengthened specifically.

[0133] Please refer to Figure 7 , this application provides an embodiment of an extended method for a single-class detection model, including: 701. Expand a multi-classification head on the single-classification detection model for detecting the probability distribution of sub-category pseudo-labels.

[0134] 702. Determine the gray-scale discrimination weight value for each sub-category pseudo-label on the new display screen.

[0135] 703. Generate the gray-scale error weight value according to the curvature parameter and micro-circuit reflection parameter of the new display screen.

[0136] 704. Construct a multi-classification cross-entropy function according to the gray-scale discrimination weight value and the gray-scale error weight value.

[0137] 705. Construct the sub-category classification target.

[0138] The terminal expands a multi-classification head on the single-classification detection model for detecting the probability distribution of sub-category pseudo-labels, and then the terminal determines the gray-scale discrimination weight value for each sub-category pseudo-label on the new display screen. Specifically, it is necessary to determine the detection degree of different defect labels (each sub-category pseudo-label) in the folding area of the folding screen. At this time, the curvature parameter and the circuit reflection parameter also need to be considered. The gray-scale discrimination weight value is as follows:

[0139] is the gray-scale discrimination weight value for the i-th sub-category pseudo-label, is the detection probability distribution of the defect corresponding to the i-th sub-category pseudo-label under the current curvature parameter, is the detection probability distribution of the defect corresponding to the i-th sub-category pseudo-label under the current circuit reflection degree (micro-circuit reflection parameter). Both of these probability distributions can be determined by the historical detection model.

[0140] Then, the terminal constructs a multi-classification cross-entropy function according to the gray-scale discrimination weight value and the gray-scale error weight value. The function is as follows:

[0141] Finally, the terminal constructs the sub-category classification target. By incorporating the gray-scale error weight value, which is generated according to the curvature parameter and micro-circuit reflection parameter of the new display screen, the calculated multi-classification loss value can be made more in line with the new folding screen, and the single-classification detection model can be better updated.

[0142] Please refer to Figure 8 , an embodiment of a device for single-classification anomaly detection of a display screen provided by this application includes: The first construction unit 801 is used to construct a single-classification detection model and a single-classification loss function.

[0143] The first acquisition unit 802 is configured to acquire training images of the new display screen labeled with single-category labels, and input the training images into the single-classification detection model.

[0144] Optionally, the first acquisition unit 802 includes: Acquire the captured images of the new display screen, and divide the captured images into abnormal images and non-abnormal images.

[0145] Combine the abnormal images with the non-abnormal images, and expand the data volume through image enhancement technology to generate the first abnormally augmented image.

[0146] Identify the reference area in the non-abnormal image, replace the reference area with the corresponding abnormal area in the abnormal image, and perform assimilation processing on the replaced abnormal area to generate the second abnormally augmented image, where the reference area is the area where the permission influence weight reaches the preset value.

[0147] Perform scrambling processing on the non-abnormal image to generate the third abnormally augmented image.

[0148] Label the abnormal images, non-abnormal images, the first abnormally augmented image, the second abnormally augmented image, and the third abnormally augmented image with single-category labels, and determine them as training images.

[0149] Input the training images into the single-classification detection model.

[0150] The first generation unit 803 is configured to perform shallow feature enhancement processing on the training images according to the single-category labels to generate feature maps.

[0151] Optionally, the new display screen is a folding screen with a microcircuit layer provided in the folding area, the single-classification detection model is a folding screen single-classification detection model, the training images are the captured images of the folding area of the folding screen, and the single-category labels are the folding area single-category labels.

[0152] The first generation unit 803 includes: Determine that the gain data of the training images according to the folding area single-category labels are the folding screen curvature and the folding screen microcircuit reflection.

[0153] Determine the curvature parameters of the folding area of the training images.

[0154] Set the first folding screen and the second folding screen according to the curvature parameters, and project light with a preset brightness value onto the first folding screen and the second folding screen through an external light source. The first folding screen is a folding screen without a microcircuit layer provided in the folding area, and the second folding screen is a folding screen with a microcircuit layer provided in the folding area.

[0155] Collect the images of the first folding screen and the second folding screen respectively, and calculate the microcircuit reflection parameters according to the folding areas of the two images.

[0156] Construct a shallow feature gain convolution kernel according to the curvature parameter and the circuit reflection parameter.

[0157] Perform shallow feature enhancement processing on the training image through the shallow feature gain convolution kernel to generate a feature map.

[0158] A second generation unit 804, configured to respectively extract the shallow features of the training image and the feature map, perform feature fusion on the two sets of shallow features, then extract deep features, and generate a binary classification prediction result for the deep features according to a single classification head.

[0159] A first training unit 805, configured to calculate a first loss value according to the single classification loss function and the binary classification prediction result, and update the single classification detection model according to the first loss value until the initial training target is reached.

[0160] A third generation unit 806, configured to perform clustering processing on the deep features to generate a sub-category pseudo label for each training image.

[0161] A second construction unit 807, configured to expand a multi-classification head on the single classification detection model, and construct a multi-classification loss function and a sub-category classification target.

[0162] Optionally, the second construction unit 807 includes: Expand a multi-classification head on the single classification detection model for detecting the probability distribution of the sub-category pseudo labels.

[0163] Determine the gray discrimination weight value of each sub-category pseudo label on the new type of display screen.

[0164] Generate a gray error weight value according to the curvature parameter and the microcircuit reflection parameter of the new type of display screen.

[0165] Construct a multi-classification cross-entropy function according to the gray discrimination weight value and the gray error weight value.

[0166] Construct a sub-category classification target.

[0167] A second training unit 808, configured to jointly train the feature extractor, the two classification heads, and the sub-category pseudo labels in the single classification detection model through the deep features, the multi-classification head, the multi-classification loss function, and the sub-category classification target until the single classification head reaches the training target.

[0168] A second acquisition unit 809, configured to remove the expanded multi-classification head to obtain an unlabeled image in the new scenario.

[0169] A perturbation unit 810, configured to copy the unlabeled image and perform perturbation processing on the copied unlabeled image.

[0170] Optionally, the new display screen is a folding screen with a microcircuit layer provided in the folding area, the single-class detection model is a folding screen single-class detection model, and the unlabeled image is a real-time captured image of the folding area of the folding screen on the production line.

[0171] The perturbation unit 810 includes: Copy the unlabeled image.

[0172] Introduce a light source perturbation into the copied unlabeled image according to the curvature parameter of the folding screen in the folding area.

[0173] Introduce noise perturbation to the unlabeled image with perturbation through cropping and rotation.

[0174] The fourth generation unit 811 is used to perform shallow feature enhancement processing on the unlabeled image to generate a scene feature map.

[0175] The fifth generation unit 812 is used to extract shallow features from the unlabeled image, the unlabeled image after perturbation processing, and the scene feature map, and generate the first feature, the second feature, and the third feature respectively.

[0176] The sixth generation unit 813 is used to fuse the first feature and the third feature to generate a first fusion feature, and then perform depth feature extraction to generate a first depth feature.

[0177] The seventh generation unit 814 is used to fuse the second feature and the third feature to generate a second fusion feature, and then perform depth feature extraction to generate a second depth feature.

[0178] The eighth generation unit 815 is used to input the first depth feature and the second depth feature into the single-class head to generate a conventional prediction result and a perturbation prediction result.

[0179] The ninth generation unit 816 is used to construct a consistency loss function, and generate a consistency loss according to the conventional prediction result, the perturbation prediction result, and the consistency loss function.

[0180] The update unit 817 is used to update the new scene applicable parameters of the single-class detection model according to the consistency loss.

[0181] The annotation unit 818 is used to precisely annotate this part of the unlabeled image when the conventional prediction result and the perturbation prediction result show that there is an unlabeled image at the classification boundary.

[0182] The optimization unit 819 is used to determine the precisely annotated unlabeled image as an un-trained image, and optimize the single-class detection model with the un-trained image during the production process.

[0183] Please refer to Figure 9, this application provides a device for single-class anomaly detection of a display screen, including: A processor 901, a memory 902, an input / output unit 903, and a bus 904.

[0184] The processor 901 is connected to the memory 902, the input / output unit 903, and the bus 904.

[0185] The memory 902 stores a program, and the processor 901 calls the program to execute the methods as described in Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 .

[0186] This application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, it executes the methods as described in Figure 1 , Figure 2 and Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 .

[0187] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0188] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0189] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.

[0190] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0191] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.

Claims

1. A method for single-category anomaly detection of a display screen, characterized in that, Including: Construct a single-class detection model and a single-class loss function; Obtain training images of the new display screen labeled with single-class labels, and input the training images into the single-class detection model; Perform shallow feature enhancement processing on the training images according to the single-class labels to generate feature maps; Extract the shallow features of the training images and the feature maps respectively, perform feature fusion on the two sets of shallow features, then extract deep features, and generate binary classification prediction results for the deep features according to the single-class head; Calculate the first loss value according to the single-class loss function and the binary classification prediction results, and update the single-class detection model according to the first loss value until the initial training target is reached; Perform clustering processing on the deep features to generate sub-class pseudo-labels for each training image; Expand a multi-class head on the single-class detection model, and construct a multi-class loss function and a sub-class classification target; Jointly train the feature extractor, the two classification heads and the sub-class pseudo-labels in the single-class detection model through the deep features, the multi-class head, the multi-class loss function and the sub-class classification target until the single-class head reaches the training target.

2. The method according to claim 1, characterized in that, After the step of jointly training the feature extractor, the two classification heads and the sub-class pseudo-labels in the single-class detection model through the deep features, the multi-class head, the multi-class loss function and the sub-class classification target, the method further includes: Remove the expanded multi-class head to obtain unlabeled images in the new scenario; Copy the unlabeled images, and perform perturbation processing on the copied unlabeled images; Perform shallow feature enhancement processing on the unlabeled images to generate scene feature maps; Extract the shallow features of the unlabeled images, the perturbed unlabeled images and the scene feature maps respectively to generate the first feature, the second feature and the third feature; Fuse the first feature and the third feature to generate the first fused feature, and then extract deep features to generate the first deep feature; Fuse the second feature and the third feature to generate the second fused feature, and then extract deep features to generate the second deep feature; Input the first deep feature and the second deep feature into the single-class head to generate a normal prediction result and a perturbed prediction result; Construct a consistency loss function, and generate a consistency loss according to the normal prediction result, the perturbed prediction result and the consistency loss function; Update the parameters of the single-class detection model applicable to the new scenario according to the consistency loss.

3. The method according to claim 2, characterized in that, After the step of updating the parameters of the single-class detection model applicable to the new scenario according to the consistency loss, the method further includes: When the normal prediction result and the perturbed prediction result show that there are unlabeled images at the classification boundary, perform precise annotation on these unlabeled images; Determine the precisely annotated unlabeled images as untrained images, and use the untrained images to optimize the single-class detection model during the production process.

4. The method according to claim 2, characterized in that The new display screen is a folding screen with a microcircuit layer in the folding area. The single-class detection model is a folding screen single-class detection model. The unlabeled image is a real-time captured image of the folding area of the folding screen on the production line; The step of copying the unlabeled image and performing perturbation processing on the copied unlabeled image includes: Copy the unlabeled image; Introduce light source perturbation into the copied unlabeled image according to the curvature parameter of the folding screen in the folding area; Introduce noise perturbation into the unlabeled image with perturbation introduced by means of cropping and rotation.

5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining the training image of the new display screen labeled with a single-class label and inputting the training image into the single-class detection model includes: Obtain the captured image of the new display screen and divide the captured image into abnormal images and non-abnormal images; Combine the abnormal images and the non-abnormal images, and expand the data volume through image enhancement technology to generate the first abnormal augmented image; Identify the reference area in the non-abnormal image, replace the reference area with the corresponding abnormal area in the abnormal image, and perform assimilation processing on the replaced abnormal area to generate the second abnormal augmented image. The reference area is the area where the permission influence weight reaches the preset value; Perform perturbation processing on the non-abnormal image to generate the third abnormal augmented image; Label the single-class label for the abnormal images, the non-abnormal images, the first abnormal augmented image, the second abnormal augmented image, and the third abnormal augmented image, and determine them as training images; Input the training images into the single-class detection model.

6. The method according to any one of claims 1 to 4, characterized in that, The new display screen is a folding screen with a microcircuit layer in the folding area. The single-class detection model is a folding screen single-class detection model. The training image is the captured image of the folding area of the folding screen. The single-class label is the folding area single-class label; Perform shallow feature enhancement processing on the training image according to the single-class label to generate a feature map, including: Determine that the gain data of the training image according to the folding area single-class label is the folding screen curvature and the folding screen microcircuit reflection; Determine the curvature parameter of the folding area of the training image; Set the first folding screen and the second folding screen according to the curvature parameter, and project light rays with a preset brightness value onto the first folding screen and the second folding screen through an external light source. The first folding screen is a folding screen without a microcircuit layer in the folding area, and the second folding screen is a folding screen with a microcircuit layer in the folding area; Collect the images of the first folding screen and the second folding screen respectively, and calculate the microcircuit reflection parameter according to the folding areas of the two images; Construct a shallow feature gain convolution kernel according to the curvature parameter and the circuit reflection parameter; Perform shallow feature enhancement processing on the training image through the shallow feature gain convolution kernel to generate a feature map.

7. The method according to claim 6, wherein The step of expanding a multi-classification head on the single-class detection model and constructing a multi-classification loss function and a sub-classification target includes: Expand a multi-classification head on the single-class detection model for detecting the probability distribution of the sub-class pseudo-labels; Determine the gray discrimination weight value of each sub-class pseudo-label on the new display screen; Generate a grayscale error weight value based on the curvature parameter and microcircuit reflection parameter of the novel display screen; Construct a multi-class cross-entropy function according to the grayscale discrimination weight value and the grayscale error weight value; Construct a sub-category classification target.

8. An apparatus for single-classification anomaly detection of a display screen, characterized in that, It includes: The first construction unit is used to construct a single-class detection model and a single-class loss function; The first acquisition unit is used to acquire a training image labeled with a single-class label of the novel display screen and input the training image into the single-class detection model; The first generation unit is used to perform shallow feature enhancement processing on the training image according to the single-class label to generate a feature map; The second generation unit is used to extract the shallow features of the training image and the feature map respectively, perform feature fusion on the two groups of shallow features, then extract deep features, and generate a binary classification prediction result for the deep features according to the single-class head; The first training unit is used to calculate a first loss value according to the single-class loss function and the binary classification prediction result, and update the single-class detection model according to the first loss value until the initial training target is reached; The third generation unit is used to perform clustering processing on the deep features to generate a sub-category pseudo label for each training image; The second construction unit is used to expand a multi-class head on the single-class detection model and construct a multi-class loss function and a sub-category classification target; The second training unit is used to jointly train the feature extractor, the two classification heads and the sub-category pseudo label in the single-class detection model through the deep features, the multi-class head, the multi-class loss function and the sub-category classification target until the single-class head reaches the training target.

9. The device according to claim 8, characterized in that, After the second training unit, the device further includes: The second acquisition unit is used to remove the expanded multi-class head and acquire an unlabeled image in the new scenario; The perturbation unit is used to copy the unlabeled image and perform perturbation processing on the copied unlabeled image; The fourth generation unit is used to perform shallow feature enhancement processing on the unlabeled image to generate a scene feature map; The fifth generation unit is used to extract the shallow features of the unlabeled image, the perturbed unlabeled image and the scene feature map respectively to generate a first feature, a second feature and a third feature; The sixth generation unit is used to fuse the first feature and the third feature to generate a first fused feature, and then perform deep feature extraction to generate a first deep feature; The seventh generation unit is used to fuse the second feature and the third feature to generate a second fused feature, and then perform deep feature extraction to generate a second deep feature; The eighth generation unit is used to input the first deep feature and the second deep feature into the single-class head to generate a normal prediction result and a perturbed prediction result; The ninth generation unit is used to construct a consistency loss function and generate a consistency loss according to the normal prediction result, the perturbed prediction result and the consistency loss function; The update unit is used to update the parameters applicable to the new scenario of the single-class detection model according to the consistency loss.

10. A computer-readable storage medium, characterized in that, A program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unsupervised anomaly detection method based on multi-task auto-encoder and exponential loss

    CN118378187A

  • Transform-based image classification method, system and device, and storage medium

    CN118799651A

  • Generalized deep forgery detection method based on multi-classification guidance

    CN119445679A

  • Defect detection using multiple models

    US20210209414A1

Cited By

  • LED chip detection method, system and equipment

    CN121391725A