OLED display screen picture defect detection method based on machine learning
By combining correlation analysis of brightness and chromaticity data, multi-view image acquisition, and convolutional neural networks, the problems of inaccurate detection results and insufficient positioning accuracy in OLED display defect detection have been solved, achieving more efficient defect identification and positioning.
Patent Information
- Application Number
- CN202511131414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies cannot fully utilize the correlation between brightness and chromaticity for OLED display defect detection, resulting in inaccurate detection results; they fail to fully consider the texture and spatial distribution characteristics of defects, limiting their ability to identify complex defects; and their defect localization accuracy is insufficient under multi-view image acquisition, making effective neighborhood analysis impossible.
By acquiring luminance and chromaticity distribution data, calculating the spatial difference between the luminance distribution and the standard luminance template, determining the geometric center coordinates of the defect area, and using a convolutional neural network to extract texture maps, combined with multi-view image acquisition and spatial coordinate transformation, defect feature matching and localization are performed.
It improves the comprehensiveness and reliability of defect detection, accurately locates defect areas, enhances the ability to identify defects in complex backgrounds, and reduces false positives and false negatives.
Smart Images

Figure CN121032945A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of display panel quality detection, and more particularly, to an OLED display screen picture defect detection method based on machine learning. BACKGROUND
[0002] In today's display technology field, OLED display screens are widely used due to their excellent contrast, color performance and response speed. However, various picture defects such as bright spots, dark spots, color spots, etc. may occur during the production process of OLED display screens, which will seriously affect the display quality and user experience. Traditional defect detection methods mainly rely on manual visual inspection or simple image processing technology. Manual visual inspection is not only inefficient, but also easily affected by subjective factors, leading to inconsistent detection results. While simple image processing technology can achieve a certain degree of automated detection, its detection accuracy is limited and it is difficult to accurately identify complex defect types, especially when the defect characteristics are not obvious or have small differences with normal pictures, which may lead to misjudgment or omission.
[0003] With the continuous development of machine learning technology, its application in image recognition and analysis field has gradually attracted attention. Machine learning algorithms can automatically extract features and classify through learning a large amount of image data, thereby achieving more accurate detection of OLED display screen picture defects. However, existing defect detection methods based on machine learning have some limitations. For example, some methods only rely on single brightness or chroma information for defect detection, ignoring the correlation between brightness and chroma, resulting in incomplete detection results. Some methods do not fully consider the texture and spatial distribution characteristics of defects when extracting defect features, making it difficult to identify complex defects. In addition, the accuracy of defect positioning in existing technologies needs to be improved, especially in the case of multi-view image acquisition, how to accurately determine the geometric center coordinates of the defect and effectively analyze the neighborhood is still a problem to be solved.
[0004] There are at least the following problems or defects in the prior art: the prior art cannot fully utilize the correlation between brightness and chroma for defect detection, resulting in inaccurate detection results; the prior art does not fully consider the texture and spatial distribution characteristics of defects when extracting defect features, limiting the recognition ability of complex defects; the prior art lacks accuracy in defect positioning in the case of multi-view image acquisition, and cannot effectively analyze the neighborhood to determine whether the defect region contains normal pixel units. SUMMARY
[0005] This invention provides a machine learning-based method for detecting defects in OLED displays. The OLED display includes several pixel units located on the same detection plane. An image acquisition device is disposed on the detection plane. The image acquisition device is used to acquire brightness distribution data and chromaticity distribution data when the OLED display is displaying an image. The OLED display defect detection method includes:
[0006] In response to the luminance distribution data, acquire the chromaticity distribution data associated with the luminance distribution data;
[0007] The degree of abnormality in the defect area is calculated by the spatial difference between the brightness distribution data and the standard brightness template.
[0008] The geometric center coordinates of the defect area are determined based on the image acquisition device.
[0009] Determine whether the neighborhood of the geometric center coordinates contains normal pixel units;
[0010] If not, then a texture map based on defect features is extracted from the brightness distribution data using a convolutional neural network;
[0011] Analyze whether the texture map matches a preset defect type;
[0012] If so, defect location information and quality warning instructions are generated and sent to the external testing terminal.
[0013] Furthermore, in response to the brightness distribution data, acquiring chromaticity distribution data associated with the brightness distribution data includes: real-time acquisition of frame sequence images of the OLED display screen;
[0014] The brightness quantization value and color quantization value of each pixel unit are obtained by the image acquisition device based on a preset sampling frequency.
[0015] A three-dimensional feature space is constructed with pixel coordinates as the horizontal axis and brightness and chromaticity values as the vertical axis.
[0016] The luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface;
[0017] Determine whether the gradient between the brightness value of the current pixel and the brightness values of its neighboring pixels exceeds a preset tolerance;
[0018] If so, it is determined to be a potential defect area;
[0019] Extract the continuous region with the largest gradient change in the brightness surface and mark it as candidate defect data;
[0020] Within the pixel range corresponding to the candidate defect data, obtain the chromaticity distribution data with the most significant chromaticity change trend.
[0021] Further, the luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface. Then, the process includes transmitting the three-dimensional feature space, the luminance surface, and the chrominance surface to an external visualization analysis platform.
[0022] Further, determining the geometric center coordinates of the defect area based on the image acquisition device includes: setting at least two image acquisition devices with different perspectives on the detection plane;
[0023] A screen coordinate system is established with the optical axis of the first image acquisition device as the reference axis;
[0024] Mark the installation locations of other image acquisition devices in the screen coordinate system;
[0025] Calculate the boundary coordinates of the defect area captured by each image acquisition device;
[0026] The spatial coordinate transformation is performed on the capture results of each image acquisition device to generate multiple sets of projection boundaries;
[0027] The center point of the overlapping region of multiple sets of projected boundaries is the coordinate of the geometric center.
[0028] Furthermore, the center point of the overlapping area of multiple sets of projection boundaries is obtained as the geometric center coordinates. Then, the screen coordinate system, the positions of each image acquisition device, and the geometric center coordinates are superimposed and displayed on the external three-dimensional reconstruction interface.
[0029] Furthermore, the texture map based on defect features in the brightness distribution data is extracted by a convolutional neural network, including: dividing the brightness distribution data into multiple detection sub-regions using a sliding window;
[0030] Multi-scale feature extraction is performed in parallel on each detection sub-region to generate a primary feature map;
[0031] By fusing primary feature maps of different scales through the residual connection module, an intermediate feature map is formed.
[0032] An attention mechanism is applied to the intermediate feature map to calculate the weight coefficients of each channel;
[0033] The intermediate feature map is reconstructed by weighting coefficients to obtain the high-level feature map;
[0034] The high-level feature map is input into the classification head network, and the defect type confidence score is output.
[0035] The feature response regions with confidence levels higher than a threshold are mapped back to the original image space to generate the texture map.
[0036] Further, it is determined whether the neighborhood of the geometric center coordinates contains normal pixel units, and then the process includes: if so, marking the defective region as a repairable defect;
[0037] Record the geometric parameters, occurrence frame number, and duration of the repairable defect, and generate a repair priority list;
[0038] Synchronize the repair priority list to the external maintenance management system.
[0039] Further, the multi-scale feature extraction includes: inputting the detection sub-region into a parallel processing branch; the first branch performs dilated convolution using a 3×3 convolution kernel with a dilation rate of 2 to generate a first feature map containing long-distance correlation features; the second branch performs downsampling convolution using a 5×5 convolution kernel with a stride of 2 to generate a second feature map containing local detail features; the third branch compresses the feature dimension through a max pooling layer followed by a 1×1 convolution kernel to perform channel dimensionality reduction to generate a third feature map containing global semantic features; and the first feature map, the second feature map, and the third feature map are concatenated along the channel dimension to form the primary feature map.
[0040] Furthermore, the spatial coordinate transformation includes: inputting the boundary coordinates of the defect area captured by each image acquisition device into a perspective projection matrix, which is pre-calibrated using a checkerboard calibration plate during the installation phase; establishing a mapping relationship between the original pixel coordinates and the coordinates after radial distortion correction, wherein the radial distortion correction uses bilinear interpolation to compensate for the distortion coefficients; and transforming the corrected two-dimensional coordinates in the coordinate system of each image acquisition device to the screen coordinate system using a homogeneous coordinate transformation matrix to generate a set of projection boundary coordinates corresponding to the viewpoint of each device.
[0041] Furthermore, the classification head network includes: inputting a high-level feature map into an encoder consisting of three fully connected layers; the first fully connected layer compresses the feature dimension to 1 / 4 of its original size and outputs an embedding vector; the second fully connected layer applies L2 regularization constraints to the embedding vector and outputs latent features; the third fully connected layer maps the latent features to a classification dimension equal to the number of preset defect types; inputting the classification dimension vector into a Softmax classifier with a temperature coefficient, wherein the temperature coefficient is dynamically adjusted through adversarial training to enhance noise robustness; and finally, a noise robustness discrimination module calibrates the classification results to output the final defect type confidence.
[0042] The embodiments of the present invention have at least the following beneficial effects:
[0043] 1. By combining the spatial difference calculation of brightness distribution data and standard template with the correlation analysis of chromaticity distribution data, defect areas of OLED displays can be accurately identified, effectively solving the problem of missed detection caused by insufficient correlation between brightness anomalies and chromaticity shifts in traditional detection methods, and improving the comprehensiveness and reliability of defect detection.
[0044] 2. By employing a multi-view image acquisition device and spatial coordinate transformation technology, the geometric center coordinates of the defect area are accurately calculated. The neighboring pixel judgment mechanism distinguishes between repairable and unrepairable defects, solving the problem of insufficient defect positioning accuracy in traditional detection and providing accurate spatial reference for subsequent repair or quality assessment.
[0045] 3. Based on a multi-scale feature extraction and classification head network of convolutional neural networks, combined with attention mechanism and noise robustness optimization, it can adaptively extract defect texture features and accurately match preset defect types, effectively overcoming the misjudgment problem caused by small defects or noise interference in complex backgrounds, and improving the accuracy and stability of defect classification. Attached Figure Description
[0046] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of the invention are illustrated in the drawings by way of example and not limitation, wherein:
[0047] Figure 1 This is a flowchart illustrating a machine learning-based OLED display screen defect detection method according to an embodiment of the present invention. Detailed Implementation
[0048] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make the invention more thorough and complete, and to fully convey the scope of the invention to those skilled in the art.
[0049] Those skilled in the art will understand that embodiments of the present invention can be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0050] It should be noted that the number of any elements in the accompanying drawings is for illustrative purposes only and not as a limitation, and any naming is for distinction only and has no limiting meaning.
[0051] The following is for reference. Figure 1, Figure 1 This is a flowchart illustrating a machine learning-based defect detection method for OLED displays, provided in an embodiment of the present invention. Figure 1 As shown, a machine learning-based method for detecting defects in OLED display screens includes:
[0052] S1. In response to the luminance distribution data, acquire the chromaticity distribution data associated with the luminance distribution data;
[0053] S2. The degree of abnormality index of the defect area is calculated by the spatial difference between the brightness distribution data and the standard brightness template.
[0054] S3. Determine the geometric center coordinates of the defect area based on the image acquisition device;
[0055] S4. Determine whether the neighborhood of the geometric center coordinates contains normal pixel units;
[0056] S5. If not, then extract the texture map based on defect features from the brightness distribution data using a convolutional neural network;
[0057] S6. Analyze whether the texture map matches the preset defect type;
[0058] S7. If so, generate defect location information and quality warning instructions and send them to the external testing terminal.
[0059] It should be noted that this invention proposes a machine learning-based method for detecting defects in OLED displays. The OLED display is composed of several pixel units, which together constitute the basic unit of the displayed image. The image acquisition device is a key device for acquiring brightness and chromaticity distribution data of the OLED display. Brightness distribution data reflects the brightness of different locations in the image, while chromaticity distribution data describes the color of different locations. Acquiring associated chromaticity distribution data in response to brightness distribution data means that brightness and chromaticity data are interrelated during the detection process, and this correlation helps to more comprehensively identify image defects. The anomaly index is obtained by calculating the spatial difference between the brightness distribution data and a standard brightness template; it quantifies the degree of difference between the defective area and the normal area. The geometric center coordinates refer to the center position of the defective area on the detection plane. Determining these coordinates using the image acquisition device allows for precise location of the defect. Determining whether the neighborhood of the geometric center coordinates contains normal pixel units further confirms the validity of the defective area. If no normal pixel units are present in the neighborhood, a texture map based on defect features is extracted from the brightness distribution data using a convolutional neural network. The texture map is a visual representation of the defect features, aiding in subsequent defect type analysis. Analyzing whether the texture map matches a preset defect type is crucial for determining the specific defect type. If a match is found, defect location information and a quality warning instruction are generated and sent to an external detection terminal. This process automates and intelligentizes defect detection, improving efficiency and accuracy.
[0060] Specifically, the pixel unit of an OLED display is the smallest unit that makes up the displayed image, and each pixel unit has its own brightness and chromaticity values. The image acquisition device can be a high-resolution camera or other imaging equipment used to capture the image information of the display screen. Brightness distribution data refers to the set of brightness values possessed by pixel units at different locations on the display screen; these brightness values can be directly measured by the image acquisition device. Chromaticity distribution data describes the color characteristics of pixel units, typically including parameters such as hue and saturation. A standard brightness template is a predefined reference model representing the brightness distribution of a normal display image. By comparing it with the actually measured brightness distribution data, potential defect areas can be identified. The geometric center coordinates are calculated using image processing algorithms; they represent the center position of the defect area on the display plane. The neighborhood range refers to the set of pixel units within a certain area centered on the geometric center coordinates. This range can be set according to actual needs; for example, it can be a square or circular area centered on the geometric center coordinates. A convolutional neural network is a deep learning model that automatically extracts image features by learning from large amounts of image data, used to identify and classify different image patterns. Texture maps are extracted using convolutional neural networks. They contain texture features of defective regions, such as edges, lines, and spots. These features are matched with preset defect types to determine the specific type of defect. Preset defect types are defined based on the types of defects that may occur during actual production and use, such as bright spots, dark spots, and discoloration. Defect location information includes parameters such as the defect's location and size. A quality warning command is a signal used to notify external inspection terminals of the presence of a defect so that appropriate corrective measures can be taken.
[0061] Preferably, the image acquisition device uses a high-resolution industrial camera, whose sampling frequency can be set according to the refresh rate of the display screen and the detection accuracy requirements, for example, it can be set to 60 frames per second or higher. When constructing the three-dimensional feature space, pixel coordinates are used as the horizontal axis, and brightness and chromaticity values are used as the vertical axis, which can intuitively represent the brightness and chromaticity information of each pixel unit. When calculating the spatial difference between the brightness distribution data and the standard brightness template, Euclidean distance or other similarity measurement methods can be used to quantify the degree of difference. The construction of the convolutional neural network includes multiple convolutional layers, pooling layers, and fully connected layers. The input parameter is the brightness distribution data. Through multi-layer feature extraction and learning, the texture map of defect features is finally output. When analyzing the matching of the texture map with the preset defect type, support vector machine or other classification algorithms can be used to determine whether there is a match based on the similarity between the feature vector of the texture map and the feature vector of the preset defect type. If the match is successful, defect location information containing information such as defect location, size, and type is generated, and a quality warning command is sent to the external detection terminal through the network or other communication methods for timely subsequent processing and repair.
[0062] In some embodiments, in response to the luminance distribution data, acquiring chromaticity distribution data associated with the luminance distribution data includes: real-time acquisition of frame sequence images of the OLED display screen;
[0063] The brightness quantization value and color quantization value of each pixel unit are obtained by the image acquisition device based on a preset sampling frequency.
[0064] A three-dimensional feature space is constructed with pixel coordinates as the horizontal axis and brightness and chromaticity values as the vertical axis.
[0065] The luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface;
[0066] Determine whether the gradient between the brightness value of the current pixel and the brightness values of its neighboring pixels exceeds a preset tolerance;
[0067] If so, it is determined to be a potential defect area;
[0068] Extract the continuous region with the largest gradient change in the brightness surface and mark it as candidate defect data;
[0069] Within the pixel range corresponding to the candidate defect data, obtain the chromaticity distribution data with the most significant chromaticity change trend.
[0070] It should be noted that the process of acquiring associated chromaticity distribution data in response to luminance distribution data in this invention is achieved by real-time acquisition of frame sequence images from the OLED display screen. Frame sequence images refer to a series of images displayed on the screen within a continuous time interval, containing luminance and chromaticity information of the display screen at different points in time. The luminance quantization value and chromaticity quantization value of each pixel unit are acquired by an image acquisition device at a preset sampling frequency. The luminance quantization value converts the continuous luminance signal into discrete values for computer processing; similarly, the chromaticity quantization value is a digital representation of color information. A three-dimensional feature space is constructed with pixel coordinates as the horizontal axis and luminance and chromaticity values as the vertical axis, allowing for a visual representation of the luminance and chromaticity information of each pixel unit in space. Luminance and chromaticity surfaces are formed in the three-dimensional feature space, describing the distribution of luminance and chromaticity on the display screen, respectively. Determining whether the gradient between the current pixel's luminance value and the luminance values of adjacent pixels exceeds a preset tolerance is to detect abnormal luminance changes; areas with gradients exceeding the preset tolerance may be potential defect areas. Extracting the continuous region with the largest gradient change in the luminance surface is to determine the range of areas that may contain defects and mark them as candidate defect data. Within the pixel range corresponding to the candidate defect data, the chromaticity distribution data with the most significant chromaticity change trend is obtained in order to further analyze the chromaticity information and identify defects more accurately.
[0071] Specifically, real-time acquisition of frame sequence images from an OLED display means that the image acquisition device needs to continuously capture images from the display to obtain complete dynamic information. The preset sampling frequency refers to the frequency at which the image acquisition device acquires images. This frequency can be set according to the display's refresh rate and detection accuracy requirements; for example, it can be set to 60 frames per second. The acquisition of luminance and chrominance quantization values is accomplished by the sensors of the image acquisition device. The sensors convert the received light signals into electrical signals, and then the analog signals are converted into digital signals by an analog-to-digital converter. Pixel coordinates are the positional identifiers of each pixel unit on the display; they are two-dimensional coordinates used to determine the specific location of the pixel unit on the display. Luminance and chrominance values are used as the vertical axis to construct a three-dimensional feature space, combining the luminance and chrominance information of each pixel unit with its position information to form a complete feature space. Luminance and chrominance surfaces are formed in this feature space, reflecting the distribution of luminance and chrominance on the display. Gradient refers to the rate of change of luminance or chrominance in space; the preset tolerance is a threshold used to determine whether gradient changes are abnormal. The specific range of the preset tolerance can be adjusted according to the actual application scenario and the characteristics of the OLED display. For example, in some high-precision inspection scenarios, the preset tolerance can be set to within 5% of the brightness value to ensure sensitive detection even for minute brightness changes; while in scenarios with relatively lower accuracy requirements, the preset tolerance can be appropriately relaxed to 10%. This adjustment is based on the brightness fluctuation range of the OLED display under normal operating conditions and the sensitivity requirements for defect detection. If the display's brightness fluctuation is small and high-sensitivity detection is required, the preset tolerance should be set smaller; conversely, if the display's brightness fluctuation is large and the sensitivity requirements are not high, the preset tolerance can be set larger. This ensures that the inspection method can effectively identify defects while avoiding misjudgments caused by normal brightness fluctuations.
[0072] Candidate defect data is obtained by extracting continuous regions with the largest gradient changes; these regions are likely where defects are located. The chromaticity distribution data with the most significant chromaticity change trends refers to the data showing the most obvious chromaticity changes within the candidate defect regions; this data is crucial for identifying defect types.
[0073] Preferably, the image acquisition device can employ a high-resolution industrial camera, whose sampling frequency can be set according to the refresh rate and detection accuracy requirements of the display screen, for example, it can be set to 60 frames per second or higher. When constructing the three-dimensional feature space, computer graphics methods can be used to map the brightness and chromaticity values of each pixel unit into the space, forming brightness and chromaticity surfaces. To determine whether the gradient exceeds the preset tolerance, image processing algorithms, such as the Sobel operator or the Canny edge detection algorithm, can be used to calculate the brightness and chromaticity gradients. The preset tolerance can be adjusted according to the actual display screen characteristics and defect detection requirements, for example, it can be set to 10% of the brightness value change or 15% of the chromaticity value change. When extracting candidate defect data, a region growing algorithm or a connected component analysis algorithm can be used, starting from the point with the largest gradient change and gradually expanding to surrounding pixel units until the gradient change is less than the preset tolerance. When obtaining the chromaticity distribution data with the most significant chromaticity change trend, principal component analysis (PCA) or other statistical analysis methods can be used to analyze the chromaticity data within the candidate defect region and extract the chromaticity distribution data with the most significant change trend. These data can serve as the basis for subsequent defect type analysis.
[0074] In some embodiments, the luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface. Then, the method includes transmitting the three-dimensional feature space, the luminance surface, and the chrominance surface to an external visualization analysis platform.
[0075] It should be noted that after constructing the luminance and chromaticity surfaces, this invention transmits these data to an external visualization and analysis platform. This allows technicians to more intuitively observe and analyze the luminance and chromaticity distribution of the display screen. The external visualization and analysis platform is a software system specifically designed for data visualization and analysis. It receives luminance and chromaticity data acquired from image acquisition devices and displays it graphically. In this way, technicians can more easily identify potential defect areas, further analyze the characteristics of defects, and thus improve the accuracy and efficiency of defect detection.
[0076] Specifically, the three-dimensional feature space is constructed using pixel coordinates as the horizontal axis and luminance and chrominance values as the vertical axis. It is a multi-dimensional data space containing luminance and chrominance information for all pixel units. Luminance and chrominance surfaces are formed within this space, representing the distribution of luminance and chrominance on the display screen, respectively. Transmitting this data to an external visualization analysis platform means sending the constructed three-dimensional feature space, luminance surface, and chrominance surface data to specialized analysis software via network or other data transmission methods. This process involves data format conversion and transmission protocol settings. For example, TCP / IP protocol can be used for data transmission, and the data format can be converted to CSV or JSON format so that the visualization analysis platform can correctly read and parse it.
[0077] Preferably, to ensure the accuracy and efficiency of data transmission, a dedicated data transmission module can be used to handle data sending and receiving. This module can be set to automatic mode, automatically initiating the data transmission process once the three-dimensional feature space, brightness surface, and chromaticity surface are constructed. During data transmission, a data verification mechanism can be set, such as CRC checksum or MD5 checksum, to ensure that no errors occur during transmission. Simultaneously, to improve data transmission efficiency, data compression can be performed, for example, using the GZIP compression algorithm, to reduce the data transmission size. After the external visualization analysis platform receives the data, automatic parsing and display functions can be set up. Once data transmission is complete, the brightness and chromaticity surfaces are immediately displayed graphically on the platform, facilitating real-time analysis and diagnosis by technicians. Furthermore, the visualization analysis platform can also provide interactive functions, such as allowing technicians to select specific areas for zooming in or out using a mouse to observe the brightness and chromaticity changes in specific areas in more detail.
[0078] In some embodiments, determining the geometric center coordinates of the defect region based on the image acquisition device includes: setting at least two image acquisition devices with different perspectives on the detection plane;
[0079] A screen coordinate system is established with the optical axis of the first image acquisition device as the reference axis;
[0080] Mark the installation locations of other image acquisition devices in the screen coordinate system;
[0081] Calculate the boundary coordinates of the defect area captured by each image acquisition device;
[0082] The spatial coordinate transformation is performed on the capture results of each image acquisition device to generate multiple sets of projection boundaries;
[0083] The center point of the overlapping region of multiple sets of projected boundaries is the coordinate of the geometric center.
[0084] It should be noted that the process of determining the geometric center coordinates of the defect area based on the image acquisition device in this invention is achieved by setting up at least two image acquisition devices with different viewing angles on the detection plane. The image acquisition devices capture images from the display screen from different angles, thereby obtaining the boundary coordinates of the defect area. By establishing a screen coordinate system with the optical axis of the first image acquisition device as the reference axis and marking the installation positions of the other image acquisition devices, image data from different viewing angles can be unified into the same coordinate system. After calculating the boundary coordinates of the defect area captured by each image acquisition device, a spatial coordinate transformation is performed on the capture results of each image acquisition device to generate multiple sets of projected boundaries. Finally, by solving for the center point of the overlapping area of the multiple sets of projected boundaries, the geometric center coordinates of the defect area are determined. This process not only improves the accuracy of defect location but also provides accurate location information for subsequent defect analysis and repair.
[0085] Specifically, the detection plane refers to the plane on which the OLED display is located. Image acquisition devices are installed at different positions on this plane to acquire image information from the display. At least two image acquisition devices with different viewing angles mean that these devices capture images of the display from different angles, providing data from multiple perspectives and helping to more comprehensively locate defects. The optical axis of the first image acquisition device serves as the reference axis to establish the screen coordinate system, which is a reference frame used to unify image data from different viewing angles. The installation positions of other image acquisition devices are marked in the screen coordinate system so that the image data they capture can be transformed into spatial coordinates. Spatial coordinate transformation is the process of converting image data from different perspectives into the same coordinate system, which involves image projection and transformation algorithms. The projection boundary refers to the boundary of the defect area captured by each image acquisition device in the screen coordinate system. By calculating the center point of the overlapping area of these projection boundaries, the geometric center coordinates of the defect area can be determined; this coordinate system is a precise representation of the defect location.
[0086] Preferably, to achieve more accurate defect localization, high-precision image acquisition devices, such as industrial cameras with high resolution and low distortion, can be used. The installation position and angle of these cameras need to be precisely measured and calibrated to ensure the accuracy and reliability of the acquired image data. When establishing the screen coordinate system, a checkerboard calibration board can be used for camera calibration. The calibration process determines the camera's intrinsic and extrinsic parameters, thereby achieving accurate spatial coordinate transformation. When calculating the projection boundary, edge detection algorithms, such as the Canny edge detection algorithm, can be used to accurately extract the boundary of the defect area. For spatial coordinate transformation, a perspective projection matrix can be used. This matrix, obtained through the calibration process, is used to convert image data from different perspectives to the screen coordinate system. When solving for the center point of the overlapping area, geometric algorithms, such as calculating the intersection center of polygons, can be used to determine the coordinates of the geometric center. Furthermore, to improve the robustness of localization, an error correction mechanism can be introduced during the calculation process, such as reducing the impact of random errors through multiple measurements and averaging.
[0087] In some embodiments, the center point of the overlapping area of multiple sets of projection boundaries is obtained as the geometric center coordinates. Then, the screen coordinate system, the positions of each image acquisition device, and the geometric center coordinates are superimposed and displayed on the external three-dimensional reconstruction interface.
[0088] It should be noted that after solving for the coordinates of the center points (geometric centers) of the overlapping areas of multiple sets of projected boundaries, this invention superimposes and displays the screen coordinate system, the positions of each image acquisition device, and the coordinates of the geometric center on the external 3D reconstruction interface. This process is to provide intuitive visual feedback, helping technicians to more clearly understand the specific location and distribution of the defect area on the display screen. The external 3D reconstruction interface is a visualization platform capable of displaying data in 3D form, allowing users to view and analyze data from different angles, thereby better understanding the geometric features and spatial relationships of the defects.
[0089] Specifically, the screen coordinate system is established based on the optical axis direction of the first image acquisition device, serving as a reference frame to unify image data from different perspectives. The positions of each image acquisition device refer to their specific installation locations marked in the screen coordinate system; this positional information was obtained through the initial calibration process. The geometric center coordinates are obtained by calculating the center point of the overlapping area of multiple sets of projected boundaries, representing the precise location of the defect area on the display screen. Overlaying this information onto the external 3D reconstruction interface means integrating this data graphically into a three-dimensional space, allowing users to intuitively see the location of the defect area and its surrounding environment. This process involves data visualization processing, including data format conversion and graphic rendering.
[0090] Preferably, to achieve a more intuitive visualization, different colors or markers can be used on the external 3D reconstruction interface to distinguish the screen coordinate system, the image acquisition device position, and the geometric center coordinates. For example, the screen coordinate system can be represented by grid lines, the image acquisition device position can be represented by dots or icons of different colors, and the geometric center coordinates can be highlighted with a prominent marker, such as a red cross. During data overlay, 3D graphics libraries, such as OpenGL or Vulkan, can be used to achieve efficient graphics rendering. Furthermore, to improve the user experience, interactive functions can be provided on the 3D reconstruction interface, such as allowing users to rotate, zoom, and drag the view using the mouse to observe the defect area from different angles. Simultaneously, an automatic update mechanism can be set up so that once the new geometric center coordinates are calculated, the display on the 3D reconstruction interface is immediately updated, ensuring that the user always sees the latest data.
[0091] In some embodiments, extracting texture maps based on defect features from the brightness distribution data using a convolutional neural network includes: dividing the brightness distribution data into multiple detection sub-regions using a sliding window;
[0092] Multi-scale feature extraction is performed in parallel on each detection sub-region to generate a primary feature map;
[0093] By fusing primary feature maps of different scales through the residual connection module, an intermediate feature map is formed.
[0094] An attention mechanism is applied to the intermediate feature map to calculate the weight coefficients of each channel;
[0095] The intermediate feature map is reconstructed by weighting coefficients to obtain the high-level feature map;
[0096] The high-level feature map is input into the classification head network, and the defect type confidence score is output.
[0097] The feature response regions with confidence levels higher than a threshold are mapped back to the original image space to generate the texture map.
[0098] It should be noted that the process of extracting texture maps based on defect features from brightness distribution data using convolutional neural networks in this invention is implemented based on deep learning technology. A convolutional neural network is a deep learning model specifically designed for processing image data, capable of automatically extracting features from images. In this process, a sliding window is first used to divide the brightness distribution data into multiple detection sub-regions. Then, multi-scale feature extraction is performed in parallel for each sub-region to generate a primary feature map. The primary feature maps of different scales are fused through a residual connection module to form an intermediate feature map. An attention mechanism is applied to the intermediate feature map to calculate the weight coefficients of each channel, and the intermediate feature map is reconstructed based on these weight coefficients to obtain a high-level feature map. Finally, the high-level feature map is input into a classification head network, which outputs a defect type confidence score. Feature response regions with confidence scores higher than a threshold are mapped back to the original image space to generate a texture map. This process not only improves the accuracy of defect detection but also enables the identification of complex defect types.
[0099] Specifically, a convolutional neural network (CNN) is a deep learning architecture containing multiple convolutional layers, pooling layers, and fully connected layers for automatically extracting image features. A sliding window is a method that divides an image into multiple small blocks, each called a detection sub-region, allowing for local feature extraction. Multi-scale feature extraction refers to extracting image features at different scales to capture defect features of varying sizes. Primary feature maps are extracted through convolutional layers and contain the image's basic features. Residual connection modules are a network structure used to address the vanishing gradient problem in deep networks by directly adding feature maps from different levels through skip connections. Intermediate feature maps are obtained by fusing primary feature maps at different scales and contain richer feature information. Attention mechanisms are models used to calculate the importance of each channel in the feature map, thus highlighting important features. High-level feature maps are obtained by weighted reconstruction of intermediate feature maps and contain higher-level feature information. The classification head network is the final part of the CNN, used to classify the high-level feature maps into different defect types and output confidence scores. Texture maps are obtained by mapping feature response regions with confidence levels above a threshold back to the original image space, and they visually represent the location and features of defects.
[0100] Preferably, the convolutional neural network can be constructed using popular deep learning frameworks such as TensorFlow or PyTorch. The network architecture can include multiple convolutional and pooling layers; for example, four convolutional layers can be set, each followed by a max-pooling layer. The kernel size of the convolutional layers can be set to 3×3 with a stride of 1 to capture local features. The pooling window size of the pooling layers can be set to 2×2 with a stride of 2 to reduce the spatial dimensionality of the feature maps. In multi-scale feature extraction, different kernel sizes and strides can be set; for example, one branch can use a 3×3 kernel with a stride of 1, while another branch can use a 5×5 kernel with a stride of 2. The residual connection module can add feature maps from different levels through skip connections to enhance feature propagation. The attention mechanism can be implemented by calculating the weight coefficients of each channel in the feature map; for example, the channel weights can be normalized using the Softmax function. The classification head network can include three fully connected layers. The first layer compresses the feature dimension to 1 / 4 of the original size, the second layer performs L2 regularization on the features, and the third layer outputs the same classification dimension as the preset number of defect types. Finally, a Softmax classifier outputs the confidence score of the defect type, and the feature response regions with confidence scores higher than the threshold are mapped back to the original image space to generate a texture map.
[0101] In some embodiments, it is determined whether a normal pixel unit is contained within the neighborhood of the geometric center coordinates, and then: if so, the defective region is marked as a repairable defect;
[0102] Record the geometric parameters, occurrence frame number, and duration of the repairable defect, and generate a repair priority list;
[0103] Synchronize the repair priority list to the external maintenance management system.
[0104] It should be noted that, after determining whether the neighborhood of the geometric center coordinates contains normal pixel units, this invention marks the defective area as a repairable defect if normal pixel units are found. This process distinguishes between the severity and repairability of defects, thus providing guidance for subsequent repair work. Recording the geometric parameters, occurrence frame number, and duration of repairable defects generates a repair priority list, allowing maintenance personnel to arrange the repair sequence based on the severity and duration of the defects. Synchronizing the repair priority list to an external maintenance management system automates and efficiently manages defect repair work, ensuring that defects are repaired promptly.
[0105] Specifically, the neighborhood of the geometric center coordinates refers to the set of pixel units within a certain area centered on the geometric center coordinates. This range can be set according to actual needs, for example, it can be a square area or a circular area centered on the geometric center coordinates. Normal pixel units refer to pixel units whose brightness and chromaticity are within the normal range within the neighborhood. The presence of these pixel units indicates that the defective area may only be partially damaged, and therefore can be marked as a repairable defect. Geometric parameters include information such as the size and shape of the defective area. These parameters help maintenance personnel understand the specific situation of the defect. The occurrence frame number refers to the frame sequence number of the first occurrence of the defect, and the duration refers to the duration from the first occurrence of the defect to the present. This information helps maintenance personnel assess the severity of the defect. The repair priority list is a list containing information on all repairable defects. It sorts defects according to their severity and duration so that maintenance personnel can repair them according to priority. The external maintenance management system is a software system used to manage and schedule maintenance work. It can receive the repair priority list and arrange maintenance tasks according to the list.
[0106] Preferably, to more accurately determine whether a normal pixel unit is contained within the neighborhood, a threshold range for brightness and chromaticity can be set. A pixel unit is considered a normal pixel unit only when both its brightness and chromaticity values fall within this threshold range. For example, the brightness threshold can be set to 90% to 110% of normal brightness, and the chromaticity threshold can be set to ±10% of normal chromaticity. When recording the geometric parameters of repairable defects, image processing algorithms can be used to accurately measure the size and shape of the defect area, such as using contour detection algorithms to determine the defect boundaries. Frame numbers and durations can be recorded using the timestamp function of the image acquisition device. Each image acquisition is accompanied by a timestamp, and the duration of the defect can be calculated by comparing the timestamps. The repair priority list can be sorted according to factors such as defect size and duration. For example, it can be sorted from largest to smallest defect area, and for cases with the same area, it can be sorted from longest to shortest duration. When synchronizing the repair priority list to an external maintenance management system, the list data can be sent to the maintenance management system via a network interface, such as an API interface. After receiving the data, the maintenance management system can automatically generate maintenance tasks and assign them to the corresponding maintenance personnel.
[0107] In some embodiments, the multi-scale feature extraction includes: inputting the detection sub-region into a parallel processing branch; the first branch performs dilated convolution using a 3×3 convolution kernel with a dilation rate of 2 to generate a first feature map containing long-distance correlation features; the second branch performs downsampling convolution using a 5×5 convolution kernel with a stride of 2 to generate a second feature map containing local detail features; the third branch compresses the feature dimension through a max pooling layer followed by a 1×1 convolution kernel to perform channel dimensionality reduction to generate a third feature map containing global semantic features; and the first feature map, the second feature map, and the third feature map are concatenated along the channel dimension to form the primary feature map.
[0108] It should be noted that the multi-scale feature extraction process mentioned in this invention is a key step in convolutional neural networks for extracting image features. By inputting the detection sub-region into a parallel processing branch, feature maps containing different features can be extracted simultaneously, thereby capturing the features of defects more comprehensively. The first branch uses a 3×3 convolutional kernel with a dilation rate of 2 to perform dilated convolution, which can capture long-distance correlation features; the second branch uses a 5×5 convolutional kernel with a stride of 2 to perform downsampling convolution, which can extract local detail features; the third branch compresses the feature dimension through a max pooling layer followed by a 1×1 convolutional kernel to perform channel dimensionality reduction, which can obtain global semantic features. These three feature maps are concatenated along the channel dimension to form a primary feature map, providing rich feature information for subsequent feature fusion and defect recognition.
[0109] Specifically, multi-scale feature extraction refers to extracting image features at different scales to better capture various features in the image. The dilated convolution in the first branch is a special convolution operation that expands the receptive field of the convolution kernel by introducing holes, thus capturing features over a wider range. A dilation rate of 2 means that elements in the convolution kernel are sampled once every other pixel in the input image. The downsampling convolution in the second branch reduces the spatial dimensionality of the feature map using a larger kernel and stride, while extracting local detail features. A stride of 2 means that the convolution kernel moves two pixels at a time in the input image. The max-pooling layer in the third branch is used to reduce the spatial dimensionality of the feature map, reducing computation while preserving the most important features. A 1×1 convolution kernel is used for channel dimensionality reduction, i.e., reducing the number of channels in the feature map while integrating features. Concatenating the three feature maps along the channel dimension means merging them into a larger feature map so that subsequent network layers can process these features at different scales simultaneously.
[0110] Preferably, to achieve more efficient multi-scale feature extraction, the architecture of the convolutional neural network can be optimized. For example, in the first branch, the dilation rate of the 3×3 convolutional kernel can be set to 2, meaning that the kernel samples every other pixel in the input image, thereby expanding the receptive field of the kernel and enabling it to capture long-range correlated features over a larger area. In the second branch, the stride of the 5×5 convolutional kernel can be set to 2, which helps extract local detail features while reducing the spatial dimension of the feature map. In the third branch, a max-pooling layer can be used to reduce the spatial dimension of the feature map, followed by a 1×1 convolutional kernel to reduce the number of channels in the feature map, thereby reducing computational cost and integrating features. When concatenating feature maps, depth concatenation can be used to merge feature maps from different branches along the channel dimension to form a primary feature map containing rich feature information. Furthermore, to further improve the feature extraction effect, a batch normalization layer can be added after each convolutional layer to stabilize the training process and improve the model's generalization ability.
[0111] In some embodiments, the spatial coordinate transformation includes: inputting the boundary coordinates of the defect area captured by each image acquisition device into a perspective projection matrix, the perspective projection matrix being pre-calibrated using a checkerboard calibration plate during the installation phase; establishing a mapping relationship between the original pixel coordinates and the coordinates after radial distortion correction, wherein the radial distortion correction uses bilinear interpolation to compensate for the distortion coefficients; and transforming the corrected two-dimensional coordinates in the coordinate system of each image acquisition device to the screen coordinate system using a homogeneous coordinate transformation matrix to generate a set of projection boundary coordinates corresponding to the viewpoint of each device.
[0112] It should be noted that the spatial coordinate transformation mentioned in this invention is the process of converting the boundary coordinates of the defect area captured by each image acquisition device to a unified screen coordinate system. This process involves the use of a perspective projection matrix, which is pre-calibrated using a checkerboard calibration plate during the installation phase. The perspective projection matrix is used to convert the two-dimensional coordinates captured by the image acquisition device into two-dimensional coordinates in the screen coordinate system. During the conversion process, radial distortion correction is required for the original pixel coordinates to compensate for coordinate deviations caused by lens distortion. Radial distortion correction uses bilinear interpolation to compensate for the distortion coefficients. Finally, the corrected coordinates are transformed to the screen coordinate system using a homogeneous coordinate transformation matrix, generating a set of projected boundary coordinates corresponding to the viewpoints of each device. This process ensures that image data from different viewpoints can be compared and analyzed in a unified coordinate system, thereby improving the accuracy of defect localization.
[0113] Specifically, spatial coordinate transformation refers to the process of converting two-dimensional coordinates captured by different image acquisition devices into a unified screen coordinate system. The perspective projection matrix is a mathematical model used to describe the geometric relationship between the image captured by the image acquisition device and the actual scene. A checkerboard calibration board is a commonly used calibration tool used to determine the intrinsic and extrinsic parameters of the image acquisition device during the installation phase. Radial distortion correction refers to the process of compensating for distortion in the image captured by the image acquisition device; distortion is caused by image deformation due to the non-ideal optical characteristics of the lens. Bilinear interpolation is a commonly used interpolation method used to estimate pixel values in image processing. A homogeneous coordinate transformation matrix is a mathematical tool used to transform two-dimensional coordinates into another coordinate system. Through these steps, image data from different perspectives can be unified into the screen coordinate system, thereby achieving accurate defect localization.
[0114] Preferably, to achieve more accurate spatial coordinate transformation, the following steps can be adopted: First, each image acquisition device is calibrated using a checkerboard calibration board, and the perspective projection matrix and radial distortion parameters are obtained through the calibration process. During the calibration process, the checkerboard calibration board is placed on the display screen, the image acquisition device captures the image of the calibration board, and then the perspective projection matrix and radial distortion parameters are calculated using computer vision algorithms. Second, for the boundary coordinates of the defect area captured by each image acquisition device, the perspective projection matrix is used to transform them into the screen coordinate system. During the transformation process, radial distortion correction is performed on the original pixel coordinates, and bilinear interpolation is used to compensate for the distortion coefficients to reduce the impact of distortion on the coordinates. Finally, the corrected coordinates are transformed into the screen coordinate system using a homogeneous coordinate transformation matrix to generate a set of projected boundary coordinates corresponding to the viewpoint of each device. To improve the accuracy of the transformation, multiple checkerboard calibration boards in different positions can be used during the calibration process to obtain more accurate perspective projection matrices and radial distortion parameters. In addition, high-precision image acquisition devices and calibration tools can be used to reduce errors.
[0115] In some embodiments, the classification head network includes: inputting a high-level feature map into an encoder consisting of three fully connected layers; the first fully connected layer compresses the feature dimension to 1 / 4 of its original size and outputs an embedding vector; the second fully connected layer applies L2 regularization constraints to the embedding vector and outputs latent features; the third fully connected layer maps the latent features to a classification dimension equal to the number of preset defect types; inputting the classification dimension vector into a Softmax classifier with a temperature coefficient, the temperature coefficient being dynamically adjusted through adversarial training to enhance noise robustness; and finally, calibrating the classification results with confidence through a noise robustness discrimination module and outputting the final defect type confidence.
[0116] It should be noted that the classification head network mentioned in this invention is the final part of the convolutional neural network, used to classify high-level feature maps into different defect types and output defect type confidence scores. The classification head network includes an encoder consisting of three fully connected layers. The first fully connected layer compresses the feature dimension to 1 / 4 of its original size and outputs an embedding vector. The second fully connected layer applies L2 regularization constraints to the embedding vector and outputs latent features. The third fully connected layer maps the latent features to a classification dimension equal to the preset number of defect types. Finally, the classification dimension vector is input into a Softmax classifier with a temperature coefficient, which is dynamically adjusted through adversarial training to enhance noise robustness. A noise robustness discrimination module calibrates the classification results to output the final defect type confidence score. This process ensures the accuracy and reliability of defect classification, maintaining high classification performance even in the presence of noise.
[0117] Specifically, the classification head network is the classification module in a convolutional neural network. It receives a high-level feature map as input and outputs a confidence score for each defect type. A fully connected layer is a neural network layer where each neuron is connected to all neurons in the previous layer, transforming input features into output features. The first fully connected layer compresses the feature dimension of the high-level feature map to one-quarter of its original size to reduce computation and extract a more compact feature representation; the output embedding vector is the compressed feature representation. The second fully connected layer applies L2 regularization constraints to the embedding vector. This is a technique to prevent overfitting by limiting the size of the weights to improve the model's generalization ability; the output latent features are the regularized features. The third fully connected layer maps the latent features to a classification dimension equal to the number of preset defect types. This means each neuron corresponds to one defect type, and the output classification dimension vector represents the score for each defect type.
[0118] The convolutional neural network is trained on a large dataset of labeled OLED display images. The dataset sources include, but are not limited to, images of normal displays and displays containing various defect types collected from actual production lines. These image data undergo preprocessing, such as normalization, cropping, and rotation, to enhance data diversity and the model's generalization ability. A supervised learning method is used for training, optimizing network parameters by minimizing the loss function between the predicted defect type and the actual labeled defect type. During training, data augmentation techniques, such as random cropping, rotation, and flipping, are used to increase the model's adaptability to different image transformations. Early stopping is employed to prevent overfitting; training stops when the loss on the validation set no longer decreases. Furthermore, a learning rate decay strategy can be used, gradually decreasing the learning rate as the number of training epochs increases to ensure the model can finely adjust parameters in later training stages.
[0119] The Softmax classifier is a commonly used classification algorithm that transforms a classification dimension vector into a probability distribution, where each element represents the confidence level for the corresponding defect type. The temperature coefficient is a hyperparameter used to adjust the sharpness of the probability distribution output by Softmax. Dynamically adjusting the temperature coefficient through adversarial training can enhance the model's robustness to noise. The temperature coefficient is dynamically determined based on the loss function value during adversarial training. During training, the initial temperature coefficient can be set to 1.0. As training progresses, if the model's classification accuracy on noisy data decreases, the temperature coefficient can be increased appropriately to smooth the probability distribution and improve the model's tolerance to noise. Conversely, if the model's classification accuracy on clean data is high and its robustness to noise is sufficient, the temperature coefficient can be decreased appropriately to improve classification accuracy. The specific range of the temperature coefficient can be adjusted based on the actual performance during training, and it is generally recommended to choose between 0.5 and 2.0. For example, when the model's classification accuracy on noisy data is below a set threshold (e.g., 90%), the temperature coefficient can be increased by 0.1; when the model's classification accuracy on clean data reaches a high level (e.g., above 98%), the temperature coefficient can be decreased by 0.1.
[0120] Preferably, to construct a more efficient classification head network, the following steps can be taken: First, input the high-level feature map into the first fully connected layer, setting the output dimension of this layer to 1 / 4 of the dimension of the high-level feature map. For example, if the dimension of the high-level feature map is 1024, the output dimension of the first fully connected layer can be set to 256. Next, apply L2 regularization constraints to the embedding vector in the second fully connected layer. The regularization parameter can be adjusted according to the validation set performance during training; for example, the regularization parameter can be set to 0.01. Then, input the regularized latent features into the third fully connected layer. The output dimension of this layer should be the same as the preset number of defect types; for example, if there are 5 defect types, the output dimension of the third fully connected layer is 5. Finally, input the classification dimension vector into a Softmax classifier with a temperature coefficient. The temperature coefficient can be dynamically adjusted through adversarial training; for example, the initial temperature coefficient can be set to 1.0 and dynamically adjusted during training according to the loss function of adversarial training. To further improve classification accuracy, data augmentation techniques, such as random pruning, rotation, and flipping, can be used during training to increase the model's generalization ability. In addition, cross-validation can be used to evaluate the model's performance, and the network structure and hyperparameters can be adjusted based on the validation results.
[0121] The above embodiments of the present invention have the following beneficial effects:
[0122] 1. By combining the spatial difference calculation of brightness distribution data and standard template with the correlation analysis of chromaticity distribution data, defect areas of OLED displays can be accurately identified, effectively solving the problem of missed detection caused by insufficient correlation between brightness anomalies and chromaticity shifts in traditional detection methods, and improving the comprehensiveness and reliability of defect detection.
[0123] 2. By employing a multi-view image acquisition device and spatial coordinate transformation technology, the geometric center coordinates of the defect area are accurately calculated. The neighboring pixel judgment mechanism distinguishes between repairable and unrepairable defects, solving the problem of insufficient defect positioning accuracy in traditional detection and providing accurate spatial reference for subsequent repair or quality assessment.
[0124] 3. Based on a multi-scale feature extraction and classification head network of convolutional neural networks, combined with attention mechanism and noise robustness optimization, it can adaptively extract defect texture features and accurately match preset defect types, effectively overcoming the misjudgment problem caused by small defects or noise interference in complex backgrounds, and improving the accuracy and stability of defect classification.
[0125] Furthermore, the storage medium in the embodiments of this application stores program instructions capable of implementing all the above methods. These program instructions can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0126] The above description is merely a selection of preferred embodiments of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention as described in the embodiments is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A method for detecting defects in an OLED display screen based on machine learning, wherein the OLED display screen includes a plurality of pixel units located on the same detection plane, an image acquisition device is disposed on the detection plane, and the image acquisition device is used to acquire brightness distribution data and chromaticity distribution data when the OLED display screen displays an image, characterized in that, The OLED display screen defect detection method includes: In response to the luminance distribution data, acquire the chromaticity distribution data associated with the luminance distribution data; The degree of abnormality in the defect area is calculated by the spatial difference between the brightness distribution data and the standard brightness template. The geometric center coordinates of the defect area are determined based on the image acquisition device. Determine whether the neighborhood of the geometric center coordinates contains normal pixel units; If not, then a texture map based on defect features is extracted from the brightness distribution data using a convolutional neural network; Analyze whether the texture map matches a preset defect type; If so, defect location information and quality warning instructions are generated and sent to the external testing terminal.
2. The OLED display screen defect detection method based on machine learning according to claim 1, characterized in that, In response to the brightness distribution data, acquire chromaticity distribution data associated with the brightness distribution data, including: real-time acquisition of frame sequence images of the OLED display screen; The brightness quantization value and color quantization value of each pixel unit are obtained by the image acquisition device based on a preset sampling frequency. A three-dimensional feature space is constructed with pixel coordinates as the horizontal axis and brightness and chromaticity values as the vertical axis. The luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface; Determine whether the gradient between the brightness value of the current pixel and the brightness values of its neighboring pixels exceeds a preset tolerance; If so, it is determined to be a potential defect area; Extract the continuous region with the largest gradient change in the brightness surface and mark it as candidate defect data; Within the pixel range corresponding to the candidate defect data, obtain the chromaticity distribution data with the most significant chromaticity change trend.
3. The OLED display screen defect detection method based on machine learning according to claim 2, characterized in that, The luminance quantization value and chrominance quantization value of each pixel unit are mapped in the three-dimensional feature space to form a luminance surface and a chrominance surface. Then, the three-dimensional feature space, the luminance surface and the chrominance surface are transmitted to an external visualization analysis platform.
4. The OLED display screen defect detection method based on machine learning according to claim 1, characterized in that, Determining the geometric center coordinates of the defect region based on the image acquisition device includes: setting up at least two image acquisition devices with different perspectives on the detection plane; A screen coordinate system is established with the optical axis of the first image acquisition device as the reference axis; Mark the installation locations of other image acquisition devices in the screen coordinate system; Calculate the boundary coordinates of the defect area captured by each image acquisition device; The spatial coordinate transformation is performed on the capture results of each image acquisition device to generate multiple sets of projection boundaries; The center point of the overlapping region of multiple sets of projected boundaries is the coordinate of the geometric center.
5. The OLED display screen defect detection method based on machine learning according to claim 4, characterized in that, The center point of the overlapping area of multiple sets of projection boundaries is the geometric center coordinate. Then, the screen coordinate system, the positions of each image acquisition device, and the geometric center coordinate are superimposed and displayed on the external three-dimensional reconstruction interface.
6. The OLED display screen defect detection method based on machine learning according to claim 1, characterized in that, Extracting texture maps based on defect features from the brightness distribution data using a convolutional neural network includes: dividing the brightness distribution data into multiple detection sub-regions using a sliding window; Multi-scale feature extraction is performed in parallel on each detection sub-region to generate a primary feature map; By fusing primary feature maps of different scales through the residual connection module, an intermediate feature map is formed. An attention mechanism is applied to the intermediate feature map to calculate the weight coefficients of each channel; The intermediate feature map is reconstructed by weighting coefficients to obtain the high-level feature map; The high-level feature map is input into the classification head network, and the defect type confidence score is output. The feature response regions with confidence levels higher than a threshold are mapped back to the original image space to generate the texture map.
7. The OLED display screen defect detection method based on machine learning according to claim 1, characterized in that, Determine whether the neighborhood of the geometric center coordinates contains normal pixel units, and then: if so, mark the defective region as a repairable defect; Record the geometric parameters, occurrence frame number, and duration of the repairable defect, and generate a repair priority list; Synchronize the repair priority list to the external maintenance management system.
8. The OLED display screen defect detection method based on machine learning according to claim 6, characterized in that, The multi-scale feature extraction includes: The detection sub-region is input into the parallel processing branch. The first branch performs dilated convolution using a 3×3 convolution kernel with a dilation rate of 2 to generate a first feature map containing long-distance correlation features. The second branch uses a 5×5 convolution kernel with a stride of 2 to perform downsampling convolution, generating a second feature map containing local detailed features; The third branch compresses the feature dimension through a max pooling layer and then performs channel dimensionality reduction using a 1×1 convolutional kernel to generate a third feature map containing global semantic features. The first feature map, the second feature map, and the third feature map are spliced together along the channel dimension to form the primary feature map.
9. The OLED display screen defect detection method based on machine learning according to claim 4, characterized in that, The spatial coordinate transformation includes: The boundary coordinates of the defect area captured by each image acquisition device are input into the perspective projection matrix, which is pre-calibrated during the installation phase using a checkerboard calibration plate. A mapping relationship is established between the original pixel coordinates and the coordinates after radial distortion correction, wherein the radial distortion correction uses bilinear interpolation to compensate for the distortion coefficients; The two-dimensional coordinates of each image acquisition device in the corrected coordinate system are transformed to the screen coordinate system through a homogeneous coordinate transformation matrix to generate a set of projection boundary coordinates corresponding to the viewpoint of each device.
10. The OLED display screen defect detection method based on machine learning according to claim 6, characterized in that, The classification head network includes: The high-level feature map is input into an encoder consisting of three fully connected layers. The first fully connected layer compresses the feature dimension to 1 / 4 of the original size and outputs the embedding vector. The second fully connected layer applies L2 regularization constraints to the embedding vector and outputs latent features; The third fully connected layer maps the latent features to a classification dimension that is the same as the number of preset defect types. The classification dimension vector is input into a Softmax classifier with a temperature coefficient, which is dynamically adjusted through adversarial training to enhance noise robustness. Finally, the classification results are calibrated using the noise robustness discrimination module, and the final defect type confidence score is output.
Citation Information
Cited By
Optical flux sheet defect detection method and system based on coordinate transformation
CN121236067A
Method, system and device for detecting quality of bright spots of cinema LED screen by using unmanned aerial vehicle, and storage medium
CN121384414A