Compensation method and device of display panel, computer equipment, readable storage medium and program product

By combining a classification model and a compensation strategy agent, the compensation action for the display panel is automatically selected, solving the accuracy and efficiency problems of color shift defects in the display panel and achieving fast and accurate display panel compensation.

CN121708874APending Publication Date: 2026-03-20HYC (CHENGDU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511791766.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, display panels may exhibit color shift defects when displaying images, and the compensation algorithm requires multiple adjustments based on human experience, resulting in poor flexibility, low efficiency, and an inability to guarantee accuracy and reliability.

Method used

By acquiring image data from the display panel, a classification model is used to output color shift state features, which are then input into the compensation strategy agent to output the target compensation action. The working parameters of the compensation strategy agent are adjusted based on the adjusted image data and execution time to achieve automated and accurate compensation.

Benefits of technology

It achieves fast and accurate display panel compensation, reduces labor costs and time resources, dynamically selects the optimal compensation action, and improves the reliability and accuracy of the compensation strategy agent, making it suitable for various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708874A_ABST
    Figure CN121708874A_ABST
Patent Text Reader

Abstract

The invention relates to a compensation method and device of a display panel, computer equipment, a readable storage medium and a program product. The method comprises the steps of obtaining first image data of a display panel, inputting the first image data into a classification model, and outputting color cast state features corresponding to the display panel through the classification model; inputting the color cast state characteristics into a compensation strategy agent, and outputting a target compensation action corresponding to the display panel through the compensation strategy agent; and adjusting display parameters of the display panel according to the target compensation action to obtain an adjusted display panel, and adjusting working parameters of the compensation strategy intelligent agent according to second image data of the adjusted display panel and execution duration of executing the target compensation action. By adopting the method, the accuracy and the efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for compensating a display panel. Background Technology

[0002] During the manufacturing process of display panels, due to uneven vapor deposition, thin film stress, or other abnormalities in the panel manufacturing process or driving circuit, color deviation defects such as bluish or pinkish tints may appear when displaying images. Typically, these defects can be corrected using compensation algorithms to achieve a more uniform display.

[0003] In related technologies, when using compensation algorithms for compensation, the parameters of the compensation algorithm need to be determined after multiple adjustments based on human experience. This results in poor flexibility and low efficiency, and cannot guarantee the accuracy and reliability of the compensation. Summary of the Invention

[0004] Therefore, it is necessary to provide a compensation method, apparatus, computer device, computer-readable storage medium, and computer program product for a display panel that can improve accuracy and efficiency in response to the above-mentioned technical problems.

[0005] In a first aspect, this application provides a compensation method for a display panel, comprising:

[0006] First image data of the display panel is acquired and input into a classification model. The classification model outputs the color shift state features corresponding to the display panel. The classification model is trained based on the correspondence between image data samples and color shift state feature labels.

[0007] The color shift state features are input into the compensation strategy agent, and the compensation strategy agent outputs the target compensation action corresponding to the display panel.

[0008] The display parameters of the display panel are adjusted according to the target compensation action to obtain the adjusted display panel, and the working parameters of the compensation strategy agent are adjusted according to the second image data of the adjusted display panel and the execution time of the target compensation action.

[0009] In one embodiment, adjusting the working parameters of the compensation strategy agent based on the adjusted second image data of the display panel and the execution duration of the target compensation action includes:

[0010] Acquire the second image data of the adjusted display panel and the execution time of the target compensation action;

[0011] Based on the first image data and the second image data, determine the reward value corresponding to the target compensation action; based on the execution duration, determine the penalty value corresponding to the target compensation action.

[0012] A target reward value is determined based on the reward value and the penalty value, and the working parameters of the compensation strategy agent are adjusted based on the target reward value.

[0013] In one embodiment, determining the reward value corresponding to the target compensation action based on the first image data and the second image data includes:

[0014] Based on the second image data, determine the brightness uniformity and color uniformity of the display panel;

[0015] Determine the correlation between the first image data and the second image data;

[0016] The reward value corresponding to the target compensation action is determined based on the brightness uniformity, the chromaticity uniformity, and the correlation.

[0017] In one embodiment, determining the reward value corresponding to the target compensation action based on the brightness uniformity, the chromaticity uniformity, and the correlation includes:

[0018] Based on the first weight corresponding to the brightness dimension, the second weight corresponding to the chromaticity dimension, and the third weight corresponding to the correlation dimension, the brightness uniformity, the chromaticity uniformity, and the correlation are weighted and summed to obtain the reward value corresponding to the target compensation action.

[0019] In one embodiment, the method for determining the classification model includes:

[0020] Obtain an image data sample set, wherein the image data sample set includes an image data sample set labeled with color deviation state feature tags;

[0021] Construct an initial classification model, wherein training parameters are set in the initial classification model;

[0022] The image data sample set is input into the initial classification model, and the initial classification model outputs the classification result.

[0023] Based on the difference between the classification result and the corresponding color deviation state feature label, the training parameters are adjusted until the adjusted difference meets the preset requirements, thus obtaining the classification model.

[0024] In one embodiment, the compensation policy agent includes a policy network and a value network, and the method further includes:

[0025] The color cast state features, the target compensation action, and the target reward value are input into the value network, and the value network outputs a predicted reward value. The predicted reward value is used to characterize the reward value accumulated within a preset time period after the target compensation action is performed under the color cast state features.

[0026] Based on the predicted reward value, the parameters of the policy network are adjusted to obtain the adjusted compensation policy agent.

[0027] Secondly, this application also provides a compensation device for a display panel, comprising:

[0028] The acquisition module is used to acquire the first image data of the display panel and input the first image data into the classification model. The classification model outputs the color deviation state features corresponding to the display panel. The classification model is trained based on the correspondence between image data samples and color deviation state feature labels.

[0029] The output module is used to input the color shift state features into the compensation strategy agent, and the compensation strategy agent outputs the target compensation action corresponding to the display panel.

[0030] The adjustment module is used to adjust the display parameters of the display panel according to the target compensation action to obtain the adjusted display panel, and to adjust the working parameters of the compensation strategy agent according to the second image data of the adjusted display panel and the execution time of the target compensation action.

[0031] In one embodiment, the adjustment module is further configured to:

[0032] Acquire the second image data of the adjusted display panel and the execution time of the target compensation action;

[0033] Based on the first image data and the second image data, determine the reward value corresponding to the target compensation action; based on the execution duration, determine the penalty value corresponding to the target compensation action.

[0034] A target reward value is determined based on the reward value and the penalty value, and the working parameters of the compensation strategy agent are adjusted based on the target reward value.

[0035] In one embodiment, the adjustment module is further configured to:

[0036] Based on the second image data, determine the brightness uniformity and color uniformity of the display panel;

[0037] Determine the correlation between the first image data and the second image data;

[0038] The reward value corresponding to the target compensation action is determined based on the brightness uniformity, the chromaticity uniformity, and the correlation.

[0039] In one embodiment, the adjustment module is further configured to:

[0040] Based on the first weight corresponding to the brightness dimension, the second weight corresponding to the chromaticity dimension, and the third weight corresponding to the correlation dimension, the brightness uniformity, the chromaticity uniformity, and the correlation are weighted and summed to obtain the reward value corresponding to the target compensation action.

[0041] In one embodiment, the apparatus further includes a module for determining the classification model, configured to:

[0042] Obtain an image data sample set, wherein the image data sample set includes an image data sample set labeled with color deviation state feature tags;

[0043] Construct an initial classification model, wherein training parameters are set in the initial classification model;

[0044] The image data sample set is input into the initial classification model, and the initial classification model outputs the classification result.

[0045] Based on the difference between the classification result and the corresponding color deviation state feature label, the training parameters are adjusted until the adjusted difference meets the preset requirements, thus obtaining the classification model.

[0046] In one embodiment, the compensation policy agent includes a policy network and a value network, and the device is further configured to:

[0047] The color cast state features, the target compensation action, and the target reward value are input into the value network, and the value network outputs a predicted reward value. The predicted reward value is used to characterize the reward value accumulated within a preset time period after the target compensation action is performed under the color cast state features.

[0048] Based on the predicted reward value, the parameters of the policy network are adjusted to obtain the adjusted compensation policy agent.

[0049] Thirdly, embodiments of this disclosure also provide a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in any one of the embodiments of this disclosure.

[0050] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of this disclosure.

[0051] Fifthly, embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in any one of the embodiments of this disclosure.

[0052] The aforementioned display panel compensation method, apparatus, computer equipment, computer-readable storage medium, and computer program product, when performing display panel compensation, acquire the first image data of the display panel, and output the color shift state features corresponding to the display panel using a classification model. These color shift state features are then input into a compensation strategy intelligent agent, which outputs a target compensation action. Based on this action, the display parameters of the display panel are adjusted to obtain the adjusted display panel. This allows for rapid and accurate compensation and correction of the display panel, ensuring its display effect. Compensation based on a classification model and a compensation strategy intelligent agent dynamically and automatically selects the optimal compensation action for different color shift problems, reducing labor costs and time resources, and achieving precise, efficient, and adaptive compensation for display panel defects. During the compensation adjustment process, the compensation strategy intelligent agent is iteratively optimized based on the image data of the display panel before and after adjustment, as well as the execution time of the action. This effectively improves the reliability and accuracy of the compensation actions output by the intelligent agent while ensuring execution efficiency, making it suitable for more application scenarios. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a flowchart illustrating a compensation method for a display panel in one embodiment;

[0055] Figure 2 This is a flowchart illustrating how the compensation strategy agent is determined in one embodiment;

[0056] Figure 3 This is a flowchart illustrating the compensation method for the display panel in another embodiment. Figure 1 ;

[0057] Figure 4This is a flowchart illustrating the compensation method for the display panel in another embodiment. Figure 2 ;

[0058] Figure 5 This is a structural block diagram of the compensation device for the display panel in one embodiment;

[0059] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] In related technologies, color shift in display panels (such as pinking, bluish tint, gradient color shift, periodic color shift, etc.) is a common mura defect in display panels, mainly caused by panel manufacturing processes (such as uneven vapor deposition, thin film stress) or abnormal driving circuits. The current mainstream correction technology is Demura technology, the core of which is to compensate for differences in panel brightness and color through algorithms to make the displayed image more uniform.

[0062] In one embodiment, such as Figure 1 As shown, a compensation method for a display panel is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0063] Step S110: Obtain the first image data of the display panel and input the first image data into the classification model. The classification model outputs the color shift state features corresponding to the display panel. The classification model is trained based on the correspondence between image data samples and color shift state feature labels.

[0064] For example, first image data of the display panel is acquired. In some examples, the display panel can be lit up, driving it to display a multi-grayscale RGB test image. Display data of the display panel is acquired through a preset image acquisition device, and processed to obtain the first image data. The preset image acquisition device may include, but is not limited to, a high-precision CCD camera. In some examples, the image acquisition device is used to acquire brightness data at each grayscale level, and the acquired data is normalized and converted into an RGB pseudo-color image to obtain the first image data.

[0065] Optionally, the first image data is input into a classification model, which outputs color cast state features corresponding to the display panel. In some examples, the color cast state features can be used to characterize the probability distribution of the color cast type corresponding to the display panel. Color cast types can include, but are not limited to, pinkish, bluish, gradient color cast, and periodic color cast. In some examples, the classification model can be built based on the YOLOv11 network, incorporating wavelet convolution WTConv2d to enhance multi-scale feature extraction capabilities. In some possible implementations, the image data samples can be obtained by processing historical data.

[0066] In some possible implementations, when training the classification model, a color cast classification dataset (corresponding to a set of image data samples) is first constructed. This involves using a camera to photograph the surface of a display panel, obtaining pre-processed CSV data at different grayscale levels from multiple display panels, including R, G, and B channel data. The R, G, and B data are then normalized and synthesized into RGB color images (corresponding to image data samples), which are then manually divided into different color cast categories and labeled to construct the color cast classification dataset. The dataset is then randomly divided into training and test sets according to a certain ratio. The training and test sets are then fed into the initial classification model for training. In some examples, a pre-trained model on the ImageNet-1k dataset can be used. The initial learning rate is... The final learning rate is 0.01. The learning rate is 0.0001, the momentum value is 0.9, the number of training epochs is 1000, and a linear learning rate decay strategy is used, the formula of which is: Train for 1000 rounds until the test set accuracy is >98.5%.

[0067] Step S120: Input the color shift state features into the compensation strategy agent, and output the target compensation action corresponding to the display panel through the compensation strategy agent;

[0068] For example, color shift state features are input into the compensation policy agent, which then outputs the target compensation action corresponding to the display panel. In some examples, the compensation policy agent determines the target compensation action based on a policy library. The target compensation action may include, but is not limited to, compensation algorithms and algorithm parameters, which can be determined according to the actual application scenario. The policy library may include various compensation algorithms, including but not limited to gradient color shift correction algorithms based on edge-preserving filtering, periodic color shift suppression algorithms based on Fourier transform, and adaptive correction algorithms based on local LUTs. Different algorithms correspond to different types of algorithm parameters, which can be determined according to the actual application scenario.

[0069] Optionally, the compensation strategy agent can be constructed based on the DDPG algorithm, comprising four neural networks for action output, value evaluation, and agent optimization.

[0070] Step S130: Adjust the display parameters of the display panel according to the target compensation action to obtain the adjusted display panel, and adjust the working parameters of the compensation strategy agent according to the second image data of the adjusted display panel and the execution time of the target compensation action.

[0071] For example, the display parameters of the display panel are adjusted according to the target compensation action to obtain the adjusted display panel. In some examples, the display parameters may include, but are not limited to, parameters such as brightness, chromaticity, and contrast. In some examples, the compensation algorithm to be executed and the corresponding algorithm parameters can be determined based on the target compensation action, and the display parameters are adjusted based on the compensation algorithm to obtain the adjusted display panel.

[0072] Optionally, the adjusted second image data of the display panel is acquired, and the execution time of the target compensation action is recorded when the target compensation action is performed. In some examples, the second image data can be acquired in the same way as the first image data.

[0073] In some examples, the display effect of the adjusted display panel can be determined based on the first image data, and the fidelity after adjustment can be determined based on the difference between the first and second image data. Simultaneously, the efficiency of the compensation action is evaluated by combining the execution time of the target compensation action. The shorter the execution time and the better the effect, the higher the overall performance of the compensation action.

[0074] For example, during the adjustment process, the operating parameters of the classification model and the compensation strategy agent can be dynamically updated according to actual needs. For instance, when a new type of color bias is detected or the compensation effect is unsatisfactory, the adaptability and accuracy of the system can be improved by retraining the classification model or adjusting the neural network parameters of the compensation strategy agent.

[0075] In this embodiment, during display panel compensation, first image data of the display panel is acquired, and a classification model is used to output the color shift state features corresponding to the display panel. These color shift state features are then input into a compensation strategy agent, which outputs a target compensation action. Based on this action, the display parameters of the display panel are adjusted to obtain the adjusted display panel. This allows for rapid and accurate compensation and correction of the display panel, ensuring its display quality. Compensation based on a classification model and a compensation strategy agent dynamically and automatically selects the optimal compensation action for different color shift problems, reducing labor costs and time resources, and achieving precise, efficient, and adaptive compensation for display panel defects. During the compensation adjustment process, the compensation strategy agent is iteratively optimized based on the image data of the display panel before and after adjustment, as well as the execution time of the action. This effectively improves the reliability and accuracy of the compensation actions output by the agent while ensuring execution efficiency, making it suitable for a wider range of application scenarios.

[0076] In one embodiment, adjusting the working parameters of the compensation strategy agent based on the adjusted second image data of the display panel and the execution duration of the target compensation action includes:

[0077] Acquire the second image data of the adjusted display panel and the execution time of the target compensation action;

[0078] Based on the first image data and the second image data, determine the reward value corresponding to the target compensation action; based on the execution duration, determine the penalty value corresponding to the target compensation action.

[0079] A target reward value is determined based on the reward value and the penalty value, and the working parameters of the compensation strategy agent are adjusted based on the target reward value.

[0080] For example, the second image data of the adjusted display panel and the execution time required to perform the target compensation action are obtained. The display effect of the adjusted display panel can be determined based on the first image data. The display fidelity can also be determined based on the first and second image data. A reward value corresponding to the target compensation action is determined based on the display effect and display fidelity; the better the display effect and the higher the fidelity, the higher the reward value. A penalty value for the target compensation action is determined based on the execution time; the longer the execution time, the higher the penalty value.

[0081] Optionally, the reward value and penalty value are combined to calculate the target reward value. The target reward value is used to evaluate the overall performance of the target compensation action and serves as an important basis for adjusting the working parameters of the compensation strategy agent. In some examples, a higher reward value results in a higher target reward value, and a higher penalty value results in a lower target reward value.

[0082] For example, the operating parameters of the compensation strategy agent are adjusted based on the target reward value. In some possible implementations, the target reward value can be compared with a preset reward value. If the target reward value is lower than the preset reward value, it can be considered that the target reward value is low and the compensation effect is poor, and the operating parameters of the compensation strategy agent are adjusted. If the target reward value is higher than the preset reward value, it can be considered that the target reward value is high and the compensation effect is good, and there is no need to adjust the operating parameters of the compensation strategy agent. Alternatively, the policy network can be optimized and adjusted based on the target reward value and the value network in the compensation strategy agent.

[0083] This embodiment of the disclosure can effectively evaluate the actual effect of the target compensation action and dynamically adjust the working parameters of the compensation strategy agent by combining reward and penalty values. This not only improves the accuracy of the compensation action but also further optimizes the agent's learning ability, allowing it to gradually approach the optimal solution through multiple iterations. Simultaneously, by considering the execution time, the efficiency of the compensation process is ensured, avoiding the impact on overall performance due to excessive time consumption.

[0084] In one embodiment, determining the reward value corresponding to the target compensation action based on the first image data and the second image data includes:

[0085] Based on the second image data, determine the brightness uniformity and color uniformity of the display panel;

[0086] Determine the correlation between the first image data and the second image data;

[0087] The reward value corresponding to the target compensation action is determined based on the brightness uniformity, the chromaticity uniformity, and the correlation.

[0088] For example, the brightness uniformity and color uniformity of the display panel are calculated based on the second image data. In some examples, brightness uniformity can be determined by analyzing the brightness distribution in different areas of the second image data, while color uniformity can be obtained by evaluating the degree of color deviation in different areas. In some examples, the brightness uniformity and color uniformity indicators can reflect the overall visual effect and consistency of the display panel after compensation.

[0089] Optionally, the correlation between the first image data and the second image data can be determined. In some implementations, structural similarity algorithms, such as the SSIM algorithm, or other image quality assessment methods can be used to measure the similarity between the two sets of data. The higher the correlation, the higher the image fidelity.

[0090] In some possible implementations, the reward value can be obtained by directly adding the brightness uniformity, chromaticity uniformity, and correlation. Alternatively, the reward value for the target compensation action can be determined through a weighted calculation based on these three metrics, depending on the specific application scenario. In some examples, different weights can be assigned to each metric according to actual needs. For instance, considering that brightness uniformity and chromaticity uniformity have a greater impact on display quality, they can be given higher weights.

[0091] This disclosure embodiment comprehensively measures the effectiveness of target compensation actions by integrating brightness uniformity, color uniformity, and image data correlation. While focusing on the visual performance of the display panel after adjustment, it also incorporates data comparison before and after adjustment to ensure maximum actual benefit of the compensation actions. Quantitative analysis of these key indicators provides a scientific basis for evaluating target compensation actions, further enhancing the reliability and adaptability of the compensation strategy.

[0092] In one embodiment, determining the reward value corresponding to the target compensation action based on the brightness uniformity, the chromaticity uniformity, and the correlation includes:

[0093] Based on the first weight corresponding to the brightness dimension, the second weight corresponding to the chromaticity dimension, and the third weight corresponding to the correlation dimension, the brightness uniformity, the chromaticity uniformity, and the correlation are weighted and summed to obtain the reward value corresponding to the target compensation action.

[0094] For example, luminance uniformity, chromaticity uniformity, and correlation are weighted and summed based on a first weight corresponding to the luminance dimension, a second weight corresponding to the chromaticity dimension, and a third weight corresponding to the correlation dimension. In some examples, the first, second, and third weights can be dynamically adjusted according to the actual application scenario to ensure that the evaluation results can more accurately reflect the actual effect of the target compensation action. For example, in scenarios with high requirements for luminance consistency, the value of the first weight can be appropriately increased; while in scenarios with higher requirements for color accuracy, the proportion of the second weight can be increased.

[0095] Optionally, the reward value can be calculated using a weighted summation formula, which can be expressed as: Reward Value = First Weight × Brightness Uniformity + Second Weight × Chromaticity Uniformity + Third Weight × Correlation. In some possible implementations, the weight values ​​can be set between 0 and 1 to ensure the reasonableness and standardization of the calculation results.

[0096] For example, the calculated reward value is used as the evaluation basis for the target compensation action, and the working parameters of the compensation strategy agent are further optimized by combining other indicators such as execution time.

[0097] This embodiment of the disclosure, by introducing a weighting mechanism, allows for more flexible adjustment of the importance of different dimensions in reward value calculation, further improving the accuracy of the evaluation. This enables the compensation strategy to be personalized according to actual needs, increasing the flexibility of display panel compensation. Optimizing the weighting parameters effectively enhances the learning efficiency and adaptability of the compensation strategy agent, thereby achieving better compensation results in complex and ever-changing real-world environments.

[0098] In one embodiment, the method for determining the classification model includes:

[0099] Obtain an image data sample set, wherein the image data sample set includes an image data sample set labeled with color deviation state feature tags;

[0100] Construct an initial classification model, wherein training parameters are set in the initial classification model;

[0101] The image data sample set is input into the initial classification model, and the initial classification model outputs the classification result.

[0102] Based on the difference between the classification result and the corresponding color deviation state feature label, the training parameters are adjusted until the adjusted difference meets the preset requirements, thus obtaining the classification model.

[0103] For example, an image data sample set is obtained, wherein the set contains image data samples labeled with color cast state feature tags. In some examples, the image data samples may be derived from historical data or obtained through actual acquisition, and the image data samples may be obtained after preprocessing the initial data to ensure data quality. In some examples, the preprocessing steps may include, but are not limited to, noise removal, image enhancement, and format standardization.

[0104] Optionally, a pre-defined network architecture can be selected as the base framework when constructing the initial classification model. In some possible implementations, existing deep learning models can be improved, for example, by introducing a wavelet convolution WTConv2d module to enhance multi-scale feature extraction capabilities. In some examples, training parameters, including but not limited to learning rate, optimizer type, and loss function, are set in the initial classification model.

[0105] For example, after inputting a set of image data samples into an initial classification model, the model outputs the corresponding classification result. Since the initial classification model is untrained, the model's output may differ significantly from the actual result, i.e., the color cast state feature label. Therefore, iterative adjustments to the model's training parameters are necessary. In some examples, the loss value can be calculated by comparing the classification result with the true color cast state feature label of the image data samples. In other examples, the cross-entropy loss function or other suitable evaluation metrics can be chosen to quantify this difference.

[0106] Optionally, the training parameters in the initial classification model can be adjusted based on the calculated loss value. In some possible implementations, backpropagation combined with gradient descent can be used to update the parameters. Through multiple iterations of training, the difference between the classification result and the true label is gradually reduced until a preset requirement is met. This preset requirement may include, but is not limited to, conditions such as the loss value being below a certain threshold or the classification accuracy reaching a specific standard. During training, a validation set can also be introduced to evaluate model performance and prevent overfitting. In some examples, early stopping strategies or regularization methods can be used to further improve the model's generalization ability.

[0107] In this embodiment, a classification model is obtained by training a sample set, which improves the efficiency and accuracy of obtaining the classification model, and ensures the accuracy and generalization ability of the classification model, making it more adaptable to various display panels.

[0108] In one embodiment, the compensation policy agent includes a policy network and a value network, and the method further includes:

[0109] The color cast state features, the target compensation action, and the target reward value are input into the value network, and the value network outputs a predicted reward value. The predicted reward value is used to characterize the reward value accumulated within a preset time period after the target compensation action is performed under the color cast state features.

[0110] Based on the predicted reward value, the parameters of the policy network are adjusted to obtain the adjusted compensation policy agent.

[0111] For example, the color-skew state features, the target compensation action, and the target reward value are input into a value network, which then outputs a predicted reward value. In some examples, this predicted reward value is used to characterize the potential accumulated reward over a predetermined period after performing the target compensation action under specific color-skew state features.

[0112] Optionally, the parameters of the policy network can be adjusted based on the difference between the predicted reward value and the actual target reward value. In some implementations, the error between the two can be calculated, and the parameters of the policy network can be updated using an optimization algorithm (such as gradient descent).

[0113] For example, when adjusting the parameters of the policy network, the learning rate can be introduced as a control parameter to balance the magnitude of each update, avoiding model instability due to excessively large step sizes or slow convergence speed due to excessively small step sizes. Furthermore, other techniques, such as momentum optimization or adaptive learning rate methods, can be combined to further improve training efficiency and stability.

[0114] Figure 2 This is a flowchart illustrating a method for determining a compensation strategy agent according to an exemplary embodiment. (Refer to...) Figure 2 As shown, a historical dataset is constructed. For display panels with various known color cast types, such as pinkish, bluish, grayish, pink on top and bluish on the bottom, bluish on the left and pink on the right, multiple different Demura algorithms are systematically run, and the results are recorded. The dataset D consists of several empirical tuples, each tuple being... .in, (State) is the classification feature vector extracted from the image captured from the display panel at time t. (Action) refers to the Demura action actually performed by the system under the current state (such as using algorithm A to adjust parameter P1 by +5%). (Reward) is for performing the action. The immediate reward obtained is automatically calculated by the reward function, taking into account both the corrective effect and the computational cost. (Next state) is after the action has been completed. Then, the new state feature vector of the display panel is displayed.

[0115] The agent is initialized. The DDPG algorithm is used to initialize the neural network. The policy network (π) takes the state s as input and outputs a suggested action a. It consists of multiple fully connected layers. The input layer receives the color-biased classification feature vector from the feature extraction and classification module. After nonlinear transformation and feature combination by multiple fully connected layers, it outputs a suggested action a in the action space. The value network (Q) is used to evaluate the value of the action output by the policy network. It takes the state s and action a as input and outputs a Q value (corresponding to the predicted reward value) through a series of calculations and analyses. The Q value represents the estimated long-term cumulative reward that can be obtained by performing the action in the current state. The value network also consists of multiple fully connected layers. By processing the input state and action information and combining historical experience and reward feedback, it evaluates the value of the action.

[0116] For example, the policy network (π) selects actions based on the current environmental state, aiming to learn a policy that maximizes expected rewards. The value network evaluates the value of states or state-action pairs and provides a stable target value for calculating the loss function and updating the parameters of the current value network. The agent also includes a target policy network and a target value network. During training, the parameters of the target policy network and target value network are updated at a slower rate, providing a relatively stable reference for the training of the policy network and value network and avoiding drastic fluctuations during training.

[0117] In some examples, the reward function is ,in, For the target reward value, Brightness uniformity is measured by calculating the reciprocal of the variance of the brightness of the corrected display panel. The smaller the variance, the more uniform the brightness distribution, and the higher the brightness uniformity reward. Color uniformity is evaluated by calculating the reciprocal of the variance of ΔE in the chromaticity of the corrected display panel. The smaller the variance of ΔE, the more uniform the chromaticity distribution, and the higher the reward for color uniformity. SSIM (Structural Similarity, or Correlation) measures the similarity between the original image with Mura and the corrected image in high-frequency details. Its range is between 0 and 1; the closer it is to 1, the more similar the two images are, and the higher the fidelity reward. T is the processing time (or execution time). The longer the processing time, the higher the efficiency penalty, thus incentivizing the agent to learn to adopt a lighter and faster Demura strategy, improving the overall efficiency of the system. As the first weight, As the second weight, As the third weight, The fourth weight corresponds to the execution duration.

[0118] For example, during training, a minibatch of N experiences is randomly sampled from the historical dataset D. For each sampled experience, a more stable target Q-value is computed using the target network. The mean squared error loss (TD error) between the predicted and target values ​​is calculated, and this loss is minimized using gradient descent to update the parameters of the Critic's current network. The Actor's policy gradient is computed, and the Actor network is updated. Using a soft update policy, the parameters of the target network are slowly moved closer to the current network. This process is repeated until the Critic's loss function converges and the Actor's policy performance no longer significantly improves on an independent validation dataset.

[0119] From the offline training environment, the parameters and structure of the finally converged Actor network are extracted. Using the exact same algorithm flow as the training phase, feature vectors for excellent partial classification are computed in real time from the input RGB image and fed into the Actor network. The Actor network performs one forward propagation, outputting a decision action within milliseconds. After receiving the action instruction, the system passes it to the Demura execution module. This module, based on the instructions, calls the corresponding correction algorithm from the pre-built algorithm library or directly loads a new set of correction parameters into the display pipeline. Subsequently, the system generates or updates the Demura lookup table of the display panel and applies the correction data to all subsequent display screens.

[0120] In this embodiment, the predicted reward value output by the value network can effectively evaluate the long-term benefits of performing the target compensation action under the corresponding color shift state characteristics. Based on this prediction result, the parameters of the policy network are dynamically adjusted, enabling the compensation policy agent to optimize its decision-making ability through continuous learning. Simultaneously, the accurate prediction of the cumulative reward value further enhances the scientific rigor and rationality of the compensation action selection, providing a reliable guarantee for the accurate calibration of the display panel.

[0121] Figure 3 This is a flowchart illustrating a compensation method for a display panel according to an exemplary embodiment. (Refer to...) Figure 3 As shown, the Demura Off brightness map (corresponding to the first image data) before Demura is obtained, and features are extracted and classified to obtain the color cast category probability (corresponding to the color cast classification feature). The color cast category probability is input into the reinforcement learning agent (corresponding to the compensation policy agent). The reinforcement learning agent adopts an appropriate Demura policy (corresponding to the target compensation action) from the action space based on the policy network, and the program is written into the module to compensate the display panel, obtaining the Demura On brightness map (corresponding to the second image data). In some examples, the reward value and penalty value can be calculated based on the second image data and the execution time to obtain the target reward value. The value network of the reinforcement learning agent adjusts and optimizes the working parameters based on the target reward value.

[0122] Figure 4 This is a flowchart illustrating a compensation method for a display panel according to an exemplary embodiment. (Refer to...) Figure 4As shown, the display panel is illuminated, driving the display of a multi-grayscale RGB test image. Brightness data at each grayscale level is acquired using a high-precision CCD camera. The acquired brightness data is normalized and converted into an RGB pseudo-color image. Color cast features are extracted using a classification model trained with an improved YOLOv11 network, outputting a probability distribution of color cast types (corresponding to color cast classification features). The classification model employs an improved YOLOv11 classification network, introducing wavelet convolution WTConv2d to enhance multi-scale feature extraction capabilities. A reinforcement learning agent (corresponding to a compensation policy agent) is constructed based on the Deep Deterministic Policy Gradient (DDPG) algorithm. The optimal Demura policy (action output, corresponding to the target compensation action) is selected based on the color cast classification results (state input, corresponding to color cast classification features), with the objective of maximizing long-term cumulative reward. The Demura policy library pre-stores various targeted correction algorithms (corresponding to compensation algorithms), such as gradient color cast correction based on edge-preserving filtering, periodic color cast suppression based on Fourier transform, and adaptive correction based on local LUTs. The reward function evaluates the Demura results (corrected luminance data) output by the agent, providing an immediate reward value for agent training. The reward function consists of three parts: uniformity reward, fidelity reward, and efficiency penalty. The uniformity reward includes luminance uniformity and chromaticity uniformity; the fidelity reward prevents overcorrection by comparing the original Mura-treated image with the corrected image in high-frequency details to ensure the content itself is not distorted; the efficiency penalty is imposed by calculating the time spent on the agent's actions, encouraging the agent to learn a lighter and faster Demura strategy.

[0123] In this embodiment of the disclosure, images can be directly classified based on deep learning, and then a suitable post-processing algorithm can be automatically matched, saving a lot of manpower and time resources in pre-processing and post-processing; the classification network uses wavelet convolution WTConv, which effectively increases the receptive field of convolution with only a few parameters, effectively improving the model's feature representation ability.

[0124] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0125] Based on the same inventive concept, this application also provides a display panel compensation device for implementing the above-described display panel compensation method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more display panel compensation device embodiments provided below can be found in the limitations of the display panel compensation method described above, and will not be repeated here.

[0126] In one exemplary embodiment, such as Figure 5 As shown, a compensation device 500 for a display panel is provided, comprising:

[0127] The acquisition module 510 is used to acquire first image data of the display panel and input the first image data into the classification model, and output the color deviation state features corresponding to the display panel through the classification model. The classification model is trained based on the correspondence between image data samples and color deviation state feature labels.

[0128] Output module 520 is used to input the color shift state features into the compensation strategy agent, and output the target compensation action corresponding to the display panel through the compensation strategy agent;

[0129] The adjustment module 530 is used to adjust the display parameters of the display panel according to the target compensation action to obtain the adjusted display panel, and to adjust the working parameters of the compensation strategy agent according to the second image data of the adjusted display panel and the execution time of the target compensation action.

[0130] In one embodiment, the adjustment module is further configured to:

[0131] Acquire the second image data of the adjusted display panel and the execution time of the target compensation action;

[0132] Based on the first image data and the second image data, determine the reward value corresponding to the target compensation action; based on the execution duration, determine the penalty value corresponding to the target compensation action.

[0133] A target reward value is determined based on the reward value and the penalty value, and the working parameters of the compensation strategy agent are adjusted based on the target reward value.

[0134] In one embodiment, the adjustment module is further configured to:

[0135] Based on the second image data, determine the brightness uniformity and color uniformity of the display panel;

[0136] Determine the correlation between the first image data and the second image data;

[0137] The reward value corresponding to the target compensation action is determined based on the brightness uniformity, the chromaticity uniformity, and the correlation.

[0138] In one embodiment, the adjustment module is further configured to:

[0139] Based on the first weight corresponding to the brightness dimension, the second weight corresponding to the chromaticity dimension, and the third weight corresponding to the correlation dimension, the brightness uniformity, the chromaticity uniformity, and the correlation are weighted and summed to obtain the reward value corresponding to the target compensation action.

[0140] In one embodiment, the apparatus further includes a module for determining the classification model, configured to:

[0141] Obtain an image data sample set, wherein the image data sample set includes an image data sample set labeled with color deviation state feature tags;

[0142] Construct an initial classification model, wherein training parameters are set in the initial classification model;

[0143] The image data sample set is input into the initial classification model, and the initial classification model outputs the classification result.

[0144] Based on the difference between the classification result and the corresponding color deviation state feature label, the training parameters are adjusted until the adjusted difference meets the preset requirements, thus obtaining the classification model.

[0145] In one embodiment, the compensation policy agent includes a policy network and a value network, and the device is further configured to:

[0146] The color cast state features, the target compensation action, and the target reward value are input into the value network, and the value network outputs a predicted reward value. The predicted reward value is used to characterize the reward value accumulated within a preset time period after the target compensation action is performed under the color cast state features.

[0147] Based on the predicted reward value, the parameters of the policy network are adjusted to obtain the adjusted compensation policy agent.

[0148] Each module in the compensation device of the aforementioned display panel can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0149] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores image data and other data involved in the methods described in this embodiment. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a compensation method for a display panel.

[0150] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0151] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0152] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0153] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0155] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0156] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0157] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A compensation method for a display panel, characterized in that, The method includes: First image data of the display panel is acquired and input into a classification model. The classification model outputs the color shift state features corresponding to the display panel. The classification model is trained based on the correspondence between image data samples and color shift state feature labels. The color shift state features are input into the compensation strategy agent, and the compensation strategy agent outputs the target compensation action corresponding to the display panel. The display parameters of the display panel are adjusted according to the target compensation action to obtain the adjusted display panel, and the working parameters of the compensation strategy agent are adjusted according to the second image data of the adjusted display panel and the execution time of the target compensation action.

2. The method according to claim 1, characterized in that, The step of adjusting the working parameters of the compensation strategy agent based on the adjusted second image data of the display panel and the execution time of the target compensation action includes: Acquire the second image data of the adjusted display panel and the execution time of the target compensation action; Based on the first image data and the second image data, determine the reward value corresponding to the target compensation action; based on the execution duration, determine the penalty value corresponding to the target compensation action. A target reward value is determined based on the reward value and the penalty value, and the working parameters of the compensation strategy agent are adjusted based on the target reward value.

3. The method according to claim 2, characterized in that, Based on the first image data and the second image data, the reward value corresponding to the target compensation action is determined, including: Based on the second image data, determine the brightness uniformity and color uniformity of the display panel; Determine the correlation between the first image data and the second image data; The reward value corresponding to the target compensation action is determined based on the brightness uniformity, the chromaticity uniformity, and the correlation.

4. The method according to claim 3, characterized in that, The step of determining the reward value corresponding to the target compensation action based on the brightness uniformity, the chromaticity uniformity, and the correlation includes: Based on the first weight corresponding to the brightness dimension, the second weight corresponding to the chromaticity dimension, and the third weight corresponding to the correlation dimension, the brightness uniformity, the chromaticity uniformity, and the correlation are weighted and summed to obtain the reward value corresponding to the target compensation action.

5. The method according to claim 1, characterized in that, The method for determining the classification model includes: Obtain an image data sample set, wherein the image data sample set includes an image data sample set labeled with color deviation state feature tags; Construct an initial classification model, wherein training parameters are set in the initial classification model; The image data sample set is input into the initial classification model, and the initial classification model outputs the classification result. Based on the difference between the classification result and the corresponding color deviation state feature label, the training parameters are adjusted until the adjusted difference meets the preset requirements, thus obtaining the classification model.

6. The method according to claim 2, characterized in that, The compensation strategy agent includes a policy network and a value network, and the method further includes: The color cast state features, the target compensation action, and the target reward value are input into the value network, and the value network outputs a predicted reward value. The predicted reward value is used to characterize the reward value accumulated within a preset time period after the target compensation action is performed under the color cast state features. Based on the predicted reward value, the parameters of the policy network are adjusted to obtain the adjusted compensation policy agent.

7. A compensation device for a display panel, characterized in that, The device includes: The acquisition module is used to acquire the first image data of the display panel and input the first image data into the classification model. The classification model outputs the color deviation state features corresponding to the display panel. The classification model is trained based on the correspondence between image data samples and color deviation state feature labels. The output module is used to input the color shift state features into the compensation strategy agent, and the compensation strategy agent outputs the target compensation action corresponding to the display panel. The adjustment module is used to adjust the display parameters of the display panel according to the target compensation action to obtain the adjusted display panel, and to adjust the working parameters of the compensation strategy agent according to the second image data of the adjusted display panel and the execution time of the target compensation action.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.