An intelligent display effect optimization method and system

By building an AI image quality enhancement model in the cloud, using screen parameters to convert the dynamic range and color enhancement of video content, the problem of poor display effect on mobile terminals is solved, and the video data is automatically adapted to the screen characteristics on mobile terminals, improving the display quality.

CN116405730BActive Publication Date: 2025-07-22北京广播电视台 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310394997.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-07-22
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the video display effect on mobile terminals such as mobile phones. Due to the computing power and cost, complex image quality enhancement algorithms cannot be implemented, resulting in poor display effect.

Method used

By building an AI image quality enhancement model in the cloud, using screen parameters to convert the dynamic range and color enhancement of video content, generate video data that adapts to screen characteristics, and play it on mobile terminals.

Benefits of technology

It realizes automatic adaptation of video data based on screen characteristics on mobile terminals, improves display effect, avoids hardware costs and computing complexity, and improves the display quality of videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405730B_ABST
    Figure CN116405730B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent display effect optimization system, including: a dynamic range and color enhancement conversion network 13, which extracts picture features from an HDR video in the HLG standard, and obtains the screen parameter PA by the mobile phone 100, fuses the screen parameter as the conversion target into the picture features to generate composite features, and then completes the enhancement conversion of the dynamic range, color gamut, and color of the video content based on the composite features to generate intermediate HDR video data in the HDR10 standard; a picture quality enhancement conversion network 14, which extracts picture features from the intermediate HDR video data in the HDR10 standard, fuses the screen parameter of the conversion target into the picture features again to generate composite features, and completes the picture quality enhancement of the video content based on the composite features to generate the final HDR video data in the HDR10 standard. Therefore, it is possible to send a video that matches the screen parameters for a specific mobile phone screen, improving the display effect when the mobile phone 100 plays a video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video playback technology, and in particular to an intelligent display effect optimization method and system. Background Art

[0002] With the need to play high-quality videos such as HDR on mobile terminals such as mobile phones, in order to improve the video display effect, the video is usually processed at the video source (for example, adjusting contrast, brightness, dynamic range, etc.) to enhance the display effect during video playback (for example, CN113902651B). Since different screens have different display-related characteristics such as color gamut, brightness range, and resolution, the actual display effect is restricted by the screen characteristics, and the effect of high-quality videos cannot be fully exerted.

[0003] Another method to improve the video display effect is to uniformly adjust the video signal at the screen end. For example, some smart TVs have picture quality enhancement functions such as noise reduction, super resolution, and color enhancement. Since the screen end can directly obtain the performance parameters of the screen, such as resolution, frame rate, color space, brightness, etc., the received video signal can be converted and enhanced according to the characteristics of the screen, so that the video adapts to the parameters of these screens and exerts the maximum effect of the screen (for example, CN108810649B).

[0004] However, this method requires the installation of an algorithm chip in the screen device and is mainly applicable to high-end TVs. For mobile terminals such as mobile phones, due to the limited computing power (due to the requirements of thinness and lightness and cost limitations, it is impossible to have high computing power), only algorithms with extremely low complexity can be used, and the effect of the algorithms is also limited. Complex algorithms not only consume too much power, but may also cause problems such as device heating and freezing.

[0005] In addition, technologies that automatically adjust parameters such as the brightness, color, and contrast of the screen according to the physical characteristics of the TV screen and the user's viewing environment to improve the visual experience of the screen (for example, CN113055752A) also require additional hardware to obtain environmental parameters and run algorithms, resulting in high costs, complex structures, and complex operations, making it difficult to popularize on mobile terminals such as mobile phones.

[0006] The purpose of the present invention is to provide an intelligent display effect optimization method and system that can automatically adapt to screen characteristics without complex calculations. Summary of the Invention

[0007] The first technical solution of the present invention is an intelligent display effect optimization method, characterized by including the following steps:

[0008] Screen parameter reading step (S20), where the screen parameter reading module (12) reads screen parameters (PA) from the screen end (100); dynamic enhancement processing step (S30), where the dynamic range and color enhancement conversion network (13) extracts the picture features from the video provided to the screen end (100), fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the dynamic range and color enhancement conversion of the video content based on the composite features to generate intermediate video data with fused screen parameters.

[0009] Picture quality enhancement processing step (S40), where the intermediate video data with fused screen parameters is input into the picture quality enhancement conversion network (14), the picture quality enhancement conversion network (14) extracts its picture features, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the picture quality enhancement of the video content based on the composite features to generate the final video data; video distribution step (S50), where the final video data is distributed to the screen end (100) through the network (50) for playing or storing.

[0010] Preferably, the video provided to the screen end (100) is an HDR video in the HLG standard.

[0011] In the dynamic enhancement processing step (S30), the dynamic range and color enhancement conversion network (13) extracts the picture features from the HDR video in the HLG standard, fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the dynamic range and color enhancement conversion of the video content based on the composite features to generate intermediate HDR video data in the HDR10 standard.

[0012] In the picture quality enhancement processing step (S40), the picture quality enhancement conversion network (14) extracts the picture features from the intermediate HDR video data in the HDR10 standard, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the picture quality enhancement of the video content based on the composite features to generate the final HDR video data in the HDR10 standard.

[0013] Preferably, the screen parameters include any one or a combination of more of the brightness range, color gamut, color depth, dynamic curve, and resolution.

[0014] Preferably, it further includes a preprocessing step (S10), where the data preprocessing module (11) preprocesses the HDR video data in the HLG standard, and as a result, the HDR video data in the HLG standard is input into the dynamic range and color enhancement conversion network (13) in the form of normalized data.

[0015] The second technical solution is an intelligent display effect optimization system, which is characterized by including the following modules: a screen parameter reading module (12) that reads screen parameters (PA) from a screen terminal (100); a dynamic range and color enhancement conversion network (13) that extracts picture features from a video provided to the screen terminal (100), fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the enhancement conversion of the dynamic range and color of the video content based on the composite features to generate intermediate video data with fused screen parameters.

[0016] An image quality enhancement conversion network (14) extracts picture features from the intermediate video data with fused screen parameters, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the enhancement of the video content's image quality based on the composite features to generate final video data. The screen parameter reading module (12), the dynamic range and color enhancement conversion network (13), and the image quality enhancement conversion network (14) are installed on a platform that provides video services.

[0017] Preferably, the video provided to the screen terminal (100) is an HDR video in HLG standard. The dynamic range and color enhancement conversion network (13) extracts picture features from the HDR video in HLG standard, fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the enhancement conversion of the dynamic range and color of the video content based on the composite features to generate intermediate HDR video data in HDR10 standard.

[0018] The image quality enhancement conversion network (14) extracts picture features from the intermediate HDR video data in HDR10 standard, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the enhancement of the video content's image quality based on the composite features to generate final HDR video data in HDR10 standard.

[0019] Preferably, the screen parameters include any one or a combination of more of the brightness range, color gamut, color depth, dynamic curve, and resolution.

[0020] Preferably, it further includes a data preprocessing module (11) that pre-normalizes the HDR video data in HLG standard entering the dynamic range and color enhancement conversion network (13) so that the HDR video data in HLG standard is input into the dynamic range and color enhancement conversion network (13) in the form of normalized data.

[0021] As the screen terminal (100), a mobile terminal can be used, including a mobile phone and a tablet computer. Description of the Drawings

[0022] Figure 1Instruction diagram for video playback with an AI image quality enhancement model installed in the cloud;

[0023] Figure 2 Flowchart for the cloud to process the downloaded video;

[0024] Figure 3 Structural diagram for the dynamic range and color enhancement conversion network;

[0025] Figure 4 Training architecture for the dynamic range and color enhancement conversion network;

[0026] Figure 5 Structural diagram for the image quality enhancement conversion network;

[0027] Figure 6 Diagram for the discriminant network used in the training of the image quality enhancement conversion network;

[0028] Figure 7 Training architecture for the image quality enhancement conversion network. Detailed implementation mode

[0029] The following further elaborates on the detailed implementation mode of the present invention in conjunction with the accompanying drawings.

[0030] Aiming at the problems existing in the prior art (see the background art), the present invention provides a method that can enhance the video image quality according to the screen characteristics of the playback terminal, fully utilize the screen performance, improve the display effect of the screen, and play high-quality videos according to the screen characteristics without adding additional hardware for complex calculations at the screen end.

[0031] 1. Basic process

[0032] i. Collect the characteristic parameters of different screens, and let professional colorists design color correction and image quality optimization for different video samples according to the screen characteristics. The optimization results are used as the training data (learning samples) of the AI image quality enhancement model.

[0033] ii. Build an AI image quality enhancement model. While receiving video data, this model inputs the screen parameter PA of the screen end (video playback terminal), and performs image quality enhancement processing on the video according to the screen parameter PA to obtain a video adapted to the screen parameter PA. The requirements for the AI image quality enhancement model are simple calculation, fast speed, and good matching with the screen.

[0034] iii. The trained AI image quality enhancement model is installed on the cloud that provides video services. When a video is sent down, the app on the screen terminal obtains the screen characteristic parameter data and sends it back to the cloud. The AI image quality enhancement model on the cloud generates a video with the optimal image quality effect under the current screen characteristics in real time according to the screen parameters, and then sends it down to the screen terminal for playback. The following takes the playback of a high-quality HLG-standard HDR video on a mobile phone as an example for illustration.

[0035] Figure 1 It is an illustration diagram for video playback with an AI image quality enhancement model installed on the cloud.

[0036] The cloud 10, as a platform for providing video services, is installed with a data preprocessing module 11, a screen parameter reading module 12, a dynamic range and color enhancement conversion network 13, and an image quality enhancement conversion network 14. The dynamic range and color enhancement conversion network 13 converts the dynamic range and color of the HLG-standard HDR video to the dynamic range and color of the intermediate HDR video in the HDR10 standard. The image quality enhancement conversion network 14 enhances the image quality based on the intermediate HDR video in the HDR10 standard. While the dynamic range and color enhancement conversion network 13 and the image quality enhancement conversion network 14 are processing the video data as the AI image quality enhancement model, the screen parameter PA of the mobile phone 100 is input, and a video with the optimal image quality effect under the current screen characteristics is generated in real time.

[0037] The mobile phone 100 as the screen terminal is installed with a video playback APP. In addition to the usual playback module 110, the video playback APP is added with a screen parameter acquisition module 120. The screen parameter acquisition module 120 acquires the screen parameter PA of the mobile phone 100 and sends it back to the cloud 10 when the playback module 110 downloads a video.

[0038] Figure 2 It is a flowchart for the cloud to process the sent-down video.

[0039] Preprocessing step S10, the data preprocessing module 11 preprocesses the sent-down HLG-standard HDR video data. For example, all the image values from 0 - 255 are divided by 255 and normalized to 0 - 1. As a result of the processing, the HLG-standard HDR video data is input into the dynamic range and color enhancement conversion network 13 in the form of normalized data. Through the normalized data processing, the model can converge better.

[0040] Screen parameter reading step S20, the screen parameter reading module 12 reads the screen parameter PA of the mobile phone 100 from the mobile phone (screen terminal) 100, and the screen parameter PA and the HLG-standard HDR video data are simultaneously input into the dynamic range and color enhancement conversion network 13.

[0041] Dynamic enhancement and color processing step S30: The dynamic range and color enhancement conversion network 13 extracts the frame features of the HDR video in the HLG standard, and fuses parameters such as the brightness range, color gamut, color depth, dynamic curve, and resolution of the mobile phone screen, which is the conversion target, into the frame features to generate composite features. Then, based on the composite features, the dynamic range and color of the video content are enhanced and converted to generate intermediate HDR video data in the HDR10 standard.

[0042] Image quality enhancement processing step S40: The intermediate HDR video data in the HDR10 standard is input into the image quality enhancement conversion network 14. The image quality enhancement conversion network 14 extracts its frame features, and once again fuses parameters such as the brightness range, color gamut, color depth, dynamic curve, and resolution of the mobile phone screen, which is the conversion target, into the frame features to generate composite features. Based on the composite features, the image quality of the video content is enhanced to generate the final HDR video data in the HDR10 standard.

[0043] As screen parameters, one or more can be arbitrarily selected from the brightness range, color gamut, color depth, dynamic curve, and resolution according to needs.

[0044] Video distribution step S50: The generated HDR video data in the HDR10 standard is distributed to the mobile phone 100 through the network 50 and played or stored by the playback module 110.

[0045] Since the video data received by the mobile phone 100 is processed by dynamic and color enhancement and image quality enhancement based on the screen parameters PA of the mobile phone 100 itself, the characteristics of the screen can be maximally utilized during playback, and the display effect of the high-quality HDR video in the HLG standard on the mobile phone can be fully exerted.

[0046] The enhancement processing in the present invention refers to the processing that can match the video with the screen, enabling the screen to fully exert its characteristics and improving the display effect of the video, which is not the same concept as simply increasing the dynamic range or resolution.

[0047] The dynamic range and color enhancement conversion network 13 will be described below.

[0048] Figure 3 It is a structural illustration of the dynamic range and color enhancement conversion network. The dynamic range and color enhancement conversion network 13 includes a frame feature extraction module 131, a feature fusion module 132, and a generation module 133. The frame feature extraction module 131 uses a simple convolutional layer with a kernel of 1 to extract the features of each frame of the video. The feature fusion module 132 consists of a splicing algorithm (splicing operation), a simple convolutional layer with a kernel of 1, and an activation layer, and is used to splice the frame features and the screen parameters into a combined feature.

[0049] The generation module 133 uses a third-order residual fully convolutional network composed of simple convolutional layers with a kernel size of 1. The three residual groups 131a are connected in series, and the output features are added to the input combined features (addition operation), and then the features are extracted by a convolutional layer with a kernel size of 1.

[0050] The residual group 131a is composed of three residual units A connected in series. The input of the residual group 131a and the output of the last residual unit A are added (addition operation), and the result is used as the output of the residual group 131a.

[0051] Each residual unit A includes two convolutional layers with a kernel size of 1, two activation layers, and one self-attention layer, and is connected in series in the order of convolutional layer, activation layer, convolutional layer, activation layer, self-attention layer, and activation layer. The features output by the self-attention layer are activated by the activation layer and then added to the input features (addition operation), and the result is used as the output of the residual unit A.

[0052] The dynamic range and color enhancement conversion network 13 and the image quality enhancement conversion network 14 are trained based on YUV format data. The training process incorporates artificial prior knowledge, and the mean squared error loss (mse) is used as the loss function, which takes into account both the effect and speed while ensuring no color deviation.

[0053] The training steps are as follows:

[0054] Step 1, training data preparation.

[0055] Collect a batch of high-quality HDR video data in the HLG standard as input data, and have professional colorists perform color grading on it in multiple versions according to different screen parameters. While recording the screen parameters, output HDR video data in the HDR10 standard in multiple versions as target data. Convert the input data and target data into a unified YUV format respectively, and extract frames one by one to form paired training data from (input data + screen parameters) to (target data).

[0056] For example, play an ultra-high-definition HDR video with 8K resolution, high bitrate, BT2020 color gamut, and HLG high dynamic range, which has a very good display effect on an 8K large TV, on a small screen terminal such as a mobile phone. After professional colorists optimize the color grading of the video, obtain an ultra-high-definition HDR video with 8K resolution, low bitrate, P3-D65 color gamut, and HDR10 standard (PQ high dynamic range) supported by the mobile phone screen. That is, obtain an ultra-high-definition HDR video in the HDR10 standard with a very good playback effect on the mobile phone through the color grading of professional colorists.

[0057] Figure 4 It is the training architecture of the dynamic range and color enhancement conversion network.

[0058] Step 2, during training, the dynamic range and color enhancement conversion network 13 processes the input data stream and outputs the result. First, the input video data is input into the frame feature extraction module 131 to extract frame features. Then, the screen parameters and the frame features are input into the feature fusion module 132 to generate composite features. Finally, the result generated by inputting the composite features into the generation module 133 is used as the network output.

[0059] Step 3, iterate the weights of the dynamic range and color enhancement conversion network 13 according to the loss function of the network output and the target data. Input the output of the previous step and the target data into the loss function calculation module 100 to calculate the loss value, and then use the backpropagation technique to iteratively optimize the network weights. The loss function is as follows:

[0060] ;

[0061] where, I Gen represents the result of the network output data, I GT represents the target data, and N is the total number of elements of the data.

[0062] Repeat Step 2 and 3 until the loss function calculated by the loss function calculation module 100 no longer decreases, that is, the network converges.

[0063] Figure 5 is the structural diagram of the image quality enhancement conversion network.

[0064] The image quality enhancement conversion network 14 includes a frame feature extraction module 141, a feature fusion module 142, and a generation module 143. The frame feature extraction module 141 uses a simple convolution with a kernel of 3. The feature fusion module 142 consists of a splicing algorithm and a simple convolution with a kernel of 3. The generation module 143 uses a third-order residual fully convolutional network composed of simple convolutions with a kernel of 3.

[0065] Except for using convolutional layers with a kernel of 3, the network structure of the image quality enhancement conversion network 14 is the same as that of the dynamic range and color enhancement conversion network 13. For details, refer to the dynamic range and color enhancement conversion network 13, which will not be elaborated here. The image quality enhancement conversion network 14 is trained based on the generative adversarial method.

[0066] Figure 6 is the structural diagram of the discriminator network for training the image quality enhancement conversion network.

[0067] The discriminator network uses the discriminator of the U-shaped network and adds a decoding architecture D enc on the basis of the conventional encoding architecture D dec .

[0068] The encoder 200 includes a structure of 4 serially-connected feature downsampling layers 210. Each time of downsampling doubles the number of feature channels and halves the feature resolution. Finally, through a fully-connected layer 220, the discrimination result of the encoder 200 is output. The decoder 300 includes a structure of 4 serially-connected feature upsampling layers 310, and fuses the features of different layers of the encoder 300 through skip connections. Finally, through a convolutional layer 320 with a kernel of 3, the discrimination result of the decoder 300 is output.

[0069] This structure of dual discrimination of encoding and decoding helps to provide pixel-level information feedback while maintaining global context information.

[0070] The model is trained based on YUV format data, fuses artificial prior knowledge, and the loss function of the generation network uses L mse (mean squared error loss), L lpips (perceptual loss), L adv (adversarial loss) combination. The loss function of the discrimination network uses the combination of L Denc (encoder loss) and L Ddec (decoder loss) to form L adv (adversarial loss).

[0071] The training steps are as follows:

[0072] Step 1, training data production.

[0073] Collect HDR video data in HDR10 standard output by the dynamic range and color enhancement conversion network as input data. Professional colorists optimize the picture quality of multiple versions according to different screen parameters. While recording the screen parameters, output HDR video data in HDR10 standard of multiple versions as target data. Respectively convert the input data and the target data into a unified YUV format and extract frame by frame to form paired training data from (input data + screen parameters) to (target data).

[0074] Figure 7 is the training architecture of the picture quality enhancement conversion network.

[0075] Step 2, the picture quality enhancement conversion network 14 processes the data stream of the input video and outputs the result. That is, first input the input data into the picture feature extraction module 141 to extract picture features, then input the screen parameters and the picture features into the feature fusion module 142 to generate composite features, and finally input the composite features into the generation module 143 to generate the result as the network output.

[0076] Step 3, iterate to generate the weights of the network according to the loss function of the network output and the target data. The loss function of the generation network uses L mse (mean squared error loss), Llpips(Perceptual loss), L adv (Adversarial loss), where L mse and Llpips only need to input the output of the previous step and the target data into the direct loss calculation module 500 to calculate the loss value. For the adversarial loss, the output of the previous step and the target data need to be input into the discriminative network 400 respectively. After obtaining the discriminative result, the adversarial loss calculation module 600 calculates the loss. Finally, in the loss aggregation module 700, the three loss functions are added according to specific weights to obtain the final loss value. Using the backpropagation technique, the network weights are iteratively optimized. The three loss functions of the generation function are as follows:

[0077] ;

[0078] where, I Gen represents the result of the output data of the generation network, I GT represents the target data, and N is the total number of elements of the data.

[0079] ;

[0080] where, I Gen represents the result of the output data of the generation network, I GT represents the target data, ϕ represents the feature extraction function, and τ is used to transform the feature difference into the LPIPS score.

[0081] ;

[0082] where, I Gen represents the result of the output data of the generation network, D enc represents the encoder of the discriminative network, D dec represents the decoder of the discriminative network, i and j respectively represent the row and column coordinates of the pixel, and E is the expectation.

[0083] Step 4, iterate the weights of the discriminative network according to the loss function of the network output and the target data. The loss function of the discriminative network is as follows:

[0084] ;

[0085] where, I Gen represents the result of the output data of the generation network, I GT represents the target data, D enc represents the encoder of the discriminative network, D dec represents the decoder of the discriminative network, i and j respectively represent the row and column coordinates of the pixel, and E is the expectation.

[0086] Repeat steps 2 to 4 until the loss function no longer decreases, that is, the network converges. After the network converges, the generation network is saved separately as the image quality enhancement conversion network 14.

[0087] When training the dynamic range and color enhancement conversion network 13 and the picture quality enhancement conversion network 14, pre-training was respectively carried out on automatically generated large datasets, and finally transfer fine-tuning training was carried out on the small data made by colorists. For example, the method of the dynamic range and color enhancement conversion network 13 is to collect a large number of HDR10-annotated videos, use simple traditional algorithms to adjust the color saturation and contrast of the videos, and then output videos in the HLG standard, so that the network learns the mapping between the HLG-annotated videos obtained and the original HDR10 standard videos. In this way, the pre-trained model has the basic ability of dynamic range and color enhancement. After that, only a small amount of data is needed to make the model transfer to a specific dynamic range and color style.

[0088] The method of the picture quality enhancement conversion network 14 is similar. Simple blurring and compression are carried out on a large number of collected HDR10 standard videos, and then the model is used to learn the mapping from the blurred videos to the original videos, obtaining a pre-trained model with the basic ability of picture quality enhancement. After that, only a small amount of data is needed to make the model transfer to a specific type of picture quality enhancement.

[0089] The following is an example to illustrate the model input and output.

[0090] 1. The input image is in YUV444 data. The shape is: 4320x7680x3

[0091] 2. The input screen parameters are tensors represented by encoding. If there are C kinds of parameters, the shape is 4320x7680xC.

[0092] a. The screen parameters include but are not limited to color gamut, maximum brightness, and can also be color depth, dynamic curve, resolution, etc.

[0093] b. Examples of encoding methods:

[0094] i. Non-numerical conditions need to be encoded. For example, there are several color gamuts such as BT709, P3, BT2020, etc., which are respectively represented by 0, 1, 2. When inputting the P3 color gamut, a 4320x7680 tensor is filled with 1, with a shape of 4320x7680x1 and all values being 1.

[0095] ii. Numerical conditions do not need to be encoded. For example, the maximum brightness is 1000 nits. When inputting a video with the maximum brightness, a 4320x7680 tensor is filled with 1000, with a shape of 4320x7680x1 and all values being 1000.

[0096] iii. When multiple conditions are input, they are spliced with each other. When inputting two conditions of color gamut and maximum brightness, it is a 4320x7680x2 tensor.

[0097] 3. Process of inputting screen parameters into the model:

[0098] a. After the input image (data) passes through the first layer of convolution (screen feature extraction module 131), a feature map with a shape of 4320x7680x16 is obtained;

[0099] b. After the input parameter tensor and the feature map of 4320x7680x16 are concatenated, a tensor with a shape of 4320x7680x18 is obtained;

[0100] c. Subsequent model calculations are performed on the tensor with a shape of 4320x7680x18.

[0101] 4. The output image is in YUV444 data. The shape is: 4320x7680x3.

[0102] The present invention has specific beneficial effects

[0103] a. As long as an APP capable of uploading screen parameters is installed on the screen side of the video player, a video corresponding to the screen characteristics can be obtained, fully leveraging the characteristics of the screen and improving the display effect of the video. Since only an APP for uploading screen parameters needs to be installed on the screen side, compared with the method of installing hardware for processing on high-end TVs, etc., it is especially suitable for mobile terminals with low computing power and limited power supply such as mobile phones and tablets.

[0104] b. Method for producing training data: Let professional colorists produce data so that through AI algorithms, the presentation effect of the video on the screen can be close to the effect of manual optimization by professional colorists on professional software, and the display effect is greatly improved.

[0105] c. Structure design, loss function design, and training strategy design of the AI model:

[0106] i. It can accept input of screen characteristic parameters, an AI model, and obtain different output effects, with high efficiency and convenient use.

[0107] ii. The model is designed according to the characteristics of color and image quality, enabling the AI model to learn the changes in color and image quality well while taking into account the inference speed.

[0108] iii. The training strategy enables the color enhancement and image quality enhancement of the AI model to not interfere with each other, ensuring the overall effect, performing end-to-end video conversion, and accelerating the convergence speed.

[0109] iv. Only a small number of samples are required for training, without the need for a large amount of data, reducing the workload of colorists.

[0110] d. Overall process: It can automatically adapt to the screen characteristics of the screen end, perform targeted picture quality optimization, and give full play to the performance of the screen.

[0111] As training data, its production method can be modified to obtain effects of different styles or output different target video parameters (different resolutions, HDR standards, etc.).

[0112] The AI model can also adopt other designs as long as it can input the characteristic parameters of the screen and fuse them with the video to obtain different output results.

[0113] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claims.

Claims

1. An intelligent display effect optimization method, characterized in that, It includes the following steps: Screen parameter reading step (S20), where the screen parameter reading module (12) reads the screen parameters (PA) from the screen terminal (100); Dynamic enhancement processing step (S30), where the dynamic range and color enhancement conversion network (13) extracts the picture features from the video provided to the screen terminal (100), fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the dynamic range and color enhancement conversion of the video content based on the composite features to generate intermediate video data integrated with the screen parameters; Picture quality enhancement processing step (S40), where the intermediate video data integrated with the screen parameters is input into the picture quality enhancement conversion network (14), the picture quality enhancement conversion network (14) extracts its picture features, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the picture quality enhancement of the video content based on the composite features to generate the final video data; Video distribution step (S50), where the final video data is distributed to the screen terminal (100) through the network (50) for playback or storage.

2. The intelligent display effect optimization method according to claim 1, characterized in that The video provided to the screen terminal (100) is an HDR video in HLG standard, In the dynamic enhancement processing step (S30), the dynamic range and color enhancement conversion network (13) extracts the picture features from the HDR video in HLG standard, fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the dynamic range and color enhancement conversion of the video content based on the composite features to generate intermediate HDR video data in HDR10 standard; In the picture quality enhancement processing step (S40), the picture quality enhancement conversion network (14) extracts the picture features from the intermediate HDR video data in HDR10 standard, fuses the screen parameters (PA) to be converted into the picture features again to generate composite features, and completes the picture quality enhancement of the video content based on the composite features to generate the final HDR video data in HDR10 standard.

3. The intelligent display effect optimization method according to claim 2, wherein, The screen parameters (PA) include any one or a combination of multiple of brightness range, color gamut, color depth, dynamic curve, and resolution.

4. The intelligent display effect optimization method according to claim 2 or 3, characterized in that It includes a preprocessing step (S10), where the data preprocessing module (11) preprocesses the HDR video data in HLG standard, and as a result, the HDR video data in HLG standard is input into the dynamic range and color enhancement conversion network (13) in the form of normalized data.

5. Intelligent display effect optimization system, characterized in that, It includes the following modules Screen parameter reading module (12), which reads the screen parameters (PA) from the screen terminal (100); Dynamic range and color enhancement conversion network (13), which extracts the picture features from the video provided to the screen terminal (100), fuses the screen parameters (PA) to be converted into the picture features to generate composite features, and then completes the dynamic range and color enhancement conversion of the video content based on the composite features to generate intermediate video data integrated with the screen parameters; The picture quality enhancement conversion network (14) extracts picture features from intermediate video data that incorporates screen parameters, and then incorporates the screen parameters (PA) that are the conversion target into the picture features again to generate composite features. Based on the composite features, the picture quality of the video content is enhanced to generate the final video data. The screen parameter reading module (12), the dynamic range and color enhancement conversion network (13), and the picture quality enhancement conversion network (14) are installed on a platform that provides video services.

6. The intelligent display effect optimization system according to claim 5, wherein The video provided to the screen terminal (100) is an HDR video in the HLG standard. The dynamic range and color enhancement conversion network (13) extracts picture features from the HDR video in the HLG standard, and incorporates the screen parameters (PA) that are the conversion target into the picture features to generate composite features. Then, based on the composite features, the dynamic range and color of the video content are enhanced and converted to generate intermediate HDR video data in the HDR10 standard. The picture quality enhancement conversion network (14) extracts picture features from the intermediate HDR video data in the HDR10 standard, incorporates the screen parameters (PA) that are the conversion target into the picture features again to generate composite features, and based on the composite features, the picture quality of the video content is enhanced to generate the final HDR video data in the HDR10 standard.

7. The intelligent display effect optimization system according to claim 5, wherein, The screen parameters (PA) include any one or a combination of multiple of the brightness range, color gamut, color depth, dynamic curve, and resolution.

8. The intelligent display effect optimization system according to claim 6 or as described above, characterized in that, It includes a data preprocessing module (11) that pre-normalizes the HDR video data in the HLG standard that enters the dynamic range and color enhancement conversion network (13), so that the HDR video data in the HLG standard is input into the dynamic range and color enhancement conversion network (13) in the form of normalized data.

9. The intelligent display effect optimization system according to any one of claims 5 to 8, characterized in that The screen terminal (100) is a mobile terminal, including mobile phones and tablet computers.

Citation Information

Patent Citations

  • Image quality adjustment methods, smart TVs and storage media

    CN108810649B

  • A video quality enhancement system based on deep learning

    CN113902651B