Video processing method, display device and storage medium
Patent Information
- Application Number
- CN202380084734.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-09-04
- Publication Date
- 2025-07-18
AI Technical Summary
Existing technology is difficult to effectively improve the display effect of standard dynamic range (SDR) video on display devices with strong brightness capabilities, resulting in video content that cannot fully utilize the brightness capabilities of the display device, resulting in a waste of resources.
By determining the brightness information of the original video and the brightness capability of the display device, the scalable dynamic range of each video frame image is calculated, and tone mapping is performed to improve the dynamic range of the video, thereby improving the display effect.
It is possible to display the video with enhanced dynamic range on a display device with strong brightness capability, which improves the user experience and display effect and fully utilizes the brightness capability of the display device.
Smart Images

Figure CN120345237A_ABST
Abstract
Description
Video processing method, display device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 9, 2022, with application number 202211585804.3 and application name “Method, display device and storage medium for processing video”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing technology, and in particular to a method for processing video, a display device, and a storage medium. Background Art
[0003] Compared with standard dynamic range (SDR) video, high-dynamic range (HDR) video has clearer light and dark levels and richer image details. It can reproduce real scenes more realistically and provide users with better video viewing experience.
[0004] With the development of HDR technology, the brightness capabilities of display devices are becoming increasingly higher to better play HDR videos. The brightness capabilities of a display device are usually expressed in terms of its dynamic range. The higher the dynamic range of a display device, the stronger its brightness capabilities.
[0005] In the course of video technology development, a large number of SDR videos have been accumulated. However, as more and more display devices with strong brightness capabilities are available, the display effects of these SDR videos are not good on display devices with strong brightness capabilities.
[0006] Summary of the Invention
[0007] The present application provides a method for processing video, a display device, and a storage medium. By determining the brightness information of the original video and combining it with the brightness capability of the current display device, the actual scalable dynamic range is determined, and tone mapping is performed based on the scalable dynamic range, thereby truly improving the dynamic range.
[0008] In a first aspect, the present application provides a method for processing a video, the method being applied to a display device, the method comprising:
[0009] Obtain a decoded video of an original video, where the decoded video includes multiple video frame images; determine brightness information corresponding to each video frame image; obtain the brightness capability of a display device and the dynamic range of the original video, and calculate an expandable dynamic range of each video frame image based on the brightness capability of the display device and the dynamic range of the original video; perform tone mapping on each video frame image using the brightness information corresponding to each video frame image and the expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image, where the dynamic range of the enhanced image is greater than the dynamic range of the video frame image.
[0010] Optionally, the display device provided in the embodiments of the present application may include a mobile phone, a tablet computer, a wearable device, a television, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), various camera devices, etc., or may be other devices or apparatuses capable of performing image processing. The embodiments of the present application do not impose any restrictions on the specific type of the display device.
[0011] Optionally, the original video may include an SDR video. The SDR video may be a video played online by a display device.
[0012] The video processing method provided in the first aspect fully considers the brightness capability of the display device and the dynamic range of the original video when determining the expandable dynamic range of each video frame image. The expandable dynamic range is then used to perform tone mapping on each video frame image, thereby truly improving the dynamic range of the video frame image. That is, the enhanced image is the image with the enhanced dynamic range. Therefore, the dynamic range of the video composed of multiple enhanced images is also improved, thereby improving the display effect and user experience when the video with the enhanced dynamic range is displayed on the display device.
[0013] In one possible implementation, the method for processing video provided in the present application may further include generating an enhanced video based on multiple enhanced images. The enhanced video is an extended dynamic range video, that is, the enhanced video is a video with an extended dynamic range. The enhanced images corresponding to each video frame image are arranged in chronological order to obtain an enhanced video, which can be played on a display device. In this implementation, after processing a video frame image, the enhanced image corresponding to the video frame image is displayed on the display device. Since the display device processes very quickly, for the user, it is like watching a video smoothly, and the video being watched is a video with an enhanced dynamic range, which improves the user experience.
[0014] Optionally, in a possible implementation, the brightness information may include a brightness value, and determining the brightness information corresponding to each video frame image includes:
[0015] Determine the pixel format of each video frame image; if the pixel format is YUV, obtain the Y value of each video frame image; or, if the pixel format is RGB, calculate the brightness value of each video frame image using a preset formula. The Y value represents the brightness value.
[0016] Optionally, in another possible implementation manner, the brightness information may include a brightness histogram, and determining the brightness information corresponding to each video frame image includes: generating a brightness histogram for each video frame image.
[0017] In one possible implementation, tone mapping is performed on each video frame image using the brightness information corresponding to each video frame image and the expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image, including: using the brightness information corresponding to each video frame image to determine the standard dynamic area and the extended dynamic area corresponding to each video frame image; determining a first coefficient corresponding to the standard dynamic area of each video frame image and a second coefficient corresponding to the extended dynamic area of each video frame image based on the expandable dynamic range of each video frame image; tone mapping is performed on the pixel points in the standard dynamic area according to the first coefficient, and tone mapping is performed on the pixel points in the extended dynamic area according to the second coefficient to obtain an enhanced image corresponding to each video frame image.
[0018] Optionally, the standard dynamic region includes a plurality of pixels in the video frame image, and the standard dynamic region also includes a plurality of pixels in the video frame image, wherein the brightness information of the pixels in the standard dynamic region is less than a preset threshold, and the brightness information of the pixels in the extended dynamic region is greater than or equal to the preset threshold.
[0019] Optionally, the preset threshold may be set by a user, may be calculated by a grayscale ratio, or may be determined by a machine learning model.
[0020] Optionally, the first coefficient is smaller than the second coefficient.
[0021] In this implementation, the brightness information corresponding to each video frame is used to determine the adjustable areas in each video frame, namely the standard dynamic area and the extended dynamic area. Different coefficients are determined for the standard dynamic area and the extended dynamic area, and the pixel values of the pixels in the different areas (such as the standard dynamic area and the extended dynamic area) are adjusted based on the different coefficients (such as the first coefficient and the second coefficient). This effectively improves the dynamic range of the final enhanced image. As a result, the dynamic range of the video composed of multiple enhanced images is also improved, thereby improving the display effect and user experience when the video with the enhanced dynamic range is displayed on a display device.
[0022] In one possible implementation, tone mapping is performed on pixels in a standard dynamic area according to a first coefficient, and tone mapping is performed on pixels in an extended dynamic area according to a second coefficient to obtain an enhanced image corresponding to each video frame image, including: calculating a first product of an original pixel value of each pixel in the standard dynamic area and the first coefficient, and updating the pixel value of each pixel in the standard dynamic area according to the first product; calculating a second product of the original pixel value of each pixel in the extended dynamic area and the second coefficient, and updating the pixel value of each pixel in the extended dynamic area according to the second product; and generating an enhanced image corresponding to each video frame image according to each updated pixel in the standard dynamic area and each updated pixel in the extended dynamic area.
[0023] In this implementation, the pixel values of the pixels in different areas (such as the standard dynamic area and the extended dynamic area) are adjusted separately according to different coefficients (such as the first coefficient and the second coefficient). When the first coefficient is less than the second coefficient and the second coefficient is 1, the pixel values of the pixels in the extended dynamic area are maintained, and the pixel values of the pixels in the standard dynamic area are reduced, so that the dynamic range of the enhanced image generated is truly improved. Alternatively, when the first coefficient is less than the second coefficient and the first coefficient is 1, the pixel values of the pixels in the standard dynamic area are maintained, and the pixel values of the pixels in the extended dynamic area are increased, so that the dynamic range of the enhanced image generated is truly improved.
[0024] In one possible implementation, tone mapping is performed on each video frame image using the brightness information corresponding to each video frame image and the expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image, including: using the brightness information corresponding to each video frame image to determine the low grayscale area, medium grayscale area and high grayscale area corresponding to each video frame image; and adjusting the brightness of at least one of the low grayscale area, the medium grayscale area and the high grayscale area according to the expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image.
[0025] For example, the current brightness of the low grayscale area and the middle grayscale area is maintained, and the brightness of the high grayscale area is increased. For another example, the current brightness of the middle grayscale area and the high grayscale area is maintained, and the brightness of the low grayscale area is reduced.
[0026] Alternatively, if the brightness information corresponding to each video frame image is determined by a brightness histogram, the brightness histogram of each video frame image can be divided into low grayscale areas, medium grayscale areas, and high grayscale areas using T1 and T2. The values of T1 and T2 can be set by the user, can be set based on the ratio of the brightness grayscale distribution, or can be determined by a machine learning model.
[0027] In this implementation, the brightness information corresponding to each video frame image is used to determine the adjustable areas in each video frame image, namely the low grayscale area, the medium grayscale area, and the high grayscale area. Then, based on the expandable dynamic range of each video frame image, the brightness of at least one of the low grayscale area, the medium grayscale area, and the high grayscale area is adjusted. For example, the current brightness of the low grayscale area and the medium grayscale area is maintained, while the brightness of the high grayscale area is increased. For another example, the current brightness of the medium grayscale area and the high grayscale area is maintained, while the brightness of the low grayscale area is reduced. This can effectively improve the dynamic range of the final enhanced image.
[0028] In one possible implementation, the brightness of at least one area among the low grayscale area, the medium grayscale area and the high grayscale area is adjusted according to the expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image, including: determining the adjusted grayscale value of the pixel points in the low grayscale area according to the expandable dynamic range of each video frame image; determining the adjusted grayscale value of the pixel points in the medium grayscale area; determining the adjusted grayscale value of the pixel points in the high grayscale area; generating an enhanced image corresponding to each video frame image according to the grayscale values of each adjusted pixel point in the low grayscale area, the medium grayscale area and the high grayscale area.
[0029] In this implementation, the brightness information corresponding to each video frame is used to determine the adjustable regions within each video frame, namely, the low, medium, and high grayscale regions. The grayscale values of each pixel in each of these regions are then determined based on the scalable dynamic range of each video frame, thereby generating an enhanced image corresponding to each video frame. Because the grayscale values of each pixel in each region are determined based on the scalable dynamic range of each video frame, the grayscale values of each pixel are adjusted to varying degrees, significantly improving the dynamic range of the resulting enhanced image.
[0030] In one possible implementation, the adjusted grayscale values of the pixels in the low grayscale area are determined based on the expandable dynamic range of each video frame image, including: determining a third coefficient corresponding to the low grayscale area based on the expandable dynamic range of each video frame image; performing tone mapping on the pixels in the low grayscale area based on the third coefficient to obtain the adjusted grayscale values of the pixels in the low grayscale area.
[0031] Optionally, the third coefficient may be determined according to the dynamic range of the original video and the scalable dynamic range of each video frame image.
[0032] Optionally, the grayscale value of each pixel in the low grayscale area after adjustment and the grayscale value before adjustment may satisfy the following relationship:
[0033] y1=(A / B)*k*x1, where x1 represents the grayscale value of the pixel in the low grayscale area before adjustment, y1 represents the grayscale value of the pixel in the low grayscale area after adjustment, A / B represents the third coefficient, A represents the dynamic range of the original video, B represents the expandable dynamic range of each video frame image, and k represents a constant.
[0034] Alternatively, the grayscale value of each pixel in the low grayscale area after adjustment and the grayscale value before adjustment may satisfy the following relationship:
[0035] y1=M*(A / B)*k*x1, (M>1), where x1 represents the grayscale value of the pixel in the low grayscale area before adjustment, y1 represents the grayscale value of the pixel in the low grayscale area after adjustment, M*(A / B) represents the third coefficient, A represents the dynamic range of the original video, B represents the expandable dynamic range of each video frame image, and M and k represent constants.
[0036] Optionally, the grayscale value of each pixel in the medium grayscale area after adjustment and the grayscale value before adjustment may satisfy the following relationship:
[0037] y2 = k1 * x2 + b1, where x2 represents the grayscale value of the pixel in the mid-grayscale region before adjustment, y2 represents the grayscale value of the pixel in the mid-grayscale region after adjustment, and k1 and b1 represent constants. The values of k1 and b1 can be set based on the ratio of the brightness grayscale distribution or determined by a machine learning model.
[0038] Optionally, the grayscale value of each pixel in the high grayscale area after adjustment and the grayscale value before adjustment may satisfy the following relationship:
[0039] y3 = k2 * x3 + b2, where x3 represents the grayscale value of the pixel in the high grayscale area before adjustment, y3 represents the grayscale value of the pixel in the high grayscale area after adjustment, and k2 and b2 represent constants. The values of k2 and b2 can be set based on the ratio of the brightness grayscale distribution or determined by a machine learning model.
[0040] In this implementation, the grayscale values of each pixel in the low, medium, and high grayscale regions are adjusted based on the expandable dynamic range of each video frame, ultimately generating an enhanced image corresponding to each video frame. Because the grayscale values of each pixel in each region are determined based on the expandable dynamic range of each video frame, the grayscale values of each pixel are adjusted to varying degrees, effectively improving the dynamic range of the resulting enhanced image.
[0041] In a second aspect, the present application provides a display device, which is included in a display device, and has the function of implementing the display device behavior in the first aspect and the possible implementation of the first aspect. The function can be implemented by hardware, or the corresponding software can be implemented by hardware. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a first acquisition module or unit, a determination module or unit, a second acquisition module or unit, a processing module or unit, etc.
[0042] In a third aspect, the present application provides a display device, comprising: a processor, a memory, and an interface; the processor, the memory, and the interface cooperate with each other so that the display device executes any one of the methods in the technical solution provided in the first aspect.
[0043] In a fourth aspect, the present application provides a chip including a processor, wherein the processor is configured to read and execute a computer program stored in a memory to perform the method of the first aspect and any possible implementation thereof.
[0044] Optionally, the chip also includes a memory, and the memory is connected to the processor via circuits or wires.
[0045] Optionally, the chip also includes a communication interface.
[0046] In a fifth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor executes any one of the methods in the technical solution of the first aspect.
[0047] In a sixth aspect, the present application provides a computer program product, which includes: a computer program code, which, when executed on a display device, enables the display device to execute any one of the methods in the technical solution of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a diagram illustrating a frame of an SDR video when playing on a display device in the related art according to an exemplary embodiment of the present application;
[0049] FIG2 is a diagram showing a single frame of an SDR video after processing and played on a display device, according to an exemplary embodiment of the present application;
[0050] FIG3 is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0051] FIG4 is a schematic diagram of an enhanced image in another application scenario provided by an embodiment of the present application;
[0052] FIG5 is a flow chart of a method for processing a video according to an embodiment of the present application;
[0053] FIG6 is a diagram showing brightness information provided in an embodiment of the present application;
[0054] FIG7 is another brightness information display diagram provided in an embodiment of the present application;
[0055] FIG8 is a schematic diagram of a process for generating an enhanced image corresponding to a video frame image according to an embodiment of the present application;
[0056] FIG9 is another schematic diagram of a process for generating an enhanced image corresponding to a video frame image according to an embodiment of the present application;
[0057] FIG10 is a schematic diagram of area division provided in an embodiment of the present application;
[0058] FIG11 is a schematic diagram of brightness area division provided in an embodiment of the present application;
[0059] FIG12 is a schematic diagram of tone mapping provided in an embodiment of the present application;
[0060] FIG13 is a schematic diagram of an implementation process shown in an exemplary embodiment of the present application;
[0061] FIG14 shows a schematic structural diagram of a display device provided in an embodiment of the present application;
[0062] FIG15 is a schematic structural diagram of a display device provided in an embodiment of the present application;
[0063] FIG16 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] The technical solution in this application will be described below with reference to the accompanying drawings.
[0065] In the description of the embodiments of this application, unless otherwise specified, " / " represents or. For example, A / B can represent A or B. "And / or" in this article is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "plurality" means two or more than two.
[0066] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0067] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0068] First, some of the terms used in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0069] 1. Dynamic Range
[0070] In the field of image processing, dynamic range refers to the ratio of the brightest light intensity to the darkest light intensity in an image. Dynamic range is one of the most important dimensions for image quality evaluation.
[0071] Natural scenes have a relatively large dynamic range, typically reaching 10^9. The human eye also has a wide dynamic range, generally considered to be at least 10^6. However, because image pixel values are typically recorded using 8-bit integer data, traditional images and video content can only distinguish 256 different brightness levels. When displayed on a typical monitor, the brightness dynamic range is approximately 10^3.
[0072] Therefore, from the perspective of dynamic range, there is a significant gap between traditional images and videos and the real scenes seen by the human eye, which limits the richness of video content. This is also an important reason why traditional images and videos always look different from the real world.
[0073] 2. High-Dynamic Range (HDR) video
[0074] HDR, also known as "High Dynamic Range Rendering," is designed to make the display image closer to the real world as observed by the human eye. As we know, video is actually a sequence of static images. Compared to ordinary images, High-Dynamic Range Images (HDRI) can provide a wider dynamic range and image detail. Based on Low-Dynamic Range (LDR) images with different exposure times, the final HDR image is synthesized using the LDR image with the best detail corresponding to each exposure time, which can better reflect the visual effects in the real environment. HDR video is composed of these HDRI frames, making the image more vivid and textured than ordinary video.
[0075] 3. Standard Dynamic Range (SDR)
[0076] SDR is a traditional technology for processing brightness and color values in images and is a common color display method. Traditional videos are generally referred to as SDR videos. To display the original scene within a limited brightness range, SDR videos severely compress the brightness range of natural scenes, resulting in a significant lack of contrast in the video content.
[0077] RGB (Red, Green, Blue) color space
[0078] Also known as the RGB domain, it refers to a color model related to the structure of the human visual system. According to the structure of the human eye, all colors are considered to be different combinations of red, green, and blue.
[0079] 5. Pixel Value
[0080] Refers to the set of color components corresponding to each pixel in a color image in the RGB color space. For example, each pixel corresponds to a set of three primary color components, where the three primary color components are red (R), green (G), and blue (B).
[0081] 6. YUV color space
[0082] YUV domain, also known as the YUV domain, refers to a color encoding method. Y represents luminance (or luma), while U and V represent chrominance (or chroma). While the RGB color space focuses on the human eye's perception of color, the YUV color space emphasizes visual sensitivity to brightness. RGB and YUV color spaces are convertible.
[0083] 7. Brightness
[0084] Brightness is the luminous flux emitted by a light source unit in a given direction, per unit area, and per unit solid angle. The symbol for brightness is L, and the unit is nit.
[0085] 8. Grayscale
[0086] Generally speaking, each dot on an LCD screen, known as a pixel, is composed of three sub-pixels: red, green, and blue (RGB). Each sub-pixel can display different brightness levels depending on the light source behind it. Grayscale represents the range of brightness levels from darkest to brightest. The more grayscale levels there are, the more detailed the image can be. For example, a typical screen can display 256 brightness levels, which we call 256 grayscales.
[0087] 9. Tone Mapping
[0088] Tone mapping is a computer graphics technique for approximating the display of high dynamic range images on media with limited dynamic range.
[0089] 10. Brightness histogram
[0090] A brightness histogram is a quantitative tool used to measure image brightness. Image brightness is typically categorized on a scale of 0 to 255. The horizontal axis of the brightness histogram represents the brightness (brightness value) between 0 and 255, while the vertical axis represents the number of pixels in the image with the corresponding brightness. The leftmost value on the horizontal axis is 0, while the rightmost value is 255. 0 represents the darkest area, and 255 represents the brightest area. The values in between represent shades of gray, with larger values indicating brighter colors.
[0091] The above is a brief introduction to the nouns involved in the embodiments of this application, and no further details will be given below.
[0092] High-Dynamic Range (HDR) video, compared to Standard Dynamic Range (SDR) video, offers clearer gradations of light and dark, richer image details, and a more realistic reproduction of real scenes. In other words, HDR video displays video content with a higher dynamic range, providing users with a better viewing experience.
[0093] With the development of HDR technology, the brightness capabilities of display devices are becoming increasingly higher to better play HDR videos. The brightness capabilities of a display device are usually expressed in terms of its dynamic range. For example, the higher the dynamic range of a display device, the stronger its brightness capabilities.
[0094] While more and more display devices have strong brightness capabilities, a large number of SDR videos have accumulated over the course of video technology development. This means that currently, mainstream videos are still SDR videos (or, in other words, the dynamic range of mainstream videos is still SDR). As a result, these SDR videos cannot realistically reproduce real scenes on display devices with strong brightness capabilities, nor can they display video content with a dynamic range higher than that of SDR. This results in poor display quality on these SDR videos on display devices with strong brightness capabilities.
[0095] At the same time, for display devices with strong brightness capabilities, the dynamic range of the display device exceeds the dynamic range of SDR video. Using this display device to display SDR video will result in the inability to fully utilize the brightness capability of the display device, resulting in a waste of display resources. For example, SDR video is usually stored with a data width of 8 bits (binary digit, bit), and the dynamic range it can express is about 100 nits. For display devices with strong brightness capabilities, the dynamic range that the display device can express can reach more than 1000 nits. Therefore, using such a display device to display SDR video will result in the inability to fully utilize the brightness capability of the display device, resulting in a waste of display resources.
[0096] In related technologies, the contrast of SDR videos is sometimes improved by increasing the color contrast of video frames, but this still fails to improve the dynamic range of SDR videos, so the above-mentioned problem still exists. To this end, an embodiment of the present application provides a method for processing videos. This method determines the actual scalable dynamic range of the SDR video by determining the brightness information of the SDR video and combining it with the brightness capability of the current display device. The SDR video is tone-mapped based on the scalable dynamic range, thereby truly improving the dynamic range of the SDR video. This improves the display effect when the SDR video with the improved dynamic range is displayed on the display device.
[0097] Currently, using display devices to play videos has become a daily routine. For example, mobile phones have increasingly higher brightness capabilities, making them better suited for playing HDR videos. However, the mainstream video format is still SDR. These SDR videos cannot realistically reproduce real scenes on display devices with high brightness capabilities, nor can they display video content with a dynamic range higher than that of SDR. Consequently, these SDR videos display poorly on these devices.
[0098] Please refer to FIG1 , which is a diagram illustrating a frame of an SDR video played on a display device in a related art according to an exemplary embodiment of the present application.
[0099] As shown in Figure 1, when an SDR video is played on a display device, the frame image has low brightness and low dynamic range, and cannot realistically reproduce the real scene, resulting in poor display effect.
[0100] The video processing method provided by the embodiments of the present application can, for example, determine the brightness information of the SDR video and combine it with the brightness capability of the current display device (such as a mobile phone) to determine the actual scalable dynamic range of the SDR video. Based on this scalable dynamic range, tone mapping is performed on the SDR video, which can truly improve the dynamic range of the SDR video. This improves the display effect when the SDR video with the increased dynamic range is displayed on the display device (such as a mobile phone).
[0101] Please refer to FIG. 2 , which is a diagram showing a frame of an SDR video after processing when it is played on a display device according to an exemplary embodiment of the present application.
[0102] As shown in Figure 2, when the processed / enhanced SDR video is played on a display device, the brightness and dynamic range of this frame image are significantly improved compared to the image in Figure 1, which can realistically reproduce the real scene and improve the display effect.
[0103] First, the application scenarios of the embodiments of the present application are briefly described.
[0104] Please refer to Figure 3, which is a schematic diagram of an application scenario provided by an embodiment of the present application. The video processing method provided in this application can be used to improve the dynamic range of SDR video.
[0105] In one example, the display device is a mobile phone. When the display device detects that a user has clicked a video application icon on an application interface, the video application can be launched and the video application interface can be displayed. The user can click a favorite video in the video application interface and select full-screen playback mode. The display device then displays a graphical user interface (GUI) as shown in FIG3(a). This GUI can be referred to as a video playback interface.
[0106] As shown in (b) of Figure 3, in the video playback interface, the user can use a touch object (such as a user's finger or a stylus, etc.) to tap the left side of the screen (only the left side is shown in this embodiment, in a possible implementation, the user can also tap the right side of the screen) and slide upwards, and the display device responds to the user's touch operation and displays a brightness control as shown in (c) of Figure 3 on the video playback interface. At the same time, the screen brightness of the display device is enhanced (not shown in (c) of Figure 3). At the same time, the frame image displayed on the video playback interface is processed by the method for processing video provided in an embodiment of the present application, such as determining the brightness information corresponding to the frame image, the dynamic range corresponding to the frame image, and determining the dynamic range of the current display device. According to the brightness information and dynamic range corresponding to the frame image and the dynamic range of the current display device, the frame image is tone mapped to obtain an enhanced image corresponding to the frame image.
[0107] For ease of understanding, please refer to Figure 4, which is a schematic diagram of an enhanced image in another application scenario provided by an embodiment of the present application. The enhanced image displayed on the video playback interface shown in Figure 4 has a significantly improved dynamic range compared to the image displayed on the video playback interface in Figure 3, and can realistically reproduce the real scene, improving the display effect.
[0108] In another example, the display device is still used as a mobile phone for illustration. When the display device detects that the user clicks the icon of the video application on the application interface, the video application can be started and the video application interface can be displayed. The user can click on his favorite video in the video application interface and select the full-screen playback mode. At this time, the display device displays a graphical user interface (GUI) as shown in (a) of Figure 3. The GUI can be called a video playback interface. The video playback interface may include an HDR control. When the display device detects that the user clicks the HDR control, the HDR mode is turned on and the video is played in the HDR mode.
[0109] In the process of playing the video in HDR mode, each frame image in the video is processed by the method for processing the video provided by the embodiment of the present application. For example, the brightness information corresponding to each frame image, the dynamic range corresponding to each frame image, and the dynamic range of the current display device are determined. According to the brightness information and dynamic range corresponding to each frame image and the dynamic range of the current display device, tone mapping is performed on each frame image to obtain an enhanced image corresponding to each frame image. Multiple frames of enhanced images constitute an enhanced video, which is a video with improved dynamic range. Therefore, when the enhanced video is played on the display device (such as a mobile phone), the display effect is improved. For example, compared with ordinary videos, the enhanced video has more vivid and vivid colors and clearer details, can provide more dynamic range and image details, greatly improve the light and dark contrast of the picture details, and better reflect the visual effects in the real environment.
[0110] In another example, the display device is still taken as a mobile phone for illustration. If the user has turned on the automatic brightness adjustment function in the display device in advance, the display device will automatically adjust the brightness when playing the video. Whenever the display device detects a change in brightness, each frame image at that brightness is processed by the method for processing video provided in an embodiment of the present application. For example, the brightness information corresponding to each frame image at that brightness, the dynamic range corresponding to each frame image at that brightness, and the dynamic range of the current display device are determined. According to the brightness information, dynamic range and dynamic range corresponding to each frame image at that brightness and the dynamic range of the current display device, tone mapping is performed on each frame image at that brightness to obtain an enhanced image corresponding to each frame image at that brightness. Multiple frames of enhanced images at that brightness constitute an enhanced video at that brightness, and the enhanced video at that brightness is a video after improving the dynamic range. Therefore, when the enhanced video at that brightness is played on the display device (such as a mobile phone), the display effect is improved.
[0111] It should be understood that the scenarios shown in Figures 3 and 4 are examples of application scenarios and do not limit the application scenarios of this application. The method for processing videos provided in the embodiments of this application can be applied but not limited to the following scenarios:
[0112] Video calls, video conferencing applications, long and short video applications, live video applications, online video courses, smart camera applications, system camera recording function for video, system camera shooting function for photo taking, video surveillance, smart cat-eye and other scenarios.
[0113] The following is a detailed introduction to the method for processing video provided in the embodiments of the present application in conjunction with the drawings in the specification.
[0114] The video processing method provided in the embodiments of the present application can be used in video-related application scenarios, where the video-related application scenarios may include a display device playing a video online; or a display device recording a video; or a display device performing a live video broadcast, etc. This is merely an example and is not intended to be limiting.
[0115] Exemplarily, the method for processing video provided in the embodiment of the present application is described by taking an application scenario related to video including online video playback on a display device as an example.
[0116] Please refer to Figure 5, which is a flow chart of a method for processing a video provided by an embodiment of the present application. As shown in Figure 5, the method for processing a video includes the following steps S101-S104.
[0117] S101: Obtain a decoded video of an original video.
[0118] Exemplarily, the original video may include SDR video, video in a call, video in a conference application, video in a live broadcast application, video in a long and short video application, video in surveillance, video recorded by a system camera recording function, etc.
[0119] An original video is obtained and decoded to obtain a decoded video of the original video, wherein the decoded video includes a plurality of video frame images.
[0120] The decoded video may include two or more video frame images. There is no limitation on the number of video frame images in the embodiment of the present application.
[0121] The embodiment of the present application is described by taking the original video including the SDR video as an example. For example, the SDR video can be a video played online by a display device. The SDR video is decoded to obtain a decoded video of the SDR video. For example, the SDR video can be decoded by a graphics processing unit (GPU) in the display device to obtain a decoded video of the SDR video. For another example, the SDR video can be decoded by a video processing card in the display device to obtain a decoded video of the SDR video. For another example, the SDR video can be decoded by a video transcoding server to obtain a decoded video of the SDR video. This is merely an exemplary description and is not intended to be limiting.
[0122] S102: Determine brightness information corresponding to each video frame image.
[0123] Exemplarily, the decoded video is processed, for example, to determine the brightness information of each video frame image contained in the decoded video. It is understood that there are multiple ways to determine the brightness information corresponding to each video frame image, and each way is described below.
[0124] In one possible implementation, the pixel format of the video frame images contained in the decoded video can be first determined, and then the brightness information corresponding to each video frame image can be determined based on the pixel format. The pixel format can include a YUV format and an RGB format. It is understood that for the same decoded video, the pixel format of each video frame image contained in the decoded video is consistent.
[0125] If the pixel format of the video frame image is YUV format, and Y represents the brightness of the video frame image, the Y value corresponding to each video frame image can be directly obtained. The Y value represents the brightness information corresponding to a video frame image.
[0126] If the pixel format of the video frame image is RGB format, the brightness information of the video frame image can be determined using the following formula (1). L = 0.299*R + 0.567*G + 0.114*B, (1)
[0127] In the above formula (1), L represents the brightness information of the video frame image, R represents the value of the red component of the video frame image, G represents the value of the green component of the video frame image, and B represents the value of the blue component of the video frame image.
[0128] If the pixel format of the video frame image is in YUV format, in order to facilitate the display device to determine the brightness information of the video frame image, the pixel format of the video frame image can also be converted from YUV format to RGB format, and then the brightness information of the video frame image can be determined by the above formula (1). Among them, the pixel format of the video frame image is converted from YUV format to RGB format, which can be achieved by the following formula (2). R = Y + 1.4075 * VG = Y - 0.3455 * U - 0.7169 * VB = Y + 1.779 * U, (2)
[0129] In the above formula (2), R represents the value of the red component of the video frame image, G represents the value of the green component of the video frame image, B represents the value of the blue component of the video frame image, Y represents the brightness of the video frame image, and U and V represent the chromaticity of the video frame image.
[0130] It should be understood that in this implementation, the brightness information can be digitized into 0-255, that is, the brightness level of each pixel in the video frame image can be represented by 0-255.
[0131] To intuitively display the brightness information corresponding to a video frame image, please refer to Figure 6, which is a brightness information display diagram provided in an embodiment of the present application. Figure 6 (a) shows an arbitrary video frame image, and Figure 6 (b) shows a display diagram of the brightness information corresponding to the video frame image.
[0132] In another possible implementation, a brightness histogram can be used to determine the brightness information corresponding to each video frame image. For example, a preset code can be run to determine the brightness histogram corresponding to each video frame image, and statistics can be performed on each brightness histogram to calculate the histogram distribution of each video frame image. The histogram distribution of each video frame image is used to represent the brightness information corresponding to each video frame image.
[0133] It should be understood that in this implementation, the brightness level of each pixel in the video frame image can be represented by the statistical frequency of each grayscale.
[0134] Please refer to Figure 7, which is another brightness information display diagram provided by an embodiment of the present application. Figure 7 (a) shows a brightness information histogram in one form of expression, and Figure 7 (b) shows a brightness information histogram in another form of expression. The horizontal axes in Figure 7 (a) and Figure 7 (b) both represent brightness (i.e., brightness value), and the vertical axes both represent the number of pixels in the video frame image corresponding to the brightness of the horizontal axis. Among them, the leftmost side of the horizontal axis represents the darkest part of the video frame image, and the rightmost side of the horizontal axis represents the brightest part of the video frame image. The values in the middle represent grays of different brightness, and the larger the value, the brighter it is.
[0135] It is worth noting that, in order to improve efficiency, if the display device detects that multiple consecutive video frame images are similar, it determines the brightness information of one of the multiple consecutive video frame images, and uses the brightness information of the video frame image as the brightness information of the remaining similar video frame images. For example, if the display device detects that four consecutive video frame images are similar (such as the similarity is greater than or equal to a preset similarity threshold), the display device determines the brightness information of the first video frame image among the four video frame images, and uses the brightness information of the first video frame image as the brightness information of the other three video frame images. This is only an exemplary description and is not limited to this.
[0136] It should be understood that the display device plays videos online by displaying images frame by frame in chronological order. When the display device determines the brightness information corresponding to each video frame image, it also determines the brightness information corresponding to each video frame image frame by frame in chronological order.
[0137] S103: Obtain the brightness capability of the display device and the dynamic range of the original video, and calculate the scalable dynamic range of each video frame image according to the brightness capability of the display device and the dynamic range of the original video.
[0138] For example, the brightness capability of the display device is obtained. The brightness capability can be represented by the dynamic range of the display device. It should be understood that each dynamic range corresponds to a peak brightness. A higher peak brightness indicates a higher dynamic range, that is, a stronger brightness capability of the display device.
[0139] It should be understood that the brightness capability of a display device is adjustable, for example, the current brightness capability of the display device can be flexibly adjusted according to different user needs. In other words, the brightness capability of a display device may be the same or different in different scenarios. For example, when a display device plays a video online, the brightness capability of the display device may be the same or different when playing different scenes.
[0140] Optionally, in one possible implementation, obtaining the brightness capability of a display device may involve obtaining the maximum brightness capability of the display device, that is, obtaining the maximum brightness capability that the display device can achieve. For example, if the dynamic range that a display device can express reaches 1000 nits, then the maximum brightness capability of the display device is obtained as 1000 nits, or in other words, the brightness value of the dynamic range of the display device is obtained as 1000 nits. This is merely an example and is not intended to be limiting.
[0141] Optionally, in a possible implementation, if the user has turned on the automatic brightness adjustment function in the display device in advance, the display device can automatically adjust the brightness. For example, the display device will automatically adjust the brightness when playing a video online. In this implementation, obtaining the brightness capability of the display device may be to obtain the current brightness capability of the display device in real time. For example, in one example, the display device automatically adjusts the brightness to 800nit when playing a video online, and the current brightness capability of the display device obtained in real time is 800nit. For another example, in another example, the display device automatically adjusts the brightness to 400nit when playing a video online, and the current brightness capability of the display device obtained in real time is 400nit. This is only an exemplary description and is not limited to this.
[0142] Optionally, in one possible implementation, the user can adjust the brightness according to their needs. For example, when the display device plays a video online, in the current video playback interface, the user can tap the left side of the screen (in one possible implementation, the user can also tap the right side of the screen) with a touch object (such as a user's finger or a stylus, etc.) and slide upwards, and the display device responds to the user's touch operation, and the screen brightness is enhanced. If the user taps the left side (or right side) of the screen with a touch object and slides downwards, the display device responds to the user's touch operation, and the screen brightness dims. In this implementation, obtaining the brightness capability of the display device may be to obtain the brightness capability of the display device after adjusting the brightness. For example, the display device adjusts the brightness to 600nit in response to the user's touch operation, and the current brightness capability of the display device is obtained to be 600nit. This is only an exemplary description and is not limited to this.
[0143] Exemplarily, the dynamic range of the original video is obtained. In one example, obtaining the dynamic range of the original video refers to obtaining the dynamic range corresponding to the original video. In another example, in order to improve the accuracy of the scalable dynamic range and thus improve the final display effect, the actual brightness statistical level of each video frame image can be considered. Therefore, obtaining the dynamic range of the original video can be to obtain the dynamic range of the original video in the scene corresponding to each video frame image. In an embodiment of the present application, the dynamic range of the SDR video can be obtained, or the dynamic range of the SDR video in the scene corresponding to each video frame image can be obtained.
[0144] It should be understood that the display device provided in the embodiment of the present application has the ability to automatically obtain the dynamic range of the original video in the scene corresponding to each video frame image, that is, the display device can automatically obtain the dynamic range in the scene corresponding to each video frame image. In one possible implementation, the display device obtains that the dynamic range of the SDR video in the scene corresponding to the first video frame image is 200nit, or in other words, the display device obtains that the brightness value of the dynamic range in the scene corresponding to the first video frame image is 200nit. The first video frame image is used to represent any one of the multiple video frame images contained in the SDR video. This is only an exemplary description and is not limited to this.
[0145] For example, the scalable dynamic range of each video frame image is calculated using the brightness capability of the display device and the dynamic range of the original video. It should be understood that the acquired brightness capability of the display device and the dynamic range of the original video can both be represented by brightness values, and the difference between the two brightness values is calculated and used as the scalable dynamic range of the video frame image.
[0146] For example, the current brightness capability of the display device is obtained to be 600nit, the brightness value of the dynamic range in the scene corresponding to the first video frame image is 200nit, the difference between 600nit and 200nit is calculated to be 400nit, and the difference, i.e. 400nit, is used as the expandable dynamic range of the first video frame image.
[0147] It is worth noting that, in order to improve efficiency, if the display device detects that multiple consecutive video frame images are similar, the scalable dynamic range of one of the multiple consecutive video frame images is calculated, and the scalable dynamic range of the video frame image is used as the scalable dynamic range of the remaining similar video frame images. For example, if the display device detects that three consecutive video frame images are similar (such as the similarity is greater than or equal to the preset similarity threshold), the display device calculates the scalable dynamic range of the first video frame image among the three video frame images, and uses the scalable dynamic range of the first video frame image as the scalable dynamic range of the other three video frame images. This is only an exemplary description and is not limited to this.
[0148] It should be understood that when a display device plays a video online, it displays images frame by frame in chronological order. When the display device calculates the scalable dynamic range of each video frame image, it also calculates the scalable dynamic range of each video frame image frame by frame in chronological order.
[0149] S104 , performing tone mapping on each video frame image using brightness information corresponding to each video frame image and an expandable dynamic range of each video frame image, to obtain an enhanced image corresponding to each video frame image.
[0150] Exemplarily, the luminance information corresponding to each video frame image is used to determine an extended region for each video frame image. This extended region includes different areas in different implementations and is used for tone mapping, which will be described in detail in the following embodiments. Based on the scalable dynamic range of each video frame image, tone mapping is performed on the extended region of each video frame image. Based on the mapping results, an enhanced image corresponding to each video frame image is generated.
[0151] Compared with the corresponding video frame image, the enhanced image has brighter and more vivid colors and clearer details. It can provide more dynamic range and image details, greatly improve the light and dark contrast of the picture details, and better reflect the visual effects in the real environment.
[0152] The method for processing video provided in an embodiment of the present application determines the brightness information corresponding to each video frame image in the decoded video of the original video; calculates the scalable dynamic range of each video frame image through the brightness capability of the display device and the dynamic range of the original video; and then uses the brightness information corresponding to each video frame image and the scalable dynamic range of each video frame image to perform tone mapping on each video frame image to obtain an enhanced image corresponding to each video frame image. When determining the scalable dynamic range of each video frame image, the brightness capability of the display device and the dynamic range of the original video are fully considered, and then uses the scalable dynamic range to perform tone mapping on each video frame image, so that the dynamic range of the video frame image is truly improved, that is, the enhanced image is the image after the dynamic range is improved. It can be seen from this that the dynamic range of the video composed of multiple enhanced images is also improved, so that when the video after the dynamic range is improved is displayed on the display device, the display effect is improved and the user experience is improved.
[0153] Optionally, in a possible implementation, the method for processing a video provided in the embodiment of the present application may further include S105 on the basis of including S101 to S104, as follows:
[0154] S105: Generate an enhanced video according to the multiple enhanced images.
[0155] The enhanced video is an extended dynamic range video, that is, the enhanced video is a video with an extended dynamic range. The enhanced image corresponding to each video frame is arranged in chronological order to obtain an enhanced video, which can be played on a display device. In this embodiment, after processing a video frame, the enhanced image corresponding to the video frame is displayed on the display device. Due to the high processing speed of the display device, the user can experience a smooth video viewing experience with an enhanced dynamic range, thereby improving the user experience.
[0156] The following describes in detail a process for generating an enhanced image corresponding to a video frame image.
[0157] Please refer to Figure 8, which is a schematic diagram of a process for generating an enhanced image corresponding to a video frame image according to an embodiment of the present application. As shown in Figure 8, the above S104 may include S1041-S1043.
[0158] S1041 : Determine a standard dynamic range and an extended dynamic range corresponding to each video frame image by using brightness information corresponding to each video frame image.
[0159] Exemplarily, the extended area in the above S104 in this implementation includes a standard dynamic area and an extended dynamic area.
[0160] Among them, the extended dynamic area is an area divided in a video frame image, and the extended dynamic area includes several pixel points in the video frame image. The pixel values of the pixel points in the extended dynamic area do not need to be adjusted, that is, the pixel values of the pixel points in the extended dynamic area maintain the original pixel values.
[0161] The standard dynamic area is another area divided in the video frame image. The standard dynamic area also contains several pixel points in the video frame image. Unlike the pixel points in the extended dynamic area, the pixel values of the pixel points in the standard dynamic area need to be adjusted.
[0162] For example, if the pixel format of the video frame image contained in the decoded video is first determined, and then the brightness information corresponding to each video frame image is determined based on the pixel format, then in this implementation, the brightness information can be digitized to 0-255, that is, the brightness level of each pixel in the video frame image can be represented by 0-255. Thus, the pixels of each video frame image can be divided according to the preset threshold and the brightness information of each video frame image. For example, the pixels of each video frame image are divided into pixels whose brightness information is less than the preset threshold, and pixels whose brightness information is greater than or equal to the preset threshold. It should be understood that the area composed of pixels whose brightness information is less than the preset threshold is the standard dynamic area, and the area composed of pixels whose brightness information is greater than or equal to the preset threshold is the extended dynamic area.
[0163] The preset threshold value may be set by the user, calculated by grayscale ratio, or determined by a machine learning model. This is merely an example and is not intended to be limiting.
[0164] In this implementation, a preset threshold is used to divide the two areas into a standard dynamic area and an extended dynamic area. In order to make the final improved dynamic range effect better, in a possible implementation, a finer division can be achieved through multiple preset threshold ranges, that is, the pixels of each video frame image are divided by multiple preset threshold ranges and the brightness information of each video frame image. For example, for the pixels of each video frame image, the pixels whose brightness information belongs to the first preset threshold range are divided into pixels of the standard dynamic area, and the pixels whose brightness information belongs to the second preset threshold range are divided into pixels of the extended dynamic area. The first preset threshold range and the second preset threshold range are different. It should be understood that the brightness information of the pixels belonging to the first preset threshold range is less than the brightness information of the pixels belonging to the second preset threshold range.
[0165] S1042: Determine a first coefficient corresponding to the standard dynamic range and a second coefficient corresponding to the extended dynamic range.
[0166] Exemplarily, the first coefficient is smaller than the second coefficient.
[0167] The first coefficient corresponding to the standard dynamic range is determined based on the dynamic range of the original video and the expandable dynamic range of each video frame image. For example, let the dynamic range of the original video be denoted as A, and the expandable dynamic range of each video frame image be denoted as B. The quotient between A and B is calculated, and this quotient is used as the first coefficient corresponding to the standard dynamic range. Typically, the value of A is smaller than the value of B. Therefore, the first coefficient corresponding to the standard dynamic range is less than 1.
[0168] It is worth noting that, in the embodiment of the present application, the dynamic range of the original video may include the dynamic range of the SDR video in the scene corresponding to each video frame image.
[0169] For example, since the pixel values of the pixels in the extended dynamic area do not need to be adjusted, the second coefficient corresponding to the extended dynamic area can be set to 1.
[0170] It should be understood that when the overall brightness of the display device is increased, the brightness of the video frame image is also increased as a whole. At this time, the contrast or dynamic range of the video frame image is not improved. When the bright areas in the video frame image are retained and the brightness of other areas is reduced, the dynamic range of the video frame image can be improved. Therefore, the second coefficient corresponding to the extended dynamic area can be set to 1, so that the pixel values of the pixels in the extended dynamic area can maintain the original pixel values. The first coefficient corresponding to the standard dynamic area is less than 1, so that the pixel values of the pixels in the standard dynamic area can be reduced.
[0171] Optionally, in one possible implementation, the pixel values of pixels in the standard dynamic range do not need to be adjusted, so the first coefficient corresponding to the standard dynamic range can be set to 1. The pixel values of pixels in the extended dynamic range are increased, so the second coefficient corresponding to the extended dynamic range can be set to a value greater than 1. For example, illustratively, the dynamic range of the original video is recorded as A, and the expandable dynamic range of each video frame image is recorded as B. The quotient between B and A is calculated, and the quotient is used as the second coefficient corresponding to the standard dynamic range. Typically, the value of A is smaller than the value of B. Therefore, the second coefficient corresponding to the extended dynamic range is greater than 1.
[0172] It should be understood that when the overall brightness of the display device is increased, the brightness of the video frame image is also increased as a whole. At this time, the contrast or dynamic range of the video frame image is not improved. When the brightness of other areas in the video frame image is retained, the brightness of the bright areas in the video frame image is increased, which can improve the dynamic range of the video frame image. Therefore, the second coefficient corresponding to the extended dynamic area is greater than 1, which can increase the pixel value of the pixel points in the extended dynamic area. The first coefficient corresponding to the standard dynamic area is set to 1, which can keep the original pixel value of the pixel points in the standard dynamic area.
[0173] S1043 : Generate an enhanced image corresponding to each video frame image according to the standard dynamic range, the first coefficient, the extended dynamic range, and the second coefficient of each video frame image.
[0174] Tone mapping is performed on the pixels in the standard dynamic range according to the first coefficient, and tone mapping is performed on the pixels in the extended dynamic range according to the second coefficient, so as to generate an enhanced image corresponding to each video frame image.
[0175] Exemplarily, for each pixel point in the standard dynamic range, the product between the original pixel value of the pixel point and the first coefficient is calculated, and the product is used as the new pixel value of the pixel point.
[0176] Optionally, in a possible implementation, for each pixel point in the extended dynamic area, the product between the original pixel value of the pixel point and the second coefficient is calculated, and the product is used as the new pixel value of the pixel point.
[0177] Optionally, in a possible implementation, since the second coefficient corresponding to the extended dynamic area is 1, for each pixel point in the extended dynamic area, the original pixel values of these pixel points do not need to be processed, and the original pixel values of these pixel points can be saved.
[0178] For each video frame, an enhanced image corresponding to the video frame is generated based on the pixel points after the pixel values are adjusted. It should be understood that in this embodiment, if the pixel values of all pixels in the video frame are adjusted, the enhanced image corresponding to the video frame is generated based on all the pixels after the pixel values are adjusted. If the pixel values of some pixels in the video frame are adjusted, the enhanced image corresponding to the video frame is generated based on both the pixels after the pixel values are adjusted and the pixels whose pixel values are not adjusted. This is merely an example and is not intended to be limiting.
[0179] It should be understood that in this implementation, tone mapping is performed on each video frame image, which means adjusting the pixel values of pixels in different areas (such as the standard dynamic area and the extended dynamic area) according to different coefficients (such as the first coefficient and the second coefficient).
[0180] In this implementation, the brightness information corresponding to each video frame image is used to determine the adjustable areas in each video frame image, namely the standard dynamic area and the extended dynamic area; different coefficients are determined for the standard dynamic area and the extended dynamic area, and then the pixel values of the pixels in different areas (such as the standard dynamic area and the extended dynamic area) are adjusted according to the different coefficients (such as the first coefficient and the second coefficient). While maintaining the pixel values of the pixels in the extended dynamic area, the pixel values of the pixels in the standard dynamic area are reduced. Only in this way can the dynamic range of the enhanced image generated be truly improved. Then, the dynamic range of the video composed of multiple enhanced images is also improved, so that when the video with the improved dynamic range is displayed on a display device, the display effect is improved and the user experience is improved.
[0181] Optionally, if the second coefficient corresponding to the extended dynamic area is greater than 1 and the first coefficient corresponding to the standard dynamic area is 1, then for each pixel point in the extended dynamic area, the product between the original pixel value of the pixel point and the second coefficient is calculated, and the product is used as the new pixel value of the pixel point.
[0182] For each pixel in the standard dynamic range, the product of the original pixel value and the first coefficient is calculated, and the product is used as the new pixel value of the pixel. Alternatively, since the first coefficient corresponding to the standard dynamic range is 1, the original pixel values of each pixel in the standard dynamic range do not need to be processed and can be saved.
[0183] In this implementation, the brightness information corresponding to each video frame image is used to determine the adjustable areas in each video frame image, namely the standard dynamic area and the extended dynamic area; different coefficients are determined for the standard dynamic area and the extended dynamic area, and then the pixel values of the pixels in different areas (such as the standard dynamic area and the extended dynamic area) are adjusted separately according to the different coefficients (such as the first coefficient and the second coefficient). While maintaining the pixel values of the pixels in the standard dynamic area, the pixel values of the pixels in the extended dynamic area are increased, so that the dynamic range of the enhanced image generated is truly improved. Then, the dynamic range of the video composed of multiple enhanced images is also improved, so that when the video with the increased dynamic range is displayed on a display device, the display effect is improved, and the user experience is improved.
[0184] In the above implementation, the pixel values of pixels in different areas (such as the standard dynamic area and the extended dynamic area) are adjusted according to different coefficients (such as the first coefficient and the second coefficient). In a possible implementation, the coefficient corresponding to each pixel can also be calculated one by one according to the actual brightness level of each pixel in the video frame image. For each pixel, the pixel value of the pixel is adjusted according to the coefficient corresponding to the pixel. For example, the product between the original pixel value of the pixel and the coefficient is calculated, and the product is used as the new pixel value of the pixel. After all the pixels are adjusted, an enhanced image corresponding to the video frame image is generated based on all the pixels after the pixel values are adjusted. In this implementation, each pixel is adjusted individually in a targeted manner, which can better improve the dynamic range of the enhanced image generated in the end, while improving the quality of the enhanced image generated in the end.
[0185] Another process for generating an enhanced image corresponding to a video frame image is described in detail below.
[0186] Please refer to Figure 9, which is another schematic diagram of a process for generating an enhanced image corresponding to a video frame image according to an embodiment of the present application. As shown in Figure 9, the above-mentioned S104 may include S1044-S1046. It is worth noting that S1044-S1046 are parallel to the above-mentioned S1041-S1043. Depending on the actual situation, either S1041-S1043 or S1044-S1046 can be selected for execution, and S1044-S1046 is not necessarily executed after S1041-S1043.
[0187] S1044: Determine a first region, a second region, and a third region corresponding to each video frame image by using brightness information corresponding to each video frame image.
[0188] Illustratively, the extended area in the above S104 in this implementation includes a first area, a second area, and a third area.
[0189] The first region may represent a low grayscale region, the second region may represent a medium grayscale region, and the third region may represent a high grayscale region.
[0190] At least one of the first, second, and third regions is an area that needs to be re-tone mapped. In one possible implementation, the current brightness of the first and second regions can be maintained, and the third region can be used as the area that needs to be re-tone mapped. For example, the current brightness of the low grayscale region and the medium grayscale region is maintained, and the brightness of the high grayscale region is improved. In another possible implementation, the first, second, and third regions are all used as areas that need to be re-tone mapped. For example, the brightness of the low grayscale region, the medium grayscale region, and the high grayscale region is all improved. It should be understood that the magnitude of the brightness improvement for different regions may be the same or different. In another possible implementation, the current brightness of the second and third regions can be maintained, and the brightness of the first region can be reduced. For example, the current brightness of the medium grayscale region and the high grayscale region is maintained, and the brightness of the low grayscale region is reduced.
[0191] For example, if the brightness information corresponding to each video frame image is determined by a brightness histogram, the brightness histogram of each video frame image can be divided by T1 and T2. For ease of understanding, please refer to Figure 10, which is a schematic diagram of region division provided in an embodiment of the present application. As shown in Figure 10, the brightness distribution represented by the brightness histogram is divided into three regions by T1 and T2, such as the first region: [0, T1), the second region [T1, T2], and the third region (T2, 255).
[0192] Among them, the values of T1 and T2 can be set by the user or determined by a machine learning model. During the training process, the machine learning model learns the grayscale distribution of the pixels in the brightness histogram of the video frame image and determines two numerical values based on the distribution, which are the values of T1 and T2 respectively. For example, in this embodiment, the brightness histogram of the video frame image can be input into the machine learning model, and the machine learning model analyzes and processes the brightness histogram and outputs the values of T1 and T2. This is only an exemplary description and is not limited to this.
[0193] In order to more intuitively show the division of brightness areas, please refer to Figure 11, which is a schematic diagram of brightness area division provided in an embodiment of the present application. As shown in Figure 11, the light gray area (including the unevenly distributed light gray points shown in Figure 11) is divided into a first area, namely a low grayscale area, the black area (including all black areas shown in Figure 11) is divided into a second area, namely a medium grayscale area, and the dark gray area is divided into a third area, namely a high grayscale area. It is worth noting that the colors shown in Figure 11 are only for distinguishing the various brightness areas and do not represent the true colors of the video frame image.
[0194] In this implementation, T1 and T2 are used to divide the low grayscale area, the medium grayscale area and the high grayscale area. In order to make the final improved dynamic range effect better, in a possible implementation, a finer division can be achieved through multiple values (such as values similar to T1 and T2), that is, the brightness distribution represented by the brightness histogram is divided into multiple areas through multiple values.
[0195] S1045: Determine a first adjustment strategy corresponding to the first area, a second adjustment strategy corresponding to the second area, and a third adjustment strategy corresponding to the third area.
[0196] Please refer to Figure 12, which is a schematic diagram of tone mapping provided in an embodiment of the present application. The horizontal axis in Figure 12 represents the value after normalization to 255 grayscale, and the vertical axis represents the grayscale value after tone mapping. For example, SDR is generally a linear mapping, such as the black straight line shown in Figure 12. The function corresponding to the black straight line can be: y = k * x.
[0197] It should be understood that the first, second, and third regions correspond to different lines in Figure 12. As shown in Figure 12, the first region corresponds to line GH, the second region corresponds to line HI, and the third region corresponds to line IJ. The function corresponding to line GH is: y = k*x.
[0198] The first adjustment strategy can be used to determine the adjusted grayscale values of pixels in the low grayscale region. For example, a third coefficient corresponding to the low grayscale region is determined based on the scalable dynamic range of each video frame image; tone mapping is performed on the pixels in the low grayscale region based on the third coefficient to obtain the adjusted grayscale values of the pixels in the low grayscale region.
[0199] For the first region, the first adjustment strategy can be to ensure that the grayscale values of each pixel in the low grayscale region after adjustment and the grayscale values before adjustment satisfy y1 = (A / B) * k * x1. For example, the slope of the straight line GH is adjusted. The function corresponding to the adjusted straight line GH in the first region is: y1 = (A / B) * k * x1.
[0200] Among them, x 1 Indicates the grayscale value of the pixel in the low grayscale area before adjustment, y 1 represents the adjusted grayscale value of the pixel in the low grayscale area, A / B represents the third coefficient, A represents the dynamic range of the original video, B represents the expandable dynamic range of each video frame image, and k represents a constant.
[0201] Optionally, in one possible implementation, to improve the overall video brightness, the function corresponding to the adjusted straight line GH in the first region may also be: y1 = M*(A / B)*k*x1, (M>1). That is, the grayscale value of each pixel in the low grayscale region after adjustment and the grayscale value before adjustment may also satisfy y1 = M*(A / B)*k*x1, (M>1).
[0202] x 1 Indicates the grayscale value of the pixel in the low grayscale area before adjustment, y 1 represents the adjusted grayscale value of the pixel in the low grayscale area, M*(A / B) represents the third coefficient, A represents the dynamic range of the original video, B represents the expandable dynamic range of each video frame image, and M and k represent constants.
[0203] It should be understood that the value of M can be set by the user according to actual conditions and is not limited thereto.
[0204] The second adjustment strategy can be used to determine the adjusted grayscale values of the pixels in the mid-grayscale area.
[0205] The second adjustment strategy for the second region can be to ensure that the grayscale values of each pixel in the mid-grayscale region after adjustment satisfy the relationship y2 = k1 * x2 + b1. For example, the line HI corresponding to the second region can be adjusted to the line HK, thereby expanding the original linear tone mapping corresponding to the second region. For example, if the function corresponding to the line HI in the second region is y = k * x, the function corresponding to the adjusted line HK in the second region is y2 = k1 * x2 + b1.
[0206] Wherein, x2 represents the grayscale value of the pixel point in the middle grayscale area before adjustment, y2 represents the grayscale value of the pixel point in the middle grayscale area after adjustment, and k1 and b1 represent constants.
[0207] The adjusted grayscale values of the pixels in the high grayscale area can be determined by the third adjustment strategy.
[0208] For the third region, a third adjustment strategy can be implemented to ensure that the grayscale values of each pixel in the high-grayscale region after adjustment and before adjustment satisfy y3 = k2 * x3 + b2. For example, the line IJ corresponding to the third region can be adjusted to the line KL, thereby expanding the original linear tone mapping corresponding to the third region. For example, if the function corresponding to the line IJ in the third region is y = k*x, the function corresponding to the adjusted line KL in the third region is y3 = k2 * x3 + b2.
[0209] Wherein, x3 represents the grayscale value of the pixel point in the high grayscale area before adjustment, y3 represents the grayscale value of the pixel point in the high grayscale area after adjustment, and k2 and b2 represent constants.
[0210] It is worth noting that the values of k1, b1, k2, and b2 can be set based on the ratio of the brightness grayscale distribution, or can be determined by a machine learning model. For example, during training, the machine learning model learns the grayscale distribution of pixels in the brightness histogram of video frame images under different scenes and predicts the values of k1, b1, k2, and b2 based on this distribution. This is only an example and is not intended to be limiting.
[0211] S1046 : Generate an enhanced image corresponding to each video frame image according to the first region, the first adjustment strategy, the second region, the second adjustment strategy, the third region, and the third adjustment strategy of each video frame image.
[0212] For example, in S1045, the functions corresponding to the adjusted first, second, and third regions are determined. For any function, an input x value corresponds to an output y value, where the y value represents the grayscale value after tone mapping, i.e., the adjusted brightness value is obtained. The brightness of the video frame image is adjusted according to the adjusted brightness value to obtain an enhanced image corresponding to the video frame image. In layman's terms, the original grayscale value of a pixel is x, and the grayscale value corresponding to the adjusted pixel is determined by these functions as y. The brightness of the pixel in the video frame image is adjusted according to the y value to obtain an enhanced image corresponding to the video frame image.
[0213] For example, for the third region, the y3 value corresponding to the x3 value of each pixel in the third region is determined according to y3=k2*x3+b2. In the video frame, the x3 value of each pixel in the third region is adjusted to the y3 value. A similar method is used to make corresponding adjustments to each pixel in the first and second regions to obtain an enhanced image corresponding to the video frame.
[0214] In this implementation, the brightness information corresponding to each video frame is used to determine the adjustable areas within each video frame, namely the low grayscale area, the medium grayscale area, and the high grayscale area. Different adjustment strategies are then determined for the low grayscale area, the medium grayscale area, and the high grayscale area. The grayscale values of the pixels in these areas are then adjusted based on these different adjustment strategies, effectively improving the dynamic range of the resulting enhanced image. Consequently, the dynamic range of the video composed of multiple enhanced images is also improved, resulting in an enhanced display quality and user experience when the video with the enhanced dynamic range is displayed on a display device.
[0215] The following describes the method for processing video provided by an embodiment of the present application again in conjunction with Figure 13. Please refer to Figure 13, which is a schematic diagram of an implementation process shown in an exemplary embodiment of the present application. Exemplarily, after obtaining the SDR video, the SDR video is decoded to obtain a decoded video of the SDR video, which may include two or more video frame images. Figure 13 shows a video frame image in the decoded video of the SDR video.
[0216] Determine the brightness information corresponding to the video frame image. Figure 13 shows the brightness information corresponding to the video frame image in the form of a brightness histogram. Use the brightness information corresponding to the video frame image to divide the brightness area into low grayscale areas, medium grayscale areas, and high grayscale areas.
[0217] Combining the brightness capability of the display device and the dynamic range of the SDR video, tone mapping is performed on the low-grayscale area, medium-grayscale area, and high-grayscale area corresponding to the video frame image to obtain an enhanced image corresponding to the video frame image. The dynamic range of the video frame image is truly improved, that is, the enhanced image is the image with the improved dynamic range. It can be seen that the dynamic range of the video composed of multiple enhanced images (i.e., the extended dynamic range video shown in Figure 13) is also improved, so that when the video with the improved dynamic range is displayed on the display device, the display effect is improved and the user experience is improved.
[0218] The above description describes in detail the video processing method provided by the embodiment of the present application in conjunction with Figures 1 to 13. The following description describes in detail the hardware system, device, and chip of the display device to which the present application is applicable in conjunction with Figures 14 to 16. It should be understood that the hardware system, device, and chip in the embodiment of the present application can execute the various video processing methods provided by the aforementioned embodiments of the present application. That is, the specific working processes of the various products below can refer to the corresponding processes in the aforementioned method embodiments.
[0219] The method for processing video provided in the embodiment of the present application can be applicable to various display devices. Correspondingly, the display device provided in the embodiment of the present application can be a display device in various forms.
[0220] In some embodiments of the present application, the display device may be various camera devices such as SLR cameras and compact cameras, mobile phones, tablet computers, wearable devices, televisions, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc., or may be other devices or apparatuses capable of performing image processing. The embodiments of the present application do not impose any restrictions on the specific type of the display device.
[0221] The following takes a mobile phone as an example of a display device, and FIG14 shows a schematic structural diagram of a display device provided in an embodiment of the present application.
[0222] The display device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0223] It should be noted that the structure shown in FIG14 does not constitute a specific limitation on the display device 100. In other embodiments of the present application, the display device 100 may include more or fewer components than those shown in FIG14, or the display device 100 may include a combination of some of the components shown in FIG14, or the display device 100 may include sub-components of some of the components shown in FIG14. The components shown in FIG14 may be implemented in hardware, software, or a combination of software and hardware.
[0224] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0225] The controller may be the nerve center and command center of the display device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0226] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0227] In an embodiment of the present application, the processor 110 can run the software code of the method for processing video provided in the embodiment of the present application, thereby effectively improving the dynamic range of the original video.
[0228] 14 is merely a schematic illustration and does not limit the connection relationship between the modules of the display device 100. Alternatively, the modules of the display device 100 may adopt a combination of the multiple connection modes described in the above embodiments.
[0229] The wireless communication function of the display device 100 can be implemented through components such as the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor.
[0230] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in display device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0231] The display device 100 can implement display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. In an embodiment of the present application, the GPU can be used to decode the original video (such as SDR video) to obtain a decoded video of the original video (such as SDR video). The GPU can also be used to perform mathematical and pose calculations for graphics rendering, etc. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0232] Exemplarily, in an embodiment of the present application, the following steps can be executed in the processor 110: obtaining a decoded video of the original video; determining the brightness information corresponding to each video frame image; obtaining the brightness capability of the display device and the dynamic range of the original video, and calculating the scalable dynamic range of each video frame image based on the brightness capability of the display device and the dynamic range of the original video; and using the brightness information corresponding to each video frame image and the scalable dynamic range of each video frame image to perform tone mapping on each video frame image to obtain an enhanced image corresponding to each video frame image.
[0233] Display screen 194 can be used to display images or videos. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, display device 100 may include one or N display screens 194, where N can be a positive integer greater than 1.
[0234] In an embodiment of the present application, the display screen 194 can be used to display original video, multiple video frame images included in the decoded video, enhanced images, enhanced video composed of enhanced images (i.e., extended dynamic range video), etc.
[0235] The display screen 194 in the embodiment of the present application may be a touch screen. A touch sensor 180K may be integrated into the display screen 194. The touch sensor 180K may also be referred to as a "touch panel". That is, the display screen 194 may include a display panel and a touch panel, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it, such as when a user lightly presses the left side of the screen of the display device 100 with a touch object (such as a user's finger or a stylus, etc.) and slides upward. After the touch operation detected by the touch sensor 180K, it can be passed to the upper layer by the kernel layer driver (such as the TP driver) to determine the type of touch event. Visual output related to the touch operation can be provided by the display screen 194. In other embodiments, the touch sensor 180K may also be provided on the surface of the display device 100, at a different position from that of the display screen 194.
[0236] The display device 100 can realize shooting and recording functions through the ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0237] The ISP processes data fed back by camera 193. For example, when shooting or recording a video, light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can perform algorithmic optimization on image noise, brightness, and color. It can also optimize parameters such as exposure and color temperature of the recorded scene. In some embodiments, the ISP can be located within camera 193.
[0238] The camera 193 is used to capture images or videos. It can be triggered to start by application instructions to realize shooting and recording functions. For example, in scenes such as live video broadcast, video conferencing, video calls, and video surveillance, videos can be recorded and acquired. The camera may include components such as an imaging lens, an optical filter, and an image sensor. The light emitted or reflected by the object enters the imaging lens, passes through the optical filter, and finally converges on the image sensor. The image sensor is mainly used to converge the light emitted or reflected by all objects in the recording angle of view to form an image; the optical filter is mainly used to filter out excess light waves in the light (for example, light waves other than visible light, such as infrared); the image sensor is mainly used to perform photoelectric conversion on the received light signal, convert it into an electrical signal, and input it into the processor 110 for subsequent processing. Among them, the camera 193 can be located in front of the display device 100 or on the back of the display device 100. The specific number and arrangement of the cameras can be set according to needs, and this application does not impose any restrictions.
[0239] Exemplarily, in an embodiment of the present application, the camera 193 can obtain recorded video.
[0240] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the display device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0241] A video codec is used to compress or decompress digital video. The display device 100 may support one or more video codecs. In an embodiment of the present application, the video codec may be used to decode an original video (e.g., an SDR video) to obtain a decoded video of the original video (e.g., an SDR video).
[0242] The gyroscope sensor 180B can be used to determine the motion posture of the display device 100. In some embodiments, the angular velocity of the display device 100 around three axes (i.e., the x-axis, the y-axis, and the z-axis) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for shooting and recording anti-shake. For example, when shooting or recording, the gyroscope sensor 180B detects the angle of the display device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the display device 100 through reverse motion to achieve anti-shake. The gyroscope sensor 180B can also be used in scenarios such as navigation and somatosensory games.
[0243] Accelerometer 180E can detect the magnitude of acceleration of display device 100 in various directions (typically the x-axis, y-axis, and z-axis). When display device 100 is stationary, it can detect the magnitude and direction of gravity. Accelerometer 180E can also be used to identify the posture of display device 100, which can be used as an input parameter for applications such as landscape / portrait switching and pedometers.
[0244] The distance sensor 180F is used to measure distance. The display device 100 can measure distance using infrared or laser. In some embodiments, such as in a shooting or recording scene, the display device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0245] Ambient light sensor 180L is used to sense ambient light brightness. Display device 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether display device 100 is in a pocket to prevent accidental touches.
[0246] The fingerprint sensor 180H is used to collect fingerprints. The display device 100 can use the collected fingerprint characteristics to implement functions such as unlocking, accessing application locks, taking photos, and answering calls.
[0247] The pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a parallel plate comprising at least two conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The display device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the display device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The display device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions.
[0248] The button 190 includes a power button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The display device 100 can receive button input and generate key signal input related to the user settings and function control of the display device 100. In an embodiment of the present application, when the display device 100 plays a video online, if the left side of the screen is used to control the screen brightness, then the right side of the screen is used to control the volume. If the right side of the screen is used to control the screen brightness, then the left side of the screen is used to control the volume. For example, the user taps the left side of the screen of the display device 100 with a touch object (such as a user's finger or a stylus, etc.) and slides up or down, and the display device 100 responds to the touch operation to increase or dim the screen brightness. The user taps the right side of the screen of the display device 100 with a touch object and slides up or down, and the display device 100 responds to the touch operation to increase or decrease the volume of the display device 100.
[0249] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. Indicator 192 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the display device 100. The display device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.
[0250] The methods in the above embodiments can all be implemented in the display device 100 having the above hardware structure.
[0251] FIG15 is a schematic diagram of the structure of a display device provided in an embodiment of the present application. As shown in FIG15 , the display device 200 includes a first acquisition module 210 , a determination module 220 , a second acquisition module 230 , and a processing module 240 .
[0252] The display device 200 can perform the following schemes:
[0253] A first acquisition module 210 is used to obtain a decoded video of the original video;
[0254] A determination module 220 is configured to determine brightness information corresponding to each video frame image;
[0255] A second acquisition module 230 is configured to acquire the brightness capability of the display device and the dynamic range of the original video, and calculate an expandable dynamic range of each video frame image based on the brightness capability of the display device and the dynamic range of the original video;
[0256] The processing module 240 is configured to perform tone mapping on each video frame image by utilizing the brightness information corresponding to each video frame image and the scalable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image.
[0257] It should be noted that the display device 200 is implemented in the form of a functional module. The term "module" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0258] For example, a "module" may be a software program, a hardware circuit, or a combination of the two that implements the aforementioned functionality. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functionality.
[0259] Therefore, the modules of each example described in the embodiments of this application can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0260] The embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions; when the computer-readable storage medium is run on a display device, the display device executes the method shown above. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more media integrated therein. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium (e.g., a solid state disk (SSD)).
[0261] An embodiment of the present application further provides a computer program product including computer instructions. The computer program product includes: computer program code. When the computer program code runs on a display device, the display device can execute the technical solution shown above.
[0262] FIG16 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip shown in FIG16 can be a general-purpose processor or a dedicated processor. The chip includes a processor 301. The processor 301 is used to support the display device in executing the technical solution shown above.
[0263] Optionally, the chip further includes a transceiver 302 , which is configured to accept control of the processor 301 and to support the display device in executing the aforementioned technical solution.
[0264] Optionally, the chip shown in FIG16 may further include: a storage medium 303 .
[0265] It should be noted that the chip shown in Figure 16 can be implemented using the following circuits or devices: one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits that can perform the various functions described throughout this application.
[0266] The display device, display apparatus, computer storage medium, computer program product, and chip provided in the above-mentioned embodiments of the present application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the corresponding beneficial effects of the method provided above, and will not be repeated here.
[0267] It should be understood that the above is only to help those skilled in the art better understand the embodiments of the present application, and is not intended to limit the scope of the embodiments of the present application. Based on the above examples given, those skilled in the art can obviously make various equivalent modifications or changes. For example, certain steps in each embodiment of the above detection method may be unnecessary, or certain new steps may be added. Or a combination of any two or any multiple embodiments described above. Such modifications, changes, or combined solutions also fall within the scope of the embodiments of the present application.
[0268] It should also be understood that the above description of the embodiments of the present application focuses on emphasizing the differences between the various embodiments. The same or similar points that are not mentioned can be referenced with each other. For the sake of brevity, they will not be repeated here.
[0269] It should also be understood that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0270] It should also be understood that in the embodiments of the present application, "pre-setting" and "pre-definition" can be achieved by pre-saving corresponding codes, tables or other methods that can be used to indicate relevant information in a device (for example, including a display device), and the present application does not limit its specific implementation method.
[0271] It should also be understood that the division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features of various modes, categories, situations and embodiments can be combined without contradiction.
[0272] It should also be understood that in the various embodiments of the present application, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. Finally, it should be noted that the above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for processing a video, characterized in that: Applied to a display device, the method includes: Obtaining a decoded video of an original video, wherein the decoded video includes a plurality of video frame images; Determine the brightness information corresponding to each video frame image; Obtaining the brightness capability of the display device and the dynamic range of the original video, and calculating an expandable dynamic range of each video frame image based on the brightness capability of the display device and the dynamic range of the original video; By utilizing the brightness information corresponding to each video frame image and the scalable dynamic range of each video frame image, tone mapping is performed on each video frame image to obtain an enhanced image corresponding to each video frame image, wherein the dynamic range of the enhanced image is greater than the dynamic range of the video frame image.
2. The method according to claim 1, wherein The method of performing tone mapping on each video frame image by utilizing brightness information corresponding to each video frame image and an expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image includes: Determining a standard dynamic range and an extended dynamic range corresponding to each video frame image using brightness information corresponding to each video frame image, wherein brightness information of pixels in the standard dynamic range is less than a preset threshold, and brightness information of pixels in the extended dynamic range is greater than or equal to the preset threshold; Determining, according to the scalable dynamic range of each video frame image, a first coefficient corresponding to the standard dynamic range of each video frame image and a second coefficient corresponding to the extended dynamic range of each video frame image; Tone mapping is performed on the pixels in the standard dynamic area according to the first coefficient, and tone mapping is performed on the pixels in the extended dynamic area according to the second coefficient to obtain an enhanced image corresponding to each video frame image.
3. The method according to claim 2, wherein The step of performing tone mapping on the pixels in the standard dynamic range according to the first coefficient and performing tone mapping on the pixels in the extended dynamic range according to the second coefficient to obtain an enhanced image corresponding to each video frame image includes: Calculating a first product of an original pixel value of each pixel point in the standard dynamic area and the first coefficient, and updating a pixel value of each pixel point in the standard dynamic area according to the first product; Calculating a second product of an original pixel value of each pixel point in the extended dynamic area and the second coefficient, and updating a pixel value of each pixel point in the extended dynamic area according to the second product; An enhanced image corresponding to each video frame image is generated according to each updated pixel point in the standard dynamic area and each updated pixel point in the extended dynamic area.
4. The method according to claim 1, wherein The method of performing tone mapping on each video frame image by utilizing brightness information corresponding to each video frame image and an expandable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image includes: Using the brightness information corresponding to each video frame image, determine the low grayscale area, medium grayscale area, and high grayscale area corresponding to each video frame image; According to the scalable dynamic range of each video frame image, the brightness of at least one of the low grayscale area, the medium grayscale area and the high grayscale area is adjusted to obtain an enhanced image corresponding to each video frame image.
5. The method according to claim 4, wherein The step of adjusting the brightness of at least one of the low grayscale area, the medium grayscale area, and the high grayscale area according to the scalable dynamic range of each video frame image to obtain an enhanced image corresponding to each video frame image includes: Determining the adjusted grayscale values of the pixels in the low grayscale area according to the scalable dynamic range of each video frame image; Determining the adjusted grayscale values of the pixels in the medium grayscale area; Determining the adjusted grayscale values of the pixels in the high grayscale area; An enhanced image corresponding to each video frame image is generated according to the adjusted grayscale values of each pixel in the low grayscale area, the medium grayscale area, and the high grayscale area.
6. The method according to claim 5, wherein The step of determining the adjusted grayscale values of the pixels in the low grayscale area according to the expandable dynamic range of each video frame image includes: According to the expandable dynamic range of each video frame image, a third coefficient corresponding to the low grayscale area is determined; according to the third coefficient, tone mapping is performed on the pixel points in the low grayscale area to obtain the adjusted grayscale values of the pixel points in the low grayscale area.
7. A display device, characterized in that: The display device comprises means for performing the method according to any one of claims 1 to 6.
8. A display device, characterized in that: include: one or more processors; one or more memories; The memory stores one or more programs, and when the one or more programs are executed by the processor, the display device executes the method according to any one of claims 1 to 6.
9. A chip, characterized in that: include: A processor, configured to call and run a computer program from a memory, so that a display device equipped with the chip executes the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 6.