Video data processing method, video coding method, device, medium and apparatus
By acquiring metadata of HDR video and display device parameters, a mapping curve is determined for color correction, solving the problem of underutilization of display device performance in existing technologies. This achieves brightness level optimization and dark detail enhancement of HDR video frames, improving the display effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIYI CENTURY SCI & TECH CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-24
AI Technical Summary
Existing HDR video processing methods cannot fully utilize the performance of high-brightness display devices, thus limiting the improvement of HDR video display effects.
By acquiring metadata of high dynamic range video and display parameters of the display device, the mapping curve corresponding to the video frame is determined. This curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space. The upper limit value is determined based on the display parameters and metadata, and color correction is performed to adapt to the display capabilities of the display device.
Optimize the brightness levels of video frames to avoid overexposure of highlights while enhancing details in shadows, preserving the high dynamic range characteristics of HDR content, and improving the display effect of video frames.
Smart Images

Figure CN122457780A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of video processing technology, and in particular to a video data processing method, a video encoding and decoding method, an apparatus, a medium, and a device. Background Technology
[0002] In the field of video processing technology, HDR (High Dynamic Range) technology aims to significantly enhance the realism, immersion, and expressiveness of video content by expanding the range of details that can be presented from the darkest to the brightest areas of the image. A wider dynamic range means that display systems can more accurately reproduce the rich contrast of light intensity and the levels of brightness and darkness in the real world, thus becoming one of the key directions in the development of current video display technology.
[0003] Currently, the processing methods for HDR video specifically include: In the video production stage, the shooting and color grading processes are first completed to create a video master with a specific peak brightness; subsequently, EOTF (Electro-Optical Transfer Function) such as PQ (Perceptual Quantizer) is used to perceptually quantize and encode the video master, generating the corresponding digital encoded signal; in the video distribution stage, the digital encoded signal is encapsulated into a standard format video for distribution. During perceptual quantization and encoding, the range of the digital encoded signal is constrained by the peak brightness set by the video master (currently, the mainstream HDR master production brightness is 1000 cd / m² (candela per square meter)); the peak brightness of the distributed video is also constrained by the peak brightness set by the video master.
[0004] However, with the rapid improvement in display device performance, the peak brightness of consumer and professional-grade displays has generally reached 2000 cd / m² or even higher, far exceeding the peak brightness constraints set for HDR video in mastering, encoding, and distribution. This means that the high brightness performance of display devices cannot be fully utilized at the video playback end, limiting further improvements in HDR video display effects. Summary of the Invention
[0005] This invention provides a video data processing method, a video encoding and decoding method, an apparatus, a medium, and a device that can effectively improve the display effect of HDR video frames.
[0006] In a first aspect, embodiments of the present invention disclose a video data processing method, the method comprising: Obtain metadata of video frames in a high dynamic range video; Obtain the display parameters of the display device; Based on the display parameters and the metadata, a mapping curve corresponding to the video frame is determined; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0007] Secondly, embodiments of the present invention disclose a video decoding method, the method comprising: Acquire the YUV signals of video frames in a high dynamic range video; Convert the YUV signal into an RGB electrical signal; The source color intensity value of the video frame at the pixel point is determined based on the RGB electrical signal; The source color intensity value is mapped using the mapping curve corresponding to the video frame to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; Based on the mapping result value, the RGB light signal converted from the RGB electrical signal is color corrected so that the color-corrected RGB light signal can be displayed on a display device or encoded.
[0008] Thirdly, embodiments of the present invention disclose a video decoding method, the method being applied at a decoding end, comprising: Obtain the metadata of video frames from the bitstream of high dynamic range video sent from the encoding end; Obtain the display parameters of the display device; Based on the display parameters and the metadata, a mapping curve corresponding to the video frame is determined; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0009] Fourthly, embodiments of the present invention disclose a video encoding method, the method being applied at an encoding end, comprising: Obtain metadata of video frames in a high dynamic range video; The metadata and the high dynamic range video are encoded to obtain a bitstream; The bitstream is sent to the decoding end so that the decoding end can determine the mapping curve corresponding to the video frame based on the display parameters of the display device and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0010] Fifthly, embodiments of the present invention disclose a video data processing apparatus, the apparatus comprising: The metadata acquisition module is used to acquire metadata of video frames in high dynamic range videos. The display parameter acquisition module is used to acquire the display parameters of the display device. The mapping curve determination module determines the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extended mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0011] Sixthly, embodiments of the present invention disclose a video decoding apparatus, the apparatus comprising: The signal acquisition module is used to acquire the YUV signals of video frames in high dynamic range videos; The signal conversion module is used to convert the YUV signal into an RGB electrical signal; The source color intensity value determination module is used to determine the source color intensity value of the video frame at the pixel point based on the RGB electrical signal; The mapping module is used to map the source color intensity value using the mapping curve corresponding to the video frame to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; The color correction module is used to correct the color of the RGB light signal converted from the RGB electrical signal according to the mapping result value, so as to display the color-corrected RGB light signal on a display device or encode the color-corrected RGB light signal.
[0012] In a seventh aspect, embodiments of the present invention disclose a video decoding device, the device being applied at a decoding end, comprising: The metadata acquisition module is used to acquire the metadata of video frames from the bitstream of high dynamic range video sent from the encoding end; The display parameter acquisition module is used to acquire the display parameters of the display device. The mapping module determination module is used to determine the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0013] Eighthly, embodiments of the present invention disclose a video encoding apparatus, the apparatus being applied at an encoding end, comprising: The metadata acquisition module is used to acquire the metadata of video frames in high dynamic range videos. The encoding module is used to encode the metadata and the high dynamic range video to obtain a bitstream; The sending module is used to send the bitstream to the decoding end, so that the decoding end can determine the mapping curve corresponding to the video frame according to the display parameters of the display device and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined according to the display parameters and the metadata.
[0014] In a ninth aspect, embodiments of the present invention disclose a non-transitory computer-readable recording medium storing a bitstream generated by a method performed by an apparatus for video data processing, wherein the method includes: acquiring metadata of video frames in a high dynamic range video; and encoding the metadata and the high dynamic range video to obtain a bitstream.
[0015] In a tenth aspect, embodiments of the present invention disclose a method for storing a bitstream of video data, comprising: acquiring metadata of video frames in a high dynamic range video; encoding the metadata and the high dynamic range video to obtain a bitstream; and storing the bitstream in a non-transitory computer-readable recording medium.
[0016] Eleventhly, embodiments of the present invention disclose a method for storing a bitstream, characterized in that it includes: performing the aforementioned video data processing method to generate a bitstream; and storing the bitstream.
[0017] In a twelfth aspect, embodiments of the present invention disclose a computer-readable medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the aforementioned method.
[0018] In a thirteenth aspect, embodiments of the present invention disclose an electronic device, comprising: One or more processors; A memory for storing one or more computer programs that, when executed by one or more processors, cause the electronic device to perform the aforementioned method.
[0019] In a fourteenth aspect, embodiments of the present invention disclose a computer program product comprising a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads from and executes the computer program from the computer-readable storage medium, causing the electronic device to perform the aforementioned method.
[0020] Compared with the prior art, the embodiments of the present invention have the following advantages: In the technical solution of this invention embodiment, the mapping curve corresponding to the video frame is determined according to the display parameters of the display device and the metadata of the video frame, and the upper limit of the expansion of the mapping result value in the mapping curve is determined according to the display parameters and metadata.
[0021] Since the metadata of HDR video frames reflects the brightness information of the video frame content, and the display parameters of the display device limit the display capabilities that the display device can support, determining the upper limit of the mapping curve based on these two factors can both control the mapping result value within the range supported by the display device to prevent display anomalies, and adapt and adjust the brightness of the video frame's light signal according to the actual brightness information of the video frame. Therefore, the mapping curve corresponding to the video frame in this embodiment of the invention can optimize the brightness level of the video frame image based on the expansion of the brightness range of the video frame's light signal. For example, it can enhance the details in the shadows while avoiding overexposure of highlights, and at the same time retain the high dynamic range characteristics of HDR content, thereby effectively improving the display effect of HDR video frames. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present invention can be applied; Figure 2 This is a flowchart illustrating the steps of a video data processing method according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the steps of a video data processing method according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the steps of a video data processing method according to an embodiment of the present invention; Figure 5 This is a data flow diagram of HDR video upconversion according to an embodiment of the present invention; Figure 6 This is a flowchart of histogram statistics, iterative binary search, and metadata processing at the generation end of an embodiment of the present invention; Figure 7This is a flowchart of pixel-by-pixel tone mapping processing at the mapping end according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a video data processing device according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a video data processing device according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a video data processing device according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a video data processing device according to an embodiment of the present invention; Figure 12 This is a schematic diagram of the structure of an electronic device 1300 according to an embodiment of the present invention. Detailed Implementation
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of embodiments of the present invention can be applied is shown.
[0025] like Figure 1 As shown, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. Figure 1 In one embodiment, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.
[0026] For example, the first terminal device 110 can encode video data (e.g., a video image stream captured by the terminal device 110) to transmit it to the second terminal device 120 via the network 150. The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device 120 can receive the encoded video data from the network 150, decode the encoded video data to recover the video data, and display video images based on the recovered video data.
[0027] In one embodiment of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded video data, such as during a video conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode video data (e.g., a video image stream captured by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded video data transmitted by the other terminal device, decode the encoded video data to recover the video data, and display the video images on an accessible display device based on the recovered video data.
[0028] exist Figure 1 In the embodiments shown, the first terminal device 110, the second terminal device 120, the third terminal device 130 and the fourth terminal device 140 may be servers or terminals, but the principles disclosed in this application are not limited to these.
[0029] Servers can be standalone physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smart voice interaction devices, smartwatches, smart home appliances, in-vehicle terminals, aircraft, etc., but are not limited to these.
[0030] Figure 1 The network 150 shown represents any number of networks, including, for example, wired and / or wireless communication networks, that transmit encoded video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140. The communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of the disclosure herein.
[0031] The video data processing method of this invention can be used to process HDR videos to improve the display effect of HDR videos on display devices.
[0032] The video signal processing method provided in this invention can be applied to electronic devices. In specific applications, these electronic devices can be set-top boxes, smart TVs, smartphones, personal computers, multimedia players, screen projection devices, in-vehicle terminals, tablet computers, laptops, smart projectors, head-mounted displays, smart cockpit hosts, video conferencing terminals, advertising playback devices, and embedded multimedia terminals, etc.
[0033] It is understandable that electronic devices with screens can be used to display video, meaning they simultaneously function as display devices. Examples include smart TVs, smartphones, and laptops. Conversely, electronic devices without screens can display video through connected display devices, such as set-top boxes, multimedia players, and desktop computers connected to a monitor.
[0034] The video data processing method of this invention will be described below through specific embodiments.
[0035] Reference Figure 2 The diagram illustrates a step-by-step flowchart of a video data processing method according to an embodiment of the present invention. The method may specifically include the following steps: Step 201: Obtain the metadata of video frames in the high dynamic range video; Step 202: Obtain the display parameters of the display device; Step 203: Determine the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0036] In the technical solution of this invention embodiment, the mapping curve corresponding to the video frame is determined according to the display parameters of the display device and the metadata of the video frame, and the upper limit of the expansion of the mapping result value in the mapping curve is determined according to the display parameters and metadata.
[0037] Since the metadata of HDR video frames reflects the brightness information of the video frame content, and the display parameters of the display device limit the display capabilities that the display device can support, determining the upper limit of the mapping curve based on these two factors can both control the mapping result value within the range supported by the display device to prevent display anomalies, and adapt and adjust the brightness of the video frame's light signal according to the actual brightness information of the video frame. Therefore, the mapping curve corresponding to the video frame in this embodiment of the invention can optimize the brightness level of the video frame image based on the expansion of the brightness range of the video frame's light signal. For example, it can enhance the details in the shadows while avoiding overexposure of highlights, and at the same time retain the high dynamic range characteristics of HDR content, thereby effectively improving the display effect of HDR video frames.
[0038] In step 201, a video frame is a still image in an HDR video. It is the basic unit that constitutes an HDR video, and numerous video frames played in sequence can form a dynamic video picture. Metadata is data used to describe the attributes, characteristics, and other information of the data. The metadata of a video frame is used to describe the content, characteristics, time position, and other related information of the video frame.
[0039] As a decoding end, the electronic device can obtain the metadata of the video frame from the bitstream or sideload metadata of the high dynamic range video sent by the encoding end. In this case, the metadata of the video frame can be obtained by the encoding end based on image processing methods and encapsulated into the bitstream or sideload metadata of the HDR video. Alternatively, the electronic device can use image processing methods to obtain the metadata of the video frame.
[0040] In a specific implementation, the metadata includes at least one of the following features: video frame extension parameters, highlight feature values, shadow feature values, average brightness value, maximum brightness value, and fusion coefficient.
[0041] Among them, the video frame expansion parameter characterizes the margin by which the brightness of the video frame can be expanded upwards. The highlight feature value characterizes the brightness distribution and highlight features of the bright areas in the video frame. The shadow feature value characterizes the brightness distribution and detail features of the dark areas in the video frame. The average brightness value characterizes the average level of the overall grayscale features of the video frame. The maximum brightness value characterizes the maximum brightness level achieved by the grayscale features of the video frame. The fusion coefficient is a weighted fusion ratio used to combine different features of the video frame.
[0042] In the specific implementation, the process by which the encoding end obtains the metadata of video frames in a high dynamic range video includes: Step A1: Construct a histogram for the grayscale features of the video frames in the PQ domain; Step A2: Determine the bisection points of the histogram, which are used to divide the histogram into dark and bright parts; Step A3: Based on the bisection point, determine at least one of the following: the bright part feature value, the dark part feature value, the average brightness value, and the maximum brightness value of the video frame.
[0043] Steps A1 to A3 are performed frame-by-frame on the HDR PQ video frames, and the execution results are written to the bitstream or sideload metadata, so that the decoding end can reconstruct the tone mapping curve when the target peak value is higher than the master peak value. Here, the target peak value can refer to the peak brightness that the target display device can achieve, and the master peak value can refer to the peak brightness of the source master used when the video content was produced.
[0044] In step A1, grayscale features can be calculated in the PQ domain for each pixel of the video frame. The grayscale features can be the maximum values of the three components of the pixel after PQ nonlinear encoding, as shown in formula (1).
[0045] S_i = max(R_i, G_i, B_i) (1) Where R_i, G_i, and B_i are the three components of the pixel after PQ nonlinear encoding. A histogram Hist_S is constructed for S within the effective grayscale range [hist_min, hist_max], where Hist_S[b] represents the pixel count falling in the b-th brightness bin.
[0046] S represents the grayscale feature of a single pixel in the PQ domain in a video frame, used to uniformly characterize the brightness-related features of that pixel.
[0047] The effective grayscale range [hist_min, hist_max] is the range of brightness values used when constructing the histogram. Hist_min represents the lower limit of the minimum value of the grayscale feature, and hist_max represents the upper limit of the maximum value of the grayscale feature. Histogram statistics are performed on the grayscale features that fall within this brightness range.
[0048] A bin (grayscale interval unit) is a discrete statistical unit obtained by dividing a continuous range of grayscale feature values. It is used to count pixels within different brightness ranges. For example, the effective range of grayscale feature values in the PQ domain [0, 1023] is divided into 256 bins. The range of each bin is 0-3, 4-7...1020-1023. The 0th bin corresponds to the brightness value of 0-3, the 1st bin corresponds to the brightness value of 4-7, and the number of pixels falling within that range is counted in each bin.
[0049] Gray-scale features can also be the luminance component y of a pixel. Gray-scale features can also be a combination of the maximum value of the three components and the luminance component, such as S_i = A* max (R_i, G_i, B_i) + (1-A)* y, where A is the weighting coefficient.
[0050] The bisection point in step A2 is used to divide the range of grayscale feature values corresponding to the histogram into dark and bright intervals, thereby distinguishing dark pixels from bright pixels in the image and providing a boundary basis for subsequent calculation of dark and bright features.
[0051] The embodiments of the present invention do not limit the specific method of determining the bisection point. For example, the bisection point can be a fixed set value, or it can be a dynamic boundary value obtained by iterative calculation based on the histogram of video frames. For example, the bin index of the midpoint in the histogram is first taken as the initial bisection point, and then the converged bisection point is obtained by calculating the weighted average of the dark part PQ and the weighted average of the linear domain of the bright part and iteratively updating.
[0052] In one alternative implementation, step A2, determining the bisection points of the histogram, specifically includes: Step A21: Take the index of the gray-level cell containing the median of the histogram in [hist_min, hist_max] (or the equivalent 50th percentile of the cumulative distribution) as the initial bisection point M, and set the maximum number of iterations to N_max.
[0053] Step A22: Within the dark region [hist_min, M), perform pixel count weighting on Hist_S, and calculate the arithmetic mean in the PQ domain using formula (2) to obtain the dark region weighted mean dark_avg: dark_avg=Σ_{b ∈[hist_min, M)} Hist_S[b] C(b) / Σ_{b ∈[hist_min,M)} Hist_S[b] (2) Where C(b) is the PQ field representative code value corresponding to the b-th bin (e.g., the bin center). The numerator of formula (2) is the sum of the products of the number of pixels in each gray-level interval unit in the dark region and the corresponding PQ field representative code value, representing the total code value contribution of all pixels in the dark region after weighting by gray-level interval units. The denominator of formula (2) is the sum of the number of pixels in each gray-level interval unit in the dark region, representing the total number of pixels participating in the statistics in the dark region.
[0054] Step A23: Within the bright region [M, hist_max], first map the representative code value of each grayscale unit to the linear optical domain using EOTF_PQ (Electro-Optical Transfer Function for PQ) to obtain the linear brightness value corresponding to each grayscale unit; then calculate the average value of these linear brightness values by weighting according to pixel count to obtain the weighted average brightness of the bright region in the linear optical domain; finally, map the weighted average brightness back to the PQ domain using OETF_PQ (Opto-Electrical Transfer Function for PQ) to obtain the weighted average value of the bright region, bright_avg, as shown in formula (3): bright_avg = OETF_PQ ( Σ_{b ∈[M, hist_max)} Hist_S [b] EOTF_PQ (C(b)) / Σ_{b ∈[M, hist_max)} Hist_S [b] ) (3) Wherein, EOTF_PQ(C(b)) is the linear brightness value obtained by mapping the representative code value C(b) of the b-th grayscale unit to the linear optical domain; Σ_{b ∈[M, hist_max)} Hist_S[b] EOTF_PQ(C(b)) is the sum of the products of the number of pixels in all grayscale units within the bright region and the corresponding linear luminance value; Σ_{b ∈[M, hist_max)} Hist_S[b] is the total number of pixels in the bright region; dividing the two yields the weighted average luminance of the bright region in the linear light domain; mapping the weighted average luminance back to the PQ domain via OETF_PQ, the final result is bright_avg.
[0055] Step A24: After converting the weighted mean of the dark area and the weighted mean of the bright area in the linear domain to the linear light domain, calculate the geometric median, and then encode it back into the PQ domain to obtain the new bisection point candidate value M_new: M_new = OETF_PQ (sqrt ( EOTF_PQ (dark_avg) EOTF_PQ (bright_avg) ) ) (4) In formula (4) “EOTF_PQ (dark_avg) The “EOTF_PQ (bright_avg)” in "It is a multiplication operation. First, dark_avg and bright_avg are mapped to the linear optical domain through EOTF_PQ and then multiplied. Then, the square root of the product is used to calculate the geometric median. Finally, it is mapped back to the PQ domain through OETF_PQ to obtain M_new.
[0056] Smoothly merge M_new with the current bisection point M, and use the result of the smoothed merge as the updated bisection point. For example, M can be updated using the rounded arithmetic mean. The smoothed merge method can include the rounded arithmetic mean: M←floor((M+Mnew) / 2) or an equivalent smoothing strategy. floor indicates rounding down, that is, taking the largest integer not greater than the result in parentheses.
[0057] Step A25: Determine if the iteration termination condition has been met. If so, stop the iteration and output the converged weighted average of dark areas (dark_avg), the weighted average of bright areas (bright_avg), and the bisection point M. The iteration termination condition includes: if the change of the bisection point M is less than one gray-level unit, or if the number of iterations reaches the maximum number of iterations N_max.
[0058] Step A3, based on the bisection point, determines at least one of the following: highlight feature value, shadow feature value, average brightness value, and maximum brightness value of the video frame, specifically including: Step A31: Based on the bisection points, perform weighted processing on the pixel brightness of the dark area of the histogram in the PQ domain to obtain the dark area feature values of the video frame; and / or Step A32: Based on the bisection point, transform the pixel brightness of the bright part of the histogram from the PQ domain to the linear domain, perform weighted processing on the pixel brightness of the bright part of the histogram in the linear domain, and then transform the corresponding weighted processing result back to the PQ domain to obtain the bright part feature value of the video frame; and / or Step A33: Convert the grayscale features of the video frame from the PQ domain to the linear domain, calculate the average, and then convert the corresponding average result back to the PQ domain to obtain the average brightness value of the video frame; and / or Step A34: Obtain the maximum brightness value of the video frame based on the maximum value of the grayscale feature in the PQ domain.
[0059] In step A31, based on the dark area intervals divided by the bisection points, a weighted operation is performed on the pixel brightness corresponding to the dark area in the histogram in the PQ domain, and the dark area feature value of the video frame is obtained through the weighted processing.
[0060] Step A32: Based on the bright area intervals divided by the bisection points, the brightness of the pixels corresponding to the bright areas in the histogram is first converted from the PQ domain to the linear domain through EOTF_PQ; after the weighted processing of the bright pixel brightness is completed in the linear domain, the weighted result is converted back to the PQ domain through OETF_PQ, and finally the bright feature value of the video frame is obtained.
[0061] Step A33 transforms the grayscale features of the entire video frame from the PQ domain to the linear domain, calculates the arithmetic mean of the pixel brightness in the linear domain, and then transforms the average result back to the PQ domain to obtain the average brightness value of the video frame, averg_pq, corresponding to formula (5): averg_pq = OETF_PQ( (1 / N)Σ_i EOTF_PQ(S_i) ) (5) Where i is the pixel index, used to traverse all pixels of the video frame; N represents the total number of pixels in the video frame Σ_i: the average brightness value summed over all pixels i. The average brightness value can be discretized using the same bin discretization method as the calculation of dark and bright feature values.
[0062] Step A34 extracts the maximum value of the grayscale feature in the PQ domain of the video frame, and uses it as the maximum brightness value of the video frame, corresponding to formula (6): max_pq = max {S_i | i traverses all pixels in the entire frame} (6) Formula (6) takes the maximum value of the PQ domain luminance value S_i of all pixels in the whole frame to obtain the maximum luminance value max_pq.
[0063] The video frame spread parameter characterizes the margin by which the brightness of a video frame can be spread upwards. The value of the video frame spread parameter can be determined by those skilled in the art based on the actual application requirements.
[0064] Alternatively, embodiments of the present invention may determine the target brightness value of a video frame and determine the video frame extension parameters based on the ratio of the target brightness value to the peak brightness of the master image corresponding to the video frame.
[0065] The target brightness value represents the maximum brightness that the human eye can tolerate, and it can be determined by the user or based on the analysis of video frames.
[0066] The fusion coefficient is used by the decoder to construct the pixel-by-pixel lookup table for the horizontal coordinate; it is written by the generator (such as the encoder) according to a strategy or by a default value agreed upon in the specification (0.5 in this example). The value is typically constrained to [0,1].
[0067] In embodiments of the present invention, the following features can be encoded as dynamic metadata (fixed-point or normalized integers, with the specific bit layout defined by the syntax). Dynamic metadata can refer to metadata that changes as the video frame or the scene in which the video frame is located: 5.1) shadow_pq carries the feature value of the dark area, that is, the weighted mean of PQ in the dark area, which represents the feature points of the dark area.
[0068] 5.2) highlight_pq carries the feature value of the bright part, that is, the weighted mean of PQ in the bright area, which represents the feature point of the bright part.
[0069] 5.3) average_pq carries the average brightness value, i.e., the average_pq obtained by formula (5).
[0070] 5.4) maximum_pq carries the maximum brightness value, which is max_pq obtained by formula (6).
[0071] 5.5) extended_headroom carries video frame extension parameters, with a lower limit clamp of 1 during encoding.
[0072] 5.6) factor_mix is the fusion coefficient, used by the decoder to construct the pixel-by-pixel lookup table for the horizontal coordinate; it is written by the generator according to the strategy or the default value is agreed upon by the specification, and the value is typically constrained to [0,1].
[0073] In summary, the process of obtaining metadata in this embodiment of the invention is as follows: 1. Construct a histogram for grayscale features in the PQ domain. Then, use an iterative bisection method combined with "updating the PQ mean of the dark area + the linear domain mean of the bright area back to PQ + updating the geometric median by taking the square root of the product of the dark area average and the bright area average" to achieve stable segmentation of the bright and dark statistical regions.
[0074] 2. The metadata carries dark area feature values, average brightness values, maximum brightness values, and video frame extension parameters, enabling the decoding end to reconstruct the brightness mapping curve without having to read back all the pixels of the frame.
[0075] 3. The fusion coefficient factor_mix is transmitted together with the above statistics as side-load information to achieve adjustable control of the pixel-by-pixel horizontal coordinate between maximum value features and perceived brightness.
[0076] Specifically, the metadata generation end is in the PQ domain, where a histogram is constructed for each pixel by a one-dimensional scalar obtained by convergence of the R / G / B components through a predetermined maximum value type; the weighted average value of the dark area (dark_avg) and the weighted average value of the bright area (bright_avg) are calculated through iterative binary search; the above results are encoded together with information such as the average brightness value, the maximum brightness value, and the brightness extension parameter, so that the mapping end (such as the decoding end or playback end of HDR video) can reproduce the shape of the brightness mapping curve without reading the full frame pixels.
[0077] In step 202, the display parameters of the display device refer to the technical characteristics of the display device used to adjust and control the image output. In specific implementations, the display parameters specifically involve technical characteristics of the display device such as brightness, contrast ratio, color gamut, and response time.
[0078] In one optional implementation, the display parameters specifically include: the peak brightness L_dst of the display device, which is in cd / m² and is used to represent the maximum screen brightness that the display device can output.
[0079] In one implementation, the display parameters include the peak brightness of the display device, and the metadata includes video frame extension parameters; the method may further include: Step B1: Determine the screen extension parameters based on the peak brightness of the master corresponding to the video frame and the peak brightness of the display device; Step B2: Select the smaller one from the video frame expansion parameter and the screen expansion parameter as the brightness expansion parameter; Step B3: Determine the upper limit of the mapping result value in the mapping curve based on the brightness expansion parameter.
[0080] In step B1, the screen extension parameter is obtained by calculating the ratio of the peak brightness L_dst of the display device to the peak brightness of the master image corresponding to the video frame. This screen extension parameter reflects the scalability of the display device relative to the master image brightness, providing a key hardware performance basis for subsequently determining the brightness extension range.
[0081] Step B2 selects the smaller of the video frame expansion parameter and the screen expansion parameter as the brightness expansion parameter. This brightness expansion parameter will neither exceed the allowed expansion of the video frame itself nor exceed the actual brightness capacity limit of the display device, thus achieving a balance between the dual constraints of content characteristics and device performance.
[0082] Step B3 can multiply the brightness extension parameter H by the master peak brightness corresponding to the video frame to obtain the upper limit value of the light space brightness in the light signal space. Then, the upper limit value of the light space brightness is converted to the PQ domain as the upper limit value of the mapping result value in the electrical signal space. This can directly correspond to the actual brightness that the display device can output, and at the same time conforms to the non-linear perception characteristics of the PQ domain.
[0083] In this embodiment of the invention, the upper limit value is used to characterize the upper limit to which the brightness of a video frame can be extended on a display device. Since the upper limit value is determined based on display parameters and metadata, different display parameters can correspond to different upper limit values, and different metadata can also correspond to different upper limit values.
[0084] In this embodiment of the invention, the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space. Therefore, the mapping curve can play a role in color correction of video frames.
[0085] The extended upper limit value is used to determine the upper limit of the mapped result value in the mapping curve. During color correction of video frames based on the mapping curve, this extended upper limit value constrains the mapped result value of the corrected video frame, optimizing the color space distribution of the corrected video frame. Furthermore, this extended upper limit value can adapt to the characteristics of different display devices, preventing excessive color enhancement in the corrected video frame that could lead to distortion. In summary, this extended upper limit value ensures that the color of the corrected video frame matches the display device and the video frame content, thereby improving the overall visual effect of the video frame.
[0086] In video technology, electrical signal space refers to the mathematical space in which pixel brightness and color information are represented by digital or analog electrical signal values. Optical signal space refers to the physical space in video technology in which visual effects are represented by light intensity and color distribution perceptible to the human eye.
[0087] In particular, the upper limit of the mapping result value in the mapping curve of this embodiment is determined based on display parameters and metadata. This upper limit ensures that the mapping result value does not exceed the capabilities of the display device while fully utilizing the brightness potential of the display device, reasonably expanding the brightness of the light signal of the video frame according to the actual needs of the video content. Therefore, the mapping curve corresponding to the video frame in this embodiment can make the brightness levels of the video frame image richer and the details clearer, thereby effectively improving the display effect of the video frame.
[0088] It should be noted that the mapping curve can be an explicit output mapping curve, where the mapping result value represents the target color intensity value. In this case, the mapping curve is used to characterize the mapping relationship between the source color intensity value and the target color intensity value within the electrical signal space. Alternatively, the mapping curve can be an intermediate control curve used to determine the output mapping relationship, where the mapping result value represents the target reference intensity value. In this case, the mapping gain can be determined based on the ratio of the mapping result value to the source color intensity value, and the target color intensity value can be obtained based on the mapping gain and the source color intensity value.
[0089] In this embodiment of the invention, the parameters of the mapping curve are calculated frame by frame, based on received metadata and the known peak brightness of the display device.
[0090] In a specific implementation, the mapping curve includes: mapping curves corresponding to at least two brightness intervals respectively; wherein, the boundary point between two adjacent brightness intervals is determined based on the metadata, and / or, the parameters of the mapping curves corresponding to the brightness intervals are obtained based on the metadata.
[0091] This invention employs a segmented mapping curve, dividing the mapping curve into at least two sub-curves corresponding to each brightness range. The boundary points between adjacent brightness ranges are determined based on metadata, and the parameters of the sub-curves corresponding to each brightness range are also obtained from metadata. This allows the boundary point settings to adapt to the brightness distribution of the content and enables the mapping parameters of each range to more specifically match the detail features within that brightness range, thereby achieving differentiated and refined adjustments in different brightness ranges. This effectively avoids the problem of a single global curve being unable to adequately address different brightness areas, significantly reducing distortion phenomena such as overexposure in highlights and loss of detail in shadows. While adapting to different display devices, it better preserves image layers and dynamic range, improving overall visual quality.
[0092] In one implementation, the display parameters include: the peak brightness of the display device; the metadata includes: the average brightness value; the at least two brightness intervals include: a first brightness interval, a second brightness interval, and a third brightness interval that are continuously distributed on the brightness axis and whose brightness increases sequentially. The process of determining the mapping curve corresponding to the video frame based on the display parameters and the metadata includes: Step C1: Determine the second parameter of the second mapping sub-curve corresponding to the video frame in the second brightness range based on the peak brightness of the display device and the average brightness value of the video frame.
[0093] The second mapping sub-curve corresponding to the second brightness range can be determined by those skilled in the art based on actual application requirements. For example, the second mapping sub-curve can be a curve provided by CUVA (China Ultra HD Video Alliance).
[0094] Formula (7) shows the form of the second mapper curve: (7) Where L represents the source color intensity value in the electrical signal space, which can be determined based on the normalized pixel value of the RGB (red, green, blue) electrical signal corresponding to the video frame. Assuming the pixel value of the RGB electrical signal is p, then L = p / 255. The function represents the second sub-mapping curve. mp represents the second parameter of the second sub-mapping curve.
[0095] Step C1, which determines the second parameter of the second mapping sub-curve corresponding to the video frame within the second brightness range, specifically includes: Step C11: Convert the peak brightness of the display device to the PQ domain to obtain the PQ domain display peak value, which is used to calculate the display brightness coefficient.
[0096] Step C12: Normalize the peak value of the PQ domain display within the set PQ domain brightness range to obtain the display brightness coefficient, which is used to determine the upper limit of mp.
[0097] Step C13: Calculate the upper limit value of the intermediate segment parameter mp based on the display brightness coefficient. This upper limit value is used for subsequent mp interpolation.
[0098] Step C14: Within the set average brightness range of the PQ domain, normalize the average brightness value to obtain the scene brightness coefficient t_avg, which is used to interpolate between 2.4 and the upper limit of mp to obtain mp.
[0099] Step C15: Interpolate the scene brightness coefficient between 2.4 and the upper limit of mp to obtain the preliminary mp.
[0100] In step C11, the peak brightness L_dst of the display device is converted to the PQ domain to obtain the PQ domain display peak value dst_pq, as shown in formula (8).
[0101] dst_pq=OETF_PQ(L_dst / L_Ref)(8) Where L_Ref represents the reference brightness, which can be set to L_Ref = 10000 cd / m².
[0102] In step C12, referring to formula (9), the peak value dst_pq of the PQ domain display is normalized within the set PQ domain brightness range [PQ_LO, PQ_HI] to obtain the display brightness coefficient t_display, which is used to determine the upper limit of mp. PQ_LO represents the lower limit of the set PQ domain brightness range, corresponding to the PQ domain value of the master peak brightness in the embodiment; PQ_HI represents the upper limit of the PQ domain brightness range, corresponding to the upper limit value of the PQ domain value in the embodiment.
[0103] t_display=clip((dst_pq PQ_LO) / (PQ_HI PQ_LO),0,1)(9) The clip function is a clipping function used to limit the calculation result to between 0 and 1; if the result is less than 0, it is set to 0, and if the result is greater than 1, it is set to 1.
[0104] In step C13, the upper limit value mp_max of the second parameter mp is calculated based on the display brightness coefficient t_display. This upper limit value is used for subsequent mp interpolation.
[0105] Formula (10) shows the process of determining the upper limit value of the second parameter.
[0106] mp_max = 2.4 + MP_DELTA_MAX × t_display(10) Where MP_DELTA_MAX is an adjustable normal value, and in this example, it can be set to 0.6.
[0107] In step C14, the formula for calculating the scene brightness coefficient is as follows: t_avg = clamp( (averg_pq AVG_LO) / (AVG_HI AVG_LO), 0, 1 )(11) AVG_LO and AVG_HI are preset PQ thresholds (the examples may correspond to the order of 10 cd / m² and 200 cd / m²).
[0108] In step C15, the scene brightness coefficient t_avg is interpolated between 2.4 and the upper limit of mp_max to obtain the preliminary mp. The interpolation formula is: mp = 2.4 + (mp_max 2.4) × t_avg(12) In extremely dark scenes, a bounded positive offset is superimposed, and finally, mp = max (mp, 2.4) is used to ensure that mp is not less than 2.4, thus obtaining the final CUVA curve shape parameter mp for the middle segment.
[0109] When the average brightness value is lower than the preset dark scene threshold, first add a fixed small positive number to the initial mp calculated by formula 12, and then use formula 13 "mp = max (mp, 2.4)" to calculate so that the final mp is not lower than 2.4; if the average brightness value is not lower than the preset dark scene threshold, substitute the result of formula (12) into formula (13) to get the final mp.
[0110] mp=max(mp,2.4)(13) Optionally, the display parameters include: the peak brightness of the display device; the metadata includes: dark area characteristic value, average brightness value, and maximum brightness value; the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially. The process of determining the mapping curve corresponding to the video frame based on the display parameters and the metadata specifically includes: Step C2: Determine the target slope of the first mapping sub-curve corresponding to the first brightness interval based on the peak brightness of the display device and the bright part feature value, dark part feature value, average brightness value, and maximum brightness value of the video frame; Step C3: Based on the target slope, search for the first boundary point between the first brightness interval and the second brightness interval on the corresponding second mapper curve within the second brightness interval; Step C4: Based on the continuity of function values and first derivatives of the first and second mapping sub-curves at the first boundary point, determine the first parameter of the first mapping sub-curve corresponding to the video frame in the first brightness range.
[0111] In step C2, the target slope k_target is the target slope of the first mapper curve, used to determine the boundary point p1 between the first brightness interval and the second brightness interval, with a value range of [SLO, 1]. SLO is the lower limit of the dark area slope determined based on the monitor's brightness. The brighter the screen, the lower the SLO. In dark scenes, k_target is closer to 1, and in bright scenes with high dark area values, k_target is closer to SLO.
[0112] The process of determining the target slope k_target includes: Calculate the lower limit of the dark area slope based on the peak brightness of the display device; The basic weights are determined based on the relationship between the brightness span of the dark area feature values and the bright area peak feature values.
[0113] Based on the average brightness value, dark area feature value, and high brightness peak feature, the basic weight is suppressed and / or increased to obtain the final weight.
[0114] Optionally, the process of adjusting the basic weight by suppressing and / or raising it specifically includes: Case 1: If the average brightness value is greater than 0.045 and the dark feature value is greater than 0.38, then calculate the weight after the first compression based on the base weight, then calculate the smoothing factor based on the average brightness value and the dark feature value, and finally calculate the final weight based on the weight after the first compression and the smoothing factor; or Case 2: If the peak feature value of the bright area is less than 0.45 and the feature value of the dark area is less than 0.40, calculate the brightness ratio based on the peak feature value of the bright area and the feature value of the dark area, then calculate the weight enhancement coefficient based on the brightness ratio, and finally calculate the final weight based on the weight after the first suppression and the weight enhancement coefficient.
[0115] It should be noted that if neither of the above two conditions is met, the final weight is equal to the basic weight. Finally, the target slope is calculated based on the lower limit of the dark slope and the final weight: Target slope = Lower limit of dark slope + (1 - Lower limit of dark slope) × Final weight.
[0116] Step C3 avoids using a fixed power function to divide the brightness range. Instead, it dynamically calculates the dividing point p1 based on the target slope. This makes the division of the brightness range more suitable for the brightness characteristics of the current video frame, thereby reducing semantic conflicts with the brightness extension parameters and the dependence on hard truncation.
[0117] Given the target slope k_target and the high brightness peak feature mp, search in the PQ domain in ascending order of brightness value to find the first brightness value L that satisfies formula (14). This L is denoted as p1, and p1 is the first dividing point between the first brightness interval and the second brightness interval.
[0118] CUVA(p1, mp) / p1≥ k_target(14) Wherein, CUVA represents the second mapper curve.
[0119] Step C4 first calculates the first slope parameter k1 and the second slope parameter k2 at the first boundary point p1 according to formulas (15) and (16). Here, k1 represents the ray slope of the second mapping sub-curve at p1, and k2 represents the first derivative of the second mapping sub-curve at p1. A first mapping sub-curve for dark area mapping is constructed at p1, such that the first mapping sub-curve and the second mapping sub-curve satisfy both function value continuity and first derivative continuity at p1.
[0120] k1 = CUVA(p1, mp) / p1(15) k2 = CUVA′(p1, mp)(16) Where CUVA′(p1, mp) represents the first derivative of CUVA with respect to L at p1.
[0121] In one example, the first mapper curve can be: a fourth-degree polynomial f(L) constructed on L ∈ [0, p1], satisfying: f(0) = 0 (17) f′(0) = k1 × tone_gain(18) f(p1) = CUVA(p1, mp)(19) f′(p1) = k2(20) The four constraints in formulas (17) to (20) uniquely determine the coefficients of f(L), ensuring that the first derivative of the first brightness interval and the second brightness interval is continuous at p1, and that the origin slope includes the mid-gray baseline gain tone_gain. The origin slope is calculated using the mid-gray baseline gain tone_gain, reflecting the influence of the brightness extension parameter on the overall brightening trend of the dark areas, and improving the consistency between dark area mapping and overall adjustment.
[0122] The formula for calculating gray baseline gain is shown in formula (21): tone_gain = sqrt(H)(21) Where H is the luminance extension parameter.
[0123] Optionally, the display parameters include: the peak brightness of the display device; the metadata includes: bright area feature values; the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially. The step of determining the mapping curve corresponding to the video frame based on the display parameters and the metadata specifically includes: Step C5: Determine the initial value of the second boundary point between the second brightness interval and the third brightness interval based on the peak brightness of the display device and the bright part feature value of the video frame; Step C6: Adjust the initial value of the second boundary point according to the linear domain constraint condition of the second boundary point relative to the first boundary point and the boundary condition of the second boundary point to obtain the adjusted second boundary point.
[0124] In step C5, the highlight feature value (highlight_feature) of the video frame is first used as the original brightness feature value lightE_raw: lightE_raw=shadow_feature(22) Then convert the peak brightness of the display device to the PQ threshold value: maxdisplay_pq=OETF_PQ(Ldst / LRef).
[0125] Then, lightE_raw is scaled according to the piecewise linear function scale(maxdisplay_pq) to obtain the initial value of the second boundary point p2.
[0126] The piecewise linear function `scale(maxdisplay_pq)` dynamically adjusts the output scaling factor based on the peak brightness `maxdisplay_pq` of the display device in the PQ domain. When the display's peak brightness is low, `maxdisplay_pq` is small, requiring a larger scaling factor to increase the initial value of `p2`, allowing more content in bright areas to be mapped into the second brightness range. When the display's peak brightness is high, `maxdisplay_pq` is large, the scaling factor decreases, and the initial value of `p2` decreases, preventing excessive compression of bright areas and thus matching the brightness display capabilities of different displays.
[0127] Step C6: Adjust the initial value of the second boundary point by constraint: Map the initial value of the second boundary point p2 and the first boundary point p1 to the linear domain via EOTF_PQ to satisfy the linear domain constraint condition of formula (23): EOTF_PQ(p2)≥min(4 EOTF_PQ(p1),1.0)(23) By simultaneously satisfying the constraints p2>p1 and p2≤1, the adjusted second boundary point p2 is finally obtained.
[0128] When the initial value of the second boundary point p2 does not satisfy the linear domain constraints or boundary conditions, it needs to be adjusted. For example, if EOTF_PQ(p2 initial value) < min(4 If EOTF_PQ(p1), 1.0), then increase the initial value of p2; if the initial value of p2 ≤ p1, then the initial value of p2 also needs to be increased; if the initial value of p2 > 1 or is lower than the lower limit of PQ, then the initial value of p2 should be decreased or increased accordingly, until the linear domain constraints and boundary conditions are satisfied. The adjusted second boundary point p2 is shared by the right end of the second mapping sub-curve and the left end of the third mapping sub-curve.
[0129] Optionally, the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially; The step of determining the mapping curve corresponding to the video frame based on the display parameters and the metadata specifically includes: Step C7: Based on the extended upper limit value, map the third mapper curve of the third brightness range to the V space; Step C8: Construct a cubic polynomial function in the V space; Step C9: Determine the third parameter of the cubic polynomial function according to the set constraints.
[0130] In step C7, based on the extended upper limit value H, a variable substitution v = output^(1 / H) is performed in the third brightness interval L∈[p2,xmax]. x_max can be the maximum brightness value.
[0131] Variable substitution can transform the brightness mapping relationship of the highlight region to the V space, making it easier for the subsequently constructed curves to satisfy the characteristics of highlight extension.
[0132] In step C8, a cubic polynomial function is constructed for the transformed function v(L) in the V space. This cubic polynomial will determine the relationship between the brightness L and the value of v in the V space. The output of the original brightness space is obtained by calculating the result of this cubic polynomial in the V space by raising it to the power of H: output(L)=v(L)^H.
[0133] Step C9 determines the third parameter of the cubic polynomial function based on the set constraints: Second boundary point function value constraint: At the second boundary point p2, the function value of the cubic polynomial function v(L) must be equal to the value of the output value of the second mapper curve at p2 after being converted to V space, to ensure that the brightness values of the two curves are continuous here; Second boundary point derivative value constraint: At p2, the derivative value of the cubic polynomial function must be equal to the derivative value of the second mapper curve at p2 after being transformed into V space, to ensure that the curve changes smoothly at the junction. Third brightness range upper limit function value constraint: At the upper limit xmax of the third brightness range, the function value of the cubic polynomial function must meet the requirement that the output gain is equal to H, to ensure that the highlight area can be expanded in brightness according to the set expansion factor; The third brightness interval upper limit derivative value constraint: at xmax, the derivative value of the cubic polynomial function must satisfy the condition that the ratio of the slope at the right end to the slope at the left end of the V space is H. This ensures that the mapping of the highlight segment is monotonically increasing and there will be no brightness reversal, while also making the expansion effect and H tightly coupled.
[0134] Within the third brightness interval, a cubic polynomial function v(x) is constructed in space V, which satisfies the following boundary conditions: v(p2) = yL^(1 / H) v'(p2) = kL / ( H * ( yL^(1 / H) )^(H-1) ) v(xmax) = (H * xmax)^(1 / H) v'(xmax) = H * v'(p2) Where yL represents the function value of the second mapper curve at the second boundary point p2; kL represents the first derivative of the second mapper curve at the second boundary point p2; H represents the brightness extension parameter; xmax represents the right-hand reference point of the third brightness range.
[0135] In summary, the embodiments of the present invention utilize metadata and the numerical representation (scaling) of the peak brightness of the display device in the PQ domain to determine the brightness extension parameter H, the mid-gray baseline gain tone_gain, and the shape parameter mp of the second mapping sub-curve.
[0136] Furthermore, in this embodiment of the invention, the target slope k_target of the first mapping sub-curve is jointly determined based on the scene statistics information contained in the metadata and the display capabilities, and the first dividing point p1 is searched on the second mapping sub-curve, so as to achieve continuity of function values and first derivatives of the first mapping sub-curve and the second mapping sub-curve at p1.
[0137] Furthermore, in this embodiment of the invention, the original bright area feature value lightE_raw is determined based on the bright area feature value. This original combination value of the bright area undergoes a scaling process that decreases as the peak brightness of the display device increases, resulting in the initial value of the second dividing point p2. Based on the relationship that the linear domain brightness of p2 relative to p1 is not less than approximately 4 times and the constraint that p2>p1, the initial value of the second dividing point p2 is adjusted to avoid the starting point of the third mapping sub-curve being too early or too late.
[0138] Furthermore, the first mapping sub-curve in this embodiment of the invention adopts a fourth-degree polynomial passing through the origin, satisfying the condition that the slope at the origin and the C0 and C1 continuity between the second mapping sub-curve and p1 are both continuous. C0 continuity means that the function values of two adjacent curve segments are equal, and C1 continuity means that the derivatives of two adjacent curve segments are equal.
[0139] Furthermore, in the highlight segment corresponding to the third brightness range, a cubic spline curve is constructed with p2 as the starting point of the x-axis. This curve, combined with the brightness capability of the display device, expands the brightness value of the highlight area. The cubic spline curve has second-order continuity characteristics, which allows the brightness value of the highlight area to change smoothly from the brightness corresponding to p2, thus fully presenting the highlight details.
[0140] Furthermore, this embodiment of the invention is applicable to upconversion scenarios where the peak brightness of the display device is higher than the peak brightness of the video master. The scene brightness structure is described using frame-level statistical metadata, and the margin of display capability relative to the master is described using a brightness extension parameter H. At the mapping end, a mapping curve is constructed based on the metadata and H, consisting of three segments: a dark segment (first brightness interval), a mid-gray segment (second brightness interval), and a highlight segment (third brightness interval). This mapping curve satisfies the characteristics of being smooth, continuous, and globally monotonically non-decreasing at the boundary points P1 and P2. Finally, a corresponding scalar gain is applied to each pixel in the linear optical domain to complete the brightness mapping.
[0141] In this embodiment of the invention, the mid-gray baseline uses tone_gain=sqrt(H), corresponding to a hierarchical relationship where the peak EV is halved. The second brightness range uses a CUVA-type curve, whose shape parameter mp is determined by the interpolation of the peak brightness of the display device in the PQ domain and the average brightness value representing the average brightness of the scene, making the brightening more sufficient for strong displays and bright scenes, and more conservative for weak displays and dark scenes.
[0142] In this embodiment of the invention, the determination of the first dividing point p1 is not based solely on the fixed power of H as the lower edge of the dark area. Instead, the target slope k_target is obtained by combining the lower edge SLO determined by the average brightness value, the dark area feature value, and the display peak value PQ. Then, p1=min{x:CUVA(x) / x≥k_target} is searched on the CUVA curve. The dark area is connected to p1 using a fourth-order polynomial, so that the slope at the origin and the value and derivative of CUVA at p1 are continuous.
[0143] Furthermore, in this embodiment of the invention, the method for determining the mid-gray and highlight boundary point p2 is as follows: first, the original bright part feature value lightE_raw is calculated, and then the initial value of p2 is obtained by scaling the value that decreases as the peak PQ value of the display increases. After that, it is also necessary to satisfy the relationship that the linear domain brightness of p2 is not less than about 4 times that of p1, p2>p1, and the full upper limit constraint, so as to avoid the highlight segment starting point being too early or too late.
[0144] Furthermore, the highlight segment of the mapping curve is constructed with a cubic spline in the space v=output^(1 / H). The left end is continuous with CUVA at p2 (C1), and the right end takes x_max=maximum_pq and TM(x_max)=H. The spline only covers the content region, and the monotonicity of this segment is achieved using the slope relationship k_R=k_L×H. The scalar gain here is reflected in the slope of the cubic spline of the highlight segment. The gain value calculated from the curve is finally applied to the pixels in the linear light domain.
[0145] Reference Figure 3 The diagram illustrates a step-by-step flowchart of a video data processing method according to an embodiment of the present invention. The method may specifically include the following steps: Step 301: Obtain the YUV signals of video frames in the high dynamic range video; Step 302: Convert the YUV signal into an RGB electrical signal; Step 303: Determine the source color intensity value of the video frame at the pixel point based on the RGB electrical signal; Step 304: Using the mapping curve corresponding to the video frame, the source color intensity value is mapped to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; Step 305: Based on the mapping result value, perform color correction on the RGB light signal converted from the RGB electrical signal, so as to display the color-corrected RGB light signal on a display device or encode the color-corrected RGB light signal.
[0146] In this embodiment of the invention, the source color intensity value of a video frame at a pixel is determined, and the source color intensity value is mapped using the mapping curve corresponding to the video frame to obtain a mapping result value. Based on the mapping result value, the color of the RGB light signal converted from the RGB electrical signal is corrected.
[0147] Since the upper limit of the mapping result value in the mapping curve is determined based on display parameters and metadata, the mapping result value can map the brightness-limited light signal in HDR video to a wider brightness range supported by both the video frame content and the display device. Therefore, embodiments of the present invention can improve the brightness level of video frames based on the expansion of the brightness range of the light signal, for example, making dark details more clearly visible and preventing highlights from being overexposed, thereby effectively improving the display effect of video frames.
[0148] In step 301, if the transmission of HDR video is involved, the decoding end can obtain the YUV (Luminance-Chrominance) signals of the video frames in the high dynamic range video from the bitstream sent by the encoding end. If the transmission of HDR video is not involved, the YUV signals of the video frames in the high dynamic range video can be obtained directly.
[0149] In step 302, the YUV signal can be converted into an RGB electrical signal according to the REC.2020 standard. Formula (24) shows an example of converting the YUV signal into an RGB electrical signal.
[0150] (twenty four) In step 303, the source color intensity value represents the color intensity of the video frame at the pixel point.
[0151] In one implementation, the source color intensity value of a video frame at a pixel can be selected from the pixel values of the RGB electrical signal in three channels. Specifically, the three channels include a red channel, a green channel, and a blue channel. For example, the maximum value can be selected from the pixel values of the RGB electrical signal in the three channels as the source color intensity value of the video frame at the pixel.
[0152] In another implementation, the source color intensity value of the video frame at a pixel can be determined based on the RGB electrical signal and the luminance component of the YUV signal. In determining the source color intensity value of the video frame at a pixel, considering both the RGB electrical signal and the luminance component of the YUV signal ensures that the source color intensity value conforms to the characteristics of an HDR image.
[0153] Optionally, determining the source color intensity value of the video frame at a pixel based on the RGB electrical signal and the luminance component of the YUV signal includes: Step D1: Select the RGB feature value of the pixel from the pixel values of the three channels of the RGB electrical signal; Step D2: Based on the fusion coefficient contained in the metadata, fuse the RGB feature values and luminance components of the pixel to obtain the source color intensity value of the video frame at the pixel.
[0154] In step D1, the maximum value of the RGB electrical signal in the three channels can be selected as the RGB feature value of the pixel.
[0155] In step D2, the specific methods for fusing RGB feature values and luminance components include weighted averaging, etc. It can be understood that the embodiments of the present invention do not limit the specific fusing methods.
[0156] Formula (24) shows an example of determining the source color intensity value RGBF of a video frame at a pixel. Here, the weight of the RGB feature value is the fusion coefficient factor_mix, and the weight of the luminance component Y is 1-factor_mix. The corresponding fusion process is as follows: RGBF = max(R,G,B)* factor_mix +Y*(1- factor_mix)(24) In step 304, formula (25) shows the process of mapping the above source color intensity values using the mapping curve TM(L), where tmK represents the mapping gain.
[0157] (25) In step 305, the process of color correction of the RGB light signal converted from the RGB electrical signal according to the mapping gain corresponding to the above mapping result value specifically includes: Step E1: If the mapping result value is not greater than 1, adjust the source pixel value of the RGB light signal at the pixel point according to the mapping result value, and use the first adjustment result as the target pixel value of the color-corrected RGB light signal at the pixel point; or, Step E2: If the mapping result value is greater than 1, the source pixel value of the RGB light signal at the pixel point is adjusted first according to the mapping result value to obtain the first adjustment result; and the first adjustment result is adjusted second according to the adjustment coefficient and the brightness value of the RGB light signal at the pixel point. The obtained second adjustment result is used as the target pixel value of the color-corrected RGB light signal at the pixel point.
[0158] Formula (26) illustrates the process of making the first adjustment to the source pixel value of the RGB light signal at the pixel point. Wherein, This represents the source pixel value of the RGB light signal in the red channel. This represents the source pixel value of the RGB light signal in the green channel. This represents the source pixel value of the RGB light signal in the blue channel. This indicates the first adjustment result of the RGB light signal in the red channel. This indicates the first adjustment result of the RGB light signal in the green channel. This represents the first adjustment result of the RGB light signal in the blue channel. The first adjustment process involves multiplying the mapped value by the source pixel value of the RGB light signal at the pixel, and the resulting product can be used as the first adjustment result.
[0159] Rt=tmK*EOTF(R) Gt=tmK*EOTF(G) Bt=tmK*EOTF(B)(26) If tmK > 1, then color correction is required: Rnew = Luma + (Rt - Luma) * adjusts Gnew = Luma + (Gt - Luma) * adjusts Bnew=Luma+(Bt-Luma)*adjusts(27) Here, adjustS is calculated as follows: (28) Where Luma represents the luminance value of the RGB light signal at a pixel, Luma = 0.2627 * Rt + 0.6780 * Gt + 0.0593 * Bt, where 0.2627, 0.6780, and 0.0593 represent the preset weighting coefficients for the red, green, and blue channels, respectively. The preset weighting coefficients for the red, green, and blue channels can be flexibly set based on historical experience or experimental data.
[0160] max(R, G, B) is the maximum pixel value among the R, G, and B channels in the RGB electrical signal, and min(R, G, B) is the minimum pixel value among the R, G, and B channels.
[0161] max(Rt, Gt, Bt) is the maximum value among the pixel values of the R, G, and B channels in the RGB light signal, and min(Rt, Gt, Bt) is the minimum value among the pixel values of the R, G, and B channels.
[0162] After the RGB light signal is corrected, the corrected RGB light signal can be directly output to a display device for display, or the corrected RGB light signal can be compressed and encoded for subsequent storage, transmission or further processing.
[0163] The video data processing method of this invention can be applied to end-to-end application scenarios.
[0164] The encoding end obtains the metadata of the video frames based on the YUV signals of the HDR video, and encodes the YUV signals and metadata of the HDR video to obtain the bitstream. This metadata can be dynamic metadata, which changes as the video frames change.
[0165] A bitstream can be a binary data stream. A bitstream is a binary data stream generated by encoding and compressing video data.
[0166] A bitstream can include a sequence of bits representing an encoded representation of video data. A bitstream can include encoded video frames and associated data. An encoded video frame is an encoded representation of a video frame. Associated data can include parameter sets for the bit sequence, parameter sets for the video frame, etc. The parameter set for the video frame can include metadata about the video frame.
[0167] After receiving the bitstream from the encoding end, the decoding end decodes the bitstream to obtain the YUV signal and metadata of the HDR video. Based on the metadata and the display parameters of the display device, the decoding end determines the mapping curve of the video frames, and based on the mapping curve and the YUV signal of the HDR video, determines the RGB light signal of the HDR video suitable for display on the display device.
[0168] Reference Figure 4 The diagram illustrates a step-by-step flowchart of a video data processing method according to an embodiment of the present invention. This method is applied at the encoding end and may specifically include the following steps: Step 401: Obtain the metadata of video frames in the high dynamic range video; Step 402: Encode the metadata and the high dynamic range video to obtain a bitstream; Step 403: Send the bitstream to the decoding end so that the decoding end can determine the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0169] Optionally, obtaining the metadata of video frames in a high dynamic range video includes: Construct a histogram for the grayscale features of video frames in the PQ domain; Determine the bisection points of the histogram, which are used to divide the histogram into dark and bright areas; Based on the bisection point, at least one of the following is determined: the bright part feature value, the dark part feature value, the average brightness value, and the maximum brightness value of the video frame.
[0170] Optionally, determining at least one of the bright area feature value, dark area feature value, average brightness value, and maximum brightness value of the video frame based on the bisection point includes: Based on the bisection points, the pixel brightness of the dark areas of the histogram is weighted in the PQ domain to obtain the dark feature values of the video frame; and / or Based on the bisection point, the pixel brightness of the bright areas in the histogram is transformed from the PQ domain to the linear domain. In the linear domain, the pixel brightness of the bright areas in the histogram is weighted, and then the corresponding weighted result is transformed back to the PQ domain to obtain the bright area feature values of the video frame; and / or The grayscale features of the video frame are transformed from the PQ domain to the linear domain and averaged. The average result is then transformed back to the PQ domain to obtain the average luminance value of the video frame; and / or The maximum brightness value of the video frame is obtained by calculating the maximum value of the grayscale feature in the PQ domain.
[0171] The encoding end of this invention can be a terminal device with an image acquisition device, including but not limited to: mobile phones, tablet computers, laptops, calculators, and smartwatches.
[0172] The scenarios to which this invention is applicable include, but are not limited to: 1. The HDR video at the encoding end is mapped to the CRT (Cathode Ray Tube) display at the decoding end for display; 2. The HDR video at the encoding end is mapped to the LCD (Liquid Crystal Display) at the decoding end for display; 3. The HDR video at the encoding end is mapped to the OLED (Organic Light Emitting Diode) display at the decoding end for display; 4. The HDR video at the encoding end is mapped to the minLED (Mini Light Emitting Diode) display at the decoding end for display.
[0173] The encoding solution provided in this invention, in conjunction with the decoding solution, can significantly improve the color distortion problem of HDR videos displayed on high-brightness screens (such as displays with a maximum brightness of 300 cd / m² or higher).
[0174] Reference Figure 5 This document illustrates a data flow diagram for HDR video upconversion according to an embodiment of the present invention. The process takes an HDR source encoded with PQ as input. At the generation end, it performs operations such as histogram statistics, dynamic metadata generation and writing, and then encapsulates the data into a bitstream or file for transmission and storage. At the mapping end, the bitstream is parsed, and an adaptive mapping curve is generated based on the parsed dynamic metadata and the display parameters of the display device. Pixel-by-pixel brightness mapping is then completed, and finally, the RGB light signal after brightness mapping is output to the display device to complete the image presentation. This achieves scene-feature-driven, display device-adaptive HDR content optimization display.
[0175] Figure 6 This is a flowchart illustrating the histogram statistics, iterative binary search, and metadata processing at the generation end. The process takes PQ domain RGB frames as input. First, it statistically processes the brightness information of the input frames to generate a PQ domain histogram with PQ brightness values on the horizontal axis and corresponding pixel frequencies on the vertical axis. Then, an iterative binary search algorithm is used to divide the brightness range of the histogram. Based on the obtained binary search points, brightness features such as bright area feature values, dark area feature values, average brightness value, and maximum brightness value are statistically calculated. These brightness features, along with video frame extension parameters and fusion coefficients, are then encapsulated to generate dynamic metadata, which is output to provide data support for accurate tone mapping that adapts to scene content at the downstream mapping end.
[0176] Figure 7 This is a flowchart of pixel-by-pixel tone mapping processing at the mapping end. The process takes the PQ domain RGB signal as input. First, the Y component of the PQ domain RGB signal is extracted, and the fusion coefficients in the metadata are obtained. Then, based on the Y component and the fusion coefficients, the source color intensity value of each pixel is calculated. Next, the source color intensity value of each pixel is input into three mapping curves to obtain the mapping result value for each pixel. Then, using this mapping result value as a reference, the original PQ domain RGB signal is converted to the linear optical domain and a linear gain is applied to complete color correction, resulting in a color-corrected linear optical domain RGB signal. The color-corrected linear optical domain RGB signal is then converted to the YCbCr (Luminance-Chrominance Color) space to adjust saturation, and then PQ encoded to obtain the PQ-encoded video signal. The PQ-encoded video signal can be displayed on HDR-supporting display devices, or used as input for compression encoding for efficient storage or transmission.
[0177] In summary, the video data processing method of this invention converts the bright half-area histogram back to the PQ domain after weighted statistics in the linear luminance domain: it alleviates the problem of luminance statistical center shift caused by PQ encoding compression, makes dynamic metadata such as brightness span more consistent with the physical luminance distribution, and improves the accuracy of parameter p2 and highlight segmentation.
[0178] The embodiments of the present invention employ the geometric median iterative bisection method to segment the dark half-region and the bright half-region: stable segmentation is achieved in the mixed dimension of perceived brightness and linear brightness, which has a smoother boundary effect compared with single-step segmentation. In this embodiment of the invention, the second parameter mp of the second mapping sub-curve is interpolated with the peak brightness of the display in the PQ domain, and combined with the scene average brightness modulation: under the same brightness expansion coefficient H, the brightness enhancement intensity of the middle section can be finely adjusted according to the display device and scene content, and the rate of change of the high brightness range conforms to the perceptual uniformity characteristics. The target slope of the first mapping sub-curve in this embodiment of the invention is jointly determined by scene features and display parameters, avoiding the defects of the lower edge of a single curve: reducing semantic conflicts related to brightness expansion parameters, reducing dependence on hard clipping, and balancing dark noise suppression and reasonable brightening of mid-gray areas. In this embodiment of the invention, the first boundary point p1 is searched analytically using the CUVA ray slope and the target coefficient k_target: the connection between the dark part and the middle section has a clear geometric meaning and achieves first-order continuity. In this embodiment of the invention, the second boundary point p2 is subject to both the display scaling factor and the linear domain multiple relative to p1: the starting point of the highlight segment moves forward as the peak brightness of the display increases, while avoiding the distance from p1 being too small. In this embodiment of the invention, the highlight segment corresponding to the third brightness interval is fitted with cubic splines in the v=output^(1 / H) space. x_max is set as the peak brightness of the material and satisfies TM(x_max)=H: This avoids extrapolating the brightness interval without effective content and ensures that the upper limit value H is fully utilized at the pixel peak. The slope coefficient on the right side of the highlight segment is the product of the slope coefficient on the left side and the brightness extension parameter, which can realize that the overall mapping curve is monotonically increasing. Furthermore, in this embodiment of the invention, the fusion coefficient factor_mix is used as metadata to drive the source color intensity value as a convex combination of the component maximum value max(R,G,B) and the luminance component Y, and color difference compression is achieved by combining the frame average gain: compared with a fixed weight allocation method, the mapping horizontal axis parameter of high saturation and neutral scenes can be adjusted according to the video sequence or single frame content to alleviate the problems of luminance gain mismatch and color oversaturation after upconversion.
[0179] Specifically, this embodiment of the invention determines the horizontal coordinate of the mapping curve for pixel-by-pixel query, which is the aforementioned source color intensity value. The process of determining the source color intensity value includes: determining the fusion coefficient factor_mix, which takes values in the range [0,1]. If the encoding end uses normalized integers for storage, the decoding end can restore it to the same range. Moreover, the fusion coefficient can be encoded together with the scene or platform strategy, supporting frame-by-frame or sequence-by-sequence updates; using formula (24), the source color intensity value RGBF of the video frame at the pixel point is determined according to the fusion coefficient.
[0180] The source color intensity value is used as the lookup x-coordinate of the three-segment mapping curve, and the corresponding mapping result value is obtained by looking up a table. A PQEOTF conversion is performed on the PQ domain RGB electrical signal, resulting in a linear luminance domain RGB signal. The linear luminance domain RGB signal is multiplied by the mapping gain corresponding to the mapping result value, resulting in a linear luminance domain RGB signal after luminance expansion or adjustment. This linear luminance domain RGB signal is converted to the YCbCr color space, resulting in a YCbCr signal containing the luminance component Y and the color difference components Cb and Cr. The color difference components Cb and Cr in the YCbCr signal are compressed according to the frame average gain, resulting in a color difference compressed YCbCr signal. Finally, the color difference compressed YCbCr signal is PQ encoded to output a PQ domain RGB electrical signal.
[0181] When the expansion factor H=1, the three-segment mapping curve appears as a straight line. Furthermore, the three-segment mapping curve is continuous and smooth at the two boundary points p1 and p2, without jumps or sharp angles. The product of the mapping gain and the input signal monotonically does not decrease with brightness; highlight expansion can be achieved. In addition, dark noise is controllable, and the calculation results have no non-numerical or infinite outliers. When adjusting factor_mix, the distribution and saturation of the source color intensity values (RGBF) appear continuous and predictable.
[0182] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the present invention.
[0183] Based on the description of the above method embodiments, the present invention also provides corresponding video data processing apparatus embodiments to implement the content described in the above method embodiments.
[0184] Reference Figure 8 The diagram illustrates the structure of a video data processing apparatus according to an embodiment of the present invention, the apparatus specifically comprising the following modules: Metadata acquisition module 801 is used to acquire metadata of video frames in high dynamic range video; Display parameter acquisition module 802 is used to acquire display parameters of the display device; The mapping curve determination module 803 determines the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0185] Reference Figure 9 The diagram illustrates the structure of a video decoding device according to an embodiment of the present invention, the device specifically including the following modules: The signal acquisition module 901 is used to acquire the YUV signals of video frames in high dynamic range video. Signal conversion module 902 is used to convert the YUV signal into an RGB electrical signal; The source color intensity value determination module 903 is used to determine the source color intensity value of the video frame at the pixel point based on the RGB electrical signal; The mapping module 904 is used to map the source color intensity value using the mapping curve corresponding to the video frame to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; The color correction module 905 is used to perform color correction on the RGB light signal converted from the RGB electrical signal according to the mapping result value, so as to display the color-corrected RGB light signal on a display device or encode the color-corrected RGB light signal.
[0186] Reference Figure 10 The diagram illustrates the structure of a video decoding device according to an embodiment of the present invention. The device is applied at the decoding end and specifically includes the following modules: Metadata acquisition module 1001 is used to acquire metadata of video frames from the bitstream of high dynamic range video sent from the encoding end; Display parameter acquisition module 1002 is used to acquire display parameters of the display device; The mapping module determination module 1003 is used to determine the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extended mapping result value in the mapping curve is determined based on the display parameters and the metadata.
[0187] Reference Figure 11The diagram illustrates a structural schematic of a video encoding apparatus according to an embodiment of the present invention. The apparatus is applied at the encoding end and specifically includes the following modules: Metadata acquisition module 1101 is used to acquire metadata of video frames in high dynamic range video; Encoding module 1102 is used to encode the metadata and the high dynamic range video to obtain a bitstream; The sending module 1103 is used to send the bitstream to the decoding end so that the decoding end can determine the mapping curve corresponding to the video frame according to the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined according to the display parameters and the metadata.
[0188] This invention discloses an electronic device, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the aforementioned method.
[0189] This invention discloses a non-transitory computer-readable storage medium storing instructions that cause a processor to execute the aforementioned method.
[0190] Non-transitory memory: This refers to components or devices in computer hardware used for non-temporary storage of data and instructions. It emphasizes the "non-transitory" nature of the storage, meaning the data will not be lost due to momentary factors such as the disappearance of electrical signals. Examples include ROM (Read-Only Memory) and the storage chips in solid-state drives (SSDs), which can store data long-term.
[0191] Non-transitory computer-readable recording media: Physical carriers that can permanently store computer data and are readable. Data storage is stable and not temporary. Examples include hard drives, USB flash drives, and optical discs. They can store various types of data such as programs, documents, and videos. Even if the device is powered off or restarted, the data is still retained and can be accessed by the computer later.
[0192] Non-transitory computer-readable storage media: These are part of the computer storage system and are media that can permanently store computer-readable data. They include hard drives, solid-state drives, and flash memory. Unlike transient storage, they can permanently or long-term retain data. Computers can read the stored instructions and data through corresponding interfaces and protocols for use in program execution, data processing, and other scenarios. For example, they store operating systems, application code, and user files, providing stable data support for the computer system.
[0193] Devices used for video data processing can be electronic devices. Figure 12A schematic diagram of an electronic device 1300 according to an embodiment of the present invention is shown. The electronic device 1300 specifically includes: one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.
[0194] Processor 1302 may include one or more single-core or multi-core processors, and processor 1302 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, electronic device 1300 can serve as a terminal device, server (cluster), or other device as described in the embodiments of the present invention.
[0195] In some embodiments, electronic device 1300 may include one or more computer-readable media (e.g., memory 1306 or non-volatile memory / storage device 1308) having instructions 1314 and one or more processors 1302 that are combined with the one or more computer-readable media and configured to execute instructions 1314 to implement modules and thus perform the actions described in this disclosure.
[0196] In one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1302 and / or any suitable device or component communicating with the control module 1304.
[0197] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0198] Memory 1306 may be used, for example, to load and store data and / or instructions 1314 for electronic device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM (Dynamic Random Access Memory). In some embodiments, memory 1306 may include dual-parameter data rate type quad synchronous dynamic random access memory.
[0199] In one embodiment, the control module 1304 may include one or more input / output controllers to provide an interface to the non-volatile memory / storage device 1308 and (one or more) input / output devices 1310.
[0200] For example, non-volatile memory / storage device 1308 may be used to store data and / or instructions 1314. Non-volatile memory / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives, one or more optical disk drives, and / or one or more digital universal optical disk drives).
[0201] The non-volatile memory / storage device 1308 may include storage resources that are physically part of a device on which the electronic device 1300 is mounted, or that can be accessed by the device without being part of the device. For example, the non-volatile memory / storage device 1308 may be accessed via a network via one or more input / output devices 1310.
[0202] One or more input / output devices 1310 may provide an interface for electronic device 1300 to communicate with any other suitable device. Input / output devices 1310 may include communication components, audio components, sensor components, etc. Network interface 1312 may provide an interface for electronic device 1300 to communicate via one or more networks. Electronic device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2-Generation wireless telephone technology), 3G (3-Generation wireless telephone technology), 4G (4-Generation wireless telephone technology), 5G (5-Generation wireless telephone technology), etc., or combinations thereof.
[0203] In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1304. In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers of the control module 1304 to form a system-in-package. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die to form a system-on-a-chip.
[0204] In various embodiments, electronic device 1300 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, touchscreen device, netbook, etc.). In various embodiments, electronic device 1300 may have more or fewer components and / or different architectures. For example, in some embodiments, electronic device 1300 includes one or more cameras, a keyboard, a liquid crystal display screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0205] This invention provides a machine-readable medium storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments.
[0206] Optionally, the machine-readable medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0207] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0208] It will be readily apparent to those skilled in the art that any combination of the above embodiments is feasible, and therefore any combination of the above embodiments constitutes an implementation of the present invention. However, due to space limitations, each embodiment will not be described in detail here. Although preferred embodiments of the present invention have been described, those skilled in the art, once they understand the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0209] The present invention has provided a detailed description of a video data processing method, video encoding and decoding method, apparatus, medium, and device. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A video data processing method, characterized in that, The method includes: Obtain metadata of video frames in a high dynamic range video; Obtain the display parameters of the display device; Based on the display parameters and the metadata, a mapping curve corresponding to the video frame is determined; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
2. The method according to claim 1, characterized in that, The display parameters include: the peak brightness of the display device; the metadata includes: video frame extension parameters; the method further includes: The screen extension parameters are determined based on the peak brightness of the master image corresponding to the video frame and the peak brightness of the display device. Choose the smaller one from the video frame expansion parameter and the screen expansion parameter as the brightness expansion parameter; Based on the brightness expansion parameter, determine the upper limit of the expansion of the mapping result value in the mapping curve.
3. The method according to claim 1, characterized in that, The mapping curves include: mapping curves corresponding to at least two brightness ranges respectively; wherein the boundary point between two adjacent brightness ranges is determined based on the metadata, and / or the parameters of the mapping curves corresponding to the brightness ranges are obtained based on the metadata.
4. The method according to claim 3, characterized in that, The display parameters include: the peak brightness of the display device; the metadata includes: the average brightness value; the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially. Determining the mapping curve corresponding to the video frame based on the display parameters and the metadata includes: Based on the peak brightness of the display device and the average brightness value of the video frame, the second parameter of the second mapping sub-curve corresponding to the video frame in the second brightness range is determined.
5. The method according to claim 3, characterized in that, The display parameters include: the peak brightness of the display device; the metadata includes: dark area characteristic value, average brightness value, and maximum brightness value; the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially. Determining the mapping curve corresponding to the video frame based on the display parameters and the metadata includes: Based on the peak brightness of the display device, and the bright part feature value, dark part feature value, average brightness value, and maximum brightness value of the video frame, determine the target slope of the first mapping sub-curve corresponding to the first brightness interval; Based on the target slope, search for the first boundary point between the first brightness interval and the second brightness interval on the corresponding second mapping sub-curve within the second brightness interval; Based on the continuity of function values and first derivatives of the first and second mapping sub-curves at the first boundary point, the first parameter of the video frame corresponding to the first mapping sub-curve within the first brightness range is determined.
6. The method according to claim 3, characterized in that, The display parameters include: the peak brightness of the display device; the metadata includes: bright area feature values; the at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially. Determining the mapping curve corresponding to the video frame based on the display parameters and the metadata includes: Based on the peak brightness of the display device and the bright part feature value of the video frame, the initial value of the second boundary point between the second brightness interval and the third brightness interval is determined; Based on the linear domain constraint condition of the second boundary point relative to the first boundary point, and the boundary condition of the second boundary point, the initial value of the second boundary point is adjusted to obtain the adjusted second boundary point.
7. The method according to claim 3, characterized in that, The at least two brightness ranges include: a first brightness range, a second brightness range, and a third brightness range that are continuously distributed on the brightness axis and whose brightness increases sequentially; Determining the mapping curve corresponding to the video frame based on the display parameters and the metadata includes: Based on the extended upper limit value, the third mapping sub-curve of the third brightness range is mapped to the V space; Construct a cubic polynomial function in the V space; The third parameter of the cubic polynomial function is determined based on the set constraints.
8. The method according to claim 7, characterized in that, The defined constraints include: The first constraint condition is set to constrain the function values of the cubic polynomial function and the second mapper curve at the second boundary point; The second constraint condition is set to constrain the derivative values of the cubic polynomial function and the second mapper curve at the second boundary point; The third constraint condition is set to constrain the function value of the cubic polynomial function at the upper limit of the third brightness range; Fourth, set constraints to constrain the derivative value of the cubic polynomial function at the upper limit of the third brightness range.
9. The method according to any one of claims 1 to 4, characterized in that, The metadata includes at least one of the following features: video frame extension parameters, highlight feature values, shadow feature values, average brightness value, maximum brightness value, and fusion coefficient.
10. The method according to claim 9, characterized in that, The acquisition of metadata for video frames in a high dynamic range video includes: Construct a histogram for the grayscale features of video frames in the PQ domain; Determine the bisection points of the histogram, which are used to divide the histogram into dark and bright areas; Based on the bisection point, at least one of the following is determined: the bright part feature value, the dark part feature value, the average brightness value, and the maximum brightness value of the video frame.
11. A video decoding method, characterized in that, The method includes: Acquire the YUV signals of video frames in a high dynamic range video; Convert the YUV signal into an RGB electrical signal; The source color intensity value of the video frame at the pixel point is determined based on the RGB electrical signal; The source color intensity value is mapped using the mapping curve corresponding to the video frame to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; Based on the mapping result value, the RGB light signal converted from the RGB electrical signal is color corrected so that the color-corrected RGB light signal can be displayed on a display device or encoded.
12. The method according to claim 11, characterized in that, The step of determining the source color intensity value of the video frame at a pixel point based on the RGB electrical signal includes: The source color intensity value of the video frame at each pixel is determined based on the RGB electrical signal and the luminance component of the YUV signal.
13. The method according to claim 12, characterized in that, Determining the source color intensity value of the video frame at a pixel based on the RGB electrical signal and the luminance component of the YUV signal includes: The RGB feature values of a pixel are selected from the pixel values of the three channels of the RGB electrical signal; Based on the fusion coefficient contained in the metadata, the RGB feature values and luminance components of the pixel are fused to obtain the source color intensity value of the video frame at the pixel.
14. The method according to claim 11, characterized in that, The step of color correction of the RGB light signal converted from the RGB electrical signal based on the mapping result value includes: If the mapping result value is not greater than 1, the source pixel value of the RGB light signal at the pixel point is adjusted according to the mapping result value, and the first adjustment result is used as the target pixel value of the color-corrected RGB light signal at the pixel point; or... If the mapping result value is greater than 1, the source pixel value of the RGB light signal at the pixel point is adjusted first according to the mapping result value to obtain the first adjustment result; and the first adjustment result is adjusted second according to the adjustment coefficient and the brightness value of the RGB light signal at the pixel point. The obtained second adjustment result is used as the target pixel value of the color-corrected RGB light signal at the pixel point.
15. A video decoding method, characterized in that, The method is applied at the decoding end and includes: Obtain the metadata of video frames from the bitstream of high dynamic range video sent from the encoding end; Obtain the display parameters of the display device; Based on the display parameters and the metadata, a mapping curve corresponding to the video frame is determined; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
16. A video encoding method, characterized in that, The method is applied to the encoding end and includes: Obtain metadata of video frames in a high dynamic range video; The metadata and the high dynamic range video are encoded to obtain a bitstream; The bitstream is sent to the decoding end so that the decoding end can determine the mapping curve corresponding to the video frame based on the display parameters of the display device and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
17. The method according to claim 16, characterized in that, The acquisition of metadata for video frames in a high dynamic range video includes: Construct a histogram for the grayscale features of video frames in the PQ domain; Determine the bisection points of the histogram, which are used to divide the histogram into dark and bright areas; Based on the bisection point, at least one of the following is determined: the bright part feature value, the dark part feature value, the average brightness value, and the maximum brightness value of the video frame.
18. The method according to claim 17, characterized in that, Determining at least one of the bright area feature value, dark area feature value, average brightness value, and maximum brightness value of the video frame based on the bisection point includes: Based on the bisection points, the pixel brightness of the dark areas of the histogram is weighted in the PQ domain to obtain the dark feature values of the video frame; and / or Based on the bisection point, the pixel brightness of the bright areas in the histogram is transformed from the PQ domain to the linear domain. In the linear domain, the pixel brightness of the bright areas in the histogram is weighted, and then the corresponding weighted result is transformed back to the PQ domain to obtain the bright area feature values of the video frame; and / or The grayscale features of the video frame are transformed from the PQ domain to the linear domain and averaged. The average result is then transformed back to the PQ domain to obtain the average luminance value of the video frame; and / or The maximum brightness value of the video frame is obtained by calculating the maximum value of the grayscale feature in the PQ domain.
19. A video data processing apparatus, characterized in that, The device includes: The metadata acquisition module is used to acquire metadata of video frames in high dynamic range videos. The display parameter acquisition module is used to acquire the display parameters of the display device. The mapping curve determination module determines the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extended mapping result value in the mapping curve is determined based on the display parameters and the metadata.
20. A video decoding device, characterized in that, The device includes: The signal acquisition module is used to acquire the YUV signals of video frames in high dynamic range videos; The signal conversion module is used to convert the YUV signal into an RGB electrical signal; The source color intensity value determination module is used to determine the source color intensity value of the video frame at the pixel point based on the RGB electrical signal; The mapping module is used to map the source color intensity value using the mapping curve corresponding to the video frame to obtain a mapping result value; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit of the mapping result value in the mapping curve is determined based on the display parameters of the display device and the metadata of the video frame; The color correction module is used to correct the color of the RGB light signal converted from the RGB electrical signal according to the mapping result value, so as to display the color-corrected RGB light signal on a display device or encode the color-corrected RGB light signal.
21. A video decoding device, characterized in that, The device is used at the decoding end and includes: The metadata acquisition module is used to acquire the metadata of video frames from the bitstream of high dynamic range video sent from the encoding end; The display parameter acquisition module is used to acquire the display parameters of the display device. The mapping module determination module is used to determine the mapping curve corresponding to the video frame based on the display parameters and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the expansion of the mapping result value in the mapping curve is determined based on the display parameters and the metadata.
22. A video encoding device, characterized in that, The device is applied to the encoding end and includes: The metadata acquisition module is used to acquire the metadata of video frames in high dynamic range videos. The encoding module is used to encode the metadata and the high dynamic range video to obtain a bitstream; The sending module is used to send the bitstream to the decoding end, so that the decoding end can determine the mapping curve corresponding to the video frame according to the display parameters of the display device and the metadata; the mapping curve is used to characterize the mapping relationship between the source color intensity value and the mapping result value in the electrical signal space; wherein, the upper limit value of the extension of the mapping result value in the mapping curve is determined according to the display parameters and the metadata.
23. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by means of video data processing, wherein the method includes: Obtain metadata of video frames in a high dynamic range video; The metadata and the high dynamic range video are encoded to obtain a bitstream.
24. A method for storing a bitstream of video data, comprising: Obtain metadata of video frames in a high dynamic range video; The metadata and the high dynamic range video are encoded to obtain a bitstream; Furthermore, the bit stream is stored in a non-transitory computer-readable recording medium.
25. A method for storing a bit stream, characterized in that, include: Perform the video data processing method according to any one of claims 16 to 18 to generate a bitstream; And, store the bit stream.
26. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 18.
27. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more computer programs that, when executed by one or more processors, cause the electronic device to perform the method of any one of claims 1 to 18.
28. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads from and executes the computer program, causing the electronic device to perform the method of any one of claims 1 to 18.