Video imaging method and apparatus, and electronic device and readable storage medium

By performing pixel interpolation on the target information area of ​​HDR video, an electro-optical conversion curve conforming to the second electro-optical conversion standard information is generated, which solves the problem that some pixel values ​​in HDR video cannot be displayed and achieves a better visual experience.

WO2025218186A1PCT designated stage Publication Date: 2025-10-23MIGU VIDEO TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136597
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2024-12-04
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Some pixel values ​​in existing HDR videos cannot be displayed correctly, resulting in the inability to fully utilize the color depth of the video.

Method used

By determining the target information region and performing pixel interpolation within the target information region of the first electro-optical conversion curve, a second electro-optical conversion curve conforming to the second electro-optical conversion standard information is generated, thereby realizing the conversion and display of video pixel values.

Benefits of technology

This allows target information areas with high relevance and high density to be displayed with more obvious brightness levels, ensuring that pixel values ​​are fully utilized and improving the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136597_23102025_PF_FP_ABST
    Figure CN2024136597_23102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of video processing. Provided are a video imaging method and apparatus, and an electronic device and a readable storage medium. The video imaging method comprises: acquiring an HDR video file, obtaining a plurality of first images on the basis of the HDR video file, and determining a first electro-optical transfer curve for the first images; by means of determining at least one target information region of the first images and performing pixel interpolation in the target information region of the first electro-optical transfer curve, obtaining a second electro-optical transfer curve that conforms to second electro-optical transfer standard information; generating third images on the basis of pixel values in the second electro-optical transfer curve; and obtaining a new HDR video file on the basis of a plurality of third images, and displaying the new HDR video file.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and readable storage medium for video imaging

[0001] Cross-reference to Related Applications

[0002] The present disclosure claims priority to Chinese Patent Application No. 202410480359.7, filed on April 19, 2024 in China, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] Embodiments of the present disclosure relate to the technical field of video processing, and in particular to a method, device, electronic device and readable storage medium for video imaging. BACKGROUND

[0004] In the related art, there are various HDR (High-Dynamic Range) standards and formats, including HDR10 and HLG (Hybrid Log-Gamma) and various other standards. Among them, different HDR standards support different EOTF (Electro-Optical Transfer Function), thereby supporting different electro-optical conversion curves.

[0005] Since the color depth of HDR10 supports 10 bits (pixel value 0-1023), the supported luminance peak value is 1000 nit, and the corresponding pixel value is 769. That is, for the HDR10 protocol, pixel values higher than 769 all correspond to a luminance of 1000 nit, thereby being unable to be normally displayed. Therefore, this part of pixel values cannot be fully utilized. SUMMARY

[0006] Embodiments of the present disclosure provide a method, device, electronic device and readable storage medium for video imaging, to solve the problem that part of pixel values in the HDR video cannot be displayed in the related art.

[0007] To solve the above technical problem, the present disclosure is implemented as follows:

[0008] In a first aspect, embodiments of the present disclosure provide a method for video imaging, comprising:

[0009] obtaining an HDR video file, obtaining a plurality of first images according to the HDR video file, and determining first electro-optical conversion standard information corresponding to the first images;

[0010] determining a first electro-optical conversion curve of the first images according to the first electro-optical conversion standard information;

[0011] determining at least one target information region of the first image, and performing pixel interpolation in the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information; wherein the pixel interpolation is to insert a pixel region greater than a maximum pixel threshold of the first electro-optical conversion standard information into the target information region;

[0012] generating a third image based on pixel values in the second electro-optical conversion curve; and obtaining a new HDR video file according to the plurality of third images, and displaying the new HDR video file.

[0013] Optionally, the target information region includes a strongly related pixel region in the first image, a high-density pixel region in the first image, and a strongly related high-density pixel region in the first image.

[0014] Optionally, the determining the at least one target information region of the first image includes:

[0015] converting a pixel value greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image;

[0016] extracting semantic correlation features of each image block in the second image, inputting the semantic correlation features of each image block into a pre-trained logistic regression model, and outputting a strongly related pixel region in the first image by the logistic regression model.

[0017] Optionally, the determining the at least one target information region of the first image includes:

[0018] converting a pixel value greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image;

[0019] extracting a high-texture-density image region in the second image, and taking the high-texture-density image region in the second image as a high-density pixel region in the first image.

[0020] Optionally, the determining the at least one target information region of the first image includes:

[0021] converting a pixel value greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image;

[0022] extracting semantic correlation features of each image block in the second image, inputting the semantic correlation features of each image block into a pre-trained logistic regression model, and outputting a strongly related pixel region in the first image by the logistic regression model.

[0023] extracting an image region with high texture density in the second image as a high-density pixel region in the first image;

[0024] performing intersection between the strongly-correlated pixel region in the first image and the high-density pixel region in the first image to obtain a strongly-correlated high-density pixel region in the first image.

[0025] Optionally, the pixel interpolation in the target information region of the first electro-optical conversion curve to obtain the second electro-optical conversion curve conforming to the second electro-optical conversion standard information further includes:

[0026] inserting, in the target information region, pixel values of a pixel range with pixel values greater than a maximum pixel threshold in the first image to obtain the second electro-optical conversion curve conforming to the second electro-optical conversion standard information;

[0027] correcting an amplitude of the second electro-optical conversion curve to obtain a corrected second electro-optical conversion curve, so that a maximum brightness of the corrected second electro-optical conversion curve does not exceed a maximum brightness of the second electro-optical conversion standard information.

[0028] In a second aspect, the embodiments of the present disclosure provide a device for video imaging, comprising:

[0029] a first processing module configured to obtain an HDR video file, obtain a plurality of first images from the HDR video file, and determine first electro-optical conversion standard information corresponding to the first images;

[0030] a second processing module configured to determine a first electro-optical conversion curve of the first images according to the first electro-optical conversion standard information;

[0031] a third processing module configured to determine at least one target information region of the first images, and perform pixel interpolation in the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to a second electro-optical conversion standard information, wherein the pixel interpolation is to insert a pixel region greater than a maximum pixel threshold of the first electro-optical conversion standard information into the target information region;

[0032] a display module configured to generate third images based on pixel values in the second electro-optical conversion curve, obtain a new HDR video file from the plurality of third images, and display the new HDR video file.

[0033] Optionally, the target information region includes a strongly-correlated pixel region in the first image, a high-density pixel region in the first image, and a strongly-correlated high-density pixel region in the first image.

[0034] Optionally, the third processing module comprises:

[0035] The first processing submodule is configured to convert pixel values greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image.

[0036] The second processing submodule is configured to extract semantic correlation features of each image block in the second image, input the semantic correlation features of each image block into a pre-trained logistic regression model, and output strongly correlated pixel regions in the first image by the logistic regression model.

[0037] Optionally, the third processing module comprises:

[0038] The third processing submodule is configured to convert pixel values greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image.

[0039] The fourth processing submodule is configured to extract image regions with high texture density in the second image, and take the image regions with high texture density in the second image as high-density pixel regions in the first image.

[0040] Optionally, the third processing module comprises:

[0041] The fifth processing submodule is configured to convert pixel values greater than a maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image.

[0042] The sixth processing submodule is configured to extract semantic correlation features of each image block in the second image, input the semantic correlation features of each image block into a pre-trained logistic regression model, and output strongly correlated pixel regions in the first image by the logistic regression model.

[0043] The seventh processing submodule is configured to extract image regions with high texture density in the second image, and take the image regions with high texture density in the second image as high-density pixel regions in the first image.

[0044] The eighth processing submodule is configured to obtain strongly correlated high-density pixel regions in the first image by performing an intersection operation on the strongly correlated pixel regions in the first image and the high-density pixel regions in the first image.

[0045] Optionally, the third processing module further comprises:

[0046] The correction submodule is configured to insert pixel values of a pixel range with pixel values greater than a maximum pixel threshold value in the first image in the target information region to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information; and correct an amplitude of the second electro-optical conversion curve to obtain a corrected second electro-optical conversion curve, so that a maximum brightness of the corrected second electro-optical conversion curve does not exceed a maximum brightness of the second electro-optical conversion standard information.

[0047] In a third aspect, an electronic device is provided, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, and the program or instructions, when executed by the processor, implement the steps in the method for video imaging according to any one of the first aspect.

[0048] In a fourth aspect, a readable storage medium is provided, which stores a program or instructions, and the program or instructions, when executed by a processor, implement the steps in the method for video imaging according to any one of the first aspect.

[0049] In a fifth aspect, a computer program product is provided, which includes computer instructions, and the computer instructions, when executed by a processor, implement the steps in the method for video imaging according to any one of the first aspect.

[0050] In the present disclosure, by determining at least one target information region of the first image, and performing pixel interpolation in the target information region of the first electro-optical conversion curve, a second electro-optical conversion curve conforming to second electro-optical conversion standard information is obtained, by shaping the electro-optical conversion curve and converting the pixel values of the video, the target information region with high correlation and high density information can be displayed with more obvious difference in brightness level, and it is ensured that the pixel values can be fully utilized, thereby providing better visual experience for users, and solving the problem that part of the pixel values in the HDR video in the related art cannot be displayed. BRIEF DESCRIPTION OF DRAWINGS

[0051] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present disclosure. Moreover, the same reference numerals in the attached drawings indicate the same or similar components. In the drawings:

[0052] FIG. 1 is a flowchart of a method for video imaging according to an embodiment of the present disclosure;

[0053] FIG. 2 is a schematic diagram of a first electro-optical conversion curve according to an embodiment of the present disclosure;

[0054] FIG. 3 is a first electro-optical conversion curve diagram of another method of video imaging provided by the embodiments of the present disclosure;

[0055] FIG. 4 is a first electro-optical conversion curve diagram of yet another method of video imaging provided by the embodiments of the present disclosure;

[0056] FIG. 5 is a second electro-optical conversion curve diagram of a method of video imaging provided by the embodiments of the present disclosure;

[0057] FIG. 6 is an image blocking diagram of a method of video imaging provided by the embodiments of the present disclosure;

[0058] FIG. 7 is an image blocking diagram of another method of video imaging provided by the embodiments of the present disclosure;

[0059] FIG. 8 is a first electro-optical conversion curve diagram of still another method of video imaging provided by the embodiments of the present disclosure;

[0060] FIG. 9 is a second electro-optical conversion curve diagram of another method of video imaging provided by the embodiments of the present disclosure;

[0061] FIG. 10 is a terminal device diagram of a method of video imaging provided by the embodiments of the present disclosure;

[0062] FIG. 11 is a structural diagram of a device of video imaging provided by the embodiments of the present disclosure;

[0063] FIG. 12 is a structural diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION

[0064] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present disclosure.

[0065] Referring to FIG. 1, the present disclosure provides a method of video imaging, comprising:

[0066] Step 11: obtaining an HDR video file, obtaining a plurality of first images according to the HDR video file, and determining first electro-optical conversion standard information corresponding to the first images;

[0067] In the embodiments of the present disclosure, the HDR standard information (HDR format) and the transfer characteristics related to the HDR standard of the HDR (High-Dynamic Range, high dynamic range) video file are read; for example, the HDR standard information is SMPTE ST 2086, which can be compatible with the HDR10 protocol, and the transfer characteristics, i.e., the electro-optical conversion characteristics, are PQ (Perceived Quality, perceived quality), and the maximum brightness is 1000 nit; according to the read HDR standard information, it is determined that the first electro-optical conversion standard information is a PQ conversion standard with a maximum brightness value of 1000 nit; so as to determine the first electro-optical conversion curve of the first image according to the first electro-optical conversion standard information, thereby shaping the first electro-optical conversion curve and converting the pixel value of the video, so that the target information area with high correlation and high density information can be displayed with more obvious difference in brightness level.

[0068] Step 12: determining the first electro-optical conversion curve of the first image according to the first electro-optical conversion standard information;

[0069] In the embodiments of the present disclosure, the PQ standard is an absolute standard rather than a relative standard, and under the PQ standard, there is a corresponding relationship between the pixel value (or gray value, level value) of the first image of the HDR video and the brightness value displayed by the terminal device as shown in FIG. 2; it can be seen that the PQ standard provides a corresponding relationship between the pixel value of 0-1023 and the brightness value of 0-10000 nit regardless of any terminal device or any HDR standard, and the conversion curve is unchanged; for example, the maximum display brightness supported by the HDR10 protocol is 1000 nit, and the corresponding maximum pixel value is 769. Since the maximum brightness value supported by the HDR10 protocol is only 1000 nit, the HDR10 image based on the PQ standard only uses a subset of the pixel value range, i.e., 0-769, and the 254 pixel values from 770 to 1023 are cut off as shown in FIG. 3, thereby causing the brightness corresponding to the pixel values from 770 to 1023 to be 1000 nit during the display process of the HDR10 protocol image, thereby failing to be normally displayed, i.e., the part of the pixel values cannot be fully utilized, and therefore it is necessary to determine the first electro-optical conversion curve of the first image, thereby shaping the first electro-optical conversion curve and converting the pixel value of the video, so that the target information area with high correlation and high density information can be displayed with more obvious difference in brightness level.

[0070] Step 13: obtaining a second electro-optical conversion curve conforming to second electro-optical conversion standard information by determining at least one target information region of the first image and performing pixel interpolation in the target information region of the first electro-optical conversion curve; wherein the pixel interpolation is inserting a pixel region greater than a maximum pixel threshold of the first electro-optical conversion standard information into the target information region;

[0071] In the embodiments of the present disclosure, by extracting semantic features and texture images in the first image, at least one target information region is determined, and pixel interpolation is performed in the target information region, as shown in FIG. 4, that is, pixel values in a pixel range r2 with a pixel value greater than a maximum pixel threshold are inserted into the target information region r1, and the second electro-optical conversion curve as shown in FIG. 5 is obtained, so that the target information region with high correlation and high density information can be displayed with more obvious difference in brightness level, and it is ensured that the pixel value can be fully utilized, thereby providing better visual experience for users.

[0072] Step 14: generating a third image based on pixel values in the second electro-optical conversion curve; and obtaining a new HDR video file according to the plurality of third images, and displaying the new HDR video file.

[0073] In the embodiments of the present disclosure, the second electro-optical conversion curve conforms to second electro-optical conversion standard information, and the second electro-optical conversion standard information can be but is not limited to HLG standard, a second electro-optical conversion function with a maximum brightness of 1000 nit is determined according to the HLG standard; a mapping table between pixel values of the first electro-optical conversion curve and the second electro-optical conversion curve is obtained according to the first electro-optical conversion curve and the second electro-optical conversion curve, a third image is obtained according to the mapping table, and pixel values of the third image cover 0-1023; and a new HDR video file is obtained according to the plurality of third images; the new HDR video file is displayed according to the second electro-optical conversion standard information, which realizes automatic identification of the HDR standard according to the HDR video file, automatic determination of the electro-optical conversion standard, display according to the determined electro-optical conversion standard, and deformation of the electro-optical conversion curve, thereby fully utilizing the color depth of the HDR image, and displaying the high correlation and high density information region with more obvious difference in brightness level.

[0074] In the embodiments of the present disclosure, at least one target information region of the first image is determined, and pixel interpolation is performed on the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to a second electro-optical conversion standard information. By shaping the electro-optical conversion curve and converting the pixel value of the video, the target information region with high correlation and high density information can be displayed in a more obvious difference in brightness level, the pixel value can be fully utilized, and better visual experience is brought to the user, thereby solving the problem that part of the pixel value in the HDR video cannot be displayed in the related art.

[0075] In an optional embodiment of the present disclosure, the target information region includes: a pixel region with strong correlation in the first image, a pixel region with high density in the first image, and a pixel region with strong correlation and high density in the first image.

[0076] In the embodiments of the present disclosure, the pixel region with strong correlation in the first image is determined by extracting semantic features in the first image; the pixel region with high density in the first image is determined by extracting texture images in the first image; at least one target information region is determined from the pixel region with strong correlation in the first image, the pixel region with high density in the first image, or the pixel region with strong correlation and high density in the first image, so that pixel interpolation is performed on the target information region, and the target information region with high correlation and high density information can be displayed in a more obvious difference in brightness level, thereby bringing better visual experience to the user.

[0077] In an optional embodiment of the present disclosure, the determination of the at least one target information region of the first image includes:

[0078] The pixel value greater than the maximum pixel threshold in the first image is converted into the maximum pixel threshold to obtain a second image.

[0079] The semantic correlation features of each image block in the second image are extracted, and the semantic correlation features of each image block are input into a pre-trained logistic regression model. The logistic regression model outputs the pixel region with strong correlation in the first image.

[0080] In the embodiments of the present disclosure, the pixel value greater than the maximum pixel threshold in the first image is converted into the maximum pixel threshold to obtain a second image, specifically including: if the pixel value of the pixel in the first image is less than or equal to the maximum pixel threshold (i.e., 769), the pixel value of the pixel is maintained unchanged; if the pixel value of the pixel in the first image is greater than the maximum pixel threshold, the maximum pixel threshold is taken as the pixel value of the pixel, thereby obtaining the second image, and the specific formula is:

[0081] Wherein, v_pix is the pixel value of a pixel in the first image, th_pix is the maximum pixel threshold; the pixel value of each pixel in the first image is between 0 and th_pix through pixel processing;

[0082] The semantic correlation features corresponding to each image block in the second image are extracted from the second image, and according to the semantic correlation features, it is determined whether each image block belongs to strong correlation or weak correlation, specifically including:

[0083] Referring to FIG. 6, the second image is divided into multiple image regions a, b, c, …, l, for example, the size of each image region can be 256*256 or 128*128, and in this embodiment, the size of each image region is taken as 128*128 as an example;

[0084] Further, referring to FIG. 7, the image region a is further divided into multiple 16*16 image blocks a1-a 64 ; By analogy, the image region b is divided into multiple 16*16 image blocks b1-b b64 ; Each image region c-l is divided into multiple 16*16 image blocks (c1-c 64 ; d1-d 64 ; …; l1-l 64 ; ) ;

[0085] The second image can be but not limited to down-sampled by using the convolution layer and the pooling layer of the convolutional neural network, so as to convert each image region a-l into 16*16 image blocks a'-l', and obtain a feature map after down-sampling; the image blocks a'-l' are sequentially input into the trained visual processing model, so as to generate semantic feature vectors corresponding to each image block a'-l' respectively; wherein, the visual processing model can but not limited to adopt ViT (Vision transformer) to extract class token corresponding to each image block as a semantic feature vector; for example, the semantic feature vector VA corresponding to the image block a', the semantic feature vector VB corresponding to the image block b', …, the semantic feature vector VL corresponding to the image block l'; the semantic feature vector VA can reflect the semantic feature of the image region a in the entire second image, that is, the global semantic feature of the image region a; the semantic feature vector VB can reflect the semantic feature of the image region b in the entire second image, that is, the global semantic feature of the image region b. By analogy, the remaining image regions are the same.

[0086] The image blocks a1-a 64 are sequentially input into the trained visual processing model, so as to generate semantic feature vectors corresponding to each image block a1-a 64Corresponding semantic feature vectors va1~va 64 Each semantic feature vector va1~va 64 Each image block a1~a 64 Local semantic features in the image block a;

[0087] For each image block a1~a 64 , its semantic feature vector va1~va 64 is calculated respectively, and the correlation between the semantic feature vector va1~va 64 and the global semantic feature VA of the image area a is calculated, so as to calculate the corresponding correlation value Ra1~Ra 64 , respectively. 64 , wherein Ra1 represents the correlation value between va1 and VA; Ra2 represents the correlation value between va2 and VA; and so on, Ra 64 represents the correlation value between va 64 and VA.

[0088] For each image block b1~b 64 , c1~c 64 and l1~l 64 , the correlation value Rb1~Rb 64 , Rc1~Rc 64 and Rl1~Rl 64 between the semantic feature vector vb1~vb 64 , vc1~vc 64 and vl1~vl 64 and the semantic feature vector VB~VL of the corresponding image area b~l is calculated.

[0089] Then, through the pre-trained logistic regression model, according to the correlation values calculated in the foregoing steps, it is determined whether each image block and the corresponding image area belong to strong correlation or weak correlation, and the image area with strong correlation is taken as the target information area. By extracting the semantic features in the first image, at least one target information area is determined, and pixel interpolation is performed in the target information area, so that the target information area with strong correlation can be displayed with more obvious difference in brightness level, thereby providing better visual experience for users.

[0090] In an optional embodiment of the present disclosure, the determination of the at least one target information area of the first image comprises:

[0091] Converting the pixel value greater than the maximum pixel threshold in the first image to the maximum pixel threshold to obtain a second image;

[0092] extracting an image region with high texture density in the second image as a high-density pixel region in the first image.

[0093] In the embodiments of the present disclosure, the pixel value greater than the maximum pixel threshold in the first image is converted into the maximum pixel threshold to obtain a second image, specifically including: if the pixel value of a pixel in the first image is less than or equal to the maximum pixel threshold (i.e. 769), the pixel value of the pixel is maintained unchanged; if the pixel value of the pixel in the first image is greater than the maximum pixel threshold, the maximum pixel threshold is taken as the pixel value of the pixel, thereby obtaining the second image, and the specific formula is:

[0094] wherein v_pix is the pixel value of a pixel in the first image, and th_pix is the maximum pixel threshold; the pixel value of each pixel in the first image is between 0 and th_pix through pixel processing;

[0095] extracting an image region with high texture density in the second image, specifically including:

[0096] performing texture feature extraction on the second image by using a texture feature extraction algorithm to obtain a texture image corresponding to the second image, wherein the texture feature extraction algorithm can be but is not limited to an LBP (Local Binary Pattern) algorithm; wherein the texture density of different regions in the same image is different, and the greater the texture density is, the more image information the region contains; the smaller the texture density is, the less image information the region contains;

[0097] inputting the texture image into a pre-trained high-density information region identification model, so as to identify a high-density information region in the texture image by using the high-density information region identification model, thereby identifying a high-density information region with rich information content in the processed second image; wherein the high-density information region identification model can be constructed based on a Faster RCNN (Faster Region-Convolutional Neural Networks) network or a target detection model of a yolo v3 network in related technologies; or texture images with high-density information regions can be pre-collected, and artificial labeling can be performed on the high-density information regions of the texture images, thereby constructing a sample set for training the high-density information region identification model, and the target detection model is trained by using the sample set, so as to identify a high-density information region from the texture image.

[0098] In the embodiments of the present disclosure, the high-density pixel region in the first image is extracted to determine at least one target information region, and pixel interpolation is performed in the target information region, so that the high-density target information region can be displayed in a more obvious difference in brightness level, thereby providing a better visual experience for the user.

[0099] In an optional embodiment of the present disclosure, the determination of the at least one target information region of the first image comprises:

[0100] Converting the pixel value greater than the maximum pixel threshold in the first image into the maximum pixel threshold to obtain a second image;

[0101] Extracting semantic correlation features of each image block in the second image, and inputting the semantic correlation features of each image block into a pre-trained logistic regression model, wherein the logistic regression model outputs a strongly correlated pixel region in the first image;

[0102] Extracting an image region with high texture density in the second image, and taking the image region with high texture density in the second image as a high-density pixel region in the first image;

[0103] Taking the intersection of the strongly correlated pixel region in the first image and the high-density pixel region in the first image to obtain a strongly correlated high-density pixel region in the first image.

[0104] In the embodiments of the present disclosure, converting the pixel value greater than the maximum pixel threshold in the first image into the maximum pixel threshold to obtain a second image specifically comprises: if the pixel value of a pixel in the first image is less than or equal to the maximum pixel threshold (i.e., 769), the pixel value of the pixel is maintained unchanged; if the pixel value of the pixel in the first image is greater than the maximum pixel threshold, the maximum pixel threshold is taken as the pixel value of the pixel, thereby obtaining the second image, and the specific formula is:

[0105] Wherein, v_pix is the pixel value of a pixel in the first image, and th_pix is the maximum pixel threshold; through pixel processing, the pixel value of each pixel in the first image is between 0 and th_pix;

[0106] Extracting semantic correlation features corresponding to each image block in the second image from the second image, and determining whether each image block belongs to strong correlation or weak correlation according to the semantic correlation features, specifically comprising:

[0107] 6 , the second image is divided into a plurality of image regions a, b, c, ... l. For example, the size of each image region may be 256*256 or 128*128. In this embodiment, the size of each image region is 128*128 as an example.

[0108] Further, referring to FIG7 , the image area a is further divided into a plurality of 16*16 image blocks a1 to a 64 ; Similarly, the image area b is divided into multiple 16*16 image blocks b1~ b64 Each image region c~l is divided into multiple 16*16 image blocks (c1~c 64 ;d1~d 64 ;...;l1~l 64 ;);

[0109] The second image can be downsampled using, but not limited to, the convolutional layer and pooling layer of a convolutional neural network, thereby converting each image region a to l into a 16*16 image block a' to l' to obtain a downsampled feature map; the image blocks a' to l' are sequentially input into a trained visual processing model to generate semantic feature vectors corresponding to each image block a' to l'; wherein the visual processing model can, but is not limited to, using ViT to extract the class token corresponding to each image block as a semantic feature vector; such as: a semantic feature vector VA corresponding to image block a', a semantic feature vector VB corresponding to image block b', ..., a semantic feature vector VL corresponding to image block l'; the semantic feature vector VA can reflect the semantic features of image region a in the entire second image, i.e., the global semantic features of image region a; the semantic feature vector VB can reflect the semantic features of image region b in the entire second image, i.e., the global semantic features of image region b. And so on, the same applies to the remaining image regions.

[0110] Then the image blocks a1~a 64 are sequentially input into the trained visual processing model to generate the image blocks a1 to a 64 The corresponding semantic feature vectors va1~va 64 , each semantic feature vector va1~va 64 Considered as each image block a1~a 64 Local semantic features in image patch a;

[0111] For each image block a1~a 64 , calculate their semantic feature vectors va1~va respectively 64 The correlation between the global semantic feature VA of the image area a and the image blocks a1 to a1 is calculated respectively. 64Corresponding correlation values Ra1~Ra 64 wherein, Ra1 represents the correlation value between va1 and VA; Ra2 represents the correlation value between va2 and VA; and so on, Ra 64 represents the correlation value between va 64 and VA;

[0112] Correlation values Rb1~Rb 64 , Rc1~Rc 64 and Rl1~Rl 64 between semantic feature vectors vb1~vb 64 , vc1~vc 64 and vl1~vl 64 of respective image blocks b1~b 64 , c1~c 64 and l1~l 64 and semantic feature vectors VB~VL of corresponding image regions b~l are calculated in the same way.

[0113] Then, according to the correlation values calculated in the foregoing steps, the strong correlation or weak correlation between each image block and the corresponding image region is determined by using a pre-trained logistic regression model, so as to obtain the strongly correlated image region in the first image;

[0114] Then, texture feature extraction is performed on the second image by using a texture feature extraction algorithm, so as to obtain a texture image corresponding to the second image, wherein the texture feature extraction algorithm may be but is not limited to LBP algorithm; wherein, in the same image, the texture density of different regions is different, and the greater the texture density is, the more image information the region contains; the smaller the texture density is, the less image information the region contains;

[0115] The texture image is input into a pre-trained high-density information region identification model, so that the high-density information region is identified in the texture image by using the high-density information region identification model, thereby identifying the high-density information region with rich information content in the processed second image, and obtaining the high-density pixel region in the first image;

[0116] The strongly correlated pixel region in the first image and the high-density pixel region in the first image are intersected, so as to obtain the strongly correlated high-density pixel region in the first image;

[0117] In the embodiments of the present disclosure, by extracting the semantic features and the texture image in the first image, at least one target information region is determined, and pixel interpolation is performed in the target information region, as shown in FIG. 4, that is, the pixel value in the pixel range r2 with a pixel value greater than the maximum pixel threshold is inserted into the target information region r1, to obtain the second electro-optical conversion curve as shown in FIG. 5, so that the target information region with high correlation and high density information can be displayed in a more obvious difference in brightness level, thereby bringing a better visual experience to the user.

[0118] In an optional embodiment of the present disclosure, the pixel interpolation in the target information region of the first electro-optical conversion curve to obtain the second electro-optical conversion curve conforming to the second electro-optical conversion standard information further includes:

[0119] inserting the pixel value in the pixel range with a pixel value greater than the maximum pixel threshold in the first image into the target information region to obtain the second electro-optical conversion curve conforming to the second electro-optical conversion standard information;

[0120] correcting the amplitude of the second electro-optical conversion curve to obtain a corrected second electro-optical conversion curve, so that the maximum brightness of the corrected second electro-optical conversion curve does not exceed the maximum brightness of the second electro-optical conversion standard information.

[0121] In the embodiments of the present disclosure, referring to FIG. 4, the pixel value in the pixel range r2 with a pixel value greater than the maximum pixel threshold is inserted into the target information region r1 to obtain the second electro-optical conversion curve as shown in FIG. 5, wherein, referring to FIG. 8, the target information region can also be a plurality of discrete ranges r 1.1 and r 1.2 , so that the pixel value ranges r 1.1 and r 1.2 are target information regions of pixel value ranges to be interpolated;

[0122] Specifically, for the curve shown in FIG. 4 or FIG. 8, the interpolation operation is performed on the pixel value range r1(r 1.1 or r 1.2 ) curve segment, and the interpolation expansion is performed along the vertical axis (i.e., the brightness axis) in equal proportion; and for other curve segments, the corresponding translation can be performed along the horizontal axis and the vertical axis based on the range of the pixel value range r1 curve segment after the interpolation expansion, without changing the shape of the curve segment; because it is considered that the maximum pixel value supported by the HDR image is only 1023, the number of pixel values that can be interpolated does not exceed the number of pixel values in the pixel value range r2, and in addition, uniform interpolation can be performed in the pixel value range r1 to obtain the second electro-optical conversion curve as shown in FIG. 5;

[0123] Please refer to FIG. 5, the width of the pixel value range r3 after interpolation is larger than the range r1 before interpolation, and the range of luminance is also correspondingly expanded relative to the curve segment of r1, so that the high correlation high density information region can be expressed by more levels of pixel values and luminance levels;

[0124] wherein the minimum pixel value e1 corresponding to the curve segment r3 is the same as the minimum pixel value m1 corresponding to the curve segment r1, after interpolation, the pixel value range and the luminance value range are expanded along the horizontal and vertical coordinate axes in the positive direction, and the curve segments after r1 are correspondingly translated, corresponding to the curve segments after the pixel value range r3 in FIG. 5;

[0125] Specifically, the pixel interpolation method is as follows:

[0126] First, the first electro-optical conversion curve as shown in FIG. 4 conforms to the following function:

[0127] E=Norm(p);

[0128] L=EOTF1(E);

[0129] wherein the function Norm(p) is a mapping function for mapping the pixel value of 0-1024 to the numerical value space range of 0-1 (i.e. non-linear color value), specifically, a linear mapping method can be adopted, for example, Norm(p)=p / 1024.

[0130] EOTF1(E) is the first light-electric conversion function of the first electro-optical conversion curve, for mapping the non-linear color value with a numerical value range of 0-1 to the luminance space range with a numerical value range of 0-10000 nit:

[0131] Please refer to FIG. 5, for the interval of pixel value 0-p1, the corresponding luminance value is calculated according to the formula E=Norm(p);L=EOTF1(E).

[0132] For the interval of pixel value p1-p3, the original curve segment p1-p2 shown in FIG. 4 is horizontally and vertically stretched by the same ratio for interpolation; specifically, for a pixel value p between p1-p3, the following steps are taken for calculation:

[0133] Calculate the difference d1=p-p1 between the pixel value p and the pixel value p1.

[0134] According to the difference d1, the pixel value p is mapped back to p' in the original pixel range space according to the ratio:

[0135] d2=(d1*r1) / r3;

[0136] p’=p1+d2;

[0137] The luminance value L corresponding to the pixel value p in FIG. 5 is calculated according to the following formula:

[0138] e' = Norm(p');

[0139] The luminance range of e1-e3 in FIG. 5 is calculated according to the pixel value range of p1-p3 by the above formula;

[0140] For the pixel value p in the pixel value range of p3-1023

[0141] According to the formula p' = p-r2, the pixel value in this pixel value range is mapped back to the pixel value range shown in FIG. 4.

[0142] The luminance L corresponding to the pixel value p is calculated according to the following formula:

[0143] e' = Norm(p');

[0144] L = EOTF(e') + e3-e2;

[0145] Where e3 is the luminance value corresponding to p3 in FIG. 5, and e2 is the luminance value corresponding to p2 in FIG. 4.

[0146] Thus, the pixel segment corresponding to the interval range of p3-1023 in FIG. 5 is obtained in this way;

[0147] In addition, for the pixel range r 1.1 and r 1.2 , interpolation can be performed respectively to expand the range of pixel values, wherein the total number of pixel values inserted into r 1.1 and r 1.2 does not exceed the number of pixel values in the pixel value range r2, and the shape of other curve portions remains unchanged except for the pixel ranges r 1.1 and r 1.2 .

[0148] As shown in FIG. 5, the luminance value of the converted second electro-optical conversion curve is slightly greater than 1000 nit, so the amplitude of the second electro-optical conversion curve needs to be adjusted. The adjustment can be, but is not limited to, a proportional shrinkage manner, so that the maximum luminance of the adjusted second electro-optical conversion curve does not exceed the maximum luminance of the second electro-optical conversion standard information, i.e., the maximum luminance is 1000 nit, so as to obtain the final converted second electro-optical conversion curve, as shown in FIG. 9. By adjusting the second electro-optical conversion curve, it is ensured that all pixel values can be fully utilized, thereby providing users with a better visual experience.

[0149] In an optional embodiment of the present disclosure, the third image is generated based on the pixel value in the second electro-optical conversion curve; and a new HDR video file is obtained based on the plurality of third images, and the new HDR video file is displayed, including:

[0150] A mapping table between the pixel values of the first electro-optical conversion curve and the second electro-optical conversion curve is obtained based on the first electro-optical conversion curve and the second electro-optical conversion curve, and the third image is obtained based on the mapping table; and a new HDR video file is obtained based on the plurality of third images.

[0151] The new HDR video file is displayed based on the second electro-optical conversion standard information.

[0152] In an embodiment of the present disclosure, a mapping table between the pixel values of the first electro-optical conversion curve and the second electro-optical conversion curve is obtained based on the first electro-optical conversion curve and the second electro-optical conversion curve, as shown in Table 1:

[0153] Table 1

[0154] As shown in FIGS. 4 and 5:

[0155] For the pixel value of 0-p1 in the first image, p'=p is taken, where p' is the interpolated pixel value;

[0156] For the pixel value of p1-p2 in the first image, p'=p1+(p-p1)*r3 / r1 is taken;

[0157] Therefore, the pixel value p1 in FIG. 4 corresponds to the pixel value p1 in FIG. 5, and the pixel value p2 in FIG. 4 corresponds to the pixel value p3 in FIG. 5;

[0158] For the pixel value of p2-769 in the first image, p'=p+r2 is taken.

[0159] Using the mapping table, the pixel value of each pixel of the second image is converted into a corresponding interpolated pixel value, thereby generating a third image with pixel values covering 0-1023, and a new HDR video file is obtained based on the plurality of third images; and the new HDR video file is displayed based on the second electro-optical conversion standard information.

[0160] In the embodiments of the present disclosure, the second electro-optical conversion curve conforms to second electro-optical conversion standard information, which can be but is not limited to HLG standard. According to the HLG standard, a second electro-optical conversion function with a maximum luminance of 1000 nit is determined. A mapping table between pixel values of the first electro-optical conversion curve and the second electro-optical conversion curve is obtained according to the first electro-optical conversion curve and the second electro-optical conversion curve. A third image is obtained according to the mapping table, and pixel values of the third image cover 0-1023. A new HDR video file is obtained according to the plurality of third images. The new HDR video file is displayed according to the second electro-optical conversion standard information, which realizes automatic identification of the HDR standard according to the HDR video file, automatic determination of the electro-optical conversion standard, and display according to the determined electro-optical conversion standard. The electro-optical conversion curve is deformed to fully utilize the color depth of the HDR image, so that the high-relevance high-density information area is displayed with more obvious difference in luminance level.

[0161] In the embodiments of the present disclosure, the method for video imaging is implemented by a terminal device 100 as shown in FIG. 10. The terminal device 100 includes an HDR video receiving module 110, an HDR standard identification module 120, an electro-optical conversion standard adaptation module 130, and an HDR video playing module 140.

[0162] The HDR video receiving module 110 is configured to receive an HDR video file and send the received HDR video file to the HDR standard identification module 120.

[0163] The HDR standard identification module 120 is configured to identify standard information of the HDR video file according to the received HDR video file, that is, to obtain a plurality of first images according to the HDR video file and determine first electro-optical conversion standard information corresponding to the first images. The identified first electro-optical conversion standard information is sent to the electro-optical conversion standard adaptation module 130.

[0164] The electro-optical conversion standard adaptation module 130 is configured to determine electro-optical conversion standard information adapted to the standard of the HDR video file according to the identified first electro-optical conversion standard information, that is, to determine a first electro-optical conversion curve of the first image according to the first electro-optical conversion standard information. A second electro-optical conversion curve conforming to second electro-optical conversion standard information is obtained by determining at least one target information region of the first image and performing pixel interpolation in the target information region of the first electro-optical conversion curve. The pixel interpolation is to insert pixel values of a pixel range with pixel values greater than a maximum pixel threshold in the first image into the target information region. The electro-optical conversion standard information is used to indicate the electro-optical conversion standard adapted to the standard of the HDR video file.

[0165] The electro-optical conversion standard adaptation module 130 comprises:

[0166] The electro-optical conversion standard determination sub-module 131 is configured to determine a first electro-optical conversion curve of the first image according to the first electro-optical conversion standard information, for example, reading the HDR format (HDR standard information) and the transfer characteristics (transfer characteristics) related to the HDR standard of the HDR video file; for example, the HDR standard information is SMPTE ST 2086, which is compatible with the HDR10 protocol, and the transfer characteristics, i.e., the electro-optical conversion characteristics, are PQ, and the maximum luminance is 1000 nit; according to the read HDR standard information, it is determined that the first electro-optical conversion standard information is the PQ conversion standard with a maximum luminance value of 1000 nit;

[0167] The luminance matching sub-module 132 is configured to obtain the maximum display luminance supported by the terminal device 100, and then compare the maximum display luminance supported by the terminal device 100 with the maximum luminance supported by the determined electro-optical conversion standard to determine whether the terminal device 100 can support the luminance range covered by the determined electro-optical conversion standard; for example, according to the determined PQ standard, it can be determined that the maximum luminance covered is 1000 nit, in this case, if the maximum display luminance supported by the terminal device 100 is greater than or equal to 1000 nit, it is determined that the terminal device 100 supports the luminance of the determined electro-optical conversion curve; otherwise, if the maximum display luminance supported by the terminal device 100 is less than 1000 nit, for example, the maximum display luminance supported by the terminal device 100 is 400 nit, it is determined that the terminal device 100 does not support the luminance covered by the electro-optical conversion curve; in the case where it is determined that the terminal device 100 can support the luminance range covered by the electro-optical conversion curve, the video image of the HDR video file is displayed based on the electro-optical conversion standard, and the information related to the electro-optical conversion standard is transmitted to the HDR video playing module 140 by the output interface sub-module 134.

[0168] The electro-optical conversion optimization sub-module 133 is configured to optimize the HDR image of the HDR10 video of the HDR10 protocol; that is, by determining at least one target information region of the first image, and performing pixel interpolation in the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information; wherein the pixel interpolation is to insert the pixel value of the pixel range in which the pixel value of the first image is greater than the maximum pixel threshold value in the target information region;

[0169] The electro-optical conversion optimization sub-module 133 comprises:

[0170] The electro-optical conversion standard adaptation unit 1331 is configured to determine whether the protocol supported by the HDR video is the HDR10 protocol.

[0171] The electro-optical conversion curve shaping unit 1332 is configured to obtain a first image from the HDR video file, and process pixel values of pixels in the first image, i.e., determine at least one target information region by extracting semantic features and texture images in the first image, and perform pixel interpolation in the target information region.

[0172] The pixel value conversion unit 1333 is configured to convert pixel values of each pixel of the second image into corresponding interpolated pixel values to obtain a third image.

[0173] The output interface sub-module 134 is configured to transmit information related to the electro-optical conversion standard to the HDR video playing module 140.

[0174] The HDR video playing module 140 is configured to display the HDR image of the HDR video file according to the electro-optical conversion standard adapted by the electro-optical conversion standard adaptation module 130.

[0175] In the embodiments of the present disclosure, by determining at least one target information region of the first image and performing pixel interpolation in the target information region of the first electro-optical conversion curve, a second electro-optical conversion curve conforming to second electro-optical conversion standard information is obtained. By shaping the electro-optical conversion curve and converting the pixel values of the video, the target information region with high correlation and high density information can be displayed with more obvious difference in brightness level, and the pixel values can be fully utilized, thereby providing better visual experience for users and solving the problem that part of the pixel values in the HDR video cannot be displayed in the related art.

[0176] Please refer to FIG. 11, the embodiments of the present disclosure provide a video imaging device, comprising:

[0177] The first processing module 111 is configured to obtain an HDR video file, obtain a plurality of first images from the HDR video file, and determine first electro-optical conversion standard information corresponding to the first images.

[0178] The second processing module 112 is configured to determine a first electro-optical conversion curve of the first image according to the first electro-optical conversion standard information.

[0179] The third processing module 113 is configured to determine at least one target information region of the first image, and perform pixel interpolation on the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to a second electro-optical conversion standard information, wherein the pixel interpolation is to insert a pixel region greater than a maximum pixel threshold of the first electro-optical conversion standard information into the target information region.

[0180] The display module 114 is configured to generate a third image based on a pixel value in the second electro-optical conversion curve, and obtain a new HDR video file according to a plurality of third images, and display the new HDR video file.

[0181] In an optional embodiment of the present disclosure, the target information region includes a strongly related pixel region in the first image, a high-density pixel region in the first image, and a strongly related high-density pixel region in the first image.

[0182] In an optional embodiment of the present disclosure, the third processing module includes:

[0183] The first processing submodule is configured to convert a pixel value greater than a maximum pixel threshold in the first image into the maximum pixel threshold to obtain a second image.

[0184] The second processing submodule is configured to extract semantic correlation features of each image block in the second image, and input the semantic correlation features of the each image block into a pre-trained logistic regression model, and the logistic regression model outputs a strongly related pixel region in the first image.

[0185] In an optional embodiment of the present disclosure, the third processing module includes:

[0186] The third processing submodule is configured to convert a pixel value greater than a maximum pixel threshold in the first image into the maximum pixel threshold to obtain a second image.

[0187] The fourth processing submodule is configured to extract an image region with high texture density in the second image, and take the image region with high texture density in the second image as a high-density pixel region in the first image.

[0188] In an optional embodiment of the present disclosure, the third processing module includes:

[0189] The fifth processing submodule is configured to convert a pixel value greater than a maximum pixel threshold in the first image into the maximum pixel threshold to obtain a second image.

[0190] The sixth processing submodule is configured to extract semantic correlation features of each image block in the second image, input the semantic correlation features of the each image block into a pre-trained logistic regression model, and output a strongly correlated pixel region in the first image from the logistic regression model;

[0191] The seventh processing submodule is configured to extract an image region with high texture density in the second image, and take the image region with high texture density in the second image as a high-density pixel region in the first image.

[0192] The eighth processing submodule is configured to obtain a strongly correlated high-density pixel region in the first image by performing an intersection operation on the strongly correlated pixel region in the first image and the high-density pixel region in the first image.

[0193] In an optional embodiment of the present disclosure, the third processing module further comprises:

[0194] The correction submodule is configured to insert, into the target information region, a pixel value of a pixel range with a pixel value greater than a maximum pixel threshold in the first image, to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information, and to correct an amplitude of the second electro-optical conversion curve to obtain a corrected second electro-optical conversion curve, so that a maximum brightness of the corrected second electro-optical conversion curve does not exceed a maximum brightness of the second electro-optical conversion standard information.

[0195] The device for video imaging provided in the embodiments of the present disclosure can implement each process implemented by the method embodiment of FIG. 1 and achieve the same technical effects. To avoid repetition, details are not described herein.

[0196] The embodiments of the present disclosure provide an electronic device 0120, as shown in FIG. 12, which is a principle block diagram of the electronic device 0120 according to an embodiment of the present disclosure, and includes a processor 0121, a memory 0122, and a program or instruction stored in the memory 0122 and executable on the processor 0121. The program or instruction is executed by the processor to implement the steps in any of the methods for video imaging of the present disclosure.

[0197] The embodiments of the present disclosure provide a readable storage medium, which stores a program or instruction. The program or instruction is executed by a processor to implement each process of the embodiments of the method for video imaging according to any of the above embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0198] The embodiments of the present disclosure also provide a computer program product, which includes computer instructions. The computer instructions are executed by a processor to implement each process of the method embodiment shown in FIG. 1 and achieve the same technical effects. To avoid repetition, details are not described herein.

[0199] Computer-readable media includes permanent and non-permanent, movable and non-movable media, and can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory PRAM (Programmable ROM, programmable read-only memory), SRAM (Static RAM, static random access memory), DRAM (Dynamic RAM, dynamic random access memory), other types of RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically EPROM, electrically erasable programmable read-only memory), flash memory or other memory technologies, CD-ROM (Compact Disc-ROM, read-only optical disc), DVD (Digital Video Disc, digital versatile disc) or other optical storage, magnetic cassette, magnetic tape storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media (readable media) such as modulated data signals and carriers.

[0200] It should be noted that in this paper, the term "includes", "contains" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0201] The above-mentioned sequence number of the embodiments of the present disclosure is only for description, and does not represent the advantages and disadvantages of the embodiments.

[0202] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, also can be through hardware, but in many cases the former is the better embodiment. Based on such understanding, the technical solutions of the disclosure essentially or say the part of the related art contribution can be embodied in the form of software product, the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disc, optical disc), including a number of instructions to make a service classification device (may be a mobile phone, computer, server, air conditioner, or network equipment, etc.) executes the method described in various embodiments of the disclosure.

[0203] The above only is the preferred embodiment of the disclosure, it should be pointed out that, for those skilled in the art, without departing from the principles of the disclosure, can make a number of improvements and refinements, these improvements and refinements also should be considered as the protection scope of the disclosure.

Claims

1. A method for video imaging, comprising: obtaining an HDR video file, obtaining a plurality of first images from the HDR video file, and determining first electro-optical conversion standard information corresponding to the first images; determining a first electro-optical conversion curve of the first images according to the first electro-optical conversion standard information; determining a second electro-optical conversion curve conforming to second electro-optical conversion standard information by determining at least one target information region of the first images and performing pixel interpolation in the target information region of the first electro-optical conversion curve, wherein the pixel interpolation is inserting a pixel region greater than a maximum pixel threshold of the first electro-optical conversion standard information into the target information region; generating a third image based on pixel values in the second electro-optical conversion curve; and obtaining a new HDR video file from the plurality of third images, and displaying the new HDR video file. 2.The method of claim 1, wherein: the target information region comprises a strongly relevant pixel region in the first images, a high-density pixel region in the first images, and a strongly relevant high-density pixel region in the first images; the determining the at least one target information region of the first images comprises: converting pixel values greater than a maximum pixel threshold in the first images to the maximum pixel threshold to obtain a second image; extracting semantic correlation features of each image block in the second image, and inputting the semantic correlation features of each image block into a pre-trained logistic regression model, wherein the logistic regression model outputs the strongly relevant pixel region in the first images; the determining the at least one target information region of the first images comprises: converting pixel values greater than a maximum pixel threshold in the first images to the maximum pixel threshold to obtain a second image; and extracting a high-texture-density image region in the second image, and taking the high-texture-density image region in the second image as the high-density pixel region in the first images; the determining the at least one target information region of the first images comprises: converting pixel values greater than a maximum pixel threshold in the first images to the maximum pixel threshold to obtain a second image; extracting semantic correlation features of each image block in the second image, and inputting the semantic correlation features of each image block into a pre-trained logistic regression model, wherein the logistic regression model outputs the strongly relevant pixel region in the first images; extracting a high-texture-density image region in the second image, and taking the high-texture-density image region in the second image as the high-density pixel region in the first images; and taking an intersection of the strongly relevant pixel region in the first images and the high-density pixel region in the first images to obtain the strongly relevant high-density pixel region in the first images; and the performing pixel interpolation in the target information region of the first electro-optical conversion curve to obtain the second electro-optical conversion curve conforming to the second electro-optical conversion standard information further comprises: ​ ​ ​ ​ ​ 3. The method of video imaging of claim 2, wherein, ​ ​ ​ 4. The method of video imaging of claim 2, wherein, ​ ​ ​ 5. The method of video imaging of claim 2, wherein, ​ ​ ​ ​ ​ 6. The method of video imaging of claim 1, wherein, ​ inserting, in the target information region of the first image, a pixel value of a pixel range whose pixel value is greater than a maximum pixel threshold value, to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information; correcting an amplitude of the second electro-optical conversion curve to obtain a corrected second electro-optical conversion curve, so that a maximum brightness of the corrected second electro-optical conversion curve does not exceed a maximum brightness in the second electro-optical conversion standard information.

7. An apparatus for video imaging, comprising: a first processing module configured to obtain an HDR video file, obtain a plurality of first images from the HDR video file, and determine first electro-optical conversion standard information corresponding to the first images; a second processing module configured to determine a first electro-optical conversion curve of the first images according to the first electro-optical conversion standard information; a third processing module configured to determine at least one target information region of the first images, and perform pixel interpolation in the target information region of the first electro-optical conversion curve to obtain a second electro-optical conversion curve conforming to second electro-optical conversion standard information, wherein the pixel interpolation is inserting a pixel region greater than a maximum pixel threshold value of the first electro-optical conversion standard information into the target information region; a display module configured to generate third images based on pixel values in the second electro-optical conversion curve, obtain a new HDR video file from the plurality of third images, and display the new HDR video file.

8. An electronic device, comprising a processor, a memory, and a program or instructions stored in the memory and executable on the processor, the program or instructions being executed by the processor to implement steps in the method for video imaging according to any one of claims 1 to 6.

9. A readable storage medium, the readable storage medium storing a program or instructions, the program or instructions being executed by a processor to implement steps in the method for video imaging according to any one of claims 1 to 6.

10. A computer program product, comprising computer instructions, the computer instructions being executed by a processor to implement steps in the method for video imaging according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and system for processing signal having high dynamic range

    CN106101679A

  • Format conversion method for HDR video

    CN108769804A

  • Image processing method and device, electronic equipment and storage medium

    CN116167950A

  • Video imaging method and device, electronic equipment and readable storage medium

    CN118381929A

  • Method and module for processing high dynamic range (HDR) image and display device using the same

    US20180122058A1