Image processing method, image processing device, electronic equipment and storage medium
By using a unified photoelectric conversion function to encode and decode the cover frame and video stream in dynamic photo processing, the problem of poor playback consistency between the cover frame and the video is solved, achieving a more consistent playback effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-17
AI Technical Summary
There is a problem with the playback consistency between the cover frame in the dynamic photo and the video.
By using the same Optical-to-Electronic Transformation Function (OETF) to encode the cover frame and video stream of the moving photo from the linear domain to the non-linear domain, and then using the corresponding transfer function to decode from the non-linear domain to the linear domain during decoding, consistency in color management between the cover frame and the video stream is ensured.
Improved the consistency of playback effect between cover frames and video streams in dynamic photos during playback.
Smart Images

Figure CN121888044A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an image processing method, an image processing apparatus, an electronic device, and a storage medium. Background Technology
[0002] In related technologies, animated photos, as a multimedia file combining a cover frame and video, can effectively express the dynamic and static information of a scene. However, because the cover frame and video belong to different standard organizations, there are differences in color management and display, resulting in differences in the effect between the cover frame and video when the animated photo is played. Therefore, in related technologies, there is a problem of poor consistency in the playback effect between the cover frame and video in animated photos. Summary of the Invention
[0003] This application provides an image processing method, an image processing apparatus, an electronic device, and a storage medium, which can solve the problem of poor consistency between the cover frame and the video in dynamic photos.
[0004] Firstly, an image processing method is provided, including:
[0005] The first metadata for encoding the motion picture is determined, the first metadata including the photoelectric conversion function OETF, the OETF being used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
[0006] Secondly, an image processing method is provided, including:
[0007] Based on the transfer function of the moving photo, the cover frame and video stream of the moving photo are decoded from the nonlinear domain to the linear domain to obtain the cover frame decoded data and the video stream decoded data.
[0008] Thirdly, an image processing apparatus is provided, comprising:
[0009] A determination module is used to determine first metadata for encoding a motion picture, the first metadata including an optoelectronic conversion function (OETF), the OETF being used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
[0010] Fourthly, an image processing apparatus is provided, comprising:
[0011] The decoding module is used to decode the cover frame and video stream of the dynamic photo from the nonlinear domain to the linear domain based on the transfer function of the dynamic photo, so as to obtain the cover frame decoding data and the video stream decoding data.
[0012] Fifthly, an electronic device is provided, including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the image processing method as described in the first aspect, or implementing the steps of the image processing method as described in the second aspect.
[0013] In a sixth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the image processing method as described in the first aspect, or implement the steps of the image processing method as described in the second aspect.
[0014] In a seventh aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the image processing method as described in the first aspect, or to implement the steps of the image processing method as described in the second aspect.
[0015] Eighthly, a computer program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the image processing method as described in the first aspect, or to implement the steps of the image processing method as described in the second aspect.
[0016] In the embodiments of this application, during the processing of dynamic photos, the cover frame and video stream of the dynamic photo are encoded from the linear domain to the nonlinear domain using the same OETF. This allows the cover frame and video stream in the dynamic photo to have unified color management information during the encoding stage, which helps to improve the consistency of the playback effect of the cover frame and video stream in the dynamic photo during playback. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the encoding / decoding system provided in an embodiment of this application;
[0018] Figure 2 This is a schematic diagram of the encoder provided in an embodiment of this application;
[0019] Figure 3 This is a schematic diagram of the decoder provided in an embodiment of this application;
[0020] Figure 4 This is one of the schematic flowcharts of the image processing method provided in the embodiments of this application;
[0021] Figure 5 This is a flowchart illustrating the acquisition process of static raw images and video raw image sequences in the embodiments of this application;
[0022] Figure 6 This is a flowchart illustrating the process of applying image effects to a sequence of original video images in an embodiment of this application.
[0023] Figure 7 This is a second schematic flowchart of the image processing method provided in the embodiments of this application;
[0024] Figure 8 This is one of the structural schematic diagrams of the image processing apparatus provided in the embodiments of this application;
[0025] Figure 9 This is a second schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application;
[0026] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0027] Figure 11 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0029] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0030] Figure 1 This is a schematic diagram of the encoding / decoding system 10 provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding and decoding (CODEC) video data (including encoding or decoding). The video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements, etc.
[0031] like Figure 1 As shown, the encoding / decoding system 10 includes a source device 100, which provides encoded video data to be decoded and displayed by the destination device 110. Specifically, the source device 100 provides video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0032] exist Figure 1 In this example, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of a video encoding device, while destination device 110 represents an example of a video decoding device. In other examples, source device 100 and destination device 110 may not include... Figure 1 Some components, or may include Figure 1 Other components besides the source device 100. For example, the source device 100 can receive video data from an external data source (such as an external camera). Similarly, the destination device 110 can interface with an external display device, without including an integrated display device. Furthermore, the memory 102 and memory 113 can be external memories.
[0033] Although Figure 1 Source device 100 and destination device 110 are illustrated as separate devices, but in some examples, they may also be integrated into a single device. In such embodiments, the same hardware or software, or separate hardware or software, or any combination thereof, may be used to implement the functionality corresponding to source device 100 and destination device 110.
[0034] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional video transmission. If it is bidirectional video transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, that is, each of source device 100 and destination device 110 includes an encoder and a decoder.
[0035] Data source 101 represents the source of video data (i.e., raw, unencoded video data) and provides encoder 200 with a series of images containing video data, which encoder 200 encodes. Data source 101 of source device 100 may include a video acquisition device (such as a video camera), a video archive containing previously acquired raw video, or a video feed interface for receiving video from a video content provider. Alternatively, data source 101 may generate computer graphics-based data as source video, or combine live video, archived video, and computer-generated video. In these cases, encoder 200 encodes the acquired, pre-acquired, or computer-generated video data. Encoder 200 may rearrange the images from the received order (sometimes referred to as the "display order") according to the encoding order. Encoder 200 may generate a bitstream including the encoded video data. Source device 100 may then output the encoded video data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0036] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw video data from data source 101, and memory 113 may store decoded video data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded video data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more video buffers, for example, to store raw, decoded, or encoded video data.
[0037] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0038] Output interface 104 may include any type of medium or device capable of transmitting encoded video data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded video data directly from source device 100 to destination device 110 in real time. The encoded video data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0039] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0040] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded video from source device 100 and provide the encoded video data to destination device 110, for example, via network transmission. The server may include, for example, a web server (for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached Storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0041] Destination device 110 can access encoded video data from a server, for example via a wireless channel (e.g., Wireless Fidelity (WIFI) connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing encoded video data stored on the server.
[0042] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 or IEEE 802.15 standard (e.g., ZigBee™ transmission mode), Bluetooth standard, or other physical components. In the example where output interface 104 and input interface 111 include a wireless component, output interface 104 and input interface 111 can be configured to transmit data, such as encoded video data, via Wi-Fi, Ethernet, cellular networks (such as 4th Generation (4G) mobile networks, Long Term Evolution (LTE), LTE Advanced, 5th Generation (5G) mobile networks, 6th Generation (6G) mobile networks, etc.).
[0043] The technology provided in this application can be applied to support video encoding and decoding in one or more multimedia applications such as video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission, digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0044] The input interface 111 of the destination device 110 receives an encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the syntax elements are used to decode the encoded data units to obtain decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0045] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0046] The encoder 200 and decoder 300 can process based on the following video codec standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding (HEVC)), H.266 (also known as Versatile Video Coding (VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video Coding Standard 1 (AVS 1), AVS2, AVS3, or next-generation video standard protocols. This application embodiment does not specifically limit the specific implementation.
[0047] Typically, encoder 200 and decoder 300 can perform block-based encoding and decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding or decoding). For example, a block can include a two-dimensional matrix of samples of luminance or chrominance data. For example, encoder 200 and decoder 300 can encode and decode video data represented in YUV format, where "Y" represents luminance (or Luma), and "U" and "V" are the two components of chrominance (or Chroma), with "U" representing the blue chrominance component (Cb) and "V" representing the red chrominance component (Cr).
[0048] See Figure 2 The figure is a schematic diagram of the structure of the encoder 200 provided in an embodiment of this application. The encoder 200 can be... Figure 1 Encoder 200 in. In Figure 2 In the example, encoder 200 includes memory 201, encoding parameter determination unit 210, residual generation unit 202, transform processing unit 203, quantization unit 204, inverse quantization unit 205, inverse transform processing unit 206, reconstruction unit 207, filter unit 208, decoded picture buffer (DPB) 209, and entropy encoding unit 220.
[0049] Memory 201 can store video data to be encoded; for example, encoder 200 can retrieve the data from the encoder. Figure 1 The data source 101 shown receives and stores video data. In some examples, the memory 201 may be on the same chip as other components of the encoder 200 (e.g., ...). Figure 2 (As shown), it can also be independent of the chip where these components are located.
[0050] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is used to obtain a first prediction block for the current block using an inter-frame prediction mode. The intra-frame prediction unit 213 is used to obtain a second prediction block for the current block using an intra-frame prediction mode. The mode selection unit 211 is used to obtain a target prediction block based on the first and second prediction blocks and determine the final prediction mode. Furthermore, the coding parameter determination unit 210 may also include other functional units, such as functional units for determining the partitioning method of coding units (CUs), functional units for determining the transformation type of the residual data of the CUs, or functional units for determining the quantization parameters of the residual data of the CUs.
[0051] For ease of description and understanding, in the embodiments of this application, the CU to be processed in the current image is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0052] Inter-frame prediction unit 212 may include a motion estimation unit and a motion compensation unit. For inter-frame prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (e.g., one or more previously encoded / decoded pictures stored in DPB 209).
[0053] The motion estimation unit can generate one or more motion vectors (MVs) representing the position of a reference block in a reference image relative to the position of the current block in the current image. The motion compensation unit can then use interpolation to obtain a predicted value with the precision indicated by the motion vectors.
[0054] The encoding parameter determination unit 210 can provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the raw uncoded video data of the current block from the memory 201 and calculates the residual between the current block and the target prediction block to obtain the residual block. In some examples, the function of the residual generation unit 202 can be implemented using one or more subtractor circuits that perform binary subtraction.
[0055] As an example, the encoding parameter determination unit 210 can provide the entropy encoding unit 220 with syntax elements representing encoding parameters for encoding. The encoding parameters include one or more of the following: the partitioning method of the CU, the final prediction mode, the transformation type of the residual data of the CU, or the quantization parameters of the residual data of the CU.
[0056] The transformation processing unit 203 transforms the residual block output by the residual generation unit 202 to obtain a transform coefficient block. This transformation may include Discrete Cosine Transform (DCT), integer transformation, direction transformation, or Karhunen-Loeve Transform, etc. In some examples, the encoder 200 may not include the transformation processing unit 203.
[0057] Quantization unit 204 can quantize the transform coefficients in the transform coefficient block according to the quantization parameter (QP) value associated with the current block to produce a quantized transform coefficient block.
[0058] The inverse quantization unit 205 and the inverse transform processing unit 206 can perform inverse quantization and inverse transform on the transform coefficient block, respectively, to obtain the reconstructed residual block. The reconstruction unit 207 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0059] Filter unit 208 can perform one or more filter operations on the reconstructed block. For example, filter unit 208 can be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, encoder 200 may not include filter unit 208.
[0060] Encoder 200 stores the reconstructed image obtained from the reconstructed blocks in DPB 209. For example, in an example where the operation of filter unit 208 is not required, reconstruction unit 207 can store the reconstructed blocks in DPB 209. In an example where the operation of filter unit 208 is required, filter unit 208 can store the filtered reconstructed blocks in DPB 209. Inter-frame prediction unit 212 retrieves the reconstructed image from DPB 209 to perform inter-frame prediction on blocks of subsequent images to be encoded. In some examples, DPB 209 can be replaced with other types of memory.
[0061] Entropy coding unit 220 can entropy code the syntax elements of other components in encoder 200 to output encoded video data. For example, entropy coding unit 220 can entropy code the quantized transform coefficient block from quantization unit 204. As another example, entropy coding unit 220 can entropy code the syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from coding parameter determination unit 210.
[0062] Understandable, Figure 2 The composition of the encoder 200 described is merely illustrative and does not constitute a limitation on the embodiments of this application.
[0063] Figure 3 This is a schematic diagram of the structure of the decoder 300 provided in the embodiments of this application. The decoder 300 can be... Figure 1 The decoder 300 is described. Figure 3 In the example, the decoder 300 includes a Coded Picture Buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0064] Entropy decoding unit 302 can receive encoded video data from CPB 301 and perform entropy decoding on the video data to obtain syntax elements. The syntax elements indicate encoding parameters, including one or more of the following: CU partitioning method, final prediction mode, transformation type of CU residual data, or quantization parameters of CU residual data.
[0065] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter-frame prediction mode, the prediction block of the current CU can be obtained through the inter-frame prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra-frame prediction mode, the prediction block of the current CU can be obtained through the intra-frame prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may also include a unit for performing prediction functions according to other prediction modes.
[0066] CPB 301 can be obtained from, for example Figure 1 The communication medium 120 shown acquires and stores encoded video data. DPB 307 is used to store decoded images. Optionally, CPB 301 and DPB 307 can be replaced with other types of memory; this application does not impose specific limitations. In some examples, CPB 301 can be on the same chip as other components of the decoder 300 (as shown), or it can be on a separate chip from those components.
[0067] Decoder 300 can perform reconstruction operations on each block individually. Entropy decoding unit 302 can entropy decode the syntax elements and transform information (e.g., QP or transform mode indication) of the quantized transform coefficients to obtain the quantized transform coefficients. Dequantization unit 303 dequantizes the quantized transform coefficients to obtain a transform coefficient block including the transform coefficients. Inverse transform processing unit 304 performs an inverse transform on the transform coefficient block to generate a residual block corresponding to the current block; this inverse transform is the reverse operation of the above transform.
[0068] Reconstruction unit 305 can reconstruct the current block based on the prediction block and the residual block. For example, reconstruction unit 305 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.
[0069] Filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the type of filter unit 306 can be referenced to the type of filter unit 208, and will not be described again here. In some examples, the operations of filter unit 306 can be skipped.
[0070] Decoder 300 can store the reconstructed image obtained from the reconstructed blocks in DPB 307. For example, in an example where filter unit 306 is not operated, reconstruction unit 305 can store the reconstructed blocks in DPB 307. In an example where filter unit 306 is operated, filter unit 306 can store the filtered reconstructed blocks in DPB 307. Decoder 300 can output the decoded image (e.g., decoded video) from DPB 307 for use with display devices (such as...). Figure 1 The subsequent presentation of the display device 114).
[0071] The following describes some aspects related to the embodiments of this application:
[0072] Live Photos: Also known as live photos, these are multimedia files that combine images and videos, effectively conveying the dynamic and static information of a scene.
[0073] Opto-Electronic Transfer Function (OETF): A transfer function used to convert scene light into electrical signals.
[0074] Electro-Optical Transfer Function (EOTF): A transfer function used to convert electrical signals into optical signals for screen display.
[0075] Opto-Optical Transfer Function (OOTF): A transfer function used to convert scene light signals into display light signals.
[0076] The dynamic photo processing method provided in this application embodiment is described below with reference to the accompanying drawings. The dynamic photo processing method provided in this application embodiment can be executed by an encoding end, for example... Figure 1 or Figure 2 The encoder 200 is shown. The dynamic photo processing method provided in this embodiment can also be executed by the decoding end, for example... Figure 1 or Figure 3 The decoder 300 is described above. The encoding and decoding ends can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding device or a video encoding device, and the decoding end can be referred to as a decoding device or a video decoding device.
[0077] Please see Figure 4 , Figure 4 This is a flowchart of a dynamic photo processing method provided in an embodiment of this application. The method can be applied to the encoding end or the decoding end. In order to better illustrate the method provided in this application, the following embodiments use the encoding end as an example to explain the dynamic photo processing method provided in this application.
[0078] like Figure 4 As shown, the dynamic photo processing method includes the following steps:
[0079] Step 401: Determine the first metadata of the encoded motion picture. The first metadata includes the photoelectric conversion function OETF, which is used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
[0080] It is understandable that the above image processing method can be applied to the generation process of moving photos. Specifically, during the capturing of a moving photo, after the camera acquires the original image data of the moving photo, the original image data can be encoded from the linear domain to the nonlinear domain based on the first metadata to obtain the moving photo. The original image data may include the original image of the cover frame and the original images of each video frame in the video stream. By using the OETF in the first metadata, each original image in the original image data can be encoded from the linear domain to the nonlinear domain.
[0081] The aforementioned first metadata may include various color management parameters during the encoding of the dynamic photo. For example, in addition to the OETF, the first metadata may also include color conversion matrix, other transfer functions, and other color management parameters.
[0082] The aforementioned OETF can be a defined or selected OETF function, which can be a standard transfer function or a custom optimization curve dynamically calculated based on the scenario content. For example, in some embodiments of this application, the OETF can be represented by the following formula:
[0083]
[0084] in, It represents the electrical signal intensity corresponding to the optical signal intensity x. n is the power exponent, representing the degree of nonlinearity, and is usually a positive number. For example, the value of n can be between 1 and 2.
[0085] It is understood that the above-mentioned encoding of the cover frame and the video stream based on the photoelectric conversion function OETF may include: encoding the cover frame based on the photoelectric conversion function OETF, and encoding each frame image in the video stream based on the photoelectric conversion function OETF.
[0086] In this embodiment, during the processing of dynamic photos, the cover frame and video stream of the dynamic photo are encoded from the linear domain to the non-linear domain using the same OETF. This allows the cover frame and video stream in the dynamic photo to have unified color management information during the encoding stage, which helps to improve the consistency of the playback effect of the cover frame and video stream in the dynamic photo during playback.
[0087] Optionally, the method further includes:
[0088] If the first metadata does not include the optical-optical conversion function OOTF, the inverse of the OETF is the transfer function of the moving image;
[0089] When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function of the moving image.
[0090] Here, the inverse of OETF is the inverse function of OETF, and the inverse of OETF can represent OETF -1 The transfer function of the animated photo can also be called EOTF. During the playback of the animated photo, the cover frame and the video stream of the animated photo need to be decoded from the nonlinear domain to the linear domain based on the EOTF to obtain the cover frame decoded data and the video stream decoded data.
[0091] In some embodiments of this application, the above-mentioned EOTF can be represented by the following formula:
[0092]
[0093] Where x represents the input electrical signal, and the value of x is a value normalized to the range [0, 1]. γ is the output light intensity, a linear brightness value; γ1 is the gamma coefficient, usually taken as 2.2 or 2.4, depending on the display standard used.
[0094] The optical-to-optical conversion function OOTF can be expressed as:
[0095]
[0096] Where x represents the input electrical signal, and the value of x is a value normalized to the range [0, 1]. γ is the output light intensity, a linear brightness value; γ2 is the gamma coefficient, usually taken as 2.2 or 2.4, depending on the display standard used.
[0097] It should be noted that when EOTF is the inverse function of OETF, OETF... -1 When OETF×EOTF = 1, OETF and EOTF are reversible, which allows for the restoration of scene lighting as much as possible.
[0098] Accordingly, when the first metadata includes OOTF, the transfer function of the moving image composed of the inverse of OETF and OOTF can mean that the transfer function of the moving image can be obtained by operating on the inverse of OETF and OOTF, wherein the type of operation can be set as needed. For example, in some embodiments of this application, when the first metadata includes OOTF, the transfer function of the moving image is the product of the inverse of OETF and OOTF. When EOTF is the product of the inverse function of OETF and the optical-optical conversion function OOTF, OETF × EOTF may not be equal to 1. In this case, it is usually based on the creator's subjective expression of the presentation effect.
[0099] It is understood that during the decoding of the animated photo, the transfer function can be determined by reading whether the first metadata includes the Optical-Optical Transfer Function (OOTF) and by reading the OETF in the first metadata. If the first metadata does not include the OOTF, the inverse of the OETF is determined to be the transfer function of the animated photo; if the first metadata includes the OOTF, the inverse of the OETF and the OOTF together constitute the transfer function of the animated photo. Then, based on the determined transfer function, the cover frame and the video stream of the animated photo are decoded from the nonlinear domain to the linear domain to obtain the cover frame decoded data and the video stream decoded data. Finally, the decoded data undergoes color conversion and is displayed.
[0100] In this embodiment, when the first metadata does not include the optical-to-optical conversion function (OOTF), the inverse of the OETF is determined as the transfer function of the moving image; when the first metadata includes the OOTF, the inverse of the OETF and the OOTF together form the transfer function of the moving image. Thus, the first metadata can include or exclude the OOTF as needed to select different transfer functions, thereby achieving the goal of restoring scene lighting as much as possible or presenting image effects according to the creator's subjective expression.
[0101] Optionally, the determination of the first metadata of the encoded dynamic photograph includes:
[0102] Determine the first metadata and second metadata of the encoded motion picture, wherein the second metadata includes a color conversion matrix used to perform color conversion on the cover frame of the motion picture and the video stream of the motion picture.
[0103] The color conversion matrix used to perform color conversion on the cover frame and video stream of the animated photo can mean that the color conversion matrix is used to perform color conversion on the cover frame and video stream of the animated photo to convert the cover frame and the animated photo into YUV format.
[0104] It is understood that the images in the cover frame and video stream mentioned above can all be in RGB format. The color conversion matrix mentioned above is an RGB2YUV conversion matrix, which can specifically be a 3×3 matrix, and the values of the elements in each row and column of the RGB2YUV conversion matrix can range from -1 to 1, which can be selected as needed.
[0105] During the capture of a dynamic photo, after the camera acquires the original image data of the dynamic photo, the original image data can first be processed with image effects. Then, based on the color conversion matrix, the image processed with image effects is converted into YUV format. Finally, based on the first metadata, the YUV format image data is encoded from the linear domain to the nonlinear domain to obtain the dynamic photo.
[0106] In this implementation, since the color conversion matrix is part of the color management information, color conversion is performed on the cover frame and the video stream based on the same color conversion matrix. This helps to further improve the consistency of the color management information used by the cover frame and the video stream during the encoding process, and thus further improves the consistency of the effect between the image and the video in the obtained dynamic photo.
[0107] Optionally, the first metadata may further include a transfer function indicator parameter;
[0108] When the transfer function indicator parameter indicates the first encoding method, the transfer function of the moving image is the inverse of the OETF;
[0109] When the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image consists of the inverse of the OETF and the OOTF in the first metadata.
[0110] Specifically, the transfer function indicator parameter can indicate different encoding methods through different parameter values. For example, when the parameter value of the transfer function indicator parameter is 0, it indicates a first encoding method; when the parameter value of the transfer function indicator parameter is 1, it indicates a second encoding method. As another example, when the parameter value of the transfer function indicator parameter is 1, it indicates the first encoding method; when the parameter value of the transfer function indicator parameter is -1, it indicates the second encoding method. The specific settings can be configured as needed.
[0111] It is understandable that during the encoding of the aforementioned cover frame and video stream, the aforementioned color conversion matrix and EOTF can be used. -1 The OOTF data is recorded as metadata for color management of moving images. An additional metadata field is added as a parameter to the pass function, where relevant personnel can use this field to select whether the moving image is in EOTF format. -1 Playback as a pass function, or as EOTF? -1 × OOTF is used as a transfer function for playback.
[0112] During the decoding process, the same decoding operation can be performed on the cover frame and the video stream according to the transfer function and color conversion matrix specified in the color management metadata.
[0113] In this embodiment, the first metadata further includes a transfer function indicator parameter; when the transfer function indicator parameter indicates a first encoding method, the transfer function of the moving image is the inverse of the OETF; when the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image is composed of the inverse of the OETF and the OOTF in the first metadata. In this way, different EOTEs can be selected according to different needs to achieve the goal of restoring scene light as much as possible or presenting image effects according to the creator's subjective expression.
[0114] Optionally, the method includes:
[0115] Obtain static raw images and video raw image sequences;
[0116] The static original image and the video original image sequence are processed with image effects respectively to obtain the cover frame and the video stream with matching image effects.
[0117] It is understood that after obtaining the cover frame and the video stream with matching image effects, the cover frame and the video stream can be encoded from the linear domain to the nonlinear domain based on the first metadata and the second metadata mentioned above.
[0118] The aforementioned static original image and video original image sequence can be image data acquired by the electronic device while it is in motion photo shooting mode. The motion photo shooting mode can be the shooting mode after the user turns on the motion photo switch in the camera. The static original image and video original image sequence are image data acquired by the electronic device in the same shooting process. For example, the video original image sequence may include: several frames of images preceding the static original image, and / or several frames of images following the static original image. This ensures that the image frames in the motion photo composed of the static original image and video original image sequence are continuously captured image frames.
[0119] The aforementioned static raw image can be a RAW image output by an image sensor in an electronic device. Correspondingly, the aforementioned video raw image sequence is a sequence composed of multiple RAW images output by an image sensor.
[0120] The above-described image effect processing of the static original image and the original video image sequence to obtain a cover frame and video stream with matching image effects can be achieved by processing the static original image and the original video image sequence with the same image effect processing method, thereby making the processed cover frame and video stream match. Here, "matching the image effects of the cover frame and the video stream" can mean that the image effects of the cover frame and the video stream are the same or approximately the same.
[0121] In this embodiment, by processing the static original image and the video original image sequence image effects separately during the dynamic photo generation process, a cover frame and video stream with matching image effects are obtained, which helps to improve the consistency of image effects between the obtained cover frame and video stream.
[0122] Optionally, the step of performing image effect processing on the static original image and the video original image sequence respectively to obtain the cover frame and the video stream with matching image effects includes:
[0123] Under the same color gamut, the static original image is processed to obtain the cover frame, and the video original image sequence is processed to obtain the video stream based on the first acquisition parameters, wherein the first acquisition parameters include the image acquisition parameters of the static original image.
[0124] The term "same color gamut" refers to all image processing performed within the same color space, which limits the range of color adjustments. For example, extreme colors outside the sRGB color gamut cannot be displayed. In some embodiments of this application, a color gamut can be selected before image effect processing. Then, within the color adjustment range indicated by the selected color gamut, image effect processing is performed on the static original image and the video original image sequence, respectively. The color gamut can be various common color gamuts, such as sRGB, BT709, P3, or BT2020.
[0125] In some embodiments of this application, image effect processing can be performed on static original images and video original image sequences under a certain color gamut using a unified effect processing flow and unified color management information. The color management information includes transfer functions and color transformation matrices, and the transfer function includes the aforementioned OETF.
[0126] The aforementioned first acquisition parameter may include various parameters of the camera when the electronic device acquires the static raw image, such as exposure, white balance, focus, gyroscope, tone, color and other processing parameters.
[0127] The above-mentioned image effect processing of the original video image sequence based on the first acquisition parameter can refer to: adjusting the color parameters of each image frame in the original video image sequence to be the same as or similar to the first acquisition parameter, thereby helping to achieve matching image effects between the cover frame and the video stream.
[0128] In some embodiments of this application, the above-described image effect processing of the static original image to obtain the cover frame may include: performing noise reduction processing on the static original image to obtain the cover frame.
[0129] In some embodiments of this application, the above-described image effect processing of the static original image to obtain the cover frame may also include: acquiring K consecutive original images of the static original image, and performing image effect processing on the K original images based on the first acquisition parameters to obtain K processed images. Specifically, this may involve adjusting the color parameters of each image frame in the K original images to be the same as or similar to the first acquisition parameters to obtain K processed images. Then, multi-frame noise reduction processing is performed on the static original image based on the K processed images to obtain the cover frame. The K original images may be K images preceding the static original image, or K images following the static original image, or K original images may be composed of at least one image preceding the static original image and at least one image following the static original image, where K is an integer greater than 1.
[0130] In this embodiment, the cover frame is obtained by performing image effect processing on the static original image under the same color gamut, and the video stream is obtained by performing image effect processing on the original video image sequence based on the first acquisition parameters. This helps to achieve matching image effects between the cover frame and the video stream.
[0131] Optionally, the static original image is an image captured by the camera in the electronic device when the electronic device receives a shooting command;
[0132] The first acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device receives the shooting command.
[0133] The aforementioned shooting instructions can be shooting instructions generated by the user clicking the shutter button or other quick shooting operations in the shooting interface.
[0134] Specifically, during the shooting process, users typically adjust shooting parameters before inputting the shooting command. Only after all shooting parameters have been adjusted will the shooting command be input. Therefore, upon receiving the shooting command, it can be determined that the image acquisition parameters at this point are satisfactory to the user. Thus, these current image acquisition parameters can be locked and used as a benchmark to process each image frame in the dynamic photo. This improves the consistency between the cover frame and the video stream in the dynamic photo, while also enhancing the consistency between the acquired parameters of the resulting dynamic photo and the user-defined acquisition parameters.
[0135] In this embodiment, by making the static original image the image captured by the camera in the electronic device when the electronic device receives the shooting command; and by making the first acquisition parameter the image acquisition parameter of the camera in the electronic device when the electronic device receives the shooting command, the consistency between the cover frame and the video stream in the dynamic photo can be improved, while also improving the consistency between the acquisition parameters of the obtained dynamic photo and the acquisition parameters set by the user.
[0136] Optionally, the original video image sequence includes: N frames of images preceding the static original image, and M frames of images following the static original image, where N and M are integers greater than 0.
[0137] In this context, the values of N and M can be the same or different. The aforementioned animated photo comprises (N+M+1) frames, namely one cover frame and (N+M) video frames. The value of (N+M+1) can be determined based on the frame rate and recording duration of the animated photo. Specifically, it can be the product of the frame rate and the recording duration. For example, when the frame rate of the animated photo is 30 frames per second and the recording duration is 3 seconds, the value of (N+M+1) is 90. In this case, the value of N can be 45, and the value of M can be 44, or the value of N can be 44, and the value of M can be 45, etc. As another example, when the frame rate of the animated photo is 24 frames per second and the recording duration is 3.5 seconds, the value of (N+M+1) is 84. In this case, the value of N can be 42, and the value of M can be 41, or the value of N can be 40, and the value of M can be 43, etc.
[0138] In this embodiment, since the original video image sequence includes N frames before the static original image and M frames after the static original image, the dynamic photo can be composed of a cover frame and multiple image frames before and after the cover frame. This makes it easier to make the image acquisition parameters of each image frame in the dynamic photo close to the image acquisition parameters of the dynamic photo, thereby improving the consistency of the image effect between the cover frame and the video stream.
[0139] Optionally, the first acquisition parameters include at least one of the following: exposure parameters, white balance parameters, focus parameters, gyroscope parameters, tone parameters, and color parameters.
[0140] The exposure, white balance, focus, gyro, tonal, and color parameters mentioned above can all be adjusted directly or simulated using post-processing software. For example, the exposure, white balance, tonal, and color parameters can all be adjusted directly in post-processing software. The focus parameter can be simulated by using a "blur and sharpen" technique. The gyro parameter can be simulated by removing motion blur using software.
[0141] In some embodiments of this application, the first acquisition parameter may simultaneously include various image acquisition parameters such as exposure parameters, white balance parameters, focus parameters, gyro parameters, tone parameters, and color parameters.
[0142] In this embodiment, by including at least one of the following in the first acquisition parameters: exposure parameters, white balance parameters, focus parameters, gyro parameters, tone parameters, and color parameters, the consistency of image effects between the cover frame and the video stream can be further improved.
[0143] Optionally, the step of performing image effect processing on the original video image sequence based on the first acquisition parameters to obtain the video stream includes:
[0144] When the scene information of the first original image frame matches the scene information of the static original image, the first original image frame is processed based on the first acquisition parameters to obtain the processed first image frame. The first original image frame is any one of the image frames in the original video image sequence, and the video stream includes the first image frame.
[0145] The aforementioned scene information may include information on various scene changes, such as at least one of brightness parameters, white balance parameters, and gyro parameters. In some embodiments of this application, the scene information of the first original image frame can be determined to match the scene information of the static original image by using the brightness parameters. For example, the average brightness of a specific area in the first original image frame and the average brightness of the same area in the static original image can be calculated. If the difference between the two average brightness values is less than a preset brightness threshold, the scene information of the first original image frame is determined to match the scene information of the static original image. If the difference between the two average brightness values is greater than or equal to the preset brightness threshold, the scene information of the first original image frame is determined to not match the scene information of the static original image.
[0146] In some embodiments of this application, the scene information of the first original image frame can be determined to match the scene information of the static original image by white balance parameters. For example, the average white balance value of a specific region in the first original image frame and the average white balance value of the same region in the static original image can be calculated. If the difference between the two average white balance values is less than a preset white balance threshold, the scene information of the first original image frame is determined to match the scene information of the static original image. If the difference between the two average white balance values is greater than or equal to the preset white balance threshold, the scene information of the first original image frame is determined to not match the scene information of the static original image.
[0147] In some embodiments of this application, the scene information of the first original image frame can be determined to match the scene information of the static original image by using the gyro parameter. For example, the average gyro value of a specific region in the first original image frame and the average gyro value of the same region in the static original image can be calculated. If the difference between the two average gyro values is less than a preset gyro threshold, the scene information of the first original image frame matches the scene information of the static original image. If the difference between the two average gyro values is greater than or equal to the preset gyro threshold, the scene information of the first original image frame does not match the scene information of the static original image.
[0148] In some embodiments of this application, the scene information of the first original image frame can be determined to match the scene information of the static original image by using two or more parameters among "brightness parameter, white balance parameter, and gyro parameter". For example, the average brightness, average white balance, and average gyro of a specific region in the first original image frame can be statistically analyzed, and then weighted and summed to obtain a first weighted sum value. Then, the average brightness, average white balance, and average gyro of the same region in the static original image can be statistically analyzed, and then weighted and summed to obtain a second weighted sum value. If the difference between the first weighted sum value and the second weighted sum value is less than a preset threshold, the scene information of the first original image frame is determined to match the scene information of the static original image. If the difference between the first weighted sum value and the second weighted sum value is greater than or equal to the preset threshold, the scene information of the first original image frame is determined to not match the scene information of the static original image. During the weighted summation process, the weights corresponding to the average brightness, the average white balance, and the average gyro can be set as needed. For example, the three can be equal or unequal, and the sum of the three is 1.
[0149] It is understood that the embodiments of this application only take the first original image frame as an example to explain how to process an image frame in the original video image sequence. In fact, each image frame in the original video image sequence can be processed according to the processing method of the first original image frame.
[0150] In some embodiments of this application, the acquisition parameters may change during the capture of dynamic photos, resulting in inconsistent image effects at different time points. Therefore, this application's embodiments record the parameters and 3A algorithm state during the capture of the static original image, locking and forcibly applying them to every frame of the entire dynamic photo video processing, fundamentally ensuring that the cover frame and video sequence originate from the same processing method. Here, "3A" typically refers to Auto Focus, Auto Exposure, and Auto White Balance.
[0151] It should be noted that if the scene information of the first original image matches the scene information of the static original image, it indicates that the scene of the first original image has not changed significantly compared to the static original image. Conversely, if the scene information of the first original image does not match the scene information of the static original image, it indicates that the scene of the first original image has changed significantly compared to the static original image. For example, if the user moves the camera quickly during shooting, the scene will change drastically. When the scene of the first original image changes significantly compared to the static original image, forcibly using the acquisition parameters of the first original image to process it may result in severe overexposure, underexposure, or color cast in the processed first original image. Therefore, in this embodiment, image effect processing is only performed on image frames in the video stream whose scene is similar to that of the static original image, according to the first acquisition parameters, thereby avoiding the problem of severe overexposure, underexposure, or color cast in some frames of the video stream.
[0152] In this embodiment, by matching the scene information of the first original image frame with the scene information of the static original image, the first original image frame is processed based on the first acquisition parameters to obtain the processed first image frame. In this way, the problem of severe overexposure, underexposure or color cast in some frames of the video stream can be avoided, which is conducive to further improving the image quality of the generated dynamic photos.
[0153] Optionally, the method further includes:
[0154] If the scene information of the first original image frame does not match the scene information of the static original image, the first acquisition parameter and the second acquisition parameter are weighted and summed to obtain the third acquisition parameter. The second acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device acquires the first original image frame. The weight of the first acquisition parameter is greater than or equal to 0, and the weight of the second acquisition parameter is greater than 0.
[0155] Based on the third acquisition parameters, the first original image frame is processed to obtain the processed first image frame, and the video stream includes the first image frame.
[0156] In some embodiments of this application, the sum of the weights of the first acquisition parameter and the second acquisition parameter can be equal to 1. When the weight of the first acquisition parameter is equal to 0, the weight of the second acquisition parameter is equal to 1, and the third acquisition parameter is the second acquisition parameter. That is, the first acquisition parameter of the cover frame is not mandatory at this time. Correspondingly, when the weight of the first acquisition parameter is greater than 0, the weight of the second acquisition parameter is greater than 0, and the sum of the weights of the first acquisition parameter and the second acquisition parameter can be equal to 1, thereby converting the acquisition parameter into a safe and restricted parameter adjustment range.
[0157] Please see Figure 5 The diagram below illustrates the process of acquiring the static original image and video original image sequence as described in some embodiments of this application. The acquisition process includes the following steps:
[0158] Step 501: The electronic device enters the dynamic photo shooting mode.
[0159] Step 502: The electronic device continues to preview and runs the 3A algorithm to adjust the preview screen in real time;
[0160] Step 503: Cache N frames of RAW image data, where the value of N is determined by the number of video frames required to capture the dynamic photo.
[0161] Step 504: Based on the scene change detection logic, record the scene information for each frame. This scene information may include the scene change status information relative to the previous frame.
[0162] Step 505: Trigger photo capture, wherein, upon receiving the shutter trigger command, it is determined that the user has performed the shooting operation.
[0163] Step 506: Determine and lock the first acquisition parameters and 3A state information of the static original image. Specifically, at the precise time point (T0) corresponding to the shutter trigger command, capture and freeze the 3A algorithm state in effect in the ISP at that moment and all its output first acquisition parameters. The first acquisition parameters include: exposure, white balance, focus, gyro, and tonal and color processing parameters, etc.
[0164] Step 507: Generate a cover frame based on the first acquisition parameters. Using one frame of raw sensor data or K-frame raw data acquired at time T0, the first acquisition parameters are applied to perform image processing to generate the final cover frame. When only one frame of raw sensor data is used, this frame is the aforementioned static raw image, and its processing may include: performing noise reduction processing on the static raw image to obtain the cover frame. When K-frame raw data is used, this K-frame raw data can be used as the aforementioned K-frame raw image, and its processing may include: performing image effect processing on the K-frame raw image based on the first acquisition parameters to obtain a K-frame processed image. Then, based on the K-frame processed image, multi-frame noise reduction processing is performed on the static raw image to obtain the cover frame.
[0165] Step 508: After the cover frame, continue to cache the M-frame RAW image data;
[0166] Step 509: Combine the above N-frame RAW image data and M-frame RAW image data to form the above original video image sequence.
[0167] Please see Figure 6 This is a flowchart illustrating the process of applying image effects to a sequence of raw video images, including the following steps:
[0168] Step 601: Obtain the original video image sequence, wherein in step 601, it can be based on... Figure 5 The method shown obtains the original image sequence of the video;
[0169] Step 602: Determine whether the scene information of each image frame in the video stream matches the scene information of the cover frame. Specifically, you can determine whether the scene information of each image frame in the video stream matches the scene information of the cover frame based on the scene information recorded in the above steps.
[0170] Depending on whether the scene matches, select the corresponding processing method for each image frame in the video stream:
[0171] Step 603: If the scene information of the first original image frame and the cover frame do not match, the first acquisition parameter and the second acquisition parameter are weighted and summed to obtain the third acquisition parameter. Based on the third acquisition parameter, the first original image frame is processed to obtain the processed first image frame. The video stream includes the first image frame.
[0172] Step 604: If the scene information of the first original image frame and the cover frame match, then perform image effect processing on the first original image frame based on the first acquisition parameters to obtain the processed first image frame.
[0173] Encapsulation and storage: The cover frame and the processed video stream are encapsulated into a live image file and stored.
[0174] It should be noted that when processing image effects on the video stream, cover frame effect parameters should be used as much as possible to achieve alignment. The noise reduction and sharpening requirements for the video stream and the cover frame can differ; therefore, the noise reduction and sharpening parameters for the video stream can be set according to individual needs, aiming for clarity as close as possible to the cover frame.
[0175] In this embodiment, when the scene information of the first original image frame does not match the scene information of the static original image, the first acquisition parameter and the second acquisition parameter are weighted and summed to obtain a third acquisition parameter. Based on the third acquisition parameter, the first original image frame is processed to obtain a processed first image frame. The video stream includes the first image frame. In this way, the problem of severe overexposure, underexposure, or color cast in some frames of the video stream can be avoided, thereby helping to further improve the image quality of the generated dynamic photos.
[0176] Please see Figure 7 This application also provides an image processing method, including:
[0177] Step 701: Based on the transfer function of the dynamic photo, decode the cover frame of the dynamic photo and the video stream of the dynamic photo from the nonlinear domain to the linear domain to obtain the cover frame decoding data and the video stream decoding data.
[0178] The above-mentioned dynamic photos can be based on the above. Figure 4 The animated photograph generated by the image processing method in the illustrated embodiment is encoded. Figure 7 The image processing method shown is the same as Figure 4 The decoding process corresponding to the illustrated embodiment is specifically implemented as follows: Figure 4 The embodiments shown correspond to each other and have corresponding beneficial effects, which will not be described again here to avoid repetition.
[0179] The image processing method can be applied to various processes that require decoding of dynamic photos, such as the display of dynamic photos.
[0180] In some embodiments of this application, the above transfer function can be expressed by the following formula:
[0181]
[0182] Where x represents the input electrical signal, and the value of x is a value normalized to the range [0, 1]. γ is the output light intensity, a linear brightness value; γ1 is the gamma coefficient, usually taken as 2.2 or 2.4, depending on the display standard used.
[0183] In this embodiment, by decoding the cover frame and video stream of the dynamic photo from the nonlinear domain to the linear domain based on the same transfer function, it is beneficial to ensure that the cover frame and video stream in the dynamic photo have unified color management information, thereby improving the consistency of the playback effect of the cover frame and video stream in the dynamic photo during playback.
[0184] Optionally, the dynamic photo is encoded based on first metadata, which includes OETF;
[0185] If the first metadata does not include OOTF, the inverse of the OETF is the transfer function;
[0186] When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function.
[0187] In this embodiment, when the first metadata does not include the optical-to-optical conversion function (OOTF), the inverse of the OETF is determined as the transfer function of the moving image; when the first metadata includes the OOTF, the inverse of the OETF and the OOTF together form the transfer function of the moving image. Thus, the first metadata can include or exclude the OOTF as needed to select different transfer functions, thereby achieving the goal of restoring scene lighting as much as possible or presenting image effects according to the creator's subjective expression.
[0188] Optionally, the dynamic photo is encoded based on the first metadata and the second metadata, wherein the second metadata includes a color conversion matrix;
[0189] After the method uses the transfer function based on the dynamic photo to decode the cover frame and video stream of the dynamic photo from the nonlinear domain to the linear domain, obtaining the cover frame decoded data and the video stream decoded data, the method further includes:
[0190] The cover frame decoding data and the video stream decoding data are color-converted based on the inverse matrix of the color conversion matrix to obtain the cover frame to be displayed and the video stream to be displayed.
[0191] In some embodiments of this application, the second metadata encoding may include not only the color conversion matrix but also its inverse. The inverse of the color conversion matrix can convert the cover frame decoding data and the video stream decoding data into RGB format. The inverse of the color conversion matrix may be a YUV 2RGB matrix.
[0192] It is understandable that after obtaining the cover frame to be displayed and the video stream to be displayed, the electronic device can directly display the cover frame to be displayed and the video stream to be displayed. In the process of displaying the cover frame to be displayed and the video stream to be displayed, the electronic device can first play the video stream to be displayed, and after playing the video stream to be displayed, the screen can be frozen on the cover frame to be displayed.
[0193] In this embodiment, since the color conversion matrix is part of the color management information, the cover frame decoding data and the video stream decoding data are color converted by the inverse matrix based on the same color conversion matrix. This helps to further improve the consistency of the color management information used by the cover frame and the video stream during the decoding process, and thus further improves the consistency of the effect between the image and the video in the obtained dynamic photo.
[0194] Optionally, the first metadata may further include a transfer function indicator parameter;
[0195] When the transfer function indicator parameter indicates a first encoding method, the transfer function is the inverse of the OETF;
[0196] When the transfer function indicator parameter indicates a second encoding method, the transfer function consists of the inverse of the OETF and the OOTF in the first metadata.
[0197] In this embodiment, the first metadata further includes a transfer function indicator parameter; when the transfer function indicator parameter indicates a first encoding method, the transfer function of the moving image is the inverse of the OETF; when the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image is composed of the inverse of the OETF and the OOTF in the first metadata. In this way, different EOTEs can be selected according to different needs to achieve the goal of restoring scene light as much as possible or presenting image effects according to the creator's subjective expression.
[0198] It should be noted that this application embodiment targets the cover frame and video stream in dynamic photos. Based on color gamut alignment, it aligns the processing modules and parameters of both during the front-end debugging process and unifies their color management information during the encoding stage, thereby achieving consistent playback effects during decoding. Specifically, a unified effect processing framework and color management process are introduced at both the generation and playback ends of the dynamic photo, ensuring that the image and video undergo the same or visually equivalent effect processing and color management, guaranteeing consistent playback effects. Simultaneously, a method for optimizing the consistency of image and video playback effects in dynamic photos is proposed. By adopting the same effect processing process and a unified color management scheme, the processing effects of both are aligned as much as possible from the generation end to the display end, thereby achieving the goal of optimizing consistent image and video playback effects. This embodiment proposes a method for optimizing the consistency of image and video effects in dynamic photos by recording scene change information and the effect parameters and 3A information of the cover frame to process the video frame, thereby achieving the goal of aligning image and video effects at the generation end.
[0199] The dynamic photo processing method provided in this application can be executed by a dynamic photo generating device. As an example, the device can be an electronic device or a component within an electronic device, such as a chip or circuit. This application uses a dynamic photo generating device executing the dynamic photo processing method as an example to illustrate the dynamic photo generating device provided in this application.
[0200] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an image processing device 800 provided in an embodiment of this application. The image processing device 800 includes:
[0201] The determination module 801 is used to determine the first metadata of the encoded motion picture. The first metadata includes the photoelectric conversion function OETF, which is used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
[0202] Optionally, if the first metadata does not include the optical-optical conversion function OOTF, the inverse of the OETF is the transfer function of the moving image;
[0203] When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function of the moving image.
[0204] Optionally, the determining module 801 is used to determine the first metadata and the second metadata of the encoded dynamic photo, wherein the second metadata includes a color conversion matrix, which is used to perform color conversion on the cover frame of the dynamic photo and the video stream of the dynamic photo.
[0205] Optionally, the first metadata may further include a transfer function indicator parameter;
[0206] When the transfer function indicator parameter indicates the first encoding method, the transfer function of the moving image is the inverse of the OETF;
[0207] When the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image consists of the inverse of the OETF and the OOTF in the first metadata.
[0208] Optionally, the device further includes:
[0209] The acquisition module is used to acquire static raw images and video raw image sequences;
[0210] The processing module is used to perform image effect processing on the static original image and the video original image sequence respectively, to obtain the cover frame and the video stream with matching image effects.
[0211] Optionally, the processing module is specifically used to perform image effect processing on the static original image under the same color gamut to obtain the cover frame, and to perform image effect processing on the video original image sequence based on the first acquisition parameters to obtain the video stream, wherein the first acquisition parameters include the image acquisition parameters of the static original image.
[0212] Optionally, the static original image is an image captured by the camera in the electronic device when the electronic device receives a shooting command;
[0213] The first acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device receives the shooting command.
[0214] Optionally, the original video image sequence includes: N frames of images preceding the static original image, and M frames of images following the static original image, where N and M are integers greater than 0.
[0215] Optionally, the first acquisition parameters include at least one of the following: exposure parameters, white balance parameters, focus parameters, gyroscope parameters, tone parameters, and color parameters.
[0216] Optionally, the processing module is specifically used to perform image effect processing on the first original image frame based on the first acquisition parameters when the scene information of the first original image frame matches the scene information of the static original image, to obtain a processed first image frame, wherein the first original image frame is any image frame in the original video image sequence, and the video stream includes the first image frame.
[0217] Optionally, the processing module is specifically used to perform a weighted summation of the first acquisition parameter and the second acquisition parameter to obtain a third acquisition parameter when the scene information of the first original image frame does not match the scene information of the static original image. The second acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device acquires the first original image frame. The weight of the first acquisition parameter is greater than or equal to 0, and the weight of the second acquisition parameter is greater than 0.
[0218] The processing module is further configured to perform image effect processing on the first original image frame based on the third acquisition parameters to obtain a processed first image frame, and the video stream includes the first image frame.
[0219] In this embodiment, during the processing of dynamic photos, the cover frame and video stream of the dynamic photo are encoded from the linear domain to the non-linear domain using the same OETF. This allows the cover frame and video stream in the dynamic photo to have unified color management information during the encoding stage, which helps to improve the consistency of the playback effect of the cover frame and video stream in the dynamic photo during playback.
[0220] The image processing apparatus 800 provided in this application embodiment can achieve... Figure 4 The various processes implemented in the method embodiments shown achieve the same technical effect, and will not be described again here to avoid repetition.
[0221] Please see Figure 9 This application also provides an image processing apparatus 900, comprising:
[0222] The decoding module 901 is used to decode the cover frame of the dynamic photo and the video stream of the dynamic photo from the nonlinear domain to the linear domain based on the transfer function of the dynamic photo, so as to obtain the cover frame decoding data and the video stream decoding data.
[0223] Optionally, the dynamic photo is encoded based on first metadata, which includes OETF;
[0224] If the first metadata does not include OOTF, the inverse of the OETF is the transfer function;
[0225] When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function.
[0226] Optionally, the dynamic photo is encoded based on the first metadata and the second metadata, wherein the second metadata includes a color conversion matrix; the device further includes:
[0227] The color conversion module is used to perform color conversion on the cover frame decoding data and the video stream decoding data based on the inverse matrix of the color conversion matrix to obtain the cover frame to be displayed and the video stream to be displayed.
[0228] Optionally, the first metadata may further include a transfer function indicator parameter;
[0229] When the transfer function indicator parameter indicates a first encoding method, the transfer function is the inverse of the OETF;
[0230] When the transfer function indicator parameter indicates a second encoding method, the transfer function consists of the inverse of the OETF and the OOTF in the first metadata.
[0231] The image processing apparatus 900 provided in this application embodiment can achieve... Figure 7 The various processes implemented in the method embodiment shown achieve the same technical effect, and will not be described again here to avoid repetition.
[0232] like Figure 10As shown, this application embodiment also provides an electronic device 1000, including a processor 1001 and a memory 1002. The memory 1002 stores a program or instructions that can run on the processor 1001. For example, when the electronic device 1000 is an encoding device, the program or instructions executed by the processor 1001 implement the various steps of the above-described image processing method embodiment and achieve the same technical effect. When the electronic device 1000 is a decoding device, the program or instructions executed by the processor 1001 implement the various steps of the above-described image processing method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here. Optionally, the memory 1002 may be... Figure 1 The processor 1001 can implement the memory 102 or memory 113 in the illustrated embodiment. Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0233] This application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the various steps of the image processing method embodiments described above. Optionally, the memory may be... Figure 1 The processing circuitry of memory 102 or memory 113 in the illustrated embodiment can implement... Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0234] This application embodiment also provides an electronic device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 11 The steps in the method embodiment shown are illustrated. This device embodiment corresponds to the above method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect.
[0235] The processors or processing circuits in the embodiments of this application may include general-purpose processors, special-purpose processors, etc., such as central processing units (CPUs), microprocessors, digital signal processors (DSPs), artificial intelligence (AI) processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), network processors (NPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The communication interfaces in the embodiments of this application may include transceivers, pins, circuits, buses, etc.
[0236] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0237] Among them, the terminal can also be called user equipment (UE), which can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication functions, such as refrigerators, televisions, washing machines or furniture, etc.), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0238] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0239] For example, the aforementioned electronic devices may include, but are not limited to, those described above. Figure 1 The type of source device 100 or destination device 110 shown.
[0240] Taking electronic devices as terminals as an example, Figure 11A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0241] The terminal 1100 includes, but is not limited to, at least some of the following components: radio frequency unit 1101, network module 1102, audio output unit 1103, input unit 1104, sensor 1105, display unit 1106, user input unit 1107, interface unit 1108, memory 1109, and processor 1110.
[0242] Those skilled in the art will understand that the terminal 1100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0243] It should be understood that, in this embodiment, the input unit 1104 may include a graphics processor 11041 and a microphone 11042. The graphics processor 11041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode, or it may process the obtained point cloud data. The display unit 1106 may include a display panel 11061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1107 includes at least one of a touch panel 11071 and other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0244] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1101 can transmit it to the processor 1110 for processing; in addition, the radio frequency unit 1101 can send uplink data to the network-side device. Typically, the radio frequency unit 1101 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0245] The memory 1109 can be used to store software programs or instructions, as well as various data. The memory 1109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1109 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1109 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0246] Processor 1110 may include one or more processing units; optionally, processor 1110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1110.
[0247] The processor 1110 is configured to: determine first metadata for encoding a motion picture, the first metadata including an optoelectronic conversion function OETF, the OETF being used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
[0248] Alternatively, the processor 1110 is configured to: decode the cover frame of the dynamic photo and the video stream of the dynamic photo from the nonlinear domain to the linear domain based on the transfer function of the dynamic photo, to obtain cover frame decoding data and video stream decoding data.
[0249] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the image processing method embodiment and achieve the same or corresponding technical effect. To avoid repetition, it will not be described again here.
[0250] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0251] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0252] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0253] It should be understood that the chips mentioned in the embodiments of this application may include system-on-a-chip (also known as system chip, chip system, or system-on-a-chip), or independent display chips, etc.
[0254] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0255] This application also provides a communication system, including: an encoding end device and a decoding end device, wherein the encoding end device can be used to perform the steps of the image processing method described above, and the decoding end device can be used to perform the steps of the image processing method described above.
[0256] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0257] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.), and the computer software product includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0258] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. An image processing method, characterized in that, include: First metadata for encoding a motion picture is determined, the first metadata including an optoelectronic conversion function OETF, the OETF being used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
2. The method according to claim 1, characterized in that, The method further includes: If the first metadata does not include the optical-optical conversion function OOTF, the inverse of the OETF is the transfer function of the moving image; When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function of the moving image.
3. The method according to claim 1, characterized in that, The first metadata for determining the encoded dynamic photograph includes: Determine first metadata and second metadata for encoding a motion picture, wherein the second metadata includes a color conversion matrix used to perform color conversion on the cover frame of the motion picture and the video stream of the motion picture.
4. The method according to any one of claims 1 to 3, characterized in that, The first metadata also includes a pass function indicator parameter; When the transfer function indicator parameter indicates the first encoding method, the transfer function of the moving image is the inverse of the OETF; When the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image consists of the inverse of the OETF and the OOTF in the first metadata.
5. The method according to any one of claims 1 to 3, characterized in that, The method includes: Obtain static raw images and video raw image sequences; The static original image and the video original image sequence are processed with image effects respectively to obtain the cover frame and the video stream with matching image effects.
6. The method according to claim 5, characterized in that, The step of performing image effect processing on the static original image and the video original image sequence respectively to obtain the cover frame and the video stream with matching image effects includes: Under the same color gamut, the static original image is processed to obtain the cover frame, and the video original image sequence is processed to obtain the video stream based on the first acquisition parameters, wherein the first acquisition parameters include the image acquisition parameters of the static original image.
7. The method according to claim 6, characterized in that, The static raw image is the image captured by the camera in the electronic device when the electronic device receives a shooting command; The first acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device receives the shooting command.
8. The method according to claim 6, characterized in that, The original video image sequence includes: N frames of images preceding the static original image, and M frames of images following the static original image, where N and M are integers greater than 0.
9. The method according to claim 6, characterized in that, The first acquisition parameters include at least one of the following: exposure parameters, white balance parameters, focus parameters, gyroscope parameters, tone parameters, and color parameters.
10. The method according to claim 6, characterized in that, The step of processing the original video image sequence based on the first acquisition parameters to obtain the video stream includes: When the scene information of the first original image frame matches the scene information of the static original image, the first original image frame is processed based on the first acquisition parameters to obtain the processed first image frame. The first original image frame is any one of the image frames in the original video image sequence, and the video stream includes the first image frame.
11. The method according to claim 10, characterized in that, The method further includes: If the scene information of the first original image frame does not match the scene information of the static original image, the first acquisition parameter and the second acquisition parameter are weighted and summed to obtain the third acquisition parameter. The second acquisition parameter is the image acquisition parameter of the camera in the electronic device when the electronic device acquires the first original image frame. The weight of the first acquisition parameter is greater than or equal to 0, and the weight of the second acquisition parameter is greater than 0. Based on the third acquisition parameters, the first original image frame is processed to obtain the processed first image frame, and the video stream includes the first image frame.
12. An image processing method, characterized in that, include: Based on the transfer function of the moving image, the cover frame and video stream of the moving image are decoded from the nonlinear domain to the linear domain to obtain the cover frame decoded data and the video stream decoded data.
13. The method according to claim 12, characterized in that, The dynamic photo is obtained based on the encoding of the first metadata, which includes OETF; If the first metadata does not include OOTF, the inverse of the OETF is the transfer function; When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function.
14. The method according to claim 12, characterized in that, The dynamic photo is encoded based on a first metadata and a second metadata, wherein the second metadata includes a color conversion matrix; After the method uses the transfer function based on the dynamic photo to decode the cover frame and video stream of the dynamic photo from the nonlinear domain to the linear domain, obtaining the cover frame decoded data and the video stream decoded data, the method further includes: The cover frame decoding data and the video stream decoding data are color-converted based on the inverse matrix of the color conversion matrix to obtain the cover frame to be displayed and the video stream to be displayed.
15. The method according to claim 13, characterized in that, The first metadata also includes a pass function indicator parameter; When the transfer function indicator parameter indicates a first encoding method, the transfer function is the inverse of the OETF; When the transfer function indicator parameter indicates a second encoding method, the transfer function consists of the inverse of the OETF and the OOTF in the first metadata.
16. An image processing apparatus, characterized in that, include: A determination module is used to determine first metadata for encoding a motion picture, the first metadata including an optoelectronic conversion function (OETF), the OETF being used to encode the cover frame of the motion picture and the video stream of the motion picture from the linear domain to the nonlinear domain.
17. The apparatus according to claim 16, characterized in that, If the first metadata does not include the optical-optical conversion function OOTF, the inverse of the OETF is the transfer function of the moving image; When the first metadata includes OOTF, the inverse of the OETF and the OOTF constitute the transfer function of the moving image.
18. The apparatus according to claim 16, characterized in that, The determining module is used to determine the first metadata and the second metadata of the encoded dynamic photo, wherein the second metadata includes a color conversion matrix, which is used to perform color conversion on the cover frame of the dynamic photo and the video stream of the dynamic photo.
19. The apparatus according to any one of claims 16 to 18, characterized in that, The first metadata also includes a pass function indicator parameter; When the transfer function indicator parameter indicates the first encoding method, the transfer function of the moving image is the inverse of the OETF; When the transfer function indicator parameter indicates a second encoding method, the transfer function of the moving image consists of the inverse of the OETF and the OOTF in the first metadata.
20. An image processing apparatus, characterized in that, include: The decoding module is used to decode the cover frame and video stream of the dynamic photo from the nonlinear domain to the linear domain based on the transfer function of the dynamic photo, so as to obtain the cover frame decoding data and the video stream decoding data.
21. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the image processing method as claimed in any one of claims 1-11, or to implement the steps of the image processing method as claimed in any one of claims 12-15.
22. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image processing method as described in any one of claims 1-11, or implement the steps of the image processing method as described in any one of claims 12-15.