Multimedia metadata encoding method and multimedia metadata decoding method
By using floating-point data format to pass multimedia metadata in the code stream, the problem of floating-point accuracy loss in the prior art is solved, the floating-point number consistency at the encoding and decoding ends is ensured, and the multimedia reconstruction quality and system compatibility are improved.
Patent Information
- Application Number
- PCT/CN2024/138885
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-24
AI Technical Summary
The prior art will lead to a loss of floating point accuracy when encoding multimedia metadata, especially when encoding high dynamic range image/video, which will cause floating point numbers to be inconsistent with the original floating point numbers at the encoding side, affecting the quality of reconstruction multimedia.
The multimedia metadata is passed in the code stream using the floating point data format to ensure that the floating point values in the memory of the encoding and decoding ends are consistent. By using the floating point storage form in the computer memory, the accuracy loss is avoided.
The floating point number obtained at the decoding end is consistent with the original floating point number at the encoding end, which improves the quality and appearance of reconstruction multimedia, especially in the display and reconstruction of high dynamic range images/videos, ensuring compatibility of different systems.
Smart Images

Figure CN2024138885_24072025_PF_FP_ABST
Abstract
Description
Multimedia metadata encoding and decoding method
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on January 16, 2024, with application number 202410064774.4 and application name “Metadata encoding and decoding method for multimedia”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of encoding and decoding, and in particular to a multimedia metadata encoding and decoding method. Background Art
[0003] As multimedia (such as audio, video, and images) develops from high definition to ultra-high definition, the amount of multimedia data also increases, and the requirements for multimedia storage and transmission bandwidth are also getting higher and higher; therefore, multimedia will be compressed (also called encoded) before storage or transmission to reduce the memory occupied by storage and transmission bandwidth.
[0004] Typically, when multimedia is encoded, metadata of the multimedia is also encoded; thereafter, the encoded multimedia and the multimedia metadata may be packaged into a file (also referred to as a code stream), and then the file (or code stream) may be stored or transmitted.
[0005] Currently, multimedia metadata in some scenarios (such as metadata for High Dynamic Range (HDR) images / videos) contains floating-point numbers. Existing solutions involve truncating floating-point numbers to a certain length during encoding and then encapsulating them in a file or stream as an American Standard Code for Information Interchange (ASCII) string. During decoding, these strings are then restored to floating-point numbers.
[0006] For example, "0.131241241241..." is truncated to "0.1312412" and then represented in the bitstream using the Extensible Metadata Platform (XMP) or eXtensible Markup Language (XML). During decoding, the string "0.1312412" is converted back to the floating-point number 0.1312412. This approach results in a loss of precision in the floating-point number in the metadata, causing the resulting floating-point number on the decoder to differ from the original floating-point number on the encoder. Summary of the Invention
[0007] In view of this, the present application provides a multimedia metadata encoding and decoding method, which can ensure that the floating-point number finally obtained by the decoding end is consistent with the original floating-point number of the encoding end without loss of precision.
[0008] It should be noted that the multimedia involved in this application may include but is not limited to: images, videos and audio.
[0009] Exemplarily, the present application can be applied to various image business scenarios (for example, cloud photo albums, etc.), various video / audio business scenarios (for example, live broadcast, on-demand, etc.), etc., and the present application does not impose any restrictions on this.
[0010] In a first aspect, the present application provides a method for encoding multimedia metadata, the method comprising: first, obtaining multimedia metadata; wherein the multimedia metadata includes first data, which is floating-point data; then, writing second data into a code stream; wherein the second data is a representation of the first data in a floating-point data format.
[0011] Since codewords of an agreed length are typically transmitted in the code stream for specific syntax elements, and since floating-point numbers are typically highly precise, they are typically fixed-pointed or converted into other formats (such as ASCII strings, which involve length truncation during the conversion process) during the encapsulation of syntax elements in existing standards, resulting in a loss of precision. However, the present application uses the form of floating-point number storage in computer memory (i.e., the floating-point data format of the second data) for transmission in the code stream, ensuring that the floating-point number values stored in the memory of the sending and receiving ends are completely consistent. In this way, the floating-point data obtained by the decoding end through the second data conversion is the same as the first data used for encoding by the encoding end, thus ensuring that the floating-point number ultimately obtained by the decoding end is consistent with the original floating-point number of the encoding end.
[0012] Secondly, since the floating-point data in the metadata of multimedia on the codec side is consistent, in scenarios where it is necessary to generate and reconstruct multimedia based on the multimedia metadata or to display and reconstruct multimedia content based on the multimedia metadata, it can improve the quality of the reconstructed multimedia or the visual experience when displaying the reconstructed multimedia content.
[0013] Exemplarily, the floating-point type includes a 32-bit float type or a 64-bit double type.
[0014] Illustratively, the floating-point data format is a data format that specifies the fields constituting floating-point data, the layout of the fields, and the arithmetic interpretation.
[0015] For example, the first data is "0.125490203", which is the mathematical expression of a floating point number, and the second data may be "0X3E008081", which is the binary expression of a floating point number in a computer.
[0016] Exemplarily, the first data may be one or more, and the second data may be one or more; based on one first data, one second data may be generated; that is, multiple first data and multiple second data correspond one to one.
[0017] It should be understood that the metadata of multimedia may also include other data in addition to the first data, and the other data in the metadata of multimedia may be encoded and decoded using existing technology methods, and this application does not impose any restrictions on this; this application describes the encoding and decoding process of floating-point data in multimedia data.
[0018] It should be understood that multimedia metadata may be image metadata, audio metadata, video metadata, etc., and this application does not impose any limitation on this.
[0019] For example, when the multimedia is an image or video, the second data may be encapsulated in an image encapsulation format to obtain a code stream. The image encapsulation format includes, but is not limited to, JPEG JFIF encapsulation format, JPEG Universal Metadata Box Format JUMBF encapsulation format, High Efficiency Image File Format (HEIF), MP4 file format, and the like, and this application does not impose any restrictions on this.
[0020] For example, the second data is encapsulated according to the JPEG JFIF encapsulation format, wherein the second data can be encapsulated into the APPN field of the JPEG JFIF encapsulation format.
[0021] For another example, the second data is encapsulated into the payload of APP11 JP in the JUMBF encapsulation format.
[0022] For another example, the second data may be encapsulated in a box in the JUMBF encapsulation format.
[0023] For another example, the second data may be encapsulated in a box of the HEIF encapsulation format, and the box may be 'idat' or others.
[0024] For another example, the second data is encapsulated into the supplemental enhancement information (SEI) of HEVC or VVC, or into a user-defined network abstraction layer (NAL) unit or a reserved NAL unit.
[0025] It should be understood that the second data can also be encapsulated into the file location indicated by the metadata location information (the metadata location information can be a type of multimedia metadata), such as the position after the EOI (end of image) of the complete JPEG file, or the end of the complete HEIF file, etc. This application does not impose any restrictions on this.
[0026] It should be understood that the present application does not impose any restrictions on the file location of the second data in the code stream.
[0027] For example, the present application does not limit the metadata structure of multimedia. For example, it may be the metadata structure corresponding to 'tmap' in 23008-12AMD 3.
[0028] For example, the second data may be written into the code stream in the form of an ASCII string; or in the form of an ASCII byte stream. It should be understood that this application does not limit the form in which the second data is written into the code stream.
[0029] According to the first aspect, the method further comprises: generating second data based on the first data.
[0030] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining a floating-point data format; generating second data based on the first data, including: processing the first data according to the floating-point data format to obtain the second data.
[0031] Since the floating-point data format has a unified standard, after the encoding end processes the first data according to the floating-point data format to obtain the second data, the decoding end processes the second data according to the floating-point data format to restore the first data; in this way, it can be ensured that the floating-point number finally obtained by the decoding end is consistent with the original floating-point number of the encoding end.
[0032] Exemplarily, the first data may be converted into the second data according to the various fields of floating-point data, the layout and arithmetic interpretation of the various fields, and the decimal-to-binary conversion method specified by the floating-point data format.
[0033] In one possible manner, the floating-point data format may be preset; thus, the preset floating-point data format may be acquired.
[0034] In one possible manner, the floating-point data format may be determined based on an application scenario, channel conditions, the number of significant bits of the first data, and the like.
[0035] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining the bit width of the second data; processing the first data according to the floating-point data format to obtain the second data, including: processing the first data according to the floating-point data format and the bit width to obtain the second data.
[0036] For example, bit width may refer to the number of bits occupied by data in a computer system. The bit width may include at least one of the following: 8 bits, 16 bits, 19 bits, 32 bits, or 64 bits, etc., and this application does not impose any limitation on this.
[0037] In one possible approach, the bit width may be preset; in this way, the preset bit width may be obtained.
[0038] In one possible approach, the bit width may be determined based on an application scenario, channel conditions, and the number of valid bits of the first data.
[0039] Specifically, the length of each field of the floating-point data specified by the floating-point data format for the bit width can be determined; then, the first data can be converted into the second data based on the various fields of the floating-point data specified by the floating-point data format, the length of each field specified for the bit width, the layout and arithmetic interpretation of each field, and the method of converting decimal to binary.
[0040] According to the first aspect, or any implementation of the first aspect above, the method further includes: writing the floating-point data format into the code stream.
[0041] Exemplarily, when the floating-point data format is determined based on the application scenario, channel conditions, and the number of valid bits of the first data, the floating-point data format can be written into the code stream; so that the decoding end can obtain the floating-point data format and convert the second data into the first data.
[0042] For example, when the floating-point data format is pre-set and the codec has pre-synchronized the floating-point data format, it is not necessary to write the floating-point data format into the bitstream. For example, when the floating-point data format is pre-set but the codec has not pre-synchronized the floating-point data format, it is also possible to write the floating-point data format into the bitstream.
[0043] According to the first aspect, or any implementation of the first aspect above, the method further includes: writing the bit width of the second data into the code stream.
[0044] Exemplarily, when the bit width is determined based on the application scenario, channel conditions, and the effective number of bits of the first data, the bit width of the second data can be written into the code stream; so that the decoding end can obtain the bit width of the second data, and then accurately divide a single second data from the multiple second data parsed from the code stream, and realize the conversion of the second data into the first data.
[0045] Exemplarily, when the bit width is pre-set and the codec has pre-synchronized the bit width, the bit width of the second data may not need to be written into the bitstream. Exemplarily, when the bit width is pre-set but the codec has not pre-synchronized the bit width, the bit width of the second data may also be written into the bitstream.
[0046] It should be noted that the present application does not limit the position of the bit width, floating point data format and the second data in the code stream. In one possible embodiment, the bit width in the code stream is located after the floating point data format, and the second data is located after the bit width.
[0047] According to the first aspect, or any implementation of the first aspect, the second data is a binary representation or a hexadecimal representation of the first data in a floating-point data format.
[0048] In one possible manner, after the first data is processed according to the floating-point data format and bit width to obtain binary data, the binary data can be directly used as the second data. In another possible manner, after the first data is processed according to the floating-point data format and bit width to obtain binary data, the binary data can be converted into hexadecimal data; and the hexadecimal data can be used as the second data.
[0049] For example, when the first data is “0.125490203” and the bit width of the second data is 32, the second data may be a hexadecimal representation of the first data in IEEE 754 format, that is, “0X3E008081”.
[0050] For example, when the first data is “0.35” and the bit width of the second data is 32, the second data may be a binary expression of the first data in IEEE 754 format, that is, “0100000001010000000000000000000000”.
[0051] According to the first aspect, or any implementation of the first aspect above, the metadata of the multimedia is metadata of a high dynamic range HDR image.
[0052] Exemplarily, the metadata of an HDR image may include dynamic metadata and static metadata, wherein dynamic metadata may refer to data that changes over time, and static metadata may refer to data that does not change over time.
[0053] According to the first aspect, or any implementation of the first aspect above, the metadata of the multimedia is metadata of enhanced data, and the enhanced data is generated based on a standard dynamic range (SDR) image and an HDR image.
[0054] The metadata of the enhanced data is generated during the process of performing a difference calculation between the base image and the first image to generate the enhanced data, and is data that affects the difference calculation result (i.e., the enhanced data). In one possible scenario, the base image is an SDR image and the first image is an HDR image. In another possible scenario, the base image is an HDR image and the first image is an SDR image.
[0055] In this case, the code stream may include an encoded basic image and encoded enhanced data. After receiving the code stream, the decoding end may decode the code stream to obtain a reconstructed image of the basic image, a reconstruction result of the enhanced data, and reconstruction metadata of the enhanced data. When the basic image is an SDR image and the first image is an HDR image, if the decoding end supports displaying HDR, the reconstructed image of the basic image and the reconstruction result of the enhanced data may be merged according to the reconstruction metadata of the enhanced data to obtain an HDR reconstructed image; if the decoding end only supports displaying SDR images, the reconstructed image of the basic image may be displayed. When the basic image is an HDR image and the first image is an SDR image, if the decoding end only supports displaying SDR, the reconstructed image of the basic image and the reconstruction result of the enhanced data may be merged according to the reconstruction metadata of the enhanced data to obtain an SDR reconstructed image; if the decoding end supports displaying HDR images, the reconstructed image of the basic image may be displayed. In this way, compatibility with formats supported by different systems can be achieved.
[0056] Due to the encoding and decoding method of the reconstructed metadata of the enhanced data in the present application, there will be no loss of accuracy of the reconstructed metadata of the enhanced data, and the reconstructed metadata of the enhanced data obtained by the decoding end is the same as the metadata of the enhanced data encoded by the encoding end, thereby ensuring the quality of the SDR image or HDR image obtained by the decoding end.
[0057] According to the first aspect, or any implementation of the first aspect above, the multimedia metadata may also be metadata of audio data, or metadata of image / video data in other formats, which is not limited in this application.
[0058] According to the first aspect, or any implementation of the first aspect above, the floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, and TensorFloat-32.
[0059] For example, the second data is encapsulated according to the JPEG JFIF encapsulation format, wherein the second data can be encapsulated into the APPN field of the JPEG JFIF encapsulation format.
[0060] According to the first aspect, or any implementation of the first aspect above, the bit width of the second data is greater than or equal to the bit width of the first data. Physical storage of floating-point numbers in computers is mostly in the form of exponent bits and decimal bits. A floating-point number format represented by fewer exponent bits and decimal bits (e.g., 16 bits) can be losslessly converted to a floating-point number format represented by more exponent bits and decimal bits (e.g., 32 bits). In the present application, by limiting the bit width of the second data to be greater than or equal to the bit width of the first data, lossless conversion from the first data to the second data can be guaranteed, thereby ensuring lossless transmission of metadata.
[0061] In a second aspect, the present application provides a code stream, which is generated according to the first aspect or any implementation manner of the first aspect.
[0062] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0063] In a third aspect, the present application provides a code stream, which includes second data, where the second data is a representation of first data in a floating-point data format, and the first data is floating-point data in multimedia metadata.
[0064] According to a third aspect, the code stream further includes a floating point data format.
[0065] According to the third aspect, the second data is obtained by processing the first data according to the floating-point data format and the bit width, and the code stream also includes the bit width.
[0066] Exemplarily, when the metadata of the multimedia is metadata of enhanced data, the code stream may further include an encoded basic image and encoded enhanced data.
[0067] The third aspect and any implementation of the third aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the third aspect and any implementation of the third aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0068] In a fourth aspect, the present application provides a method for decoding multimedia metadata, the method comprising: first, receiving a code stream; then, parsing second data from the code stream, where the second data is a representation of the first data in a floating-point data format, and the first data is floating-point data in the multimedia metadata; thereafter, generating the first data based on the second data.
[0069] Exemplarily, the second data may be parsed from an image encapsulation format in a code stream.
[0070] For example, the second data may be parsed from the JPEG JFIF encapsulation format, such as parsing the second data from the APPN field of the JPEG JFIF encapsulation format.
[0071] For another example, the second data may be parsed from a JUMBF encapsulation format. For example, the second data may be parsed from a payload of an APP11 JP in a JUMBF encapsulation format. For another example, the second data may be parsed from a box in a JUMBF encapsulation format.
[0072] For another example, the second data may be parsed from the HEIF encapsulation format, for example, the second data may be parsed from a box in the HEIF encapsulation format, or for another example, the second data may be parsed from an 'idat' box in the HEIF encapsulation format.
[0073] For another example, the second data is parsed from the SEI of HEVC or VVC or the NAL unit defined by the user or the reserved NAL unit.
[0074] It should be understood that metadata location information can also be parsed from the code stream (the metadata location information can be a type of multimedia metadata, used to indicate the file location of the multimedia metadata), and then the second data can be parsed from the file position indicated by the metadata location information in the code stream (for example, the position after the EOI (end of image) of the complete JPEG file, or the end of the complete HEIF file, etc., which is not limited in this application).
[0075] It should be understood that the decoding end can also parse the second data from other file locations in the code stream, and this application does not limit this.
[0076] According to a fourth aspect, the method further includes: obtaining a floating-point data format; and generating first data based on the second data, including: processing the second data according to the floating-point data format to obtain the first data.
[0077] According to the fourth aspect, or any implementation of the fourth aspect above, obtaining the floating-point data format includes: parsing the floating-point data format from the code stream.
[0078] In one possible approach, when the decoder fails to parse the floating point data format from the code stream, it can obtain the pre-set floating point data format.
[0079] According to the fourth aspect, or any implementation of the fourth aspect above, the method also includes: obtaining the bit width of the second data; processing the second data according to the floating-point data format to obtain the first data, including: processing the second data according to the floating-point data format and the bit width to obtain the first data.
[0080] According to the fourth aspect, or any implementation of the fourth aspect above, obtaining the bit width of the second data includes: parsing the bit width from the code stream.
[0081] In one possible approach, when the decoding end fails to parse the bit width from the code stream, the preset bit width may be obtained.
[0082] According to the fourth aspect, or any implementation of the fourth aspect, the second data is a binary representation or a hexadecimal representation of the first data in a floating-point data format.
[0083] According to the fourth aspect, or any implementation of the fourth aspect, the multimedia metadata is metadata of a high dynamic range (HDR) image.
[0084] According to the fourth aspect, or any implementation of the fourth aspect above, the metadata of the multimedia is metadata of enhanced data, and the enhanced data is determined based on the standard dynamic range SDR image and the high dynamic range HDR image.
[0085] According to the fourth aspect, or any implementation of the fourth aspect above, the floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, and TensorFloat-32.
[0086] The fourth aspect and any implementation of the fourth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0087] In a fifth aspect, the present application provides an encoding device, the encoding device comprising:
[0088] A metadata acquisition module, configured to acquire multimedia metadata; wherein the multimedia metadata includes first data, which is floating-point data;
[0089] The data writing module is used to write the second data into the code stream; wherein the second data is a representation of the first data in a floating point data format.
[0090] Illustratively, the above-mentioned encoding device can be used to execute the method in the first aspect or any possible implementation of the first aspect.
[0091] The fifth aspect and any implementation of the fifth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0092] In a sixth aspect, the present application provides a decoding device, the decoding device comprising:
[0093] A code stream receiving module, used for receiving code stream;
[0094] a parsing module, configured to parse the code stream to obtain second data, wherein the second data is a representation of the first data in a floating-point data format, and the first data is floating-point data in multimedia metadata;
[0095] The second data generating module is used to generate the first data based on the second data.
[0096] Exemplarily, the above-mentioned decoding device can be used to execute the method in the fourth aspect or any possible implementation of the fourth aspect.
[0097] The sixth aspect and any implementation of the sixth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the sixth aspect and any implementation of the sixth aspect can be referred to the technical effects corresponding to the fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0098] In the seventh aspect, the present application provides an encoder comprising: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, enables the encoder to execute the method in the first aspect or any possible implementation of the first aspect.
[0099] The seventh aspect and any implementation of the seventh aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the seventh aspect and any implementation of the seventh aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0100] In an eighth aspect, the present application provides a decoder comprising: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, enable the decoder to execute the method in the fourth aspect or any possible implementation of the fourth aspect.
[0101] The eighth aspect and any implementation of the eighth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the eighth aspect and any implementation of the eighth aspect can be referred to the technical effects corresponding to the fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0102] In a ninth aspect, the present application provides a decoder comprising: a memory and a processor, wherein the memory is coupled to the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the decoder executes the method in the first aspect or any possible implementation of the first aspect.
[0103] The ninth aspect and any implementation of the ninth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the ninth aspect and any implementation of the ninth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0104] In a tenth aspect, the present application provides a decoder comprising: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, enable the decoder to execute the method of the fourth aspect or any possible implementation of the fourth aspect.
[0105] The tenth aspect and any implementation of the tenth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the tenth aspect and any implementation of the tenth aspect can be referred to the technical effects corresponding to the fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0106] In the eleventh aspect, the present application provides a chip comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the first aspect and any one of the implementation methods of the first aspect are executed.
[0107] The eleventh aspect and any implementation of the eleventh aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the eleventh aspect and any implementation of the eleventh aspect can be referred to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, and will not be repeated here.
[0108] In the twelfth aspect, the present application provides a chip comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the fourth aspect and any one of the implementation methods of the fourth aspect are achieved.
[0109] The twelfth aspect and any implementation of the twelfth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the twelfth aspect and any implementation of the twelfth aspect can be referred to the technical effects corresponding to the above-mentioned fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0110] In the thirteenth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, it enables the computer or the processor to execute the method in the first aspect or any possible implementation of the first aspect.
[0111] The thirteenth aspect and any implementation of the thirteenth aspect respectively correspond to the first aspect and any implementation of the first aspect. The technical effects corresponding to the thirteenth aspect and any implementation of the thirteenth aspect can be referred to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, and will not be repeated here.
[0112] In the fourteenth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, it enables the computer or the processor to execute the method in the fourth aspect or any possible implementation of the fourth aspect.
[0113] The fourteenth aspect and any implementation of the fourteenth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the fourteenth aspect and any implementation of the fourteenth aspect can be referred to the technical effects corresponding to the above-mentioned fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0114] In a fifteenth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer or the processor executes the method in the first aspect or any possible implementation of the first aspect.
[0115] The fifteenth aspect and any implementation of the fifteenth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifteenth aspect and any implementation of the fifteenth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0116] In the sixteenth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer or the processor executes the method in the fourth aspect or any possible implementation of the fourth aspect.
[0117] The sixteenth aspect and any implementation of the sixteenth aspect correspond to the fourth aspect and any implementation of the fourth aspect, respectively. The technical effects corresponding to the sixteenth aspect and any implementation of the sixteenth aspect can be referred to the technical effects corresponding to the above-mentioned fourth aspect and any implementation of the fourth aspect, and will not be repeated here.
[0118] In a seventeenth aspect, the present application provides a computer-readable storage medium, which stores the code stream in the above-mentioned third aspect or any one of the implementations of the third aspect.
[0119] The seventeenth aspect and any implementation of the seventeenth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the seventeenth aspect and any implementation of the seventeenth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0120] In an eighteenth aspect, the present application provides a device for storing a code stream, the device comprising: a receiver and at least one storage medium, the receiver being used to receive the code stream; and the at least one storage medium being used to store the code stream in the above-mentioned third aspect or any one of the implementations of the third aspect.
[0121] The eighteenth aspect and any implementation of the eighteenth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the eighteenth aspect and any implementation of the eighteenth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0122] In the nineteenth aspect, the present application provides a device for transmitting a code stream, the device comprising: a transmitter and at least one storage medium, the at least one storage medium being used to store the code stream in the above-mentioned third aspect or any one of the implementation methods of the third aspect; the transmitter being used to obtain the code stream from the storage medium and send the code stream to the end-side device through the transmission medium.
[0123] The nineteenth aspect and any implementation of the nineteenth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the nineteenth aspect and any implementation of the nineteenth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0124] In the twentieth aspect, the present application provides a system for distributing code streams, the system comprising: at least one storage medium for storing the code stream in the third aspect or any one of the implementations of the third aspect, a streaming media device for obtaining a target code stream from the at least one storage medium and sending the target code stream to an end-side device, wherein the streaming media device comprises a content server or a content distribution server.
[0125] The twentieth aspect and any implementation of the twentieth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the twentieth aspect and any implementation of the twentieth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0126] FIG1A is a schematic diagram illustrating an exemplary application scenario;
[0127] FIG1B is a schematic diagram illustrating an end-to-end process of HDR video;
[0128] FIG1C is a schematic diagram illustrating an exemplary encoding process of an image and its metadata;
[0129] FIG1D is a schematic diagram illustrating an exemplary decoding process of an image and its metadata;
[0130] FIG2 is a schematic diagram of an example metadata encoding process 200;
[0131] FIG3 is a schematic diagram illustrating an exemplary metadata decoding process 300;
[0132] FIG4A is a schematic diagram illustrating an exemplary metadata encoding process 400;
[0133] FIG4B is a schematic diagram illustrating an exemplary floating-point data format;
[0134] FIG4C is a schematic diagram illustrating an exemplary floating-point data format;
[0135] FIG5A is a schematic diagram illustrating an exemplary code stream structure;
[0136] FIG5B is a schematic diagram illustrating an exemplary code stream structure;
[0137] FIG5C is a schematic diagram illustrating an exemplary code stream structure;
[0138] FIG5D is a schematic diagram illustrating an exemplary code stream structure;
[0139] FIG6 is a schematic diagram illustrating an exemplary metadata decoding process 600;
[0140] FIG. 7 is a schematic diagram illustrating an exemplary encoding process 700 of an image and its metadata;
[0141] FIG8 is a schematic diagram illustrating an exemplary decoding process 800 of an image and its metadata;
[0142] FIG9 is a schematic diagram illustrating an exemplary decoding process 900 of an image and its metadata;
[0143] FIG10 is a schematic diagram illustrating an exemplary encoding device 1000;
[0144] FIG11 is a schematic diagram illustrating an exemplary decoding device 1100;
[0145] FIG12 is a schematic block diagram of a coding and decoding system used in an embodiment of the present application;
[0146] FIG13A is a block diagram of a content providing system for implementing a content distribution service according to an embodiment of the present application;
[0147] FIG13B is a schematic diagram of an example structure of the terminal device 2106 in FIG13A ;
[0148] FIG14A is a schematic diagram of a workflow of a streaming media system used in an embodiment of the present application;
[0149] FIG14B is a schematic diagram of a streaming media system architecture used in an embodiment of the present application;
[0150] FIG15A is a schematic diagram of a possible system architecture applicable to an embodiment of the present application;
[0151] FIG15B is a schematic diagram of the structure of an image processing system provided in an embodiment of the present application;
[0152] FIG16 is a schematic structural diagram of an exemplary device. DETAILED DESCRIPTION
[0153] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0154] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0155] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0156] In the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0157] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
[0158] For example, the multimedia involved in this application may include but is not limited to: images, videos, and audios.
[0159] For example, this application can be applied to various image business scenarios (e.g., mobile phone photos, cloud albums, etc.), various video / audio business scenarios (e.g., live broadcast, on-demand, etc.), etc., and this application does not limit this. This application takes multimedia as an example to illustrate.
[0160] FIG1A is a schematic diagram illustrating an exemplary application scenario.
[0161] 1A , exemplarily, the first electronic device may include an image generating module, a metadata generating module, an encoding module (or encoder), and a transmitting module.
[0162] Exemplarily, the encoding module may be a software module or a hardware module, and this application does not impose any limitation on this.
[0163] Exemplarily, the image generation module may include but is not limited to: an image acquisition module, an image editing module, etc., and this application does not impose any restrictions on this.
[0164] It should be understood that FIG1A is only an example of the first electronic device. The first electronic device in other embodiments of the present application may have more or fewer modules than those shown in FIG1A , and the present application does not impose any limitation on this.
[0165] 1A , illustratively, the second electronic device may include a display module, a decoding module (or decoder), and a receiving module. Exemplarily, the decoding module may be a software module or a hardware module, and this application does not limit this. It should be understood that FIG1A is only an example of the second electronic device. In other embodiments of this application, the second electronic device may have more modules than those shown in FIG1A , and this application does not limit this.
[0166] Continuing to refer to Figure 1A, exemplarily, after the image generation module of the first electronic device generates an image, the generated image can be output to the encoding module and the metadata generation module. Subsequently, the metadata generation module can generate metadata for the image and output the metadata of the image to the encoding module. Afterwards, the encoding module can encode the image and the metadata of the image to obtain a code stream (wherein the code stream can also be referred to as a bit stream or bit stream or file), and output the encoded code stream to the sending module. The sending module can then send the code stream to the second electronic device. After receiving the code stream, the receiving module of the second electronic device can output the code stream to the decoding module; then, the decoding module can decode the code stream to obtain a reconstructed image and reconstructed metadata of the image, and output the reconstructed image and the reconstructed metadata of the image to the display module, which displays the reconstructed image according to the reconstructed metadata of the image.
[0167] Exemplarily, the first electronic device includes but is not limited to: a server, a PC (Personal Computer), a laptop, a tablet computer, a mobile phone, and a watch.
[0168] Exemplarily, the second electronic device includes but is not limited to: a PC, a laptop, a tablet computer, a mobile phone, and a watch.
[0169] For example, when the first electronic device is used for encoding and the second electronic device is used for decoding, the first electronic device can be called the encoding end and the second electronic device can be called the decoding end. When the second electronic device is used for encoding and the first electronic device is used for decoding, the second electronic device can be called the encoding end and the first electronic device can be called the decoding end.
[0170] In one possible approach, the image encoded by the first electronic device in FIG1A may be an HDR image. To facilitate understanding, SDR and HDR are first introduced.
[0171] Dynamic range is used in many fields to express the ratio of the maximum to minimum values of a variable. In digital images, dynamic range represents the ratio between the maximum grayscale value and the minimum grayscale value within the image display range. The dynamic range in nature is quite large. The brightness of a night scene under the starry sky is about 0.001cd / m2, and the brightness of the sun itself is as high as 1,000,000,000cd / m2. The dynamic range is 1,000,000,000 / 0.001=10 13 However, in real scenes in nature, the brightness of the sun and the brightness of the stars are not obtained at the same time. For natural scenes in the real world, the dynamic range is 10 -3 to 10 6 In most color digital images, each channel of R, G, and B is stored in an 8-bit byte. That is, the range of each channel is 0 to 255 gray levels. Here, 0 to 255 is the dynamic range of the image. In the real world, the dynamic range of the same scene is between 10 and 255. -3 to 10 6 The range is called high dynamic range (HDR), while the dynamic range of ordinary images is called low dynamic range (LDR). The imaging process of a digital camera is actually a mapping from the high dynamic range of the real world to the low dynamic range of the photograph.
[0172] Standard dynamic range images correspond to high dynamic range images. Traditionally used 8-bit images in formats such as JPEG can be considered standard dynamic range images. Before the advent of cameras capable of capturing HDR images, traditional cameras could only record captured light information within a certain range by controlling the exposure value. Since the maximum illumination information of a display device cannot match the brightness information of the real world, and we browse images through a display device, a photoelectric transfer function is required. Early display devices used cathode ray tube (CRT) displays, whose photoelectric transfer function was the gamma function. This photoelectric transfer function, based on the "gamma" function, is defined in the ITU-R Recommendation BT.1886 standard.
[0173] However, with the upgrade of display devices, the illumination range of display devices continues to increase. Existing consumer-grade HDR displays are at 600cd / m2, while high-end HDR displays can reach 2000cd / m2, far exceeding the illumination information of SDR display devices. The photoelectric conversion function in the ITU-R Recommendation BT.1886 standard cannot well express the display performance of HDR display devices. Therefore, an improved electro-optical transfer function is needed to adapt to the upgrade of display devices. The idea of the photoelectric transfer function is derived from the mapping function in the tone mapping algorithm. The photoelectric transfer function is obtained by making appropriate adjustments to the mapping function. At present, there are three common photoelectric conversion functions: perception quantization (PQ) function, gamma, logarithm (log) function, hybrid log-gamma (HLG) function, etc. These three photoelectric conversion functions are the conversion functions specified in the AVS standard.
[0174] Among them, if HDR images and videos need to obtain a better experience, it is an end-to-end process. Please refer to Figure 1B, which describes the end-to-end process of HDR video.
[0175] Because the dynamic range of an image acquisition module is limited under specific shooting conditions, to obtain an image with a higher dynamic range, multiple images with different exposures captured simultaneously are typically combined to create an image with a higher dynamic range, known as an HDR image. The resulting HDR image typically has a bit width greater than 10 bits, allowing it to accommodate scenes with a higher dynamic range.
[0176] 1B , illustratively, material production may refer to the process of synthesizing an HDR image using multiple images with different exposures captured at the same time; and synthesizing the temporally continuous HDR images into an original HDR video; the original HDR video obtained from the material production may be referred to as material.
[0177] Afterwards, the HDR image / video production module of the image generation module can perform editing / color adjustment and other processing on the material to obtain the final HDR video (that is, the Master work). Subsequently, the metadata generation module can generate dynamic metadata based on the final HDR video. Next, the encoding module can encode the HDR video and dynamic metadata to obtain a code stream and distribute / transmit the code stream through the network. After the decoding end receives the code stream, it can call the decoding module to decode the code stream to obtain the reconstructed HDR video and reconstructed dynamic metadata. Afterwards, the display module can display the reconstructed HDR video based on the reconstructed dynamic metadata.
[0178] However, some codec standards can only support the encoding and decoding of images with a bit width of up to 10 bits, making it impossible to implement the encoding and decoding of HDR images. In addition, Standard Dynamic Range (SDR) transcoding devices cannot encapsulate the metadata of HDR images into the SDR image bitstream, resulting in the terminal device receiving the SDR bitstream being unable to restore the HDR image. In view of this, the present application provides a method for encoding and decoding images and image metadata, which can be referred to in Figures 1C and 1D.
[0179] FIG1C is a schematic diagram illustrating an exemplary encoding process of an image and its metadata.
[0180] Referring to Figure 1C , for example, the encoder can use the SDR image as the base image and the HDR image as the first image; alternatively, the HDR image as the base image and the SDR image as the first image, although this application does not impose any restrictions on this. Next, a difference calculation can be performed on the base image and the first image to obtain enhanced data and metadata. The base image, enhanced data, and metadata can then be encoded to obtain a bitstream.
[0181] FIG1D is a schematic diagram showing an exemplary decoding process of an image and its metadata.
[0182] 1D , exemplarily, after receiving the code stream, the decoding end can decode the code stream to obtain the reconstructed image of the basic image, the reconstruction result of the enhanced data, and the reconstructed metadata. When the basic image is an SDR image and the first image is an HDR image, if the decoding end supports displaying HDR, the reconstructed image of the basic image and the reconstruction result of the enhanced data can be merged according to the reconstructed metadata to obtain the HDR reconstructed image; if the decoding end only supports displaying SDR images, the reconstructed image of the basic image can be displayed. When the basic image is an HDR image and the first image is an SDR image, if the decoding end only supports displaying SDR, the reconstructed image of the basic image and the reconstruction result of the enhanced data can be merged according to the reconstructed metadata to obtain the SDR reconstructed image; if the decoding end supports displaying HDR images, the reconstructed image of the basic image can be displayed.
[0183] In this way, compatibility with formats supported by different systems can be achieved.
[0184] The following describes the metadata encoding and decoding process.
[0185] FIG. 2 is a schematic diagram illustrating an example metadata encoding process 200 .
[0186] S201: Acquire multimedia metadata, where the multimedia metadata includes first data, and the first data is floating-point data.
[0187] In one possible approach, the multimedia metadata may be metadata for the enhancement data. The enhancement data metadata is generated during the process of performing a difference calculation between the base image and the first image to generate the enhancement data, and is data that affects the difference calculation result (i.e., the enhancement data). This will be described in detail in subsequent embodiments.
[0188] In one possible approach, the multimedia metadata may be metadata of an HDR image, which may include dynamic metadata and static metadata, wherein dynamic metadata may refer to data that changes over time, and static metadata may refer to data that does not change over time.
[0189] It should be understood that multimedia metadata may also be metadata of other images, audio, video, etc., and this application does not impose any limitation on this.
[0190] Exemplarily, the metadata of the multimedia may include first data, and the first data may be floating-point data (including 32-bit float type or 64-bit double type); for example, the first data is "0.125490203".
[0191] It should be understood that the metadata of multimedia may also include other data in addition to the first data, and the other data in the metadata of multimedia may be encoded and decoded using existing technology methods, and this application does not impose any restrictions on this; this application describes the encoding and decoding process of floating-point data in multimedia data.
[0192] Exemplarily, the first data may be one or more.
[0193] S202, writing second data into a code stream; wherein the second data is a representation of the first data in a floating-point data format.
[0194] Illustratively, the first data may be processed to obtain a representation of the first data in a floating-point data format (hereinafter referred to as second data).
[0195] Exemplarily, the second data may be a binary representation or a hexadecimal representation of the first data in a floating-point data format; this application does not impose any limitation thereto.
[0196] For example, a floating-point data format is a data format that specifies the fields that constitute floating-point data, the layout of each field, and the arithmetic interpretation. Floating-point data formats may include, but are not limited to, IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, and TensorFloat-32. It should be understood that other floating-point data formats may also be included, and this application does not limit this.
[0197] For example, when the first data is “0.125490203” and the bit width is 32, the second data may be a hexadecimal representation of the first data in IEEE 754 format, that is, “0X3E008081”.
[0198] For example, when the first data is “0.35” and the bit width is 32, the second data may be a binary expression of the first data in IEEE 754 format, that is, “0100000001010000000000000000000000”.
[0199] Exemplarily, the second data may be one or more, wherein one second data may be generated based on one first data; that is, the multiple first data and the multiple second data are in one-to-one correspondence.
[0200] Illustratively, after obtaining the second data, the second data can be written into the code stream. For example, the second data can be written into the code stream in the form of an ASCII string; or, for another example, the second data can be written into the code stream in the form of an ASCII byte stream. It should be understood that this application does not limit the form in which the second data is written into the code stream.
[0201] FIG3 is a schematic diagram illustrating an exemplary metadata decoding process 300. The metadata decoding process 300 in FIG3 corresponds to the metadata encoding process 200 in FIG2.
[0202] S301, receiving a code stream.
[0203] Exemplarily, after the encoding end sends the code stream to the decoding end, the decoding end may receive the code stream.
[0204] S302: parse the code stream to obtain second data, where the second data is a representation of the first data in a floating-point data format, and the first data is floating-point data in multimedia metadata.
[0205] Exemplarily, the code stream received by the decoding end may include the second data, and further, the second data may be parsed from the code stream.
[0206] S303: Generate first data based on the second data.
[0207] For example, after obtaining the second data, the second data may be processed to convert the second data into floating point data, thereby obtaining the first data. For example, the first data obtained in S303 may be part or all of the metadata reconstructed in FIG1C .
[0208] For example, the second data is "0X3E008081", and the first data generated based on the second data may be "0.125490203".
[0209] For another example, the second data is “01000000010100000000000000000000000”, and the first data generated based on the second data may be “0.35”.
[0210] Since codewords of an agreed length are typically transmitted in the code stream for specific syntax elements, and since floating-point numbers are typically highly precise, they are typically fixed-pointed or converted into other formats (such as ASCII strings, which involve length truncation during the conversion process) during the encapsulation of syntax elements in existing standards, resulting in a loss of precision. However, the present application uses the form of floating-point number storage in computer memory (i.e., the floating-point data format of the second data) for transmission in the code stream, ensuring that the floating-point number values stored in the memory of the sending and receiving ends are completely consistent. In this way, the floating-point data obtained by the decoding end through the second data conversion is the same as the first data used for encoding by the encoding end, thus ensuring that the floating-point number ultimately obtained by the decoding end is consistent with the original floating-point number of the encoding end.
[0211] Secondly, since the floating-point data in the metadata of multimedia on the codec side is consistent, in scenarios where it is necessary to generate and reconstruct multimedia based on the multimedia metadata or to display and reconstruct multimedia content based on the multimedia metadata, it can improve the quality of the reconstructed multimedia content or the visual experience when displaying the reconstructed multimedia content.
[0212] The following is a detailed description of the metadata encoding and decoding process.
[0213] FIG4A is a schematic diagram illustrating an exemplary metadata encoding process 400 .
[0214] S401: Acquire multimedia metadata, where the multimedia metadata includes first data, and the first data is floating-point data.
[0215] Exemplarily, each first data in the metadata of the multimedia may be encoded according to steps S402 to S407 .
[0216] S402, obtaining floating point data format.
[0217] In one possible manner, the floating-point data format may be preset; thus, the preset floating-point data format may be acquired.
[0218] In one possible manner, the floating-point data format may be determined based on an application scenario, channel conditions, the number of significant bits of the first data, and the like.
[0219] S403: Obtain the bit width of the second data.
[0220] For example, bit width may refer to the number of bits occupied by data in a computer system. The bit width may include at least one of the following: 8 bits, 16 bits, 19 bits, 32 bits, or 64 bits, etc., and this application does not impose any limitation on this.
[0221] In one possible approach, the bit width may be preset; in this way, the preset bit width may be obtained.
[0222] In one possible approach, the bit width may be determined based on an application scenario, channel conditions, and the number of valid bits of the first data.
[0223] For example, the combination of floating-point data format and bit width may be as shown in Table 1 below:
[0224] Table 1
[0225] It should be noted that this application does not limit the execution order of S402 and S403.
[0226] S404: Process the first data according to the floating-point data format and bit width to obtain second data.
[0227] For example, the first data may be converted according to a floating-point data format and a bit width to obtain the second data. Specifically, the length of each field of the floating-point data specified by the floating-point data format for the bit width may be determined; then, the first data may be converted to the second data according to the floating-point data format's specifications for each field of the floating-point data, the length of each field, the layout and arithmetic interpretation of each field, and the decimal-to-binary conversion method.
[0228] For example, the floating-point data format is IEEE 754, with a bit width of 16 bits. Correspondingly, the floating-point data format can be shown in FIG4B(1). From left to right, the first bit is used to represent the sign of the floating-point data, the second to sixth bits are used to represent the exponent of the floating-point data, and the seventh to sixteenth bits are used to represent the fraction of the floating-point data. According to the method of converting decimal to binary, the first data "3.25" is converted into the second data "0100000001010000", which can be shown in FIG4B(2).
[0229] For example, the floating-point data format is IEEE 754, and the bit width is 32 bits. Correspondingly, the floating-point data format can be shown in FIG4C(1). From left to right, the first bit is used to represent the sign of the floating-point data, the second to ninth bits are used to represent the exponent of the floating-point data, and the ninth to thirty-second bits are used to represent the fraction of the floating-point data. According to the method of converting decimal to binary, the first data "3.25" is converted into the second data "01000000010100000000000000000000", which can be shown in FIG4C(2).
[0230] It should be noted that, in one possible approach, after processing the first data according to the floating-point data format and bit width to obtain binary data, the binary data can be directly used as the second data. In another possible approach, after processing the first data according to the floating-point data format and bit width to obtain binary data, the binary data can be converted into hexadecimal data; and the hexadecimal data can be used as the second data.
[0231] For example, other floating-point data formats may refer to the description in the corresponding standards, which will not be described in detail here.
[0232] S405, writing the floating point data format into the code stream.
[0233] Exemplarily, when the floating-point data format is determined based on the application scenario, channel conditions, and the number of valid bits of the first data, etc., the floating-point data format can be written into the code stream; so that the decoding end can obtain the floating-point data format and convert the second data into the first data.
[0234] For example, when the floating-point data format is pre-set and the codec has pre-synchronized the floating-point data format, it is not necessary to write the floating-point data format into the bitstream. For example, when the floating-point data format is pre-set but the codec has not pre-synchronized the floating-point data format, it is also possible to write the floating-point data format into the bitstream.
[0235] S406: Write the bit width of the second data into the code stream.
[0236] Exemplarily, when the bit width is determined based on the application scenario, channel conditions, and the effective number of bits of the first data, etc., the bit width of the second data can be written into the code stream; so that the decoding end can obtain the bit width of the second data, and then accurately divide a single second data from the multiple second data parsed from the code stream, and realize the conversion of the second data into the first data.
[0237] Exemplarily, when the bit width is pre-set and the codec has pre-synchronized the bit width, the bit width of the second data may not need to be written into the bitstream. Exemplarily, when the bit width is pre-set but the codec has not pre-synchronized the bit width, the bit width of the second data may also be written into the bitstream.
[0238] It should be noted that this application does not limit the position of the bit width and floating-point data format in the code stream.
[0239] S407: Write the second data into the code stream.
[0240] For example, when the multimedia is an image or video, the second data may be encapsulated in an image encapsulation format to obtain a code stream. The image encapsulation format includes, but is not limited to, JPEG JFIF encapsulation format, JPEG Universal Metadata Box Format JUMBF encapsulation format, High Efficiency Image File Format (HEIF), MP4 file format, and the like, and this application does not impose any restrictions on this.
[0241] For example, the second data may be encapsulated in accordance with the JPEG JFIF encapsulation format. The second data may be encapsulated into the APPN field of the JPEG JFIF encapsulation format, as shown in Table 2:
[0242] Table 2
[0243] The information shown in Table 2 may be interpreted with reference to the description in ISO / IEC FDIS10918-4:2023 APPn markers (JPEG 1 part 4), which will not be repeated here.
[0244] For example, the second data is encapsulated according to the JUMBF encapsulation format.
[0245] In one possible manner, the second data may be encapsulated into the payload of the APP11 JP in the JUMBF encapsulation format, as shown in Table 3 below:
[0246] Table 3
[0247] The information shown in Table 3 may be explained with reference to the description in the JPEG general metadata box format standard, which will not be repeated here.
[0248] In one possible manner, the second data may be encapsulated in a box in the JUMBF encapsulation format.
[0249] For example, the second data may be encapsulated according to HEIF. For example, the second data may be encapsulated into a HEIF package structure. For another example, the second data may be encapsulated into 'tmap' and 'idat' of HEIF; the description of 'tmap' and 'idat' may refer to the description in ISOIEC 23008-12AMD 3 and will not be repeated here.
[0250] Exemplarily, when the multimedia is an image or video, the second data can be encapsulated into the supplemental enhancement information (SEI) of HEVC or VVC, or a user-defined network abstraction layer (NAL) unit or a reserved NAL unit.
[0251] It should be understood that the second data can also be encapsulated into the file location indicated by the metadata location information (the metadata location information can be a type of multimedia metadata), such as the position after the EOI (end of image) of the complete JPEG file, or the end of the complete HEIF file, etc. This application does not impose any restrictions on this.
[0252] It should be understood that the present application does not impose any restrictions on the file location of the second data in the code stream.
[0253] For example, the present application does not limit the metadata structure of multimedia. For example, it may be the metadata structure corresponding to 'tmap' in 23008-12AMD 3, as shown below:
[0254] Among them, SignedRational represents the metadata type, ain_map_min, ... alternate_offset are metadata, and part or all of these metadata are first data.
[0255] For another example, the metadata structure of multimedia can be described in XML or XML format, as shown below:
[0256] Among them, PerComponentMinGainmapValues, ...AlternateHDRheadroom are metadata, and part or all of these metadata are first data.
[0257] In addition, when the second data is written into the code stream in the form of an ASCII code byte stream, the second data can be written into the code stream in the form of a high bit first and a low bit last; or the second data can be written into the code stream in the form of a low bit first and a high bit last. This application does not impose any restrictions on this.
[0258] It should be noted that the present application does not limit the execution order of S405, S406, and S407. For example, the execution can be performed in the order of S405 → S406 → S407; correspondingly, the bit width in the code stream is located after the floating-point data format, and the second data is located after the bit width.
[0259] Figure 5A is a schematic diagram of an exemplary code stream structure. As shown in Figure 5A, in one possible manner, the code stream of the present application may include second data.
[0260] Figure 5B is a schematic diagram of an exemplary code stream structure. As shown in Figure 5B, in one possible manner, the code stream of the present application may include the second data and floating-point data format.
[0261] Figure 5C is a schematic diagram of an exemplary code stream structure. As shown in Figure 5C, in one possible manner, the code stream of the present application may include the second data and the bit width.
[0262] Figure 5D is a schematic diagram of an exemplary code stream structure. As shown in Figure 5D, in one possible manner, the code stream of the present application may include second data, a floating point data format, and a bit width.
[0263] It should be noted that this application does not limit the floating-point data format, bit width, and order / arrangement position of the second data in the code stream.
[0264] Exemplarily, the present application provides an apparatus for transmitting a code stream, which may include: a transmitter and at least one storage medium, wherein the at least one storage medium is used to store the code stream generated by the above-mentioned metadata encoding process 200 or the metadata encoding process 400; the transmitter is used to obtain the code stream from the storage medium and send the code stream to the end-side device via the transmission medium.
[0265] Exemplarily, an encoding end or a device transmitting a bitstream can send the bitstream to a system distributing the bitstream. The system distributing the bitstream can include: at least one storage medium and a streaming device; the storage medium is used to store the bitstream generated by metadata encoding process 200 or metadata encoding process 400; and the streaming device is used to obtain a target bitstream from the at least one storage medium and send the target bitstream to a device on the end side, wherein the streaming device includes a content server or a content distribution server.
[0266] Illustratively, the present application further provides a device for storing a codestream. The device for storing a codestream may include: a receiver for receiving the codestream; and at least one storage medium for storing the codestream. The codestream may be generated by metadata encoding process 200 or metadata encoding process 400.
[0267] Fig. 6 is a schematic diagram illustrating an exemplary metadata decoding process 600. The metadata decoding process 600 in Fig. 6 corresponds to the metadata encoding process 400 in Fig. 4A.
[0268] S601, receiving a code stream.
[0269] S602: parse the code stream to obtain the second data, the floating-point data format, and the bit width of the second data.
[0270] Illustratively, when the encoding end writes the floating-point data format and the bit width of the second data into the code stream, the decoding end can parse the second data, the floating-point data format, and the bit width from the code stream.
[0271] Exemplarily, the second data may be parsed from an image encapsulation format in a code stream.
[0272] For example, the second data may be parsed from the JPEG JFIF encapsulation format, such as parsing the second data from the APPN field of the JPEG JFIF encapsulation format.
[0273] For another example, the second data may be parsed from a JUMBF encapsulation format. For example, the second data may be parsed from a payload of an APP11 JP in a JUMBF encapsulation format. For another example, the second data may be parsed from a box in a JUMBF encapsulation format.
[0274] For another example, the second data is parsed from the SEI of HEVC or VVC or the NAL unit defined by the user or the reserved NAL unit.
[0275] It should be understood that metadata location information can also be parsed from the code stream (the metadata location information can be a type of multimedia metadata, used to indicate the file location of the multimedia metadata), and then the second data can be parsed from the file position indicated by the metadata location information in the code stream (for example, the position after the EOI (end of image) of the complete JPEG file, or the end of the complete HEIF file, etc., which is not limited in this application).
[0276] It should be understood that the decoding end can also parse the second data from other file locations in the code stream, and this application does not limit this.
[0277] Exemplarily, when the decoding end fails to parse out the floating-point data format from the code stream, a preset floating-point data format may be obtained.
[0278] Exemplarily, when the decoding end fails to parse the bit width from the code stream, the preset bit width may be obtained.
[0279] Exemplarily, multiple second data can be parsed from the code stream, and these multiple second data are spliced together; further, the multiple spliced together second data can be divided according to the floating-point data format, bit width and metadata structure to obtain single independent second data.
[0280] S603: Process the second data according to the floating-point data format and the bit width to obtain first data.
[0281] Exemplarily, each second data can be converted into floating-point data according to the floating-point data format and bit width to obtain a first data. Specifically, the process of converting the representation expressed in the floating-point data format into the floating-point data can refer to the description of the prior art and will not be repeated here.
[0282] Based on the above metadata encoding and decoding process, the following describes the generation process of the base image, enhanced data, and metadata of the enhanced data, as well as the encoding and decoding process.
[0283] Fig. 7 is a schematic diagram of an exemplary encoding process 700 of an image and its metadata. In the encoding process 700, the metadata of the image is metadata of the enhancement data.
[0284] S701, obtaining a basic image.
[0285] In one possible approach, an SDR image may be obtained and used as a basic image.
[0286] In one possible approach, an HDR image may be acquired and used as a basic image.
[0287] It should be noted that the foreground of the SDR image is the same as the foreground of the HDR image, and the background of the SDR image is the same as the background of the HDR image. In other words, the SDR image and the HDR image contain the same content (including objects, scenes, etc.).
[0288] For example, the HDR image may be obtained as follows:
[0289] For example, a camera can capture multiple SDR images with different exposure values, and then use these multiple SDR images with different exposure values to synthesize an HDR image. The camera's photosensitive element has different photoelectric characteristics at different positions, which allows it to capture multiple SDR images with different exposure values.
[0290] For example, HDR images can be captured using professional acquisition equipment; the HDR images can then be transmitted to the encoding end by sending or copying.
[0291] It should be noted that this application does not limit the method of obtaining HDR images.
[0292] It should be noted that this application does not limit the format of HDR images. For example, in terms of color space, HDR images can be YUV images or RGB images. In terms of data bit width, the bit width of HDR images can be 8 bits, 10 bits, or 12 bits, etc. The photoelectric conversion function corresponding to the HDR image can be a perception quantization (PQ) function, gamma, logarithm (log) function, hybrid log-gamma (HLG) function, etc.
[0293] Exemplarily, the process of acquiring an SDR image may be as follows:
[0294] For example, an SDR image with different light and dark characteristics from an HDR image is captured by a camera, or an SDR image in a format that adapts to the display characteristics of many existing devices is captured by a camera.
[0295] For example, SDR images are captured by professional acquisition equipment, and then transmitted to the encoding end by sending or copying.
[0296] Another example is generating an SDR image from an HDR image. For example, the SDR image can be obtained by tone mapping the HDR image. Another example is inputting the HDR image into a neural network to obtain the SDR image.
[0297] It should be noted that this application does not limit the method of obtaining SDR images.
[0298] It should be noted that this application does not limit the format of SDR images. For example, in terms of color space, SDR images can be YUV images or RGB images. In terms of data bit width, the bit width of SDR images can be 8 bits, 10 bits, or 12 bits, etc. The photoelectric conversion function of SDR images can be PQ function, gamma function, log function, HLG function, etc.
[0299] S702: Acquire a first image.
[0300] Exemplarily, when an SDR image is used as a basic image, an HDR image may be used as the first image.
[0301] Exemplarily, when the HDR image is used as the base image, the SDR image may be used as the first image.
[0302] S703: Acquire enhanced data and metadata of the enhanced data according to the basic image and the first image.
[0303] Illustratively, S703 may include S7031 to S7034:
[0304] S7031: Obtain a first intermediate image based on the basic image.
[0305] In one possible approach, the basic image may be directly used as the first intermediate image.
[0306] In one possible approach, a base image can be encoded to obtain a base image code stream. The base image code stream is then decoded to obtain a reconstructed image of the base image; the reconstructed image of the base image is used as the first intermediate image. The base image can be encoded and decoded using codecs such as JPEG, HEIF, H.264, and HEVC, but this application does not limit this.
[0307] S7032: Generate a second intermediate image based on the first intermediate image.
[0308] In one possible approach, the first intermediate image data may be used as the second intermediate image.
[0309] In one possible approach, a local or global first mapping relationship between the first intermediate image and the HDR image may be acquired; and then the first intermediate image may be mapped using the first mapping relationship to obtain a second intermediate image.
[0310] Exemplarily, multiple characteristic brightness values of the HDR image can be obtained; for example, a histogram of the HDR image is obtained, and the brightness of the peak position in the histogram is obtained as the characteristic brightness value. For another example, the average brightness value of a specific area (such as a human face, green plants, and blue sky) in the HDR image is obtained as the characteristic brightness value. Then, the pixel position of the characteristic brightness value in the HDR image is determined, and the brightness value of the pixel position in the basic image is determined. Thereafter, a first mapping relationship is determined based on the characteristic brightness value and the brightness value of the pixel position in the basic image.
[0311] For example, the first mapping relationship may include multiple forms, such as Sigmoid, cubic spline, gamma, straight line, etc. or their inverse function forms, which are not limited in this application. For example, the first mapping relationship may be represented by the following curve:
[0312] Wherein, L and L' in the formula may be normalized optical signals or electrical signals, a and b are scaling coefficients, n and m are power exponents, and p is a control parameter, which is not limited in this application.
[0313] For another example, the first mapping relationship may be an inverse curve of the above curve.
[0314] For example, the first intermediate image may be mapped using the first mapping relationship (which may be represented by TMB1()) as follows: baseAfter[i]=TMB1(base[i]), or baseAfter[i]=TMB1()*base[i]. Where base[i] is the pixel value of the i-th pixel in the first intermediate image, and baseAfter[i] is the pixel value of the i-th pixel in the second intermediate image.
[0315] It should be noted that in this application, baseAfter[i] and base[i] can be any domain, such as linear domain, PQ domain, or log domain. In addition, this application does not limit the color space of baseAfter[i] and base[i], which can be YUV, RGB, Lab, HSV, CIEXYZ, and other color spaces.
[0316] It should be noted that in one possible approach, gamut mapping can be performed on the first intermediate image, converting it from the current color gamut to the target color gamut. Subsequently, the first mapping relationship is used to map the first intermediate image after gamut mapping to obtain a second intermediate image. In another possible approach, after mapping the first intermediate image using the first mapping relationship to obtain the second intermediate image, gamut mapping can be performed on the second intermediate image, converting it from the current color gamut to the target color gamut. The current color gamut and the target color gamut include, but are not limited to, BT.2020, BT.709, DCI-P3, sRGB, etc.
[0317] Exemplarily, the parameters of the first mapping relationship may be used as metadata of the enhanced data. The parameters of the first mapping relationship may be the first data.
[0318] S7033: Generate a third intermediate image based on the first image and the second intermediate image.
[0319] In one possible approach, the third intermediate image = g(first image) / f(baseAfter[i]), where g() and f() are conversion functions in the numerical domain. For example, g() and f() can be log, an optical-electro transfer function (OETF), an electro-optical transfer function (EOTF), a linear function, or a piecewise curve. Exemplarily, g() and f() may also include processing of the second intermediate image; for example, obtaining the data range (such as the minimum and / or maximum value) of the second intermediate image, mapping the minimum value of the second intermediate image to 0 or a preset value, and / or mapping the maximum value of the second intermediate image to 1.0 or a preset value, and / or mapping the intermediate value of the second intermediate image to a certain intermediate value according to a second mapping relationship (the second mapping relationship can be determined based on the maximum and / or minimum value of the second intermediate image).
[0320] In one possible approach, the third intermediate image = g(first image) - f(baseAfter[i]), where g() and f() are conversion functions of the numerical domain. For example, g() and f() can be log, OETF, EOTF, linear functions, or piecewise curves, etc. Exemplarily, g() and f() may also include processing of the second intermediate image; for example, obtaining the data range (such as the minimum value and / or maximum value) of the second intermediate image, mapping the minimum value of the second intermediate image to 0 or a preset value, and / or mapping the maximum value of the second intermediate image to 1.0 or a preset value, and / or mapping the intermediate value of the second intermediate image to a certain intermediate value according to a second mapping relationship (the second mapping relationship can be determined based on the maximum value and / or minimum value of the second intermediate image).
[0321] Exemplarily, the parameters of the second mapping relationship may be used as metadata of the enhanced data. The parameters of the second mapping relationship may be the first data.
[0322] S7034: Process the third intermediate image to obtain enhanced data and metadata of the enhanced data.
[0323] Optionally, the third intermediate image may be downsampled to obtain a fourth intermediate image (which may be represented by enchence).
[0324] In one possible approach, the data range (e.g., minimum and maximum values) of the fourth intermediate image can be obtained, and the minimum value of the fourth intermediate image can be mapped to 0 or a preset value, and / or the maximum value of the fourth intermediate image can be mapped to 1.0 or a preset value; and / or the intermediate value of the fourth intermediate image can be mapped to an intermediate value based on a third mapping relationship (the third mapping relationship can be determined based on the maximum and / or minimum values of the fourth intermediate image) to obtain enhanced data. The maximum, minimum, and intermediate values of the fourth intermediate image can all serve as metadata, and the maximum, minimum, and intermediate values of the fourth intermediate image can also serve as first data.
[0325] Exemplarily, the parameters of the third mapping relationship may be used as metadata of the enhanced data. The parameters of the third mapping relationship may be the first data.
[0326] In one possible approach, a histogram of the fourth intermediate data can be obtained and processed using a histogram equalization method to obtain a fourth mapping relationship TMB2(). The fourth intermediate image can then be mapped according to the fourth mapping relationship to obtain enhanced data. For example, the fourth intermediate image can be mapped as follows: enchenceAfter[i] = TMB2(enchence[i]), or, enchenceAfter[i] = TMB2() * enchence[i]. enchenceAfter[i] is the enhanced data corresponding to the i-th pixel in the base image, and enchence[i] is the i-th pixel in the fourth intermediate image.
[0327] Exemplarily, the parameters of the fourth mapping relationship may be used as metadata of the enhanced data.
[0328] That is, the metadata of the enhanced data may include data that has an impact on a processing result during the process of determining the enhanced data based on the basic image and the first image.
[0329] It should be understood that the present application does not limit the manner of determining the enhanced data and metadata of the enhanced data based on the basic image and the first image.
[0330] Exemplarily, the fourth mapping relationship has various forms, including Sigmoid, cubic spline, gamma, straight line, linear, piecewise curve, etc. or their inverse function forms, which are not limited in this application.
[0331] It should be noted that in this application, enchenceAfter[i] and enchence[i] can be any domain; for example, enchenceAfter[i] and enchence[i] can be linear domain, PQ domain, or log domain. This application does not limit the color space of enchenceAfter[i] and enchence[i], which can be YUV, RGB, Lab, HSV, and other color spaces.
[0332] It should be noted that, in one possible approach, gamut mapping can be performed on the fourth intermediate image, converting the fourth intermediate image from the current color gamut to the target color gamut; then, the gamut-mapped fourth intermediate image is mapped using a fourth mapping relationship to obtain enhanced data. In another possible approach, after the fourth intermediate image is mapped using the fourth mapping relationship to obtain enhanced data, gamut mapping can be performed on the enhanced data, converting the enhanced data from the current color gamut to the target color gamut. The current color gamut and the target color gamut include, but are not limited to, BT.2020, BT.709, DCI-P3, sRGB, etc.
[0333] S704: Encode the basic image.
[0334] Illustratively, the process of encoding the basic image may refer to the image encoding process in the prior art, which will not be described in detail here.
[0335] S705, encode enhanced data.
[0336] For example, one pixel in the basic image may correspond to one enhanced data, or multiple pixels may correspond to one enhanced data, wherein the enhanced data may include multiple values, and the number of values included in the enhanced data corresponding to each pixel is the same.
[0337] Exemplarily, the enhanced data may be encapsulated into the SEI of HEVC or VVC, a user-defined NAL unit or a reserved NAL unit, or may be APP extension information encapsulated in JFIF, or a data segment encapsulated in MP4, and the like.
[0338] Exemplarily, the enhanced data may also be encapsulated into a file location indicated by metadata location information, such as a location after the EOI (end of image) of a complete JPEG file.
[0339] It should be understood that the present application does not limit the position of the enhancement data in the bitstream.
[0340] Exemplarily, the metadata of the enhancement data may be encoded with reference to the following S706 to S711 .
[0341] S706, obtain floating point data format
[0342] S707: Obtain the bit width of the second data.
[0343] S708: Process the first data according to the floating-point data format and bit width to obtain second data.
[0344] S709, writing the floating point data format into the code stream.
[0345] S710: Write the bit width of the second data into the code stream.
[0346] S711, write the second data into the code stream.
[0347] For example, S706 to S711 may refer to the description of S402 to S407 and will not be repeated here.
[0348] It should be noted that this application is not limited to encoders that encode basic images, enhancement data, and metadata of enhancement data, such as JPEG, HEIF, H.264, HEVC, etc.
[0349] It should be noted that this application does not limit the execution order of encoding the basic image, enhancement data, and metadata of the enhancement data.
[0350] It should be noted that the encoded basic image, the encoded enhancement data, and the metadata encoded for the enhancement data can be stored in one code stream / file. The encoded basic image, the encoded enhancement data, and the metadata encoded for the enhancement data can be regarded as three parts of one code stream / file.
[0351] FIG8 is a schematic diagram illustrating an exemplary decoding process 800 of an image and its metadata. The decoding process 800 in FIG8 corresponds to the encoding process 700 in FIG7 .
[0352] S801 , obtaining a reconstructed image of a basic image, a reconstruction result of enhanced data, and encoded metadata from a received code stream.
[0353] Among them, S801 can refer to the description of the following S902 to S905, and will not be repeated here.
[0354] Exemplarily, the encoded metadata may include the second data.
[0355] S802: Process the encoded metadata according to a floating-point data format to obtain reconstructed metadata.
[0356] Exemplarily, the second data may be processed according to a floating-point data format to obtain reconstructed second data, namely, the first data.
[0357] S803 : Merge the reconstructed image of the basic image and the reconstruction result of the enhanced data according to the reconstructed metadata to obtain a reconstructed first image.
[0358] Exemplarily, when the base image is an HDR image and the first image is an SDR image, if the decoding end only supports displaying SDR, the reconstructed image of the base image and the reconstruction result of the enhanced data can be merged according to the reconstructed metadata to obtain an SDR reconstructed image; if the decoding end supports displaying HDR images, the reconstructed image of the base image can be displayed.
[0359] Exemplarily, when the base image is an SDR image and the first image is an HDR image, if the decoding end supports displaying HDR, the reconstructed image of the base image and the reconstruction result of the enhanced data can be merged according to the reconstructed metadata to obtain an HDR reconstructed image; if the decoding end only supports displaying SDR images, the reconstructed image of the base image can be displayed.
[0360] FIG9 is a schematic diagram illustrating an exemplary decoding process 900 of an image and its metadata. The decoding process 900 in FIG9 is a detailed description of the decoding process 800 in FIG8 .
[0361] S901: Receive code stream.
[0362] S902: Decode the basic image from the code stream to obtain a reconstructed image.
[0363] Exemplarily, any decoder (such as HEVC, JPEG, ProRes, HEIF) can be used to decode the received code stream to obtain a reconstructed image of the basic image.
[0364] It should be noted that this application does not limit the format of the reconstructed image of the basic image. In terms of color space, the reconstructed image of the basic image can be a YUV image or an RGB image. In terms of data bit width, the reconstructed image of the basic image can have a bit width of 8 bits, 10 bits, 12 bits, etc.
[0365] S903: Decode the enhanced data from the bitstream to obtain a reconstruction result.
[0366] Exemplarily, the reconstruction result of the enhanced data can be decoded from the SEI of HEVC or VVC, the user-defined NAL unit or the reserved NAL unit, the APP extension information encapsulated in JFIF, or the data segment encapsulated in MP4.
[0367] Exemplarily, the reconstruction result of the enhanced data may also be decoded from the file position indicated by the metadata position information, such as the position after the EOI (end of image) of the complete JPEG file.
[0368] It should be understood that the reconstruction result of the enhanced data can also be decoded from other file locations, and this application does not limit this.
[0369] S904: parse the code stream to obtain the second data, the floating-point data format, and the bit width of the second data.
[0370] S905 , processing the second data according to the floating-point data format and the bit width of the second data to obtain first data.
[0371] For example, S904 to S905 may refer to the description of S502 to S503 above, which will not be repeated here.
[0372] For example, the present application does not limit the execution order of S902 , S903 , and S904 .
[0373] It should be understood that other reconstructed data (other reconstructed data is the reconstruction result of data other than the first data in the metadata of the enhanced data) can also be decoded from the code stream, and the reconstructed metadata of the enhanced data can be obtained.
[0374] S906 , merging the reconstructed image of the basic image and the reconstruction result of the enhanced data according to the reconstruction metadata of the enhanced data to obtain a reconstructed first image.
[0375] For example, the reconstructed image of the base image and the reconstruction result of the enhanced data may be merged according to the reconstruction metadata of the enhanced data to obtain an SDR reconstructed image or an HDR reconstructed image. This application is described by taking the base image as an SDR image and the first image as an HDR image as an example. The following S1 to S6 may be included:
[0376] S1, processing a reconstructed image of a basic image according to reconstruction metadata of enhanced data to obtain a first reconstructed image.
[0377] In one possible approach, the reconstructed image of the basic image is used as the first reconstructed image, that is, R_baseAfter[i]=R_Base[i]. Here, R_baseAfter[i] represents the pixel value of the i-th pixel in the first reconstructed image, and R_Base[i] represents the pixel value of the i-th pixel in the reconstructed image of the basic image.
[0378] In one possible approach, a lower limit THL of the base image is obtained from the reconstruction metadata of the enhanced data, and the first reconstructed image is obtained according to the following approach: R_baseAfter[i]=R_Base[i]+THL.
[0379] In one possible approach, parameters of the first mapping relationship and a lower limit THL of the base image are obtained from the reconstruction metadata of the enhanced data; a first mapping relationship TMB1() is obtained based on the parameters of the first mapping relationship. Thereafter, a first reconstructed image is obtained as follows: R_baseAfter[i] = TMB1(R_Base[i] + THL).
[0380] In one possible approach, parameters of the first mapping relationship are obtained from the reconstruction metadata of the enhanced data. Based on the parameters, the first mapping relationships TMB11() and TMB12() are obtained. TMB11() can be a global or local tone mapping function, and TMB12 can be in the form of a linear, spline, or piecewise curve. Subsequently, the first reconstructed image is obtained as follows: R_baseAfter[i] = TMB11(TMB12(R_Base[i])).
[0381] It should be noted that in this application, R_baseAfter[i] and R_Base[i] can be any domain; for example, R_baseAfter[i] and R_Base[i] can be linear domain, PQ domain, or log domain. In addition, this application does not limit the color space of R_baseAfter[i] and R_Base[i], and can be YUV, RGB, Lab, HSV, and other color spaces.
[0382] It should be noted that, in one possible approach, the reconstructed image of the basic image can be subjected to color gamut mapping to convert the reconstructed image of the basic image from the current color gamut to the target color gamut; thereafter, the reconstructed image of the basic image after color gamut mapping is processed according to the reconstruction metadata of the enhanced data to obtain a first reconstructed image. In another possible approach, after the reconstructed image of the basic image is processed according to the reconstruction metadata of the enhanced data to obtain the first reconstructed image, the first reconstructed image can be subjected to color gamut mapping to convert the first reconstructed image from the current color gamut to the target color gamut. The current color gamut and the target color gamut include but are not limited to BT.2020, BT.709, DCI-P3, sRGB, etc.
[0383] S2, transforming the reconstruction result of the enhanced data according to the reconstruction metadata of the enhanced data to obtain a first reconstruction result.
[0384] Exemplarily, the transformation process may be upsampling the reconstruction result of the enhanced data to obtain the first reconstruction result. For example, upsampling may be performed using a preset filter, or using a method including multiple directional interpolation methods, or using a bicubic spline method, which is not limited in this application.
[0385] S2: Process the first reconstruction result according to the reconstruction metadata of the enhanced data to obtain a second reconstruction result.
[0386] Optionally, the first reconstruction result (which may be referred to as R_enchence) may be normalized. Typically, 8-bit data ranges from 0 to 255, and all numbers in R_enchence are converted to a range of 0 to 1.0 by dividing them by 255.0, which is not limited in this application.
[0387] In one possible approach, the first reconstruction result or the normalized first reconstruction result is used as the second reconstruction result, that is, R_enchenceAfter[i] = R_enchence[i]. R_enchenceAfter[i] is the i-th pixel point of the second reconstruction result, and R_enchence[i] is the i-th pixel point of the first reconstruction result / or the normalized first reconstruction result.
[0388] In one possible approach, the parameters of the fourth mapping relationship, the upper limit THL and the lower limit THH of the enhanced data can be obtained from the reconstructed metadata of the enhanced data, and the fourth mapping relationship TMB2() can be obtained according to the parameters of the fourth mapping relationship, R_enchenceAfter[i]=TMB2(R_enchence[i]*THH+(A-R_enchence[i])*THL).
[0389] In one possible approach, the parameters of the fourth mapping relationship, the upper limit THL and the lower limit THH of the enhanced data can be obtained from the reconstructed metadata of the enhanced data, and the fourth mapping relationship TMB2() can be obtained according to the parameters of the fourth mapping relationship, R_enchenceAfter[i]=TMB2(R_enchence[i])*THH+(A-TMB2(R_enchence[i]))*THL.
[0390] It should be noted that in this application, R_enchenceAfter[i] and R_enchence[i] can be any domain, for example, R_enchenceAfter[i] and R_enchence[i] can be linear domain or PQ domain log domain, etc. This application does not limit the color space of R_enchenceAfter[i] and R_enchence[i], which can be YUV, RGB, Lab, HSV and other color spaces.
[0391] It should be noted that, in one possible approach, the first reconstruction result can be subjected to color gamut mapping to convert the first reconstruction result from the current color gamut to the target color gamut; thereafter, the first reconstruction result after color gamut mapping can be mapped using a fourth mapping relationship to obtain a second reconstruction result. In another possible approach, after the first reconstruction result is mapped using the fourth mapping relationship to obtain a second reconstruction result, the second reconstruction result can be subjected to color gamut mapping to convert the second reconstruction result from the current color gamut to the target color gamut. The current color gamut and the target color gamut include but are not limited to BT.2020, BT.709, DCI-P3, sRGB, etc.
[0392] Afterwards, the second reconstructed image recHDR can be obtained based on R_baseAfter and R_enhenceAfter: recHDR[i] = g(baseAfter[i])*f(R_enchenceAfter[i]) or recHDR[i] = R_baseAfter[i] + f(R_enchenceAfter[i]), where g() and f() are conversion functions in the numerical domain. recHDR[i] can be any component of RGB or YUV components, and f(R_enchenceAfter[i]) is the gain of any component obtained by reconstructing the enhanced data.
[0393] It should be noted that the present invention does not limit whether other image processing is performed on the second reconstructed image recHDR after it is obtained and before it is displayed.
[0394] S4, generating an HDR reconstructed image recHDRAfter according to the reconstruction metadata of the enhanced data.
[0395] In one possible approach, the second reconstructed image may be used as the HDR reconstructed image, that is, recHDRAfter[i]=recHDR[i].
[0396] In one possible approach, the lower limit THL of the second reconstructed image may be obtained from the reconstruction metadata of the enhanced data, and the HDR reconstructed image may be generated as follows: recHDRAfter[i]=recHDR[i]+THL.
[0397] In one possible approach, the upper limit THH of the second reconstructed image may be obtained from the reconstruction metadata of the enhanced data, and the HDR reconstructed image may be generated as follows: recHDRAfter[i]=TMB1(R_base[i]).
[0398] It should be noted that the present invention does not limit the form of f(x) and TMB1(), that is, recHDRAfter[i] and recHDR[i] can be in any domain, for example, recHDRAfter[i] and recHDR[i] can be in the linear domain or the PQ domain or the log domain. The present invention also does not limit the color space of recHDRAfter[i] and recHDR[i], which can be YUV, RGB, Lab, HSV and other color spaces.
[0399] Due to the encoding and decoding method of the reconstructed metadata of the enhanced data in the present application, there will be no loss of accuracy of the reconstructed metadata of the enhanced data, and the reconstructed metadata of the enhanced data obtained by the decoding end is the same as the metadata of the enhanced data encoded by the encoding end, thereby ensuring the quality of the SDR image or HDR image obtained by the decoding end.
[0400] 10 is a schematic diagram of an exemplary encoding device 1000. The encoding device 1000 can be used to execute the method of the aforementioned embodiment. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method provided above, and will not be repeated here.
[0401] 10 , illustratively, the encoding apparatus 1000 may include:
[0402] The metadata acquisition module 1001 is used to acquire multimedia metadata; wherein the multimedia metadata includes first data, and the first data is floating-point data;
[0403] The data writing module 1002 is configured to write second data into the code stream; wherein the second data is a representation of the first data in a floating-point data format.
[0404] Exemplarily, the encoding device 1000 further includes:
[0405] The first data generating module is used to generate second data based on the first data.
[0406] Exemplarily, the encoding device 1000 further includes:
[0407] A first format acquisition module is used to acquire the floating point data format;
[0408] The first data generating module is specifically configured to process the first data according to a floating point data format to obtain the second data.
[0409] Exemplarily, the encoding device 1000 further includes:
[0410] a first bit width acquisition module, configured to acquire a bit width of second data;
[0411] The first data generating module is specifically configured to process the first data according to the floating point data format and the bit width of the second data to obtain the second data.
[0412] Exemplarily, the data writing module 1002 is further configured to write the bit width into the code stream.
[0413] Exemplarily, the data writing module 1002 is further configured to write floating-point data format into the code stream.
[0414] Exemplarily, the second data is a binary representation or a hexadecimal representation of the first data in a floating point data format.
[0415] Exemplarily, the metadata of the multimedia is metadata of a high dynamic range HDR image.
[0416] Exemplarily, the metadata of the multimedia is metadata of enhanced data, and the enhanced data is generated based on a standard dynamic range SDR image and an HDR image.
[0417] Exemplarily, the floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, and TensorFloat-32.
[0418] 11 is a schematic diagram of an exemplary decoding device 1100. The decoding device 1100 can be used to execute the method of the aforementioned embodiment, and therefore, the beneficial effects achieved by the decoding device 1100 can refer to the beneficial effects of the corresponding method provided above, which will not be repeated here.
[0419] 11 , illustratively, a decoding device 1100 may include:
[0420] The code stream receiving module 1101 is used to receive the code stream;
[0421] A parsing module 1102 is configured to parse the code stream to obtain second data, where the second data is a representation of the first data in a floating-point data format, and the first data is floating-point data in multimedia metadata;
[0422] The second data generating module 1103 is configured to generate first data based on the second data.
[0423] Exemplarily, the decoding device 1100 further includes:
[0424] The second format acquisition module is used to obtain the floating point data format;
[0425] The second data generating module 1103 is specifically configured to process the second data according to a floating point data format to obtain the first data.
[0426] Exemplarily, the second format acquisition module includes: a parsing module 1102; the parsing module 1102 is further configured to parse the floating point data format from the code stream.
[0427] Exemplarily, the second bit width acquisition module is configured to acquire the bit width of the second data;
[0428] The second data generating module 1103 is specifically configured to process the second data according to the floating point data format and the bit width to obtain the first data.
[0429] Exemplarily, the second bit width acquisition module includes: a parsing module 1102; the parsing module 1102 is further configured to parse the bit width from the code stream.
[0430] Exemplarily, the second data is a binary representation or a hexadecimal representation of the first data in a floating point data format.
[0431] Exemplarily, the metadata of the multimedia is metadata of a high dynamic range HDR image.
[0432] Exemplarily, the metadata of the multimedia is metadata of enhanced data, and the enhanced data is determined according to a standard dynamic range SDR image and a high dynamic range HDR image.
[0433] Exemplarily, the floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, and TensorFloat-32.
[0434] The following describes the codec system used in the present application in conjunction with FIG12 . FIG12 is a schematic block diagram of a codec system used in an embodiment of the present application, such as a video codec system 10 (or simply codec system 10) that can utilize the techniques of the present application. The video encoder 20 (or simply encoder 20) and video decoder 30 (or simply decoder 30) in the video codec system 10 represent devices that can be used to perform various techniques according to the various examples described in this application.
[0435] As shown in FIG. 12 , the encoding and decoding system 10 includes a source device 12 , which is configured to provide encoded data 21 such as an encoded image to a destination device 14 for decoding the encoded data.
[0436] The source device 12 includes an encoder 20 and, optionally, may further include an image source 16 , an image preprocessor 18 (or a preprocessing unit), and a communication interface or communication unit 22 .
[0437] Image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the above images.
[0438] In order to distinguish the processing performed by the pre-processor 18 or the pre-processing unit 18 , the image or image data 17 may also be referred to as a raw image or raw image data 17 .
[0439] The preprocessor 18 is configured to receive (raw) image data 17 and preprocess the image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. For example, the preprocessing performed by the preprocessor 18 may include cropping, color format conversion (e.g., from RGB to YCbCr), color grading, or denoising. It will be appreciated that the preprocessor 18 may be an optional component.
[0440] The video encoder 20 is configured to receive the pre-processed image data 19 and provide encoded image data 21 .
[0441] Communication interface 22 in source device 12 may be used to receive encoded image data 21 and transmit the encoded image data 21 (or any other processed version thereof) to another device, such as destination device 14, or any other device, via communication channel 13 for storage or direct reconstruction.
[0442] Destination device 14 includes a decoder 30 (eg, video decoder 30 ) and, in addition and optionally, may include a communication interface or communication unit 28 , a post-processor 32 (or post-processing unit 32 ), and a display device 34 .
[0443] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, a storage device that is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0444] The communication interface 22 and the communication interface 28 may be used to send or receive the encoded image data 21 or the encoded data via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network and public network, or any combination thereof.
[0445] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.
[0446] The communication interface 28 corresponds to the communication interface 22 , for example, and can be used to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain encoded image data 21 .
[0447] Both the communication interface 22 and the communication interface 28 can be configured as a unidirectional communication interface as indicated by the arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14 in Figure 12, or a bidirectional communication interface, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.
[0448] The decoder 30 is configured to receive the encoded image data 21 and provide decoded image data 31 or a decoded image 31 .
[0449] The post-processor 32 in the destination device 14 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data), such as the decoded image 31, to obtain post-processed image data 33, such as the post-processed image 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color grading, cropping, or resampling, or any other processing for generating the decoded image data 31 for display on a display device 34, etc.
[0450] A display device 34 in the destination device 14 is configured to receive the post-processed image data 33 and display the image to a user or viewer. The display device 34 may be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen.
[0451] Although FIG12 shows source device 12 and destination device 14 as separate devices, device embodiments may also include both source device 12 and destination device 14 or the functionality of both source device 12 and destination device 14, that is, both source device 12 or the corresponding functionality and destination device 14 or the corresponding functionality. In these embodiments, source device 12 or the corresponding functionality and destination device 14 or the corresponding functionality may be implemented using the same hardware and / or software or through separate hardware and / or software, or any combination thereof.
[0452] Based on the description, it will be apparent to those skilled in the art that the existence and (precise) division of different units or functions in the source device 12 and / or the destination device 14 shown in FIG. 12 may vary depending on actual devices and applications.
[0453] The content providing system of the content distribution service used in the present application is described below in conjunction with Figure 13A. Figure 13A is a block diagram of a content providing system for implementing the content distribution service used in an embodiment of the present application. The content providing system 2100 includes a capture device 2102, a terminal device 2106 and an (optional) display 2126. The capture device 2102 communicates with the terminal device 2106 via a communication link 2104. The communication link may include the above-mentioned communication channel 13. The communication link 2104 includes but is not limited to WiFi, Ethernet, wired, wireless (3G / 4G / 5G), USB or any combination thereof.
[0454] Capture device 2102 generates data and can use the coding method shown in the above embodiment to encode the data. Alternatively, capture device 2102 can distribute data to a streaming media server (not shown), which encodes the data and transmits the encoded data to terminal device 2106. Capture device 2102 includes but is not limited to a camera, a smart phone or a tablet computer, a computer or a notebook computer, a video conferencing system, a PDA, a vehicle-mounted device, or any combination thereof. For example, capture device 2102 can include above-mentioned source device 12. When data includes video, the video encoder 20 in capture device 2102 can actually perform video encoding. When data includes audio (i.e., sound), the audio encoder 20 in capture device 2102 can actually perform audio encoding. In some actual scenarios, capture device 2102 distributes encoded video data and encoded audio data by multiplexing the encoded video data and the encoded audio data together. In other actual scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. Capture device 2102 distributes the encoded audio data and the encoded video data to terminal device 2106 respectively.
[0455] The terminal device 2106 in the content providing system 2100 receives and regenerates the encoded data. The terminal device 2106 can be a device with data reception and recovery capabilities, such as a smartphone or tablet computer 2108, a computer or laptop computer 2110, a network video recorder (NVR) / digital video recorder (DVR) 2112, a television 2114, a set-top box (STB) 2116, a video conferencing system 2118, a personal digital assistant (PDA) 2122, an in-vehicle device 2124, or any combination thereof, or any other device capable of decoding the encoded data. For example, the terminal device 2106 can include the destination device 14 described above. When the encoded data includes video, the video decoder 30 in the terminal device prioritizes video decoding. When the encoded data includes audio, the audio decoder in the terminal device prioritizes audio decoding. The terminal device 2106 can be a video playback application, a streaming media playback application, a streaming media playback platform, or a live broadcast platform running on the terminal device.
[0456] For end devices with displays, such as smartphones or tablets 2108, computers or laptops 2110, NVRs / DVRs 2112, TVs 2114, PDAs 2122, or in-vehicle devices 2124, the end device can send the decoded data to its display. For end devices without displays, such as STBs 2116 and video conferencing systems 3118, the device can be connected to an external display 2126 to receive and display the decoded data.
[0457] When performing encoding or decoding, each device in this system may use the image encoding device or image decoding device shown in the above embodiments.
[0458] FIG13B is a schematic diagram of the exemplary structure of the terminal device 2106 in FIG13A . After the terminal device 2106 receives the bitstream from the capture device 2102, the protocol processing unit 2202 analyzes the transmission protocol of the bitstream. The protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG Dynamic Adaptive Streaming over HTTP (MPEG-DASH), Real-time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any combination thereof.
[0459] After processing the stream, protocol processing unit 2202 generates a stream file. This file is output to demultiplexing unit 2204. Demultiplexing unit 2204 can separate the multiplexed data into coded audio data and coded video data. As mentioned above, in other practical scenarios, such as in a video conferencing system, the coded audio data and coded video data are not multiplexed. In this case, the coded data is not transmitted to video decoder 3206 and audio decoder 2208 via demultiplexing unit 2204.
[0460] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optional subtitles are generated. The video decoder 2206, including the video decoder 30 described in the above embodiment, decodes the video ES by the decoding method shown in the above embodiment to generate video frames, and sends the data to the synchronization unit 2212. The audio decoder 2208 decodes the audio ES to generate audio frames, and sends the data to the synchronization unit 2212. Alternatively, the video frames can be stored in a buffer (not shown in the figure) before being sent to the synchronization unit 2212. Similarly, the audio frames can be stored in a buffer (not shown in the figure) before being sent to the synchronization unit 2212.
[0461] The synchronization unit 2212 synchronizes the video frames and audio frames and provides the video / audio to the video / audio display 3214. For example, the synchronization unit 2212 synchronizes the presentation of the video and audio information. The information can be encoded in the syntax using the timestamps associated with the representation of the coded audio and visual data and the timestamps associated with the transmission of the data stream.
[0462] If subtitles are included in the bitstream, the subtitle decoder 2210 decodes the subtitles, synchronizes them with the video frames and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 2216.
[0463] The present invention is not limited to the above-mentioned system. The image encoding device or the image decoding device in the above-mentioned embodiments can be used in other systems such as automobiles.
[0464] The following describes a streaming media system applicable to an embodiment of the present application in conjunction with Figure 14A. Figure 14A is a schematic diagram of a workflow of a streaming media system applicable to an embodiment of the present application.
[0465] The streaming media system includes a content creation module that generates the required content data, such as video or audio. The streaming media system also includes a video encoding module that encodes the generated content through an encoder. The streaming media system also includes a video stream transmission module that transmits the encoded video in the form of a code stream. Optionally, the format of the video stream can be converted into a code stream format of a commonly used transmission protocol for OTT (over-the-top) devices, for example, the protocol includes but is not limited to Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live streaming protocol (HLS), MPEG HTTP Dynamic Adaptive Streaming over HTTP (MPEG-DASH), Real-time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP) or any combination thereof. Optionally, the video stream can be stored to store the original format of the video stream and / or the converted multiple code stream formats for easy use. Furthermore, the streaming system also includes a video stream encapsulation module for encapsulating the video stream to generate an encapsulated video stream. The encapsulated video stream can be referred to as a video streaming package. Exemplarily, the video streaming package can be generated based on a transcoded video stream or a stored video stream. Furthermore, the streaming system also includes a content distribution network (CDN) for distributing the video streaming package to multiple OTT devices, such as mobile phones, computers, tablets, and home projectors.
[0466] It should be noted that video encoding, video streaming transmission, video stream transcoding, video stream storage, video streaming package generation and content distribution network can all be implemented on cloud servers.
[0467] An exemplary streaming media system architecture of the present application is described below in conjunction with FIG14B . The streaming media system architecture includes: a client device, a content distribution network, and a cloud server.
[0468] The user on the client device sends a play or playback request to the cloud platform. Optionally, the content of the request can be the title of the movie or TV program to be played.
[0469] The cloud platform makes a decision and responds to the client, sending the client the address of the requested content on the CDN. Optionally, the content sent to the client can be a URL link (uniform resource locator). Specifically, the playback application service in the cloud platform checks user authorization and permissions, and then considers individual client characteristics and current network conditions to determine which specific files are needed to process the playback request. It should be noted that the content delivery network (CDN) regularly reports its operating status, learned routes, and available content (files) to the cache control service in the cloud platform.
[0470] The client then requests the CDN to play the content based on the address, and the CDN provides the content to the client, ultimately completing the client's request.
[0471] The following describes the system architecture applicable to the embodiment of the present application in conjunction with Figure 15A. Figure 15A is a schematic diagram of a possible system architecture applicable to the embodiment of the present application. The system architecture of the embodiment of the present application includes: front-end equipment, transmission links, and terminal display equipment.
[0472] Among them, the front-end equipment is used to capture or produce HDR / SDR content (for example, HDR / SDR video or images).
[0473] In a possible embodiment, the front-end device can also be used to extract corresponding metadata from the HDR content. The metadata may include global mapping information, local mapping information, and dynamic metadata and static metadata corresponding to the HDR content. The front-end device can send the HDR content and metadata to the terminal display device via a transmission link. Specifically, the HDR content and metadata can be transmitted in the form of one data packet, or respectively in two data packets, which is not specifically limited in the embodiment of the present application.
[0474] Optionally, the terminal display device can be used to receive metadata and HDR content, and extract the global mapping information, local mapping information, and the terminal display device information contained in the corresponding metadata according to the HDR content, obtain a mapping curve to perform global tone mapping and local tone mapping on the HDR content, convert it into display content adapted for the HDR display device or SDR device in the terminal display device, and display it. It should be understood that in different embodiments, the terminal display device may include a display device with a display capability of a lower dynamic range or a higher dynamic range than the HDR content generated by the front-end device, and this application is not limited thereto.
[0475] Optionally, the front-end device and the terminal display device in this application can be independent and different physical devices. For example, the front-end device can be a video capture device or a video production device, where the video capture device can be a video camera, a camera, an image rendering machine, etc. The terminal display device can be a device with video playback function, such as virtual reality (VR) glasses, a mobile phone, a tablet, a television, a projector, etc.
[0476] Optionally, the transmission link between the front-end device and the terminal display device can be a wireless connection or a wired connection, wherein the wireless connection can adopt technologies such as long term evolution (LTE), fifth generation (5G) mobile communications, and future mobile communications. Wireless connections can also include wireless fidelity (WiFi), Bluetooth, near field communication (NFC), and other technologies. Wired connections can include Ethernet connections, local area network connections, etc. There is no specific limitation on this.
[0477] The present application may also integrate the functions of the front-end device and the terminal display device into the same physical device, for example, a terminal device such as a mobile phone or tablet with a video capture function. The present application may also integrate some functions of the front-end device and some functions of the terminal display device into the same physical device. This is not specifically limited.
[0478] The following describes an end-to-end image processing system provided by an embodiment of the present application in conjunction with Figure 15B. This system can be applied to the system architecture shown in Figure 15A. Figure 15B is a schematic diagram of the structure of an image processing system provided by an embodiment of the present application. In Figure 14B, HDR / SDR content is exemplified by HDR video. The image processing system includes: an HDR preprocessing module, an HDR video encoding module, an HDR video decoding module, and a tone mapping module.
[0479] Among them, the HDR preprocessing module and the HDR video encoding module can be located in the front-end device shown in Figure 15A, and the HDR video decoding module and the tone mapping module can be located in the terminal display device shown in Figure 15A.
[0480] HDR pre-processing module: This module is used to extract dynamic metadata (e.g., maximum, minimum, average, and range of brightness) from HDR video, determine mapping curve parameters based on the dynamic metadata and the display capabilities of the target display device, write the mapping curve parameters into the dynamic metadata to obtain HDR metadata, and transmit the metadata. HDR video can be either captured or processed by a colorist; the display capabilities of the target display device are the brightness range that the target display device can display.
[0481] HDR video encoding module: used to encode HDR video and HDR metadata according to the video compression standard (for example, AVS or HEVC standard) (for example, embedding HDR metadata into the user-defined part of the bitstream) and output the corresponding bitstream (AVS or HEVC bitstream).
[0482] HDR video decoding module: used to decode the generated bitstream (AVS bitstream or HEVC bitstream) according to the standard corresponding to the bitstream format, and output the decoded HDR video and HDR metadata.
[0483] Tone mapping module: used to generate a mapping curve according to the parameters of the mapping curve in the decoded HDR metadata, and perform tone mapping on the decoded HDR video (i.e. HDR adaptation processing or SDR adaptation processing), and send the HDR adapted video after tone mapping to the HDR display terminal for display, or send the SDR adapted video to the SDR display terminal for display.
[0484] Exemplarily, the HDR pre-processing module may exist in a video acquisition device or a video production device.
[0485] Exemplarily, the HDR video encoding module may exist in a video acquisition device or a video production device.
[0486] Exemplarily, the HDR video decoding module may be present in a set-top box, a television display device, a mobile terminal display device, and a video conversion device for live broadcasting, online video applications, and the like.
[0487] For example, the tone mapping module may be present in a set-top box, a television display device, a mobile terminal display device, or a video conversion device for a live webcast or online video application. More specifically, the tone mapping module may be present in the form of a chip or software program in the set-top box, television display, or mobile terminal display, and may be present in the form of a software program in the video conversion device for a live webcast or online video application.
[0488] In a possible embodiment, when the tone mapping module and the HDR video decoding module are both present in a set-top box, the set-top box can complete the functions of receiving, decoding, and tone mapping the video stream. The set-top box sends the decoded video data to a display device through a high-definition multimedia interface (HDMI) for display, so that the user can enjoy the video content.
[0489] In an example, FIG16 shows a schematic block diagram of a device 1600 according to an embodiment of the present application. The device 1600 may include: a processor 1601 and a transceiver / transceiver pin 1602 , and optionally, a memory 1603 .
[0490] The various components of the device 1600 are coupled together via a bus 1604, wherein the bus 1604 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are referred to as bus 1604 in the figure.
[0491] Optionally, the memory 1603 may be used to store instructions in the aforementioned method embodiment. The processor 1601 may be used to execute the instructions in the memory 1603 and control the receiving pin to receive a signal and control the transmitting pin to send a signal.
[0492] The apparatus 1600 may be the electronic device or a chip of the electronic device in the above method embodiment.
[0493] Illustratively, the electronic device in the above embodiment may be a terminal device or a server.
[0494] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0495] This application also provides a chip including one or more interface circuits and one or more processors; the one or more processors receive or send data via the one or more interface circuits, and when the one or more processors execute computer instructions, the above-mentioned related method steps are implemented to implement the steps of the method in the above embodiment. The interface circuit is a transceiver / transceiver pin 1602.
[0496] The present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the method in the above-mentioned embodiment.
[0497] The present application also provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer executes the above-mentioned related steps to implement the method in the above-mentioned embodiment.
[0498] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the methods in the above-mentioned method embodiments.
[0499] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0500] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0501] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0502] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0503] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0504] Any content of each embodiment of this application, as well as any content of the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0505] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0506] The steps of the method or algorithm described in conjunction with the disclosure of the embodiments of the present application can be implemented in a hardware manner, or can be implemented by a processor executing a software instruction. The software instruction can be composed of corresponding software modules, and the software module can be stored in a random access memory (Random Access Memory, RAM), a flash memory, a read-only memory (Read Only Memory, ROM), an erasable programmable read-only memory (Erasable Programmable ROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), a register, a hard disk, a mobile hard disk, a read-only compact disc (CD-ROM) or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0507] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0508] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A method for encoding metadata of multimedia, characterized in that, The method includes: Obtaining metadata of multimedia; wherein, the metadata of the multimedia includes first data, and the first data is floating-point data; Writing second data into a bitstream; wherein, the second data is a representation form of the first data in a floating-point data format.
2. The method according to claim 1, wherein The method further includes: Generating the second data based on the first data.
3. The method according to claim 2, wherein The method further includes: Obtaining the floating-point data format; The generating the second data based on the first data includes: Processing the first data according to the floating-point data format to obtain the second data.
4. The method according to claim 3, characterized in that, The method further includes: Obtaining the bit width of the second data; The processing the first data according to the floating-point data format to obtain the second data includes: Processing the first data according to the floating-point data format and the bit width to obtain the second data.
5. The method according to claim 4, wherein The method further includes: Writing the bit width into the bitstream.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Writing the floating-point data format into the bitstream.
7. The method according to any one of claims 1 to 6, wherein the second data is a binary representation form or a hexadecimal representation form of the first data in a floating-point data format.
8. The method according to any one of claims 1 to 7, characterized in that, The metadata of the multimedia is metadata of a high dynamic range (HDR) image.
9. The method according to any one of claims 1 to 7, characterized in that, The metadata of the multimedia is metadata of enhanced data, and the enhanced data is generated based on a standard dynamic range (SDR) image and a high dynamic range (HDR) image.
10. The method according to any one of claims 1 to 9, characterized in that, The floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, TensorFloat-32.
11. A bitstream, characterized in that, The bitstream is generated according to the method of any one of claims 1 to 10 above.
12. A bitstream, characterized in that, The bitstream includes second data, and the second data is a representation form of first data in a floating-point data format, and the first data is floating-point data in the metadata of multimedia.
13. The bitstream according to claim 12, wherein The bitstream further includes the floating-point data format.
14. The bitstream according to claim 12 or 13, characterized in that, The second data is obtained by processing the first data according to the floating-point data format and the bit width of the second data, and the bitstream further includes the bit width.
15. A method for decoding metadata of multimedia, characterized in that, The method includes: Receiving a bitstream; Parsing second data from the bitstream, where the second data is a representation form of first data in a floating-point data format, and the first data is floating-point data in the metadata of multimedia; Generating the first data based on the second data.
16. The method according to claim 15, wherein The method further includes: Obtaining the floating-point data format; The generating the first data based on the second data includes: Processing the second data according to the floating-point data format to obtain the first data.
17. The method according to claim 16, characterized in that, The obtaining the floating-point data format includes: Parsing the floating-point data format from the bitstream.
18. The method according to claim 16 or 17, characterized in that, The method further includes: Obtaining the bit width of the second data; The processing the second data according to the floating-point data format to obtain the first data includes: Processing the second data according to the floating-point data format and the bit width to obtain the first data.
19. The method according to claim 18, characterized in that, Obtaining the bit width of the second data includes: Parsing the bit width from the bitstream.
20. The method according to any one of claims 15 to 19, wherein: The second data is a binary representation or a hexadecimal representation of the first data in floating-point data format.
21. The method according to any one of claims 15 to 20, wherein: The metadata of the multimedia is the metadata of a high-dynamic range (HDR) image.
22. The method according to any one of claims 15 to 21, characterized in that, The metadata of the multimedia is enhanced data metadata, and the enhanced data is determined based on a standard dynamic range (SDR) image and a high-dynamic range (HDR) image.
23. The method according to any one of claims 15 to 22, characterized in that, The floating-point data format includes at least one of the following: IEEE 754, FP8 (E4M3), FP8 (E5M2), Bfloat16, TensorFloat-32.
24. An encoder, characterized in that, Including: A memory and a processor, the memory being coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the encoder is caused to execute the method according to any one of claims 1 to 10.
25. A decoder, characterized in that, Including: A memory and a processor, the memory being coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the decoder is caused to execute the method according to any one of claims 15 to 23.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on a computer or a processor, the computer or the processor is caused to execute the method according to any one of claims 1 to 10, or execute the method according to any one of claims 15 to 23.
27. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a computer or a processor, the steps of the method according to any one of claims 1 to 10 are caused to be executed, or the steps of the method according to any one of claims 15 to 23 are caused to be executed.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores the bitstream according to any one of claims 11 to 14 above.
Citation Information
Patent Citations
Metadata encoding and decoding method for multimedia
CN120343268A
Method for encoding floating-point data, method for decoding floating-point data, and corresponding encoder and decoder
CN102498673A
Encoding method and decoding method for video sequence parameters, corresponding devices and code streams
CN103139555A
Data conversion method based on new short floating point type data
CN105634499A
3D audio encoding and decoding method and device thereof
CN109448741A