Media data processing method and apparatus, encoder, decoder, and system
Patent Information
- Application Number
- PCT/CN2025/148299
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-11
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025148299_17092026_PF_FP_ABST
Abstract
Description
Media data processing methods, apparatus, encoders, decoders and systems
[0001] This application claims priority to Chinese patent application filed on March 11, 2025, with application number 202510292940.0 and entitled "Media Data Processing Method, Apparatus, Encoder, Decoder and System", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of media, and more particularly to a media data processing method, apparatus, encoder, decoder, and system. Background Technology
[0003] Currently, terminal devices centrally manage and present media content. For example, they use grids, page turning, time categorization, and location categorization to display multiple images. However, the implementation methods of media content management vary greatly among different applications, terminal devices, and systems. Media content management lacks uniformity, application adaptation is complex, and the diversity of integrated presentation of various media content is poor. Summary of the Invention
[0004] This application provides a media data processing method, apparatus, encoder, decoder, and system, thereby enhancing the diversity of multi-media content fusion presentation.
[0005] Firstly, a media data processing method is provided. The method includes acquiring multiple media contents and generating a media file based on the multiple media contents. The multiple media contents include a first image and other first media content besides the first image; the first media content includes a first video. The media file includes image data of the first image, metadata, and data of the first media content. The metadata is used to associate and / or transform the first image and / or the first media content.
[0006] In one possible implementation, a first image and first media content are obtained based on a set of temporally continuous images.
[0007] In another possible implementation, the first media content is obtained based on the first group of images in a set of temporally continuous images.
[0008] In another possible implementation, the first image is obtained based on a second image group from a set of temporally continuous images.
[0009] In another possible implementation, a set of temporally consecutive images comprises multiple images taken at the same time with different exposure parameters.
[0010] In another possible implementation, the various media content also includes image format information. Format information includes SDR and / or HDR.
[0011] In another possible implementation, a media file is generated based on multiple media contents, including: encoding a first image to obtain an image file, the image file including image data; encoding first media content to obtain data of the first media content; and encoding the data and metadata of the first media content into the image file to obtain the media file.
[0012] In another possible implementation, a media file is generated based on multiple media contents, including: encoding a first image to obtain image data; encoding a first media content to obtain data of the first media content; and generating a media file based on the image data, the data of the first media content, and metadata.
[0013] Secondly, a media data processing method is provided, the method comprising: acquiring a media file, the media file including image data of a first image, metadata and data of first media content, the metadata being used to associate and / or transform the first image and / or the first media content, the first media content including a first video; and decoding the media file to obtain the first image and / or the first media content.
[0014] In one possible implementation, obtaining the first image from the media file includes: decoding the media file to obtain the first image.
[0015] In another possible implementation, obtaining the first image from the media file includes: decoding the image data in the media file to obtain the first image.
[0016] In another possible implementation, the method further includes associating and / or transforming the first image and / or the first media content based on metadata.
[0017] The technical solution provided in this application includes encoding multiple media contents into the same media file, and using metadata in the media file to indicate the association and / or transformation of the multiple media contents contained in the media file. When presenting multiple media contents in a fused form, the metadata in the media file is used to associate and / or transform the multiple media contents in the media file, thus presenting the multiple media contents in a fused form. This application improves the structure of media files and sets metadata to process multiple media contents, expanding the forms of fusion of multiple media contents. The media data processing method provided in this application is a standardized encoding method, which facilitates system implementation, achieves unified management of media content, improves the convenience of media content management, reduces application adaptation costs, and enhances the diversity of multi-media content fusion presentation and the user's cross-platform experience.
[0018] The technical solution provided in this application also includes encoding multiple media contents into different media files and using metadata in the media files to indicate that the media files together contain multiple media contents with other media files.
[0019] In another possible implementation, the metadata includes time-related metadata, which is used to indicate the temporal relationship between the first media content and the first image included in the media file.
[0020] In another possible implementation, the metadata includes time-related metadata, which is used to indicate the temporal relationship between the first media content, the second image, and the first image included in the media file.
[0021] In another possible implementation, the time-related metadata is obtained from first time-domain information and second time-domain information; the first time-domain information is obtained based on the time of acquiring the first video or image sequence of the first media content, and the second time-domain information is obtained based on the time of acquiring the first image.
[0022] In another possible implementation, the time-related metadata indicates the difference in start time between the first image and the video / image sequence; or, the time-related metadata indicates the difference in frame number between the first image and the start frame of the video / image sequence; or, the time-related metadata indicates the timestamp of the first image; or, the time-related metadata indicates the difference in frame number between the second image and the start frame of the video / image sequence; or, the time-related metadata indicates the timestamp of the second image; or, the time-related metadata indicates the start timestamp of the video / image sequence; or, the time-related metadata indicates the end timestamp of the video / image sequence; or, the time-related metadata indicates the duration of the video / image sequence; or, the time-related metadata indicates the number of frames in the video / image sequence.
[0023] Thus, time-related images and other media content (e.g., videos) can be retrieved from media files based on time-related metadata.
[0024] In another possible implementation, the method includes: obtaining the video frame corresponding to the first image from the first video based on time-related metadata.
[0025] In another possible implementation, the metadata also includes transformation-related metadata, which is used to process the first media content and / or transform the first image.
[0026] In another possible implementation, the metadata also includes transformation-related metadata, which is used to determine the transformation of the first media content and / or the first image.
[0027] In another possible implementation, the conversion-related metadata includes at least one of format information, color gamut information, resolution, or rotation.
[0028] In another possible implementation, the conversion may also include conversion between HDR and SDR, or conversion between HDR content with different dynamic ranges.
[0029] In another possible implementation, when the first image is an HDR image and the first video is an SDR video, the first video is converted based on the conversion-related metadata.
[0030] In another possible implementation, the conversion-related metadata is obtained from first format information and second format information; the first format information is obtained based on consecutive images of the first video of the first media content, and the second format information is obtained based on consecutive images of the first image.
[0031] In another possible implementation, the transformation associated metadata is obtained from the data associated with the first image.
[0032] Therefore, by fusing time-related images and other media content (such as videos) obtained from media files based on transformation-related metadata, the diversity of multi-media content fusion presentation is improved.
[0033] In another possible implementation, the media file is in the High Efficiency Image Format (HEIF) file format; metadata is located in the ftyp or meta field of the media file.
[0034] In another possible implementation, the meta field includes an iinf subfield; the metadata is located in the meta field of the media file, including: the metadata is located in the data segment corresponding to the type other than the image codec format type in the iinf subfield.
[0035] In another possible implementation, the image codec format type includes at least one of MPEG, Joint Photographic Experts Group (JPEG), Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or Versatile Video Codec (VVC).
[0036] In another possible implementation, the type, in addition to the image codec format type, includes at least one of uri, mime, it35, xml, or exif.
[0037] In another possible implementation, the meta field includes the iloc subfield, the iprp subfield, the iref subfield, and the idat subfield; the metadata is located in the meta field of the media file, including: the metadata is located in the data segment corresponding to at least one of the iloc subfield, the iprp subfield, the iref subfield, or the idat subfield.
[0038] In another possible implementation, at least one field in the field identifying time-related metadata and the field identifying transformation-related metadata are different.
[0039] In another possible implementation, the field that identifies time-related metadata is associated with the field that identifies transformation-related metadata.
[0040] In another possible implementation, the fields identifying time-related metadata and the fields identifying transformation-related metadata include at least one of the IPRP subfields or the IRef subfields; the fields identifying tag metadata are associated with at least one of the IPRP subfields or the IRef subfields.
[0041] In another possible implementation, the field that identifies time-related metadata and the field that identifies transformation-related metadata are associated through item association rules.
[0042] In another possible implementation, the data for the first media content is located in the idat subfield of the mdat field, moov field, or meta field.
[0043] In another possible implementation, the media file contains a first field or a first preset character. Metadata is located in the first field or the first preset character.
[0044] The first category refers to any one of the following abbreviations or full names related to continuous shooting and moving images: live, motion, movie photo, continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, burst mode, etc.
[0045] The first field refers to the box, fullbox, or APPn identified by the first type.
[0046] The first preset character is any one of the Chinese or English abbreviations or full names related to continuous shooting and moving images, such as live, motion, movie photo, continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, burst mode, etc.
[0047] In another possible implementation, the media file is in JPEG format; the metadata is located in APPn or MPF within the media file.
[0048] In another possible implementation, the metadata is located in the APPn or MPF of the media file, including: at least one of the tags, labels, or payloads of the APPn in the media file is of type first.
[0049] The first type uses any one of the following abbreviations or full names related to continuous shooting and moving images: live, motion, movie photo, continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, burst mode, etc.
[0050] In another possible implementation, the media file is in JPEG format; metadata is located in APPn within the media file; APPn contains data structures related to the High Efficiency Image Format (HEIF).
[0051] In another possible implementation, the APPn flag indicates HEIF.
[0052] In another possible implementation, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0053] In another possible implementation, APPn's payload also includes the Start of Image (SOI).
[0054] In another possible implementation, the metadata is located in the APPn of the media file, including: the metadata is located in the payload of the APPn in the media file.
[0055] Therefore, JPEG animated images are encapsulated in a JPEG-compatible manner and unified to HEIF, meaning that JPEG file format media files include content related to the HEIF file format. This enables unified management of media content, facilitates the expansion of various media content, and improves the convenience of media content management.
[0056] Thirdly, a method for encapsulating media files is provided, which encapsulates multiple media contents into a media file. The media file is in the JPEG format and includes APPn, which contains data structures related to HEIF.
[0057] In one possible implementation, the APPn flag indicates HEIF.
[0058] In another possible implementation, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0059] In another possible implementation, APPn's payload also includes the Start of Image (SOI).
[0060] In another possible implementation, the media file is the media file of the first or second aspect mentioned above.
[0061] In another possible implementation, the aforementioned metadata is located in the APPn of the media file.
[0062] In another possible implementation, the aforementioned metadata is located in the APPn of the media file, including: the metadata is located in the payload of the APPn in the media file.
[0063] Fourthly, an encoding apparatus is provided, the encoding apparatus including processing circuitry for performing the methods of the first aspect or any possible design of the first aspect. For example, the encoding apparatus includes a communication module and an encoding module.
[0064] The communication module is used to acquire multiple media content, including a first image and other media content besides the first image, including a first video.
[0065] An encoding module is used to generate media files based on various media content. The media files include image data of a first image, metadata, and data of the first media content. The metadata is used to associate and / or transform the first image and / or the first media content.
[0066] In one possible implementation, a first image and first media content are obtained based on a set of temporally continuous images.
[0067] In another possible implementation, the first media content is obtained based on the first group of images in a set of temporally continuous images.
[0068] In another possible implementation, the first image is obtained based on a second image group from a set of temporally continuous images.
[0069] In another possible implementation, a set of temporally consecutive images comprises multiple images taken at the same time with different exposure parameters.
[0070] In another possible implementation, the various media content also includes image format information. Format information includes SDR and / or HDR.
[0071] In another possible implementation, when the encoding module generates a media file based on multiple media contents, it is specifically used to encode a first image to obtain an image file, the image file including image data; encode the first media content to obtain the data of the first media content; and encode the data and metadata of the first media content into the image file to obtain the media file.
[0072] In another possible implementation, when the encoding module generates a media file based on multiple media contents, it is specifically used to encode a first image to obtain image data; encode a first media content to obtain data of the first media content; and generate a media file based on the image data, the data of the first media content, and metadata.
[0073] Fifthly, a decoding apparatus is provided, the decoding apparatus including processing circuitry for performing the methods of the second aspect or any possible design of the second aspect. For example, the encoding apparatus includes a communication module and a decoding module.
[0074] The communication module is used to acquire media files, which include image data of a first image, metadata, and data of first media content. The metadata is used to associate and / or transform the first image and / or the first media content, which includes a first video.
[0075] A decoding module is used to decode media files to obtain a first image and / or first media content.
[0076] In one possible implementation, when the decoding module obtains the first image based on the media file, it is specifically used to decode the media file to obtain the first image.
[0077] In another possible implementation, when the decoding module obtains the first image from the media file, it is specifically used to decode the image data in the media file to obtain the first image.
[0078] In another possible implementation, the decoding module is also used to associate and / or transform the first image and / or the first media content based on metadata.
[0079] In a sixth aspect, an encoder is provided, the encoder including at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by at least one processor, it implements the method described in the first aspect or any possible design of the first aspect.
[0080] In a seventh aspect, a decoder is provided, the decoder including at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by at least one processor, it implements the method described in the second aspect or any possible design of the second aspect.
[0081] Eighthly, a coding / decoding system is provided, the coding / decoding system comprising an encoder as described in the sixth aspect and a decoder as described in the seventh aspect.
[0082] A ninth aspect provides a chip, comprising: a processor and a power supply circuit; wherein the power supply circuit is used to supply power to the processor; the processor is used to perform operational steps of the method in the first aspect or any possible implementation of the first aspect, and to perform operational steps of the method in the second aspect or any possible implementation of the second aspect.
[0083] In a tenth aspect, a computer program product is provided, comprising a computer program or instructions, which, when executed on a processor, cause the processor to perform operational steps of the method in the first aspect or any possible implementation thereof, or to perform operational steps of the method in the second aspect or any possible implementation thereof.
[0084] Eleventhly, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a computing device, causing the computing device to perform operational steps of the method in the first aspect or any possible implementation thereof, or to perform operational steps of the method as described in the second aspect or any possible implementation thereof.
[0085] In a twelfth aspect, a media file is provided, said media file being obtained from the first aspect or any possible implementation thereof.
[0086] In a thirteenth aspect, a media file is provided, the media file including metadata, the metadata including tag metadata, and at least one of attribute metadata or location metadata.
[0087] Fourteenth aspect, a method for storing media files is provided, the method comprising: receiving a media file generated according to the first aspect or any possible implementation thereof, or a media file as described in the thirteenth aspect; and storing the media file in a storage medium.
[0088] In a fifteenth aspect, an apparatus for storing media files is provided, the apparatus being used to store media files generated according to the first aspect or any possible implementation thereof, or for storing media files as described in the thirteenth aspect. Exemplarily, the apparatus may be a computer-readable storage medium.
[0089] In a sixteenth aspect, a method for transmitting a media file is provided, the method comprising: acquiring a media file, the media file being generated by the first aspect or any possible implementation thereof, or the media file being the media file described in the thirteenth aspect; and sending the media file.
[0090] In a seventeenth aspect, an apparatus for transmitting media files is provided, the apparatus being used to acquire and transmit media files generated by the first aspect or any possible implementation thereof, or to acquire and transmit media files as described in the thirteenth aspect.
[0091] The technical effects of any of the implementation methods in aspects three through seventeen can be found in the technical effects of the corresponding implementation methods in aspects one or two, and will not be repeated here.
[0092] All possible implementations of any of the above aspects can be combined, provided that the solutions do not contradict each other. Attached Figure Description
[0093] Figure 1 is a schematic diagram of a method for presenting multiple media content in a fusion format according to this application;
[0094] Figure 2 is a schematic diagram of the structure of an encoding / decoding system provided in this application;
[0095] Figure 3 is a schematic diagram of another encoding / decoding system provided in this application;
[0096] Figure 4 is a flowchart illustrating a media data processing method provided in this application;
[0097] Figure 5 is a schematic diagram of the structure of an encoding device provided in this application;
[0098] Figure 6 is a schematic diagram of a decoding device provided in this application;
[0099] Figure 7 is a schematic diagram of the structure of an encoder provided in this application;
[0100] Figure 8 is a schematic diagram of the structure of a decoder provided in this application. Detailed Implementation
[0101] To facilitate understanding, the main terms used in this application will be explained first.
[0102] Currently, terminal devices centrally manage and present media content. Presenting multiple media content in a converged format has become a trend.
[0103] For example, multiple images can be presented in a fused manner. As shown in Figure 1(a), the terminal device presents multiple images in a grid format. When a user views an image, the terminal device receives the user's click action and displays the image viewed by the user.
[0104] As shown in Figure 1(b), the terminal device presents multiple images in a page-turning manner. When the user views an image, the terminal device receives the user's page-turning operation and presents the image viewed by the user.
[0105] As shown in Figure 1(c), the terminal device presents multiple images using categories such as time, location, and people. For example, multiple images are presented by a gallery application on the terminal device.
[0106] For example, still images can be combined with continuous images. Or, still images can be combined with video. For instance, an image can be displayed first, followed by a video to achieve an animated effect. Animated images typically refer to Graphics Interchange Format (GIF) animated images, which are image formats that create animation effects by playing a series of static images in succession.
[0107] However, the media content is centrally managed by the system, and the fusion format is simplistic, merely a simple arrangement of multiple images. The fusion implementation varies greatly across different applications, terminal devices, and system management, resulting in poor diversity in the fusion of various media content and a subpar user experience.
[0108] To address the issue of poor diversity in the fusion presentation of multiple media content, this application provides a media data processing method, namely, a multi-media content fusion scheme. The method includes encoding multiple media content into a media file, and using metadata in the media file to indicate the association and / or transformation of the multiple media content contained in the media file. When presenting multiple media content in a fused form, the metadata in the media file is used to associate and / or transform the multiple media content in the media file, thus presenting the multiple media content in a fused form. This application improves the structure of media files and sets metadata to process multiple media content, expanding the fusion forms of multiple media content. Furthermore, the media data processing method provided in this application is a standardized encoding method, facilitating system implementation, achieving unified management of media content, improving the convenience of media content management, reducing application adaptation costs, and enhancing the diversity of multi-media content fusion presentation and the user's cross-platform experience.
[0109] Media content refers to information presented and disseminated to the public in various forms. Media content includes, but is not limited to, one or more of text, images, audio, and video. For example, depending on the medium and presentation method, media content can be categorized into various types, such as news reports, feature articles, advertisements, TV dramas, movies, music, and social media content.
[0110] Subtitles are text-based representations of dialogue or narration in a video displayed at the bottom of the screen for reading. They are primarily used in situations where sound is unavailable or where understanding the content is desired without interfering with the audio, such as movie subtitles or watching videos in a silent environment. Subtitles are typically automatically synchronized with the audio content in the video and have a relatively fixed format (such as font, size, and position) to ensure smooth and accurate reading.
[0111] Text refers to static or dynamic text elements added to a video, including titles, descriptive text, and tags. It is commonly used in intros, outros, transitions, and to emphasize content, enhancing the video's visual appeal and information delivery. Text styles, sizes, colors, and animation effects can all be customized, offering high flexibility.
[0112] Subtitles are primarily used to convey dialogue or narration, while text enhances the visual appeal and information delivery of a video. Subtitles are typically located at the bottom of the screen and have a relatively fixed format. Text can be placed anywhere in the video and offers a wider variety of styles and animation effects. Subtitles are suitable for situations where understanding the dialogue is crucial, such as movie subtitles. Text is suitable for situations where emphasis or embellishment is needed, such as titles and explanatory text at the beginning and end of credits.
[0113] A media file is a document containing multimedia data. Media files are typically used to store and transmit multimedia data. Multimedia data includes various types of data such as video, audio, images, subtitles, and text. In this application, the type and number of media contained in a media file are not limited. The more types of media a media file contains, the higher the integration level of multiple media content.
[0114] The media data processing method provided in this application will be described in detail below with reference to the accompanying drawings.
[0115] Figure 2 is a schematic diagram of the structure of an encoding / decoding system provided in this application. The encoding / decoding system 200 includes a source device 210 and a destination device 220. The source device 210 is used to process various media content to generate a media file. The various media content includes images and other media content besides images. The other media content can be, for example, at least one of video, audio, subtitles, or text. The media file includes data of various media content and metadata, the metadata being used to associate and / or convert a first image and / or the first media content. The source device 210 sends the media file to the destination device 220. The destination device 220 is used to obtain images and other media content besides images from the media file. Optionally, the destination device 220 is also used to perform a fusion process on the images and other media content besides images, the fusion process including at least one of editing, display, or transcoding. For example, the destination device 220 is also used to convert the first image or the first media content.
[0116] Specifically, the source device 210 includes an image acquisition unit 211, a preprocessor 212, an encoder 213, and a communication interface 214.
[0117] Image acquisition device 211 is used to acquire raw images. Image acquisition device 211 includes or is any category of image capture device for, for example, capturing real-world images, and / or any category of image or commentary generation device (for screen content encoding, some text on the screen is also considered an image to be encoded or part of an image). For example, a computer graphics processor for generating computer-animated images, or any category of device for acquiring and / or providing real-world images, computer-animated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). Image acquisition device 211 is a camera for capturing images or a memory for storing images. Image acquisition device 211 also includes any category of (internal or external) interface for storing previously captured or generated images and / or acquiring or receiving images. If image acquisition device 211 is a camera, it is, for example, a local or integrated camera in the source device. If image acquisition device 211 is a memory, it is, for example, a local or integrated memory in the source device. When the image acquisition device 211 includes an interface, the interface may be, for example, an external interface for receiving images from an external video source. The external video source may be, for example, an external image capture device, such as a camera, external storage, or an external image generation device. The external image generation device may be, for example, an external computer graphics processor, a computer, or a server. The interface may be any type of interface according to any proprietary or standardized interface protocol, such as a wired or wireless interface, or an optical interface.
[0118] An image can be viewed as a two-dimensional array or matrix of pixels (picture elements). Pixels in an array are also called sample points. The number of sample points in an array or image along the horizontal and vertical directions (or axes) defines the image's size and / or resolution. To represent color, three color components are typically used; that is, an image can be represented as or contain three sample arrays. For example, in RBG format or color space, an image includes corresponding red, green, and blue sample arrays. However, in image coding or video coding, each pixel is typically represented in a luma / chroma format or color space. For example, for a YUV format image, this includes a luma component indicated by Y (sometimes also indicated by L) and two chroma components indicated by U and V. The luma component Y represents the brightness or grayscale level intensity (e.g., both are the same in a grayscale image), while the two chroma components U and V represent chroma or color information components. Accordingly, a YUV format image includes a luma sample array of luma sample values (Y) and two chroma sample arrays of chroma values (U and V). An RGB format image can be converted or transformed to YUV format, and vice versa; this process is also called color transformation or conversion. If the image is black and white, it may only include a luminance sampling array. In this application, the image transmitted from the image acquisition unit 211 to the encoder 213 can also be referred to as raw image data.
[0119] The preprocessor 212 receives the raw image acquired by the image acquisition unit 211 and preprocesses the raw image to obtain a preprocessed image. For example, the preprocessing performed by the preprocessor 212 includes retouching, color format conversion (e.g., from RGB format to YUV format), color adjustment, or noise reduction.
[0120] In some embodiments, source device 210 further includes an audio acquisition unit. The audio acquisition unit is used to acquire raw audio. The audio acquisition unit can be any type of audio acquisition device for capturing real-world sounds, and / or any type of audio generation device. The audio acquisition unit is, for example, a computer audio processor for generating computer audio. The audio acquisition unit can also be any type of memory or storage device for storing audio. Audio includes real-world sounds, virtual scene sounds (such as VR or augmented reality (AR) sounds), and / or any combination thereof. Preprocessor 212 is also used to receive the raw audio acquired by the audio acquisition unit and to preprocess the raw audio. For example, the preprocessing performed by preprocessor 212 includes channel conversion, audio format conversion, or noise reduction.
[0121] Encoder 213 is used to acquire various media content such as images, videos, and audio, and encodes each type of media content separately to obtain encoded data for each type of media content. For example, an image encoding algorithm is used to compress the original image to obtain encoded image data, reducing the data size of the encoded image data while maintaining the quality of the original image as much as possible. Similarly, an audio / video encoding algorithm is used to compress the original audio and video to obtain encoded audio / video data, reducing the data size of the encoded audio / video data while maintaining the quality of the original audio and video as much as possible.
[0122] Optionally, although encoders and encoders are functionally distinct, in practical applications they are often integrated together. The encoder's software or hardware has built-in encoding functions, allowing the encoded media file to be obtained directly after data encoding is completed. In some embodiments, the encoder, in addition to having the function of encoding data, also has the function of encoding the encoded data into a media file.
[0123] Encoder 213 is also used to encode various media content into media files. For example, it encodes encoded audio and video data, along with other relevant information (such as subtitles and text), into media files for easier playback, transmission, and storage of the encoded data. Common encoding formats include MP4, MOV, and MXF.
[0124] For example, encoder 213 includes encoding unit 2131 and encapsulation unit 2132. Encoding unit 2131 is used to encode multiple media contents separately to obtain encoded data of multiple media contents. Encapsulation unit 2132 is used to encode the encoded data of multiple media contents to obtain a media file.
[0125] In this application, the packaging unit 2132 and the encoder 213 can be integrated into one physical device or disposed on different physical devices, without limitation. If the encoder 213 does not include the packaging unit 2132, it means that the packaging unit 2132 and the encoder 213 are two different physical devices.
[0126] The communication interface 214 is used to receive media files generated by the encoder 213 and send the bit stream through the communication channel 230.
[0127] The target device 220 includes a display 221, a post-processor 222, a decoder 223, and a communication interface 224.
[0128] The communication interface 224 is used to receive media files and transmit them to the decoder 223.
[0129] Communication interfaces 214 and 224 can be used to communicate via a communication link between source device 210 and destination device 220, such as through wired or wireless connections, or through any type of mesh, such as wired mesh, wireless mesh or any combination thereof, any type of private network and public network or any combination thereof, to send or receive relevant data from the dynamic mesh.
[0130] Both communication interface 214 and communication interface 224 can be configured as a one-way communication interface or a two-way communication interface as indicated by the arrow pointing from the source device 210 to the corresponding communication channel 230 of the destination device 220 in FIG2, and can be used to send and receive messages, etc., to establish a connection, acknowledge and exchange any other information related to the communication link and / or data transmission such as encoded bitstream transmission, etc.
[0131] Communication interface 214 can also transfer media files to a storage device. Communication interface 224 can also communicate with the storage device to retrieve media files. For example, the storage device can be a readable storage medium, a storage server, a storage gateway, an edge server, a content delivery network (CDN), etc. The CDN can receive and store the media files and then send them.
[0132] Decoder 223 is used to extract multiple media contents from a media file using metadata in the media file, and to present the multiple media contents in a fused form.
[0133] For example, decoder 223 includes decapsulation unit 2231 and decoding unit 2232. Decapsulation unit 2231 is used to obtain encoded data of various media content from the media file using metadata in the media file. Decoding unit 2232 is used to decode the encoded data of various media content to obtain various media content.
[0134] The post-processor 222 receives multiple media contents output by the decoder 223, performs fusion processing on the multiple media contents, and presents the multiple media contents in a fused form. For example, the post-processing performed by the post-processor 222 includes color format conversion (e.g., from YUV format to RGB format), color correction, retouching or resampling, or any other processing.
[0135] Display 221 is used to display media content in a merged format. Display 221 is or includes any category of display device for presenting media content in a merged format. For example, an integrated or external display or monitor. For example, the display may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other category of display.
[0136] Both encoder 213 and decoder 223 can be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented in part in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium, and one or more processors can be used to execute the instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.
[0137] The image acquisition device (e.g., an image acquisition device, an audio / video acquisition device) and encoder 213 can be integrated into a single physical device or located on different physical devices; this is not limited. For example, as shown in Figure 2, the source device 210 includes an image acquisition device 211 and an encoder 213, indicating that the image acquisition device 211 and encoder 213 are integrated into a single physical device. In this case, the source device 210 can also be referred to as an acquisition device. The source device 210 can be, for example, a mobile phone, tablet computer, computer, laptop computer, camcorder, camera, wearable device, in-vehicle device, terminal device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, extended reality (XR) device, or other image acquisition device. If the source device 210 does not include the image acquisition device 211, it means that the image acquisition device 211 and encoder 213 are two different physical devices, and the source device 210 can acquire raw images from other devices (e.g., image acquisition devices or image storage devices).
[0138] Furthermore, the display 221 and decoder 223 can be integrated into a single physical device or located on different physical devices; this is not limited. For example, as shown in FIG2, the destination device 220 includes the display 221 and decoder 223, indicating that the display 221 and decoder 223 are integrated into a single physical device. In this case, the destination device 220 can also be called a playback device, and it has the function of decoding and displaying merged media content. The destination device 220 can be, for example, a monitor, television, digital media player, video game console, in-vehicle computer, or other image display device. If the destination device 220 does not include the display 221, it means that the display 221 and decoder 223 are two different physical devices. After decoding the media file, the destination device 220 transmits the merged media content to other display devices (such as a television or digital media player) for display.
[0139] Furthermore, Figure 2 shows that the source device 210 and the destination device 220 are integrated on a single physical device, but they can also be set on different physical devices, without limitation.
[0140] For example, as shown in Figure 3(a), a schematic diagram of an encoding / decoding system, source device 210 is a server, exemplarily a server in a cloud system. Destination device 220 is a display of various possible forms. Source device 210 acquires video of a first scene and generates fused media content based on the video and original images within the video; alternatively, source device 210 obtains video from a storage device and generates fused media content based on the video and original images within the video. Source device 210 includes an encoding module and an encapsulation module. The encoding module encodes various media contents to obtain encoded data, and the encapsulation module encapsulates the encoded data and metadata to obtain a media file, which is then transmitted through a channel.
[0141] The destination device 220 includes a decapsulation module and a decoding module. The decapsulation module uses metadata in the media file to obtain encoded data of various media contents from the media file. The decoding module decodes the encoded data of the various media contents to obtain various media contents. Optionally, the destination device 220 also performs fusion processing on the various media contents to obtain fused media content. The destination device 220 performs operations such as storing and displaying the fused media content.
[0142] As another example, as shown in Figure 3(b), a schematic diagram of an encoding / decoding system, the source device 210 and the destination device 220 are integrated into the same device, such as a smartphone, tablet, computer, laptop, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, or extended reality (XR) device. In this case, the device has the capability to encode and decode dynamic meshes. For example, the source device 210 captures video of the user's real-world scene and generates fused media content based on the video and stored images. The destination device 220 displays the real-world scene in a virtual environment.
[0143] In these embodiments, the source device 210 or its corresponding functions and the destination device 220 or its corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. As described, the presence and division of different units or functions in the source device 210 and / or destination device 220 shown in FIG2 may vary depending on the actual device and application, which will be apparent to those skilled in the art.
[0144] The structure of the above-described encoding / decoding system is only illustrative. In some possible implementations, the encoding / decoding system may also include other devices, such as end-side devices or cloud-side devices. After the source device 210 acquires the original image, it preprocesses the original image to obtain a preprocessed image; and then transmits the preprocessed image to the end-side device or cloud-side device, which performs encoding / decoding on the preprocessed image.
[0145] Next, the media data processing process will be described with reference to the accompanying drawings. Figure 4(a) shows a flowchart of a media data processing method at the encoding end provided in this application. Here, the process of performing multiple media content processing steps by the source device 210 in Figure 2 will be used as an example. As shown in Figure 4(a), the method includes the following steps.
[0146] Step 410: Obtain various media content.
[0147] The multiple media content includes at least images and also other media content. For example, the multiple media content includes a first image and other first media content besides the first image. The first image may include a single image, and may also include a main image (e.g., the original image, a 3D image, etc.) and secondary images (images associated with the main image, such as thumbnails, etc.). The first media content includes a first video. Optionally, the first media content also includes at least one of a second image, audio, subtitles, or text. Optionally, the second image is a set of images. Optionally, the second image includes one or more images. Optionally, the second image may or may not include the first image. Optionally, the content of the first image and the content of the second image are different.
[0148] This application does not limit the source of various media content. The source device may collect media content itself, obtain media content from memory or other storage, or obtain media content from network resources. As described in the above embodiments, if the source device 210 carries an image acquisition device 211, the source device 210 acquires raw images through the image acquisition device 211. Optionally, the source device 210 may receive raw images acquired by other devices; or obtain raw images from memory or other storage in the source device 210. The raw images may include at least one of real-time acquired real-world images, images stored in the device, and images synthesized from multiple images. This embodiment does not limit the method of acquiring raw images or the type of raw images.
[0149] In some embodiments, the source device acquires different media content from different sources. For example, the source device captures images and videos, and acquires at least one of audio, subtitles, or text from memory or network resources. Alternatively, the source device acquires desired images and videos from a set of images. For example, the set of images includes a set of images that are sequential in the time domain. Or, the set of images includes multiple images with different exposure parameters. Or, the set of images includes multiple images with different exposure parameters captured at the same time.
[0150] Optionally, the original media content may also include information about the images or a set of images. This information may include at least one of exposure parameters, timing parameters, or format parameters.
[0151] This application does not limit the methods of acquiring various media content. Examples of methods for acquiring the first image and video are given below.
[0152] In some examples, a first image and a video containing the first media content are obtained based on a set of temporally consecutive images. That is, a video is obtained by selecting a first image from a set of temporally consecutive images and extracting images within a certain duration from it. The video may or may not contain the first image. The video comprises a set of temporally consecutive images.
[0153] In other examples, the video comprises a subset of images from a set of temporally consecutive images. For instance, the video contained in the first media content is obtained based on a first set of images from a set of temporally consecutive images. The first set of images comprises a subset of images from a set of temporally consecutive images. As another example, a second set of temporally consecutive images is obtained from the first set of images from a set of temporally consecutive images, and the video contained in the first media content is obtained from the second set of images. For example, the first set of images is obtained by performing frame interpolation on a set of temporally consecutive images. Alternatively, the second set of images is obtained from a subset of images in the first set obtained by performing frame interpolation on a set of temporally consecutive images.
[0154] In other examples, the first image is derived from a third image group within a set of temporally consecutive images. The third image group comprises multiple images with different exposure parameters.
[0155] For example, a first image is selected from an image group according to a preset time. Exemplarily, the preset time includes at least one of a midpoint time, a smiling time, or a time when the eyes are not closed. The first image is selected from the image group according to the midpoint time of the image group. Alternatively, the first image is selected from the image group according to the smiling time of the image group, then the first image includes a smiling image. Alternatively, the first image is selected from the image group according to the time when the eyes are not closed, then the first image includes an image when the eyes are not closed. The image group includes a set of images that are consecutive in the time domain or multiple images with different exposure parameters.
[0156] For example, select the first image from the image group based on a preset format. The original image includes a raw format image or a multi-exposure image. Preset formats include HDR (see BT.2100, ISO 22028-5, etc.).
[0157] For example, you can select video from an image group according to a preset time. The preset time includes all moments, every other frame, X seconds before and Y seconds after, where X and Y are equal to at least one of 1, 1.5, or other numbers. You can select images from the image group at all moments as video. Alternatively, you can select images from the image group every other frame as video.
[0158] Optionally, the selected multiple images can be upsampled or downsampled to form a video.
[0159] In other examples, the first image is obtained from a single frame.
[0160] Step 420: Generate media files based on various media content.
[0161] In one possible implementation, a first image is encoded to obtain image data, and an image file is generated based on the image data, the image file including the image data. First media content is encoded to obtain data of the first media content; the data of the first media content is then encoded into the image file to obtain a media file.
[0162] In one embodiment, metadata is encoded into a media file, which is used to indicate data of the first media content.
[0163] For example, reusing fields in an image file to set the data and metadata of the first media content makes the media file also include the data and metadata of the first media content. Alternatively, inserting the data and metadata of the first media content into an image file makes the media file also include the data and metadata of the first media content.
[0164] In another possible implementation, the first image is encoded to obtain image data; the first media content is encoded to obtain data of the first media content; and a media file is generated based on the image data, metadata, and data of the first media content. That is, instead of generating an image file, the image data, data of the first media content, and metadata are directly encoded into a media file. In this implementation, relative to the image file, new fields are added to the media file, and the data and metadata of the first media content are located in the newly added fields; or, fields from the image file are reused, and the data and metadata of the first media content are located in the fields of the image file.
[0165] This application does not limit the encoding format of media content such as images, videos, audio, subtitles, or text, nor the format of media files. For example, image codec format types include at least one of the following: Moving Pictures Experts Group (MPEG), Joint Photographic Experts Group (JPEG), Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or Versatile Video Codec (VVC).
[0166] Audio codec format types include at least one of the following: Pulse Code Modulation (PCM), Differential Pulse Code Modulation (DPCM), Adaptive Differential Pulse Code Modulation (ADPCM), Advanced Audio Coding (AAC), or Free Lossless Audio Codec (FLAC).
[0167] PCM is an uncompressed audio encoding method that converts analog audio signals into digital pulse signals. PCM encoding produces high-quality audio, but also generates a large amount of data.
[0168] DPCM is a lossy audio compression method that reduces the amount of data by predicting the difference between the current sample and the previous sample.
[0169] ADPCM improves upon DPCM by adaptively adjusting the quantization step size to enhance coding efficiency.
[0170] AAC is a lossy audio encoding method that combines the advantages of various audio encoding technologies, resulting in high compression efficiency and sound quality.
[0171] FLAC is a lossless audio encoding method that does not lose any audio information during compression and decompression.
[0172] The subtitle encoding / decoding format type includes at least one of Unicode, UTF-8, or ANSI.
[0173] The text encoding / decoding format type includes at least one of American Standard Code for Information Interchange (ASCII), GBK, Unicode, or UTF-8.
[0174] Media file formats include the JPEG (Joint Photographic Experts Group File Interchange Format, JFIF) and the High Efficiency Image Format (HEIF). JPEG is a standard for encoding and exchanging digital images. HEIF is an image file format based on HEVC video coding technology.
[0175] For example, the image data obtained by encoding the first image according to the HEIF standard is encoded to obtain an HEIF format image file. Alternatively, the image data obtained by encoding the first image according to the JPEG or JFIF standard is encoded to obtain a JPEG format image file.
[0176] As shown in Table 1, the JPEG image file format is illustrated.
[0177] Table 1
[0178] The JPEG file format is used to store and distribute a series of JPEG images that are played in sequence to simulate an animation effect.
[0179] The Start of Image (SOI) is used to mark the beginning of JPEG image data. Every JPEG file begins with an SOI marker.
[0180] Application Marker n (APPn), where APP0 stores JFIF information, including version number, thumbnail, etc. JFIF is a common format for JPEG, used to define the structure of JPEG data at the file level.
[0181] The quantization table (DQT) defines the quantization matrix used in the JPEG compression process. Quantization is one of the key steps in converting image data from the spatial domain to the frequency domain, determining the quality of the compressed image and the file size.
[0182] Huffman tables (DHT) are used in JPEG compression with Huffman coding, a variable-length coding technique used to further reduce the size of image data. DHTs optimize based on the statistical properties of image data to improve compression efficiency.
[0183] The Restart Interval (DRI) is used to insert a restart marker during JPEG compression so that decoding can restart from a specific point during decoding, which helps to restore synchronization in case of decoding errors.
[0184] The Start of Frame (SOF) is used to mark the beginning of a JPEG image frame and contains information about the image size, number of color components, sampling factor, etc.
[0185] The Start of Scan (SOS) marker indicates the start of a JPEG image scan and specifies the color components to be scanned and the scanning mode. The actual compressed data follows the SOS marker.
[0186] Compression data is the actual data of a JPEG image, obtained after compression processes such as quantization and Huffman coding. This data is the core part of the image, containing all its information.
[0187] The End of Image (EOI) is used to mark the end of JPEG image data. Every JPEG file ends with an EOI marker.
[0188] As shown in Table 2, the HEIF format image file format is illustrated.
[0189] Table 2
[0190] The file type (ftyp) is used to identify the file's format and version, ensuring that the decoder can correctly parse the file. For example, for an HEVC-encoded video file, the ftyp field might indicate that it is an ISO Base Media File Format (isom) file, and version and compatibility flags would also be included.
[0191] The meta tag contains various information about the media file, such as author, title, date, and copyright. Optionally, the meta tag may also include technical information such as encoding details, color space, and timestamps.
[0192] The processor or processor type (HDLR) indicates the type of software or hardware processor that handles the media content. For example, it indicates that the video is processed by a specific decoder or player.
[0193] The Primary Item Box (pitm) is used to identify the main item or content in a media file. In files containing multiple media tracks (such as video, audio, subtitles, etc.), the pitm field indicates which track is the primary track or should be played first.
[0194] Data items (idat) are used within a context that may depend on their implementation. In some cases, they may be used to store additional data related to a media project.
[0195] The Item Info Box (iinf) contains detailed information about a specific item in a media project. For example, it may contain information such as the resolution, frame rate, and bitrate of a video track.
[0196] An item information entry (infe) is used to specifically describe an information item. For example, hvc1 indicates that this is an HEVC-encoded video data item.
[0197] The Item Location Box (iloc) provides the physical location information of data items in a media file. This helps the decoder quickly locate and read the required media data.
[0198] The Item Properties Box (Iprp) contains information about the properties of specific items in a media project. For example, it may contain technical details about color spaces, encoding settings, etc.
[0199] The Item Property Container Box (ipco) contains the actual item property information.
[0200] Color attributes or color information (co1r) provide information about the colors of a media item, such as color space, gamma value, etc. Under Ipco, repetition may indicate different contexts or uses.
[0201] Item references (irefs) provide referencing information between different items in a media file. For example, they might be used to indicate which video track a subtitle track is associated with.
[0202] Under `iref`, `dimg` may indicate an image reference or a similar meaning (the specific meaning may depend on the context) referencing image data in a media file.
[0203] A Media Data Box (mdat) contains the actual media data, such as video frames and audio samples. This is the most important data part of a media file.
[0204] Data1 could be a placeholder or example name. In this context, Data1 likely refers to the actual image bitstream data stored in the mdat box. This is the core data the decoder needs to decode and display the image.
[0205] In some embodiments, optionally, the metadata includes tag metadata. Tag metadata is used to indicate that the media file includes first media content.
[0206] The source device determines the tag metadata based on the content of the first media.
[0207] For example, if the source device determines that the acquired media content includes a first media content other than the first image, it determines that the media file contains tag metadata.
[0208] In other words, a media file containing tag metadata indicates that it contains media content other than the first image. Conversely, a media file without tag metadata indicates that it contains only the first image, meaning it contains no other media content besides the first image. In other words, a media file containing tag metadata indicates that it contains media content other than the directly decoded image. Conversely, a media file without tag metadata indicates that it contains only the directly decoded image, meaning it contains no other media content besides the directly decoded image.
[0209] For example, different values of the tag metadata contained in a media file indicate whether the media file contains media content other than the first image. For instance, if the tag metadata contained in a media file has a first value (e.g., a value of 1), it means that the media file contains media content other than the first image. If the tag metadata contained in a media file has a second value (e.g., a value of 0), it means that the media file contains only the first image, that is, the media file does not contain media content other than the first image.
[0210] The following describes how to set tag metadata in media files for different media file formats.
[0211] When the media file is in HEIF format, the following methods are used to include tag metadata in the media file.
[0212] Method 1: The tag metadata is located in the `ftyp` field of the media file. For example, as shown in Table 3, the tag metadata is located in the data segment corresponding to the `ftyp` field in the media file. That is, in the data segment corresponding to the `ftyp` field of the HEIF format image file shown in Table 2.
[0213] Table 3
[0214] Method 2: The tag metadata is located in the meta field of the media file. For example, as shown in Table 4, the tag metadata is located in the data segment corresponding to the meta field in the media file. That is, in the data segment corresponding to the meta field of the HEIF format image file shown in Table 2.
[0215] Table 4
[0216] Method 3: The meta field includes the iinf subfield. The tag metadata is located in the data segment corresponding to the type in the iinf subfield, excluding the image codec format type.
[0217] Among them, the image codec format type includes at least one of MPEG, JPEG, AVC, HEVC or VVC.
[0218] The types mentioned above, excluding image encoding / decoding format types, include at least one of uri, mime, it35, xml, exif, or live.
[0219] For example, the tag metadata is located in the data segment corresponding to the type URI in the iinf subfield, excluding the image codec format type.
[0220] For example, the tag metadata is located in the data segment corresponding to the type MIME in the iinf subfield, excluding the image codec format type.
[0221] For example, the tag metadata is located in the data segment corresponding to type it35 in the iinf subfield, excluding image codec format types.
[0222] For example, the tag metadata is located in the data segment corresponding to the type xml in the iinf subfield, excluding image encoding / decoding format types.
[0223] For example, the tag metadata is located in the data segment corresponding to the type exif in the iinf subfield, excluding image codec format types.
[0224] For example, as shown in Table 5, the tag metadata is located in the data segment corresponding to the first type of the iinf subfield, excluding the image codec format type, such as the data segment corresponding to "live". That is, the meta field of the HEIF format image file shown in Table 2 contains the data segment corresponding to the first type under the iinf subfield. For example, the tag metadata is located in the idat field corresponding to the item of the type other than the image codec format type in the iinf subfield.
[0225] Table 5
[0226] Method 4: The meta field includes the iloc subfield, iprp subfield, ifef subfield, and idat subfield. Tag metadata resides in at least one of the iloc, iprp, ifef, or idat subfields.
[0227] For example, the tag metadata is located in the data segment corresponding to the iloc subfield.
[0228] For example, the tag metadata is located in the data segment corresponding to the IPRP subfield.
[0229] For example, the tag metadata is located in the data segment corresponding to the ifef subfield.
[0230] For example, the tag metadata is located in the data segment corresponding to the idat subfield.
[0231] For example, as shown in Table 6, the tag metadata is located in the data segment corresponding to the idat subfield. That is, in the data segment corresponding to the idat subfield contained in the meta field of the HEIF format image file shown in Table 2.
[0232] Table 6
[0233] Method 5: In the HEIF format image file shown in Table 2, the ftyp, meta, and mdat fields are used as items in the image file. A first field is added to the media file to identify and tag metadata. For example, the media file adds an item of type 1.
[0234] When the media file is in JPEG format, the inclusion of tag metadata in the media file includes the following methods.
[0235] Method 1: The tag metadata is located in the APPn field of the media file. For example, as shown in Table 7, the tag metadata is located in the APPn field of the media file. That is, in the data segment corresponding to the APPn added to the JPEG format image file shown in Table 1.
[0236] Table 7
[0237] Method 2: The tag metadata is located in the MPF of the media file. For example, as shown in Table 8, the tag metadata is located in the MPF of the media file, specifically in the data segment corresponding to the APPn added to the JPEG format image file shown in Table 1.
[0238] Table 8
[0239] For example, the tag metadata is located in at least one of the tags, tags, or payloads of an APPn in the media file. For example, the tag metadata is located in an APPn of the first type in the media file. Or, the tag metadata is located in an APPn in the media file that contains a first preset character.
[0240] In other embodiments, optionally, the metadata also includes at least one of attribute metadata or location metadata. Attribute metadata is used to indicate the attributes of the first media content included in the media file. Attribute metadata is used to indicate what kinds of media content the media file contains. Location metadata is used to indicate the location of the first media content in the media file. Location metadata is used to indicate the location of the media content included in the media file.
[0241] In some embodiments, the metadata also includes time-related metadata. Time-related metadata is used to indicate the temporal relationship between the first media content and the first image included in the media file.
[0242] For example, time-related metadata is used to indicate the temporal relationship between the first video and the first image included in the first media content.
[0243] In some embodiments, the time-related metadata is obtained from first time-domain information and second time-domain information. The first time-domain information is obtained based on the time when the first video of the first media content was acquired. The second time-domain information is obtained based on the time when the first image was acquired.
[0244] In some embodiments, time-related metadata indicates one or more of the following:
[0245] 1) The time-related metadata indicates the difference in start time between the first image and the first video.
[0246] For example, the difference in start time between the first image and the video is determined based on the first time domain information and the second time domain information.
[0247] 2) The time-related metadata indicates the difference between the starting frame of the first image and the first video.
[0248] For example, the time-related metadata includes the frame number corresponding to the first image in the first video.
[0249] For example, the difference between the starting frame of the first image and the first video is determined based on the frame number extracted from the first time domain information and the second time domain information.
[0250] 3) Time-related metadata indicates the timestamp of the first image.
[0251] 4) Time-related metadata indicates the start timestamp of the first video.
[0252] 5) Time-related metadata indicates the duration of the first video.
[0253] Optionally, the first media content may also include a second image, with time-related metadata examples indicating the following:
[0254] 1) The time-related metadata indicates the difference in start time between the second image and the first video.
[0255] 2) The time-related metadata indicates the difference between the starting frame of the second image and the first video.
[0256] 3) Time-related metadata indicates the timestamp of the second image.
[0257] For example, the first image is an image selected from the first video according to a first preset rule. The first preset rule refers to a midpoint in the duration of the first video. Alternatively, the first preset rule may be user-specified.
[0258] For example, the second image is an image selected from the first video or the first image according to a second preset rule. The second preset rule refers to a preferred moment in the first video. Alternatively, the second preset rule means that when the user specifies a new first image, the first image that was previously selected will be used as the second image.
[0259] Optionally, the first time domain information includes a preset moment or a preset time. The preset moment includes at least one of the intermediate moment or the moment when the eyes are not closed. The preset time includes at least one of the following: all moments, frame intervals, X seconds before and Y seconds after, where X and Y are 1, 1.5, or other values.
[0260] The second time domain information includes the time when a frame of an image is acquired or the time when the first image is acquired from a frame of an image.
[0261] In other embodiments, the metadata also includes transformation-related metadata. Optionally, the transformation-related metadata is used to determine the transformation to be performed on the first media content and / or the first image. Optionally, the transformation-related metadata is used to indicate the manner in which the first media content and / or the first image is transformed.
[0262] For example, the conversion-related metadata is obtained from first format information and second format information. The first format information is obtained based on consecutive images of the video from which the first media content is acquired. The second format information is obtained based on consecutive images from which the first image is acquired.
[0263] The associated metadata is obtained from the data associated with the first image.
[0264] For example, the first format information includes relevant information from the image group, and / or preset formats, and / or format information obtained during the video extraction process.
[0265] For example, the second format information includes information related to a frame of image, and / or a preset format, and / or format information obtained during the extraction of the first image.
[0266] The conversion associated metadata can include the format information of the first image.
[0267] For example, the conversion associated metadata indicates whether the format of the first image is SDR or HDR.
[0268] For example, the conversion associated metadata indicates how the format of the first image is converted from an SDR image to an HDR image.
[0269] For example, the conversion associated metadata indicates how the format of the first image is converted from an HDR image to an SDR image.
[0270] For example, the conversion associated metadata indicates the dynamic range, peak brightness, or peak brightness to reference white ratio information of the first image as HDR.
[0271] For example, the transformation-related metadata indicates how the first image is transformed from an HDR image to an HDR image, and the two have different dynamic ranges.
[0272] The conversion-related metadata can include the format information of the first video.
[0273] For example, the conversion associated metadata indicates whether the first video is SDR or HDR.
[0274] For example, the conversion associated metadata indicates how the format of the first video is converted from SDR video to HDR video.
[0275] For example, the conversion associated metadata indicates how the format of the first video is converted from HDR video to SDR video.
[0276] For example, the conversion associated metadata indicates the dynamic range, peak brightness, or peak brightness to reference white ratio information of the first video if it is HDR.
[0277] For example, the conversion-related metadata indicates how the first video was converted from an HDR video to an HDR video.
[0278] Optionally, the conversion-related metadata includes at least one of color gamut information, resolution, or rotation.
[0279] Optionally, the conversion-related metadata may also include the conversion methods for images and videos, such as the conversion methods between HDR and SDR, or the conversion relationships between HDR content with different dynamic ranges.
[0280] Optionally, the conversion relationship between different dynamic range HDR content can be converted from one dynamic range, peak brightness, or peak brightness to reference white ratio information to another dynamic range, peak brightness, or peak brightness to reference white ratio.
[0281] Optionally, the conversion method between HDR and SDR, or the conversion relationship between HDR content with different dynamic ranges, can be achieved using parametrically recorded tone mapping curves and / or associated color conversion matrices.
[0282] Optionally, the conversion method between HDR and SDR, or the conversion relationship between HDR content with different dynamic ranges, can be described using 1D, 2D, or 3D lookup tables and / or conversion matrices.
[0283] The following describes the location of time-related metadata and transformation-related metadata in the media file.
[0284] Optionally, the source device may set at least one of time-related metadata or transformation-related metadata in the media file, depending on the method of setting tag metadata in the media file.
[0285] In some embodiments, the fields that identify tag metadata, the fields that identify time-related metadata, and the fields that identify transformation-related metadata are the same.
[0286] For example, as shown in Table 9, the tag metadata, time-related metadata, and transformation-related metadata are all located in the ftyp field of the media file.
[0287] Table 9
[0288] For example, as shown in Table 10, tag metadata, time-related metadata, and transformation-related metadata are all located in the meta field of the media file.
[0289] Table 10
[0290] For example, tag metadata, time-related metadata, and transformation-related metadata all reside in the same field within the media file, and the types of the fields identifying tag metadata, time-related metadata, and transformation-related metadata are identical. As shown in Table 11, tag metadata, time-related metadata, and transformation-related metadata are all located in the infe subfield of the iinf subfield contained in the meta field of the media file. For instance, tag metadata is located in the data segment corresponding to the type live of the infe subfield. Time-related metadata is located in the data segment corresponding to the type live of the infe subfield. Transformation-related metadata is located in the data segment corresponding to the type live of the infe subfield.
[0291] Table 11
[0292] In some embodiments, at least one of the fields identifying tag metadata, identifying time-related metadata, and transformation-related metadata is different.
[0293] For example, the fields identifying time-related metadata and transformation-related metadata are the same. The fields identifying tag metadata are different from the fields identifying time-related metadata and transformation-related metadata, respectively. As shown in Table 12, tag metadata is located in the ftyp field of the media file, while time-related metadata and transformation-related metadata are both located in the meta field of the media file.
[0294] Table 12
[0295] Optionally, the tag metadata is located in the live type subfield of the iinf field, and the time-related metadata and conversion-related metadata are located in the idat field.
[0296] For example, the fields identifying tag metadata, time-related metadata, and transformation-related metadata are all different. As shown in Table 13, tag metadata is located in the iinf subfield contained in the meta field of the media file. For instance, tag metadata is located in the data segment corresponding to the live type in the iinf subfield, excluding the image codec format type. Time-related metadata is located in the ifef subfield contained in the meta field of the media file, or, alternatively, in the iprp subfield contained in the meta field of the media file. Transformation-related metadata is located in the iloc subfield contained in the meta field of the media file.
[0297] Table 13
[0298] For example, tag metadata, time-related metadata, and transformation-related metadata all reside in the same field within the media file. However, the types of the fields identifying tag metadata, time-related metadata, and transformation-related metadata are all different. As shown in Table 14, tag metadata, time-related metadata, and transformation-related metadata are all located in the infe subfield of the iinf subfield contained in the meta field of the media file. For instance, tag metadata is located in the data segment corresponding to the type live of the infe subfield. Time-related metadata is located in the data segment corresponding to the type mime of the infe subfield. Transformation-related metadata is located in the data segment corresponding to the type xml of the infe subfield.
[0299] Table 14
[0300] For example, tag metadata, time-related metadata, and transformation-related metadata may all reside in the same field within the media file; however, at least one of the types of the fields identifying tag metadata, time-related metadata, and transformation-related metadata may differ. As shown in Table 15, tag metadata, time-related metadata, and transformation-related metadata are all located in the infe subfield of the iinf subfield contained within the meta field of the media file. For instance, tag metadata is located in the data segment corresponding to the type live of the infe subfield. Time-related metadata is located in the data segment corresponding to the type live of the infe subfield. Transformation-related metadata is located in the data segment corresponding to the type xml of the infe subfield.
[0301] Table 15
[0302] Optionally, metadata includes tag metadata, location metadata, time-related metadata, and transformation-related metadata. Metadata does not include attribute metadata. The fields identifying tag metadata are the same as those identifying location metadata, time-related metadata, and transformation-related metadata. As shown in Table 16, the ftyp fields identifying tag metadata and location metadata do not include attribute metadata.
[0303] Table 16
[0304] As shown in Table 17, the meta fields of the identifier tag metadata, identifier location metadata, time-related metadata, and transformation-related metadata do not include attribute metadata.
[0305] Table 17
[0306] The fields for identifying tag metadata differ from those for identifying time-related metadata and transformation-related metadata. As shown in Table 18, tag metadata is located in the ftyp field of the media file, while time-related metadata and transformation-related metadata are both located in the meta field of the media file.
[0307] Table 18
[0308] In some embodiments, metadata is set according to the association relationships between fields of identifier tag metadata, fields of identifier time-related metadata, and fields of identifier transformation-related metadata. The fields of identifier tag metadata, fields of identifier time-related metadata, and fields of identifier transformation-related metadata are associated. Tag metadata, time-related metadata, and transformation-related metadata are placed in the data segments corresponding to the associated fields.
[0309] For example, fields of identifier tag metadata are associated with fields of identifier time-related metadata, and fields of identifier time-related metadata are associated with fields of identifier transformation-related metadata.
[0310] For example, as shown in Table 13, the iinf field of the identifier tag metadata is associated with the iloc field of the identifier transformation-related metadata. The ifef field of the identifier transformation-related metadata is associated with the ifef field of the identifier time-related metadata.
[0311] For example, the types of identifier tag metadata can be associated with the types of time-related metadata and the types of identifier transformation-related metadata. For instance, the type refers to the type or tag of box or fullbox. For instance, the first type of item in the identifier tag metadata is associated with the IPRP of the time-related metadata. For instance, the type IPRP of the time-related metadata is associated with the type iloc of the identifier transformation-related metadata.
[0312] For example, fields identifying time-related metadata and fields identifying transformation-related metadata include at least one of the IPRP subfield or the iref subfield; fields identifying tag metadata are associated with at least one of the IPRP subfield or the iref subfield. Alternatively, fields identifying time-related metadata and fields identifying transformation-related metadata include the IPRP subfield, and fields identifying tag metadata are associated with the IPRP subfield. Or, fields identifying time-related metadata and fields identifying transformation-related metadata include the iref subfield, and fields identifying tag metadata are associated with the iref subfield. Or, fields identifying time-related metadata include the IPRP subfield, fields identifying transformation-related metadata include the iref subfield, and fields identifying tag metadata are associated with both the IPRP subfield and the iref subfield.
[0313] For example, as shown in Table 19, the fields that identify time-related metadata and the fields that identify transformation-related metadata include the IPRP subfield, and the iinf field that identifies tag metadata is associated with the IPRP subfield.
[0314] Table 19
[0315] In some embodiments, metadata is set through a single item association rule. The type of the metadata is identified as the first type of item in the iinf field. The iinf field and the iloc field are associated through the item, and the field that identifies the transformation associated metadata is iloc. The iinf field and the iprp field are associated through the item, and the field that identifies the time-related metadata is iprp.
[0316] In some embodiments, metadata is set through association rules for multiple items. The types of identifier tag metadata, identifier time-related metadata, and identifier transformation-related metadata are associated through item association rules. For example, as shown in Table 22, image data items use item items of type hvc1, and tag metadata uses item items of type first; the two items are associated through iref.
[0317] In some embodiments, metadata is set through multiple items and a single association rule. The types of identifier tag metadata, identifier time-related metadata, and identifier transformation-related metadata are associated through item association rules. For example, as shown in Table 22, image data items use item items of type hvc1, and tag metadata uses item items of type 1; the two items are associated through iref. Alternatively, image data items use item items of type hvc1, video data items use item items of type MIME, URI, or type 1, and tag metadata items use item items of a different type than video (type 1); the three items are associated through iref. Simultaneously, these image data items, video data items, and tag metadata items are associated with the iloc and iprp fields through a single item association rule.
[0318] In some embodiments, association is achieved through predefined rules for the brand included in ftyp. For example, as shown in Table 18, the tag metadata is located in the ftyp field of the media file, while the transformation-related metadata is located in the meta field of the media file. The ftyp and meta fields belong to two separate items; when these two items are associated, the field identifying the tag metadata and the field identifying the transformation-related metadata are associated. The first type of brand needs to include a specific combination of fields from the meta tag.
[0319] In some embodiments, metadata is set according to the association relationships between the types of tag metadata, tag time-related metadata, and tag transformation-related metadata. The types of tag metadata, tag time-related metadata, and tag transformation-related metadata are associated. The tag metadata, time-related metadata, and transformation-related metadata are placed in the data segment corresponding to the associated type.
[0320] For example, the type of tag metadata is associated with the type of tag time-related metadata, and the type of tag time-related metadata is associated with the type of tag transformation-related metadata.
[0321] For example, the tag tag metadata type iprp is associated with the tag time-related metadata type ifre. The tag time-related metadata type ifre is associated with the tag transformation-related metadata type iloc.
[0322] For example, the types of metadata are marked by marking the types of metadata associated with time and the types of metadata associated with transformation.
[0323] The tag metadata type is the first type in the iinf field. The iinf field is associated with the iloc field, and the tag location associated metadata type is iloc. The iinf field is associated with the iprp field, and the tag time or attribute associated metadata type is iprp. The iinf field is associated with the idat field, and the tag transformation associated metadata.
[0324] The type of the image or video tag is hvc1, mime, or uri in the iinf field. The iinf field is associated with the iloc field. The type of the metadata associated with the location tag is iloc. The iinf field is associated with the IPrp field. The type of the metadata associated with the time or attribute tag is IPrp. The iinf field is associated with the idat field. The metadata associated with the transformation tag is IPrp.
[0325] The type of the metadata tag is the first type in the iinf field. For images or videos, the type is hvc1, mime, or uri in the iinf field. The iinf field is associated with the iloc field; the type of location-related metadata is iloc. The iinf field is associated with the IPrp field; the type of time or attribute-related metadata is IPrp. The iinf field is associated with the idat field; and the conversion-related metadata is tagged. The first type uses any abbreviation or full name related to burst shooting or moving images, such as live, motion, movie photo, continuous shooting, live, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, or burst mode.
[0326] The first field can use any of the following abbreviations or full names related to continuous shooting and moving images: live, motion, movie photo, continuous shooting, live, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, burst mode, etc.
[0327] The first type of media content includes multiple media types, and the media file includes time-related metadata and transformation-related metadata for each type of media content. For example, as shown in Table 20, the tag metadata is located in the `iinf` subfield contained in the `meta` field of the media file. For instance, the tag metadata is located in the data segment corresponding to type `live` in the `iinf` subfield, excluding the image codec format type. The time-related metadata for the first type of media content is located in the `iprp` subfield contained in the `meta` field of the media file, and the transformation-related metadata is located in the `iloc` subfield contained in the `meta` field of the media file. The time-related metadata and transformation-related metadata for the second type of media content are located in the `idat` subfield contained in the `meta` field of the media file.
[0328] Table 20
[0329] The first type of media content includes multiple media types, and the media file includes time-related metadata and transformation-related metadata for each type of media content. For example, as shown in Table 21, the tag metadata for the first type of media content is located in the data segment corresponding to type `live` in the `iinf` subfield, excluding the image codec format type. The transformation-related metadata for the first type of media content is located in the `iloc` subfield contained in the `meta` field of the media file. The tag metadata, time-related metadata, and transformation-related metadata for the second type of media content are located in the `idat` subfield contained in the `meta` field of the media file.
[0330] Table 21
[0331] Optionally, the metadata for the first media content may reside in two or more fields. For example, as shown in Table 21, location metadata may reside in the iloc and idat fields.
[0332] Add the first media content to the media file. The method for setting the data of the first media content in the media file varies depending on the media file format.
[0333] When the media file is in HEIF format, the first media content can be set in the media file in the following ways.
[0334] In Method 1, the data for the first media content is located in the `mdat` field. As shown in Table 22, the first media content is video, and the video stream is located in the `mdat` field.
[0335] Table 22
[0336] Method 2: The data for the first media content is located in the idat subfield of the meta field. As shown in Table 23, the first media content is video, and the video stream is located in the idat subfield of the meta field.
[0337] Table 23
[0338] Method 3: The data of the first media content is located in the moov field. As shown in Table 24, for example, the first media content is video, and the video stream is located in the moov field. In an optional embodiment, the moov packet is obtained from the media content data (e.g., video MP4) and added to the media file. The image data of the first image is added to the mdat field of the media file. The offset metadata of moov is modified, with moov indicating the offset to start from the idat start position.
[0339] Table 24
[0340] Method 4: The data for the first media content is carried in other fields at the same level as ftyp, meta, mdat, moov, etc. As shown in Table 25, if the first media content is video, the video stream is carried in the free field. Alternatively, the video stream can be carried in any field related to any abbreviation or full name of continuous shooting or moving images, such as live, motion, movie photo, continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, continuous shooting, burst mode, etc., as shown in the figure for live.
[0341] Table 25
[0342] For example, primary media content data includes video streams.
[0343] When the media file is in JPEG format, the data containing the first media content in the media file can include the following methods.
[0344] Method 1: Add the data of the first media content to the end of the first image file, as shown in Table 26. It should be understood that APP11 is for illustrative purposes only.
[0345] Table 26
[0346] Method 2: Add the first media content from the JUMBF package, as shown in Table 27. It should be understood that APP11 is for illustrative purposes only.
[0347] Table 27
[0348] Method 3: Add the first media content to the end of the media file, as shown in Table 28.
[0349] Table 28
[0350] Method 4: Add the first media content after the independent sub-file format in the media file. See Tables 29 to 31.
[0351] Table 29
[0352] Table 30
[0353] Table 31
[0354] APPX is used to identify and tag metadata.
[0355] JUMBF (JPEG User Marker Blocks for Features) may be an application- or standard-specific term used to describe user-defined marker blocks in a JPEG image. These marker blocks are used to store various additional information or features related to the image.
[0356] APP11 JUMBF1 may represent the first JUMBF block, used to store certain characteristics or information.
[0357] APP11 JUMBF2 may represent a second JUMBF block used to store another specific feature or information.
[0358] APP11 JUMBF3 may represent the third JUMBF block, and so on.
[0359] APP11 JUMBF4 may represent the fourth JUMBF block.
[0360] For suitability, the above APP11 JUMBF1~4 are data of the primary media content.
[0361] In some embodiments, when the media file is in a JPEG format, the JPEG-formatted media file includes content related to the HEIF file format.
[0362] By encapsulating HEIF-related content within JPEG media files, unified management of media content is achieved. This facilitates the expansion of various media content formats, ensures media file standards are as compatible as possible with existing standards, and enhances the convenience of media content management.
[0363] When the media file format conforms to the JPEG file format, the media file includes a new APPn. The APPn tag indicates HEIF. An example, as shown in Table 32, is a general structure for a media file. The media file includes an image. The media file includes an APPn:heif. Optionally, the media file may also include video.
[0364] Table 32
[0365] Table 33 shows the structure of a multi-layer HDR image-related media file. The media file includes various media content. For example, a media file may include multiple images and videos.
[0366] Table 33
[0367] APPn includes APP start, APP length, tag, and payload.
[0368] The APPn marker indicates a data structure related to HEIF, meaning that the APPn includes content related to the HEIF file format. For example, the APPn marker indicates HEIF, HEIC, or MIAF, etc. Optionally, the APPn marker indicates Urn:iso:std:iso:23008:-12, indicating that the APPn includes content related to the HEIF file format.
[0369] The data structures associated with the HEIF file format include the File-type Box, Metabox, and Media Data Box Header.
[0370] As shown in Table 34, the payload of APPn includes the File-type Box, Metabox, and Mdat Box Header.
[0371] Table 34
[0372] Optionally, as shown in Table 1, JPEG format image files include SOI and EOI, and between SOI and EOI, DQI, Compression Data, etc., are also included. Therefore, the payload of APPn also includes SOI. For example, as shown in Table 34, the payload field includes SOI.
[0373] Metadata is located in the APPn of the media file. If the APPn is marked as HEIF, the metadata is located in the payload of the APPn in the media file. Optionally, the media file may contain metadata in the manner described in Tables 3 to 6 of the above embodiments.
[0374] The new APPn payload contains structures related to the HEIF file, such as the File-type Box, Metabox, and Mdat Box Header. These structures allow linking to or access to the first image and / or the first media content. For example, after removing the portion before the new APPn payload in the media file, the remaining portion can be processed as a complete HEIF structure (in other words, the new APPn payload combined with subsequent data can be processed as a complete HEIF structure). For example, an SOI can be added after the Mdat Box Header in the new APPn payload. For example, some or all of the data following the new APPn in the current media file can be considered as included in the Mdat data (data[]).
[0375] For example, as shown in Table 35, the SOI in the APPn payload includes data following the APPn in JPEG file format, marked with HEIF. The SOI in the APPn payload also includes other APPn (e.g., APPn:MPF, in this example, APPn:heif before APPn:MPF), DQT, Compressed Data, and EOI, as well as other complete files. For example, the SOI in the APPn payload also includes a complete JPEG file format media file and a complete MP4 file format media file.
[0376] Table 35
[0377] For example, as shown in Table 35, the four columns on the left show the payload of APPn in the media file and the subsequent data structure, while the three columns on the right show the result of heif encapsulation / management obtained by combining the heif-related structures in the payload of APPn with the above data structure.
[0378] This adds metadata related to photo / video and other media content to the image file. By improving the structure of media files and setting metadata, multiple media contents can be processed uniformly, expanding the forms of media content fusion. For example, as shown in Table 36.
[0379] Table 36
[0380] The storage order and quantity of the metadata in the table above are not restricted.
[0381] The above embodiments illustrate the media data processing method at the encoding end. The media data processing method at the decoding end will be described below.
[0382] The following explanation uses the example of the target device 220 in Figure 2 performing various media content processing procedures, as shown in Figure 4(b), which is a flowchart illustrating a media data processing method at a decoding end provided in this application. The method includes the following steps.
[0383] Step 430: Obtain the media file. The media file includes the image data of the first image, metadata, and the data of the first media content.
[0384] The destination device receives a media file from the source device. Alternatively, the destination device generates the media file, or retrieves the media file from its own memory or other storage. The description of the media file, the first image, the first media content, and the metadata is as explained in step 410.
[0385] Step 440: Obtain the first image from the media file.
[0386] In some embodiments, a first image is obtained by decoding a media file. The media file is directly decoded as image data to obtain the first image.
[0387] In some embodiments, image data in a media file is decoded to obtain a first image. The process involves acquiring image data from a media file, decoding the image data, and obtaining the first image.
[0388] Optionally, data blocks in the media file are deleted, and the resulting file is decoded to obtain the first image. For example, the unnecessary parts of the media file are determined based on the format of the first image (or the media file format), and then the unnecessary parts are deleted (e.g., deleting the APP-related content in Tables 19 to 21). The remaining parts are then decoded to obtain the first image. Optionally, since the positions of fields in the remaining content have changed due to the deletion of some content in the media file, the positions of the remaining content in the media file are modified, and decoding is performed based on the modified content to obtain the first image. Optionally, the data of the first media content in the media file can also be deleted, and then the remaining content is decoded to obtain the first image.
[0389] Optionally, the fields required for decoding the first image are obtained from the media file according to the format of the first image (or the media file format), and the data of the first image is decoded based on the obtained fields to obtain the first image.
[0390] For example, when the media file is in the HEIF file format, the first image is obtained from the HEIF file file in one or more of the following ways.
[0391] Method 1: Treat the entire HEIF file as a static image file, decode it according to the HEIF standard, and obtain the first image.
[0392] Method 2: Obtain the primary image and related fields (ftyp, meta, mdat) from the media file as an image file and decode it to obtain the first image.
[0393] Method 3: Obtain other types of standard files (e.g., ISO 21496-1, ISO 22028-5, HEIF tmap, etc.) from the media file as image files and decode them to obtain the first image.
[0394] Method 4: Remove all fields except ftyp, meta, and mdat from the media file and decode it as an image file to obtain the first image.
[0395] Method 5: Remove all fields except ftyp, meta, and mdat from the media file, modify some data in the meta field, and then decode it as an image file to obtain the first image.
[0396] When the media file is in JPEG format, the first image is obtained from the JPEG media file using one or more of the following methods.
[0397] Method 1: Treat the entire JPEG file as a still image file, decode it according to the JPEG standard, and obtain the first image.
[0398] Method 2: Obtain the complete JPEG structure confirmed from the media file, decode it into an image file, and obtain the first image.
[0399] Method 3: Confirm the complete JPEG structure obtained from the media file, remove the APP except for APP0 JFIF, and / or remove the data with the file extension of the complete structure file, and decode the obtained part as an image file to obtain the first image.
[0400] Method 4: Obtain other types of standard files (e.g., ISO 21496-1, ISO 22028-5, HEIF tmap, JPEG standard extension standards, etc.) from the media file as image files and decode them to obtain the first image.
[0401] Step 450: Obtain the first media content from the media file.
[0402] For example, when the media file is in HEIF format, the first media content can be obtained from the media file in the following ways.
[0403] Method 1, as shown in Table 22, has the data for the first media content located in the `mdat` field. Since the first media content is video, the video stream is obtained from the `mdat` field.
[0404] Method 2, as shown in Table 23, involves retrieving the data for the first media content from the `idat` subfield within the `meta` field. Since the first media content is video, the video stream is obtained from the `idat` subfield of the `meta` field.
[0405] Method 3, as shown in Table 24, has the data for the first media content located in the moov field. Since the first media content is video, the video stream is obtained from the moov field.
[0406] Method 4, as shown in Table 25, involves the data of the first media content being carried in the first type or first field of the same level as meta or mdat, such as free or other types.
[0407] For example, when the media file is in JPEG format, obtaining the first media content from the media file includes the following methods.
[0408] Method 1, as shown in Table 26, obtains the first media content from the end position of the image file of the first image.
[0409] Method 2, as shown in Table 27, obtains the first media content from the JUMBF package.
[0410] Method 3, as shown in Table 28, retrieves the first media content from the end of the media file.
[0411] Method 4, as shown in Tables 29 to 31, involves adding the first media content after the independent sub-file format in the media file.
[0412] Optionally, the metadata includes tag metadata. Tag metadata is used to indicate that the media file includes first media content. The first media content is retrieved from the media file based on the tag metadata.
[0413] In some embodiments, the destination device determines whether a media file contains first media content based on tag metadata. If the media file contains tag metadata, it is determined that the media file contains first media content. If the media file does not contain tag metadata, it is determined that the media file does not contain first media content.
[0414] Optionally, the relevant fields or types carried by the above-mentioned tag metadata conform to the first type or the first field.
[0415] When the media file is in HEIF format, the following methods can be used to obtain tag metadata from the media file.
[0416] Method 1, as shown in Table 3, involves retrieving the tag metadata from the `ftyp` field in the media file.
[0417] Method 2, as shown in Table 4, involves retrieving the tag metadata from the meta field of the media file.
[0418] Method 3: The meta field includes the iinf subfield. As explained in step 420 above, the tag metadata is located in the data segment corresponding to the type other than the image codec format type in the iinf subfield. The tag metadata is retrieved from the data segment corresponding to the type other than the image codec format type in the iinf subfield.
[0419] Method 4: The meta field includes the iloc subfield, iprp subfield, ifef subfield, and idat subfield. As described in step 420 above, the tag metadata is located in at least one of the iloc, iprp, ifef, or idat subfields. The tag metadata is obtained from the data segment corresponding to at least one of the iloc, iprp, ifef, or idat subfields.
[0420] Method 5: Add a new item box to the media file to identify the tag metadata. Retrieve the tag metadata from the new item box.
[0421] Optionally, the relevant fields or types carried by the above-mentioned tag metadata conform to the first type or the first field.
[0422] When the media file is in JPEG format, the following methods can be used to obtain tag metadata from the media file.
[0423] Method 1, as described in step 420 above, involves retrieving the tag metadata from the APPn file within the media file.
[0424] Method 2, as described in step 420 above, involves retrieving the tag metadata from the MPF file within the media file.
[0425] After determining that the media file contains tag metadata, the target device obtains the first media content in the media file based on the tag metadata.
[0426] In some embodiments, when the media file is in JPEG format and the JPEG-formatted media file includes content related to the HEIF file format, the media file is decoded according to the JPEG file format, and the content related to the HEIF file format is decoded according to the HEIF file format.
[0427] For example, if the media file is in JPEG format, it includes a new APPn. The APPn is marked with HEIF, and the metadata is located in the APPn within the media file. The metadata, the first image, and the first media content are then obtained from the media file according to the acquisition method described above for media files in HEIF format. The metadata includes at least one of the following: tag metadata, location metadata, time-related metadata, transformation-related metadata, or display metadata.
[0428] For example, as shown in Table 35, assume that item 1 is a tag metadata field, located in the subfield infe within the iinf subfield. The location information of the data corresponding to item 1 is located in the subfield iloc within the meta field. The data block of item 1 is obtained based on the location information of the data corresponding to item 1.
[0429] Assume item 2 is display metadata, located in the subfield 'infe' within the 'iinf' subfield. The location information for the data corresponding to item 2 is located in the subfield 'iloc' within the 'meta' field. The data block for item 2 is retrieved based on this location information.
[0430] In other embodiments, the metadata also includes attribute metadata and location metadata.
[0431] The first media content in the media file is determined by obtaining at least one of the tag metadata, attribute metadata, or location metadata from the destination device.
[0432] Method 1: Mark metadata to indicate that the media file contains first media content, analyze the data of the non-image part of the media file, and obtain the first media content.
[0433] Method 2: Mark the metadata to indicate that the media file contains the first media content, and obtain the first media content from the location indicated by the location metadata based on the location metadata.
[0434] Method 3: Mark the metadata to indicate that the media file contains the first media content, obtain the data of other media content based on the location metadata, and obtain the first media content according to the decoding method indicated by the attribute metadata.
[0435] Optionally, the relevant fields or types carried by the above-mentioned tag metadata conform to the first type or the first field.
[0436] In some embodiments, the media file is used to present merged image media. After the destination device acquires the first image and the first media content, it performs a fusion process on the first image and the first media content, the fusion process including at least one of editing, displaying, or transcoding.
[0437] In some embodiments, the target device associates and / or transforms the first image and / or the first media content based on metadata.
[0438] In some embodiments, the metadata includes at least one of time-related metadata or transformation-related metadata.
[0439] In some embodiments, after determining that the media file contains tag metadata, the destination device determines the time-related metadata and / or transformation-related metadata in the media file based on the tag metadata. This allows the destination device to process and / or display the first image and the first media content in the media file based on the time-related metadata and / or transformation-related metadata.
[0440] For descriptions of time-related metadata, conversion-related metadata, and their storage location in media files, please refer to the above description of encoding; they will not be repeated here.
[0441] Optionally, the metadata includes first format information and second format information. After obtaining the first format information and second format information, the destination device determines the conversion-related metadata based on the first format information and second format information.
[0442] Optionally, the metadata includes first format information and second format information. After obtaining the first format information and second format information, the destination device determines the conversion-related metadata based on the first format information, the second format information, the first image, and the video of the first media content.
[0443] Optionally, the metadata includes first time domain information and second time domain information. After the destination device obtains the first time domain information and the second time domain information, it determines the transformation-related metadata based on the first time domain information and the second time domain information.
[0444] Optionally, a third image is obtained from the first video of the first media content based on the first time domain information, and conversion-related metadata is obtained based on the mapping relationship between the first image and the third image.
[0445] Optionally, a third image can be obtained from the first video based on associated metadata.
[0446] Optionally, the first and / or third images can be converted based on the conversion metadata.
[0447] For example, the first and third images are converted to RGB, and then converted to the same color gamut (such as BT.2020). The luminance values of the first and third images are obtained from their RGB values. The mapping relationship is determined based on the mapping relationship between the luminance values of corresponding pixels in the first and third images. Alternatively, the color gamut of the first and third images (such as BT.2020, sRGB, BT.709, P3) can be obtained, and the conversion matrix or algorithm used when converting the first and third images to the same color gamut (such as BT.2020) can be recorded.
[0448] For example, the target device determines whether the first image is SDR or HDR based on the first image obtained from the media file and / or the format information of the first image.
[0449] For example, the target device determines that the first image is HDR based on the first image and / or the format information of the first image obtained from the media file, and obtains its dynamic range, and / or peak brightness, and / or the ratio of peak brightness to reference white.
[0450] For example, the destination device determines whether the video is SDR or HDR based on the video and / or video format information obtained from the media file. For instance, transfer function information can be obtained from the colr included in an HEIF file or the CICP in the ICC file included in a JPEG file. If the transfer function indicates the use of HLG or PQ, then it is HDR.
[0451] For example, the target device obtains conversion information between HDR and SDR based on the conversion information (such as processing information) in the format information of the acquired image. For instance, it converts the conversion information into tone mapping processing flow information according to the format of the conversion information (HDR Vivid standard, HDR10+ standard, ICC standard). For example, in the HDR Vivid standard format conversion information, the tone mapping curve for the conversion between HDR and SDR can be obtained. Simultaneously, the tone mapping flow can be obtained from the HDR Vivid standard format, showing that the process involves first converting to the luminance domain / linear domain and then performing tone mapping. Furthermore, information on saturation adjustment in the YUV domain after tone mapping can also be obtained from the HDR Vivid standard format.
[0452] For example, the target device can obtain information on the conversion between HDR based on the conversion information (such as processing information) in the format information of the acquired image.
[0453] For example, the image data acquired by the target device contains a format (such as SDR or HDR), and the processing information for the format information is obtained from the extraction process. For instance, HDR information contains corresponding peak brightness information, one with a peak brightness of 1000 nits and another with a peak brightness of 1600 nits. It is necessary to convert the 1000-nit HDR and the 1600-nit HDR to the same peak brightness. For instance, HDR information contains corresponding peak brightness / reference white ratio information, one with a peak brightness / reference white ratio of 4 and another with a peak brightness / reference white ratio of 8. It is necessary to convert two HDRs with different peak brightness / reference white ratios to the same peak brightness / reference white ratio. Alternatively, the peak brightness / reference white ratio can be multiplied by the reference white (e.g., 203 nits) to obtain the peak brightness, and then the two can be converted to the same peak brightness.
[0454] For example, the image data acquired by the target device contains data in two formats, and the processing information in the format information is obtained based on the two formats.
[0455] In some examples, the target device acquires at least one of time-related metadata or transformation-related metadata, and processes the first image and the first media content based on the time-related metadata and / or transformation-related metadata.
[0456] For example, the target device performs joint image and video processing based on conversion-related metadata and preset external states. Optionally, preset external states include display state and display device state. Display state includes editing state or normal display state. Display device state includes SDR or HDR, or different configurations of HDR. For example, conversion between HDR and SDR, or between HDR and HDR based on conversion-related metadata, includes one or more of the following.
[0457] Method 1: Convert the image from HDR to SDR based on preset external conditions.
[0458] Method 2: Convert the image from SDR to HDR based on preset external conditions.
[0459] Method 3: Convert the video from HDR to SDR based on preset external conditions.
[0460] Method 4: Convert the video from SDR to HDR based on preset external conditions.
[0461] Method 5: Based on preset external conditions, convert the image from the first dynamic range HDR to the second dynamic range HDR.
[0462] The method flow converts the video from a first dynamic range HDR to a second dynamic range HDR based on a preset external state.
[0463] For example, the target device performs joint processing of images and videos based on time-related metadata and preset external states.
[0464] Method 1: Based on a preset external state, display or process the first image and / or video by associating metadata with the time domain of the first image and video.
[0465] For example, starting from the start time of the video, the first image is displayed after the difference between the start time of the first image and the start time of the video.
[0466] For example, starting from the beginning frame of the video, the first image is displayed after the difference between the first image and the beginning frame of the video.
[0467] For example, the first image is displayed based on the timestamp of the first image indicated by the time-related metadata.
[0468] For example, the video can be displayed based on the start timestamp of the video as indicated by the time-related metadata.
[0469] For example, playing a video based on the duration indicated by the time-related metadata.
[0470] For example, after playing the first image, the video can start playing from the position in the video where the start time of the first image differs from that of the video.
[0471] For example, when displaying a video, mark the positions in the video where the start time differs from the first image and the beginning time of the video.
[0472] For example, when playback stops, the position in the video that differs from the starting time of the first image.
[0473] For example, in editing and other processing, other images at positions different from the first image and the start time of the video are extracted.
[0474] In some examples, image relationships are used to convert videos. Image formats include both SDR and HDR. Video formats include either SDR or HDR.
[0475] Based on the transformation-related metadata and preset external states, perform joint processing of the first image and video. For example, perform conversion between HDR and SDR based on the transformation-related metadata.
[0476] Method 1, with the default external state being SDR. When the first image is in HDR single-layer or dual-layer (SDR+HDR enhancement layer) format, and the video is in SDR format, playing the cover frame or default still image includes the following steps.
[0477] 1) Obtain conversion information from HDR to SDR. For example, obtain the conversion tone mapping information based on the HDR to SDR conversion information obtained from the conversion-related metadata. Alternatively, obtain the preset tone mapping information.
[0478] 2) Based on the tone mapping information, perform HDR to SDR conversion on the cover frame or default static image to obtain the processed cover frame or default static image.
[0479] 3) Display the processed cover frame or default static image, or continue with subsequent processing.
[0480] Method 2, with the preset external state being HDR. When the first image is in HDR single-layer or dual-layer (SDR+HDR enhancement layer) format, and the video is in SDR format, playing the video or consecutive images includes the following steps.
[0481] 1) Obtain SDR to HDR conversion information. For example, based on the SDR to HDR conversion information obtained from the conversion-related metadata, obtain the converted tone mapping information. Alternatively, obtain the preset tone mapping information.
[0482] 2) Based on the tone mapping information, the video is converted from SDR to HDR to obtain the processed HDR video.
[0483] 3) Display the processed HDR video or continue with further processing.
[0484] Method 3, with the default external state being SDR. When the first image is in HDR single-layer or dual-layer (SDR+HDR enhancement layer) format, and the video is in HDR format, playback includes the following steps.
[0485] 1) When playing the cover frame or default static image, if the first image is an HDR dual-layer image, and the SDR part of the image is played, if the first image is an HDR single-layer image, the first image is converted to an SDR image for display or subsequent processing according to Method 1.
[0486] 2) Play videos or continuous images. If the video is HDR, convert it to SDR video according to Method 1 for display or subsequent processing.
[0487] Method 4: The preset external state is a combination of HDR and SDR with a specific dynamic range. One type of combination refers to the ratio of SDR to HDR peak brightness to SDR peak brightness in the SDR dynamic range. Another type of combination refers to the ratio of SDR to HDR preset peak brightness to SDR preset peak brightness in the SDR dynamic range. When the first image is in HDR single-layer or dual-layer (SDR + HDR enhancement layer) format, and the video is in HDR format, playback includes the following steps.
[0488] 1) Obtain the dynamic range of SDR and HDR in HDR single-layer images or dual-layer (SDR+HDR enhancement layer) formats, as well as the dynamic range of SDR in HDR videos and HDR-to-SDR conversions.
[0489] 2) Regarding the dynamic range of SDR and HDR, in the process of the above method one, method two or method three, the first image is converted to the same dynamic range as the video, or the video is converted to the same dynamic range as the image.
[0490] Even when both the first image and video are HDR content, they may have different dynamic ranges. For example, HDR video includes HLG and PQ formats. The peak brightness of HLG format is typically 1000 nits, while the peak brightness of PQ format is 10000 nits. The SDR peak brightness and HDR peak brightness of an HDR image are usually related. The reference peak brightness or default peak brightness of SDR is usually set to 203 nits, while the HDR peak brightness may be 203 nits * N, where N can be any number greater than 1, and therefore may not be 1000 nits or 10000 nits. When switching between HDR dynamic ranges of 1000 nits, 10000 nits, and 203 nits * N, a conversion is also required.
[0491] Optionally, media content in a merged format can be displayed based on the temporal-related metadata of images and videos.
[0492] Method 1: Add the media file to the editing software. With the editing window displaying the first image and video, mark the positions in the video that have the same time relationship as the first image based on the time-related metadata, or display the time domain data of the first image and video.
[0493] Method 2: Add the media file to the editing software. When the editing window displays the first image and video, when playing the video part based on the time-related metadata, start playing from the video frame that is at the same time as or closest to the first image; or, stop playing at the video frame that is at the same time as or closest to the first image.
[0494] This application also provides a method for encapsulating media files. The method includes encapsulating various media contents into a media file, the media file being in a JPEG format. The media file includes an APPn, which contains data structures related to HEIF.
[0495] In some possible implementations, the APPn flag indicates HEIF. Alternatively, the APPn flag may also indicate HEIC or MIAF, etc.
[0496] Optionally, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0497] Optionally, APPn's payload also includes SOI.
[0498] Optionally, the metadata is located in the APPn file of the media file.
[0499] Optionally, the metadata is located in the payload of APPn in the media file.
[0500] For a description of the media files obtained by encapsulating various media content using this encapsulation method, please refer to the above embodiments.
[0501] The media files obtained by this encapsulation method can be applied to the media data processing methods described in the above embodiments.
[0502] For example, an encoder generates a media file from multiple media contents, and a decoder decodes the media file to obtain a first image and / or first media content. The media file is obtained by encapsulating the content according to this encapsulation method.
[0503] For example, the media file obtained by this encapsulation method may include the aforementioned metadata, such as associated metadata, display metadata, etc.
[0504] This application improves the structure of media files and sets metadata to process multiple media contents, expanding the forms of media content fusion. The media data processing method provided in this application is a standardized encoding method, which facilitates system implementation, enables unified management of media content, improves the convenience of media content management, reduces application adaptation costs, and enhances the diversity of multi-media content fusion presentation and user cross-platform experience.
[0505] It is understood that, in order to achieve the functions in the above embodiments, the encoder and decoder include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0506] The media data processing method provided according to this embodiment has been described in detail above with reference to Figures 1 to 4. The encoding device and decoding device provided according to this embodiment will be described below with reference to Figures 5 and 6.
[0507] Figure 5 is a schematic diagram of the possible encoding device provided in this embodiment. These encoding devices can be used to implement the encoding function of the media data processing method in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. In this embodiment, the encoding device is the encoder 213 shown in Figure 2, and can also be a module (such as a chip) applied to a terminal device or server.
[0508] As shown in Figure 5, the encoding device 500 includes a communication module 510, an encoding module 520, and a storage module 530.
[0509] The encoding device 500 is used to implement the function of the encoder in the method embodiment shown in FIG4 above.
[0510] The communication module 510 is used to acquire various media content, including a first image and other media content besides the first image, wherein the media content includes a first video. For example, the communication module 510 is used to perform step 410 in Figure 4.
[0511] Encoding module 520 is used to generate a media file based on various media content. The media file includes image data of a first image, data of the first media content, and metadata. The metadata is used to instruct the first image or the first media content to be associated and / or converted. For example, encoding module 520 is used to perform step 420 in Figure 4.
[0512] The communication module 510 is also used to send media files.
[0513] Storage module 530 is used to store media content, metadata, and media files, so that media files can be generated at the encoding end based on various media content.
[0514] Figure 6 is a schematic diagram of a possible decoding device provided in this embodiment. These decoding devices can be used to implement the decoding function of the media data processing method in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. In this embodiment, the decoding device can be the decoder 223 shown in Figure 2, or it can be a module (such as a chip) applied to a terminal device or server.
[0515] As shown in Figure 6, the decoding device 600 includes a communication module 610, a decoding module 620, and a storage module 630.
[0516] The decoding device 600 is used to implement the function of the decoder in the method embodiment shown in FIG4 above.
[0517] The communication module 610 is also used to acquire a media file, which includes image data of a first image, data of first media content, and metadata. The metadata is used to instruct the first image or the first media content to be associated and / or converted. The first media content includes a first video. For example, the communication module 610 is used to perform step 430 in FIG4.
[0518] The decoding module 620 is used to obtain a first image from the media file and to obtain first media content from the media file based on metadata. For example, the decoding module 620 is used to perform steps 440 and 450 in Figure 4.
[0519] Storage module 630 is used to store media content, metadata, and media files, so that the first image and first media content can be obtained from the media files at the decoding end.
[0520] It should be understood that the encoding device 500 and decoding device 600 in the embodiments of this application can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Alternatively, when the method shown in FIG4 is implemented in software, the encoding device 500 and decoding device 600 and their respective modules can also be software modules.
[0521] A more detailed description of the communication module, encoding module, decoding module and storage module mentioned above can be obtained directly from the relevant description in the method embodiment shown in Figure 4, and will not be repeated here.
[0522] Figure 7 is a schematic diagram of the structure of an encoder 700 provided in this embodiment. As shown in Figure 7, the encoder 700 includes a processor 710, a bus 720, a memory 730, and a communication interface 740.
[0523] It should be understood that in this embodiment, the processor 710 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0524] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0525] The communication interface 740 is used to enable communication between the encoder 700 and external devices or components. In this embodiment, the communication interface 740 is used to acquire various media content.
[0526] Bus 720 may include a pathway for transmitting information between the aforementioned components (such as processor 710 and memory 730). In addition to a data bus, bus 720 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 720 in the figure.
[0527] As an example, encoder 700 may include multiple processors. A processor may be a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or computing units used to process data (e.g., computer program instructions). Processor 710 generates media files based on various media content.
[0528] It is worth noting that Figure 7 only shows an encoder 700 including one processor 710 and one memory 730 as an example. Here, the processor 710 and the memory 730 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0529] The memory 730 can correspond to the storage medium used in the above method embodiments for storing information such as media content, metadata and media files, for example, a disk, such as a mechanical hard disk or a solid-state hard disk.
[0530] The encoder 700 described above can be a general-purpose device or a special-purpose device. For example, the encoder 700 can be an x86 or ARM-based server, or other special-purpose servers, such as a policy control and charging (PCC) server. This application does not limit the type of encoder 700.
[0531] It should be understood that the encoder 700 of this embodiment can correspond to the encoding device 500 of this embodiment, and can correspond to the corresponding subject that executes any of the methods in FIG4. The above and other operations and / or functions of each module in the encoding device 500 are respectively for implementing the corresponding processes of each method in FIG4. For the sake of brevity, they will not be described in detail here.
[0532] Figure 8 is a schematic diagram of the structure of a decoder 800 provided in this embodiment. As shown in Figure 8, the decoder 800 includes a processor 810, a bus 820, a memory 830, and a communication interface 840.
[0533] It should be understood that in this embodiment, the processor 810 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0534] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0535] The communication interface 840 is used to enable communication between the decoder 800 and external devices or components. In this embodiment, the communication interface 840 is used to acquire media files.
[0536] Bus 820 may include a pathway for transmitting information between the aforementioned components (such as processor 810 and memory 830). In addition to a data bus, bus 820 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 820 in the figure.
[0537] As an example, decoder 800 may include multiple processors. A processor may be a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or computing units used to process data (e.g., computer program instructions). Processor 810 acquires a first image based on a media file and acquires first media content from the media file based on metadata.
[0538] It is worth noting that Figure 8 only shows an example of a decoder 800 including one processor 810 and one memory 830. Here, the processor 810 and the memory 830 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0539] The memory 830 can correspond to the storage medium used to store media content, metadata and media files in the above method embodiments, such as a disk, like a mechanical hard disk or a solid-state hard disk.
[0540] The decoder 800 described above can be a general-purpose device or a dedicated device. For example, the decoder 800 can be an x86 or ARM-based server, or other dedicated servers, such as a policy control and charging (PCC) server. This application does not limit the type of decoder 800.
[0541] It should be understood that the decoder 800 according to this embodiment can correspond to the decoding device 600 in this embodiment, and can correspond to the corresponding subject that executes any of the methods in FIG4. The above and other operations and / or functions of each module in the encoding device 500 and the decoding device 600 are respectively for implementing the corresponding processes of each method in FIG4. For the sake of brevity, they will not be described in detail here.
[0542] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a network device or terminal device. Of course, the processor and storage medium can also exist as discrete components in the network device or terminal device.
[0543] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0544] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0545] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
[0546] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. In the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following or similar expressions" refers to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and / or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0547] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0548] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0549] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, throughout the specification, various embodiments do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0550] Some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current underlying solution, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in other scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated upon here.
[0551] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.
Claims
1. A media data processing method, characterized in that, include: Acquire multiple media content, the multiple media content including a first image and other first media content besides the first image, the first media content including a first video; A media file is generated based on the various media content. The media file includes image data of the first image, metadata, and data of the first media content. The metadata is used to associate and / or transform the first image and / or the first media content.
2. The method according to claim 1, characterized in that, The first image and the first media content are obtained based on a set of time-domain continuous images.
3. The method according to claim 1, characterized in that, The first media content is obtained from the first image group in a set of time-domain continuous images.
4. The method according to claim 1, characterized in that, The first image is obtained from a second image group in a set of time-domain continuous images.
5. The method according to any one of claims 2-4, characterized in that, The set of time-domain continuous images includes multiple images taken at the same time with different exposure parameters.
6. The method according to any one of claims 1-5, characterized in that, Media files are generated based on the aforementioned multiple media contents, including: The first image is encoded to obtain an image file, the image file including the image data; Encode the first media content to obtain the data of the first media content; The data of the first media content and the metadata are encoded into the image file to obtain the media file.
7. The method according to any one of claims 1-5, characterized in that, Media files are generated based on the aforementioned multiple media contents, including: The image data is obtained by encoding the first image; Encode the first media content to obtain the data of the first media content; The media file is generated based on the image data, the data of the first media content, and the metadata.
8. A media data processing method, characterized in that, include: Obtain a media file, the media file including image data of a first image, metadata and data of a first media content, the metadata being used to associate and / or convert the first image and / or the first media content, the first media content including a first video; The media file is decoded to obtain the first image and / or the first media content.
9. The method according to claim 8, characterized in that, Obtaining the first image from the media file includes: The first image is obtained by decoding the media file.
10. The method according to claim 8, characterized in that, Obtaining the first image from the media file includes: The first image is obtained by decoding the image data in the media file.
11. The method according to any one of claims 8-10, characterized in that, The method further includes: The first image and / or the first media content are associated and / or transformed based on the metadata.
12. The method according to any one of claims 1-11, characterized in that, The metadata includes time-related metadata, which is used to indicate the temporal relationship between the first media content and the first image.
13. The method according to claim 12, characterized in that, The time-related metadata indicates the difference in start time between the first image and the first video included in the first media content; or, The time-related metadata indicates the difference between the first image and the starting frame of the first video included in the first media content; or, The time-related metadata indicates the timestamp of the first image; or... The time-related metadata indicates the start timestamp of the first video included in the first media content; or, The time-related metadata indicates the duration of the first video included in the first media content.
14. The method according to claim 12 or 13, characterized in that, The method includes: obtaining video frames corresponding to the first image from the first video based on the time-related metadata.
15. The method according to any one of claims 12-14, characterized in that, The metadata also includes transformation-related metadata, which indicates the method of transformation of the first media content and / or the first image.
16. The method according to any one of claims 12-14, characterized in that, The metadata also includes transformation-related metadata, which is used to determine the transformation of the first media content and / or the first image.
17. The method according to claim 15 or 16, characterized in that, The conversion-related metadata includes at least one of format information, color gamut information, resolution, or rotation.
18. The method according to claim 17, characterized in that, The conversion also includes conversion between HDR and SDR, or conversion between HDR content with different dynamic ranges.
19. The method according to claim 18, characterized in that, The method includes converting the first video based on the conversion-related metadata when the first image is an HDR image and the first video is an SDR video.
20. The method according to any one of claims 15-19, characterized in that, The transformation-related metadata is obtained from the data associated with the first image.
21. An encoding device, characterized in that, It includes a processing circuit for implementing the method as described in any one of claims 1-7 and 12-20.
22. A decoding device, characterized in that, It includes a processing circuit for implementing the method as described in any one of claims 8-11, 12-20.
23. An encoder, characterized in that, The encoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the method as described in any one of claims 1-7, 12-20.
24. A decoder, characterized in that, The decoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the method as described in any one of claims 8-11, 12-20.
25. A codec system, characterized in that, The encoding / decoding system includes an encoder as described in claim 23 and a decoder as described in claim 24, wherein the encoder is used to perform the operation steps of the method according to any one of claims 1-7 and 12-20, and the decoder is used to perform the method according to any one of claims 8-11 and 12-20.
26. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-20.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions or programs that, when executed on a computer, implement the method as described in any one of claims 1-20.
28. A media file, characterized in that, The media file is obtained by the method according to any one of claims 1-7 and 12-20.
29. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes media files obtained by the method according to any one of claims 1-7, 12-20.