Media data processing method, apparatus, encoder, decoder and system
Patent Information
- Application Number
- CN202510620233.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-11
- Filing Date
- 2025-05-13
- Publication Date
- 2026-09-11
AI Technical Summary
但是,不同应用、不同终端设备、不同系统对媒体内容的管理的实现方式的差异很大,媒体内容的管理缺乏统一性,应用适配复杂,多种媒体内容融合呈现的多样性较差
[0073] The technical effects of any of the implementation methods in aspects three through seventeen can be found in the technical effects of the corresponding implementation methods in aspects one or two, and will not be repeated here.
Smart Images

Figure CN122741748A_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application filed on March 11, 2025, with application number 202510289464.7, entitled "Media Data Processing Method, Apparatus, Encoder, Decoder and System", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of media, and more particularly to a media data processing method, apparatus, encoder, decoder, and system. Background Technology
[0003] Currently, terminal devices centrally manage and present media content. For example, they use grids, page turning, time categorization, and location categorization to display multiple images. However, the implementation methods of media content management vary greatly among different applications, terminal devices, and systems. Media content management lacks uniformity, application adaptation is complex, and the diversity of integrated presentation of various media content is poor. Summary of the Invention
[0004] This application provides a media data processing method, apparatus, encoder, decoder, and system, thereby enhancing the diversity of multi-media content fusion presentation.
[0005] Firstly, a media data processing method is provided. The method includes acquiring multiple media contents and generating a media file based on the multiple media contents. The multiple media contents include a first image and other media content besides the first image. The first media content includes at least one of a second image or a first video. The media file includes image data of the first image, metadata, and data of the first media content. The metadata is used to indicate display information for the multiple media contents.
[0006] In one possible implementation, generating a media file based on multiple media contents includes: encoding a first image to obtain an image file, the image file including image data; encoding first media content to obtain data of the first media content; and encoding metadata and data of the first media content into the image file to obtain the media file.
[0007] In another possible implementation, a media file is generated based on multiple media contents, including: encoding a first image to obtain image data; encoding the first media content to obtain data of the first media content; and generating a media file based on the image data, metadata, and the data of the first media content.
[0008] In a second aspect, a media data processing method is provided, the method comprising acquiring a media file, the media file including image data of a first image, metadata and data of first media content, the metadata being used to indicate display information of multiple media contents, the multiple media contents including the first image and the first media content, the first media content including at least one of a second image or a first video; and decoding the media file to obtain the first image and / or the first media content.
[0009] In one possible implementation, decoding the media file to obtain the first image includes: decoding image data in the media file to obtain the first image.
[0010] In another possible implementation, the method further includes: performing display processing on the first image and / or the first media content based on metadata.
[0011] The technical solution provided in this application includes encoding multiple media contents into the same media file and using metadata in the media file to indicate the display information of the multiple media contents. When presenting multiple media contents in a fused form, the metadata in the media file is used to process the display of the multiple media contents in the media file, thus presenting the multiple media contents in a fused form. This application expands the fusion forms of multiple media contents by improving the structure of the media file and setting metadata, and processing the display of multiple media contents based on the metadata. The media data processing method provided in this application is a standardized encoding method, which facilitates system implementation, achieves unified management of media content, improves the convenience of media content management, reduces application adaptation costs, and enhances the diversity of multi-media content fusion presentation and the user's cross-platform experience.
[0012] The technical solution provided in this application also includes encoding multiple media contents into different media files and using metadata in the media files to indicate that the media files together contain multiple media contents with other media files.
[0013] In another possible implementation, the metadata includes display metadata, which indicates at least one of the display methods of the cover frame or various media content.
[0014] In another possible implementation, the display metadata indicates that at least one of the multiple images included in the multiple media content is used as the cover frame; or, the display metadata indicates that the result of processing at least one of the multiple images included in the multiple media content is used as the cover frame.
[0015] In another possible implementation, the processing result includes any one of the following: a deformation result of at least one of the plurality of images, a fusion result of at least one of the plurality of images, or a fusion result of at least one of the plurality of images with media content other than images in the first media content.
[0016] Thus, displaying a cover frame based on display metadata makes the display of various media content more attractive and enhances the diversity of the integrated presentation of various media content.
[0017] In another possible implementation, the display method includes display style and / or playback information of multiple frames of images contained in the first image and the first media content.
[0018] In another possible implementation, the playback information includes at least one of the following: playback order, starting frame of playback, or image switching method.
[0019] In another possible implementation, the display metadata also includes at least one of the following: display order, position, size, or scaling information in the display styles corresponding to various media content.
[0020] Therefore, by displaying images and videos in media files based on the display metadata, the display of various media content becomes smoother, enhancing the diversity of integrated media content presentation.
[0021] In another possible implementation, the media file is in JPEG format; metadata is located in APPn within the media file; APPn contains data structures related to the High Efficiency Image Format (HEIF).
[0022] In another possible implementation, the APPn flag indicates HEIF.
[0023] In another possible implementation, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0024] In another possible implementation, APPn's payload also includes the Start of Image (SOI).
[0025] In another possible implementation, the metadata is located in the APPn of the media file, including: the metadata is located in the payload of the APPn in the media file.
[0026] Therefore, JPEG animated images are encapsulated in a JPEG-compatible manner and unified to HEIF, meaning that JPEG file format media files include content related to the HEIF file format. This enables unified management of media content, facilitates the expansion of various media content, and improves the convenience of media content management.
[0027] In another possible implementation, the media file is in HEIF format; metadata is located in the ftyp or meta field of the media file.
[0028] In another possible implementation, the meta field includes an iinf subfield; the metadata is located in the meta field of the media file, including: the metadata is located in the data segment corresponding to the type other than the image codec format type in the iinf subfield.
[0029] In another possible implementation, the image codec format type includes at least one of MPEG, Joint Photographic Experts Group (JPEG), Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or Versatile Video Codec (VVC).
[0030] In another possible implementation, the type, in addition to the image codec format type, includes at least one of uri, mime, it35, xml, or exif.
[0031] In another possible implementation, the meta field includes the iloc subfield, the iprp subfield, the iref subfield, and the idat subfield; the metadata is located in the meta field of the media file, including: the metadata is located in the data segment corresponding to at least one of the iloc subfield, the iprp subfield, the iref subfield, or the idat subfield.
[0032] In another possible implementation, at least one of the fields identifying tag metadata, identifying attribute metadata, identifying location metadata, identifying time-related metadata, identifying transformation-related metadata, and identifying display metadata is different.
[0033] In another possible implementation, fields that identify the metadata associated with the identifier are associated with fields that identify the metadata displayed.
[0034] In another possible implementation, the field that identifies the displayed metadata includes at least one of the iprp subfield or the iref subfield; the field that identifies the tag metadata is associated with at least one of the iprp subfield or the iref subfield.
[0035] In another possible implementation, the field that identifies the tag metadata and the field that identifies the display metadata are associated through item association rules.
[0036] In another possible implementation, the data for the first media content is located in the idat subfield of the mdat field, moov field, or meta field.
[0037] In another possible implementation, the media file contains a first field or a first preset character. Metadata is located in the first field or the first preset character.
[0038] The first field refers to a box, fullbox, or APPn identified using the first type identifier.
[0039] The first category refers to any one of the following abbreviations or full names related to continuous shooting and moving images: Live, Motion, Movie Photo, Continuous Shooting, mpvd, mpic, Live / motion / movie / multiple photo, Continuousshooting, burst mode, etc.
[0040] The first preset character is any one of the Chinese or English abbreviations or full names related to continuous shooting and moving images, such as Live, motion, movie photo, mpho, Continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, Continuous shooting, burst mode, etc.
[0041] In another possible implementation, the media file is in JPEG format; metadata is located in APPn or MPF within the media file.
[0042] In another possible implementation, the metadata is located in the APPn or MPF of the media file, including: at least one of the tags, labels, or payloads of the APPn in the media file is of the first type.
[0043] The first type uses any one of the following abbreviations or full names related to continuous shooting and moving images: Live, motion, movie photo, mpho, Continuous shooting, mpvd, mpic, Live / motion / movie / multiple photo, Continuous shooting, burst mode, etc.
[0044] Thirdly, a method for encapsulating media files is provided, which encapsulates multiple media contents into a media file. The media file is in the JPEG format and includes APPn, which contains data structures related to HEIF.
[0045] In one possible implementation, the APPn flag indicates HEIF.
[0046] In another possible implementation, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0047] In another possible implementation, APPn's payload also includes the Start of Image (SOI).
[0048] In another possible implementation, the media file is the media file of the first or second aspect mentioned above.
[0049] In another possible implementation, the aforementioned metadata is located in the APPn of the media file.
[0050] In another possible implementation, the aforementioned metadata is located in the APPn of the media file, including: the metadata is located in the payload of the APPn in the media file.
[0051] Fourthly, an encoding apparatus is provided, the encoding apparatus including processing circuitry for performing the methods of the first aspect or any possible design of the first aspect. For example, the encoding apparatus includes a communication module and an encoding module.
[0052] A communication module is used to acquire multiple media content, including a first image and other media content, wherein the first media content includes at least one of a second image or a first video.
[0053] The encoding module is used to generate media files based on various media content. The media file includes image data of the first image, metadata, and data of the first media content. The metadata is used to indicate display information for the various media content.
[0054] In one possible implementation, when the encoding module generates a media file based on multiple media contents, it is specifically used to encode a first image to obtain an image file, the image file including image data; encode the first media content to obtain data of the first media content; and encode the metadata and the data of the first media content into the image file to obtain the media file.
[0055] In another possible implementation, when the encoding module generates a media file based on multiple media contents, it is specifically used to encode a first image to obtain image data; encode a first media content to obtain data of the first media content; and generate a media file based on the image data, metadata, and data of the first media content.
[0056] Fifthly, a decoding apparatus is provided, the decoding apparatus including processing circuitry for performing the methods of the second aspect or any possible design of the second aspect. For example, the encoding apparatus includes a communication module and a decoding module.
[0057] A communication module is used to acquire media files, which include image data of a first image, metadata, and data of first media content. The metadata is used to indicate display information for various media content, including the first image and the first media content itself, which includes at least one of a second image or a first video.
[0058] A decoding module is used to decode media files to obtain a first image and / or first media content.
[0059] In one possible implementation, decoding the media file to obtain the first image includes: decoding image data in the media file to obtain the first image.
[0060] In another possible implementation, the method further includes: performing display processing on the first image and / or the first media content based on metadata.
[0061] In a sixth aspect, an encoder is provided, the encoder including at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by at least one processor, it implements the method described in the first aspect or any possible design of the first aspect.
[0062] In a seventh aspect, a decoder is provided, the decoder including at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by at least one processor, it implements the method described in the second aspect or any possible design of the second aspect.
[0063] Eighthly, a coding / decoding system is provided, the coding / decoding system comprising an encoder as described in the sixth aspect and a decoder as described in the seventh aspect.
[0064] A ninth aspect provides a chip, comprising: a processor and a power supply circuit; wherein the power supply circuit is used to supply power to the processor; the processor is used to perform operational steps of the method in the first aspect or any possible implementation of the first aspect, and to perform operational steps of the method in the second aspect or any possible implementation of the second aspect.
[0065] In a tenth aspect, a computer program product is provided, comprising a computer program or instructions, which, when executed on a processor, cause the processor to perform operational steps of the method in the first aspect or any possible implementation thereof, or to perform operational steps of the method in the second aspect or any possible implementation thereof.
[0066] Eleventhly, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a computing device, causing the computing device to perform operational steps of the method in the first aspect or any possible implementation thereof, or to perform operational steps of the method as described in the second aspect or any possible implementation thereof.
[0067] In a twelfth aspect, a media file is provided, said media file being obtained from the first aspect or any possible implementation thereof.
[0068] In a thirteenth aspect, a media file is provided, the media file including metadata, the metadata including display metadata.
[0069] Fourteenth aspect, a method for storing media files is provided, the method comprising: receiving a media file generated according to the first aspect or any possible implementation thereof, or a media file as described in the thirteenth aspect; and storing the media file in a storage medium.
[0070] In a fifteenth aspect, an apparatus for storing media files is provided, the apparatus being used to store media files generated according to the first aspect or any possible implementation thereof, or for storing media files as described in the thirteenth aspect. Exemplarily, the apparatus may be a computer-readable storage medium.
[0071] In a sixteenth aspect, a method for transmitting a media file is provided, the method comprising: acquiring a media file, the media file being generated by the first aspect or any possible implementation thereof, or the media file being the media file described in the thirteenth aspect; and sending the media file.
[0072] In a seventeenth aspect, an apparatus for transmitting media files is provided, the apparatus being used to acquire and transmit media files generated by the first aspect or any possible implementation thereof, or to acquire and transmit media files as described in the thirteenth aspect.
[0073] The technical effects of any of the implementation methods in aspects three through seventeen can be found in the technical effects of the corresponding implementation methods in aspects one or two, and will not be repeated here.
[0074] All possible implementations of any of the above aspects can be combined, provided that the solutions do not contradict each other. Attached Figure Description
[0075] Figure 1 This application provides a schematic diagram of presenting multiple media content in a fusion format;
[0076] Figure 2 A schematic diagram of the structure of an encoding / decoding system provided in this application;
[0077] Figure 3 A schematic diagram of another encoding / decoding system provided in this application;
[0078] Figure 4 A flowchart illustrating a media data processing method provided in this application;
[0079] Figure 5 A schematic diagram of a cover frame provided for this application;
[0080] Figure 6 A schematic diagram illustrating the display of multiple media content provided in this application;
[0081] Figure 7 A schematic diagram of the structure of an encoding device provided in this application;
[0082] Figure 8 A schematic diagram of the structure of a decoding device provided in this application;
[0083] Figure 9 A schematic diagram of an encoder provided in this application;
[0084] Figure 10 This is a schematic diagram of the structure of a decoder provided in this application. Detailed Implementation
[0085] To facilitate understanding, the main terms used in this application will be explained first.
[0086] Currently, terminal devices centrally manage and present media content. Presenting multiple media content in a converged format has become a trend.
[0087] For example, multiple images are fused together for presentation. Examples include... Figure 1 As shown in (a), the terminal device presents multiple images in a grid format. When a user views an image, the terminal device receives the user's click action and displays the image viewed by the user.
[0088] like Figure 1As shown in (b), the terminal device presents multiple images in a page-turning manner. When the user views an image, the terminal device receives the user's page-turning operation and presents the image viewed by the user.
[0089] like Figure 1 As shown in (c), the terminal device presents multiple images using categories such as time, location, and people. For example, multiple images presented by a gallery application on the terminal device.
[0090] For example, still images can be blended with continuous images. Or, still images can be blended with video. For instance, an image can be displayed first, followed by a video to achieve an animated effect. Animated images typically refer to dynamic images in the Graphics Interchange Format (GIF), an image format that creates animation effects by playing a series of static images in succession.
[0091] However, the media content is centrally managed by the system, and the fusion format is simplistic, merely a simple arrangement of multiple images. The fusion implementation varies greatly across different applications, terminal devices, and system management, resulting in poor diversity in the fusion of various media content and a subpar user experience.
[0092] To address the issue of poor diversity in the presentation of integrated media content, this application provides a media data processing method, namely, a multi-media content fusion scheme. The method includes encoding multiple media contents into a media file and using metadata within the media file to indicate the display information of the various media contents contained within the media file. When presenting multiple media contents in a fused form, the metadata in the media file is used to process the display of the multiple media contents within the media file, thus presenting the multiple media contents in a fused form. This application expands the forms of multi-media content fusion by improving the structure of media files and setting metadata, and by processing the display of multiple media contents based on metadata. The media data processing method provided in this application is a standardized encoding method, which facilitates system implementation, enables unified management of media content, improves the convenience of media content management, reduces application adaptation costs, and enhances the diversity of multi-media content fusion presentation and the user's cross-platform experience.
[0093] Media content refers to information presented and disseminated to the public in various forms. Media content includes, but is not limited to, one or more of text, images, audio, and video. For example, depending on the medium and presentation method, media content can be categorized into various types, such as news reports, feature articles, advertisements, TV dramas, movies, music, and social media content.
[0094] Subtitles are text-based representations of dialogue or narration in a video displayed below the screen for reading. They are primarily used in situations where sound is inaudible or where it's desirable to understand the content without interfering with the audio, such as movie subtitles or watching videos in a silent environment. Subtitles are typically synchronized with the audio content in the video and have a relatively fixed format (such as font, size, and position) to ensure smooth and accurate reading.
[0095] Text refers to static or dynamic text elements added to a video, including titles, descriptive text, and tags. It is commonly used in intros, outros, transitions, and to emphasize content, enhancing the video's visual appeal and information delivery. Text styles, sizes, colors, and animation effects can all be customized, offering high flexibility.
[0096] Subtitles are primarily used to convey dialogue or narration, while text enhances the visual appeal and information delivery of a video. Subtitles are typically located at the bottom of the screen and have a relatively fixed format. Text can be placed anywhere in the video and offers a wider variety of styles and animation effects. Subtitles are suitable for situations where understanding the dialogue is crucial, such as movie subtitles. Text is suitable for situations where emphasis or embellishment is needed, such as titles and explanatory text at the beginning and end of credits.
[0097] A media file is a document containing multimedia data. Media files are typically used to store and transmit multimedia data. Multimedia data includes various types of data such as video, audio, images, subtitles, and text. In this application, the type and number of media contained in a media file are not limited. The more types of media a media file contains, the higher the integration level of multiple media content.
[0098] The media data processing method provided in this application will be described in detail below with reference to the accompanying drawings.
[0099] Figure 2 This application provides a schematic diagram of the structure of an encoding / decoding system. The encoding / decoding system 200 includes a source device 210 and a destination device 220. The source device 210 is used to process multiple media contents to generate a media file. The multiple media contents include a first image and other first media content besides the first image. The first media content besides the first image is, for example, at least one of image, video, audio, subtitle, or text. The media file includes data of the multiple media contents and metadata, the metadata being used to indicate display information of the multiple media contents, and the first image and / or the first media content being displayed according to the display information indicated by the metadata. The source device 210 sends the media file to the destination device 220. The destination device 220 is used to obtain the first image and other first media content from the media file. Optionally, the destination device 220 is also used to perform a fusion process on the first image and other first media content, the fusion process including at least one of editing, display, or transcoding.
[0100] Specifically, the source device 210 includes an image acquisition unit 211, a preprocessor 212, an encoder 213, and a communication interface 214.
[0101] Image acquisition device 211 is used to acquire raw images. Image acquisition device 211 includes or is any category of image capture device for, for example, capturing real-world images, and / or any category of image or commentary generation device (for screen content encoding, some text on the screen is also considered an image to be encoded or part of an image). For example, a computer graphics processor for generating computer-animated images, or any category of device for acquiring and / or providing real-world images, computer-animated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). Image acquisition device 211 is a camera for capturing images or a memory for storing images. Image acquisition device 211 also includes any category of (internal or external) interface for storing previously captured or generated images and / or acquiring or receiving images. If image acquisition device 211 is a camera, it is, for example, a local or integrated camera in the source device. If image acquisition device 211 is a memory, it is, for example, a local or integrated memory in the source device. When the image acquisition device 211 includes an interface, the interface may be, for example, an external interface for receiving images from an external video source. The external video source may be, for example, an external image capture device, such as a camera, external storage, or an external image generation device. The external image generation device may be, for example, an external computer graphics processor, a computer, or a server. The interface may be any type of interface according to any proprietary or standardized interface protocol, such as a wired or wireless interface, or an optical interface.
[0102] An image can be viewed as a two-dimensional array or matrix of pixels (picture elements). Pixels in an array are also called sample points. The number of sample points in an array or image along the horizontal and vertical directions (or axes) defines the image's size and / or resolution. To represent color, three color components are typically used; that is, an image can be represented as or contain three sample arrays. For example, in RBG format or color space, an image includes corresponding red, green, and blue sample arrays. However, in image coding or video coding, each pixel is typically represented in a luma / chroma format or color space. For example, for a YUV format image, this includes a luma component indicated by Y (sometimes also indicated by L) and two chroma components indicated by U and V. The luma component Y represents the brightness or grayscale level intensity (e.g., both are the same in a grayscale image), while the two chroma components U and V represent chroma or color information components. Accordingly, a YUV format image includes a luma sample array of luma sample values (Y) and two chroma sample arrays of chroma values (U and V). An RGB format image can be converted or transformed to YUV format, and vice versa; this process is also called color transformation or conversion. If the image is black and white, it includes a luminance sampling array. In this application, the image transmitted from the image acquisition unit 211 to the encoder 213 can also be referred to as raw image data.
[0103] The preprocessor 212 receives the raw image acquired by the image acquisition unit 211 and preprocesses the raw image to obtain a preprocessed image. For example, the preprocessing performed by the preprocessor 212 includes retouching, color format conversion (e.g., from RGB format to YUV format), color adjustment, or noise reduction.
[0104] In some embodiments, source device 210 further includes an audio acquisition unit. The audio acquisition unit is used to acquire raw audio. The audio acquisition unit can be any type of audio acquisition device for capturing real-world sounds, and / or any type of audio generation device. The audio acquisition unit is, for example, a computer audio processor for generating computer audio. The audio acquisition unit can also be any type of memory or storage device for storing audio. Audio includes real-world sounds, virtual scene sounds (such as VR or augmented reality (AR) sounds), and / or any combination thereof. Preprocessor 212 is also used to receive the raw audio acquired by the audio acquisition unit and to preprocess the raw audio. For example, the preprocessing performed by preprocessor 212 includes channel conversion, audio format conversion, or noise reduction.
[0105] Encoder 213 is used to acquire various media content such as images, videos, and audio, and encodes each type of media content separately to obtain encoded data for each type of media content. For example, an image encoding algorithm is used to compress the original image to obtain encoded image data, reducing the data size of the encoded image data while maintaining the quality of the original image as much as possible. Similarly, an audio / video encoding algorithm is used to compress the original audio and video to obtain encoded audio / video data, reducing the data size of the encoded audio / video data while maintaining the quality of the original audio and video as much as possible.
[0106] Optionally, although encoders and encapsulators are functionally distinct, in practice they are often integrated together. The encoder's software or hardware has built-in encapsulation functionality, allowing the encapsulated media file to be obtained directly after data encoding. In some embodiments, the encoder, in addition to encoding data, also has the function of encapsulating the encoded data into a media file.
[0107] Encoder 213 is also used to encapsulate encoded data from various media content into media files. For example, it encapsulates encoded audio and video data, along with other relevant information (such as subtitles and text), into media files for easier playback, transmission, and storage of the encoded data. Common encapsulation formats include MP4, MOV, and MXF.
[0108] For example, encoder 213 includes encoding unit 2131 and encapsulation unit 2132. Encoding unit 2131 is used to encode multiple media contents separately to obtain encoded data of multiple media contents. Encapsulation unit 2132 is used to encapsulate the encoded data of multiple media contents to obtain a media file.
[0109] In this application, the packaging unit 2132 and the encoder 213 can be integrated into one physical device or disposed on different physical devices, without limitation. If the encoder 213 does not include the packaging unit 2132, it means that the packaging unit 2132 and the encoder 213 are two different physical devices.
[0110] The communication interface 214 is used to receive media files generated by the encoder 213 and send the media files through the communication channel 230.
[0111] The target device 220 includes a display 221, a post-processor 222, a decoder 223, and a communication interface 224.
[0112] The communication interface 224 is used to receive media files and transmit them to the decoder 223.
[0113] Communication interfaces 214 and 224 can be used to communicate via a communication link between source device 210 and destination device 220, such as via wired or wireless connection, or via any type of mesh, such as wired mesh, wireless mesh or any combination thereof, any type of private network and public network or any combination thereof, to send or receive data related to media files.
[0114] Both communication interface 214 and communication interface 224 can be configured as follows: Figure 2 The arrow pointing from the source device 210 to the corresponding communication channel 230 of the destination device 220 indicates a one-way communication interface or a two-way communication interface, and can be used to send and receive messages, etc., to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded bitstream transmission, etc.
[0115] Communication interface 214 can also transfer media files to the storage device. Communication interface 224 can also communicate with the storage device to retrieve media files. For example, the storage device can be a readable storage medium, a storage server, a storage gateway, an edge server, a content delivery network (CDN), etc. The CDN can receive and store the media files and then send them.
[0116] Decoder 223 is used to extract multiple media contents from a media file using metadata in the media file, and to present the multiple media contents in a fused form.
[0117] For example, decoder 223 includes decapsulation unit 2231 and decoding unit 2232. Decapsulation unit 2231 is used to obtain encoded data of various media content from the media file using metadata in the media file. Decoding unit 2232 is used to decode the encoded data of various media content to obtain various media content.
[0118] The post-processor 222 receives multiple media contents output by the decoder 223, performs fusion processing on the multiple media contents, and presents the multiple media contents in a fused form. For example, the post-processing performed by the post-processor 222 includes color format conversion (e.g., from YUV format to RGB format), color correction, retouching or resampling, or any other processing.
[0119] Display 221 is used to display media content in a merged format. Display 221 is or includes any category of display device for presenting media content in a merged format. For example, an integrated or external display or monitor. For example, the display may include a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other category of display.
[0120] Both encoder 213 and decoder 223 can be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology is implemented in part in software, the device can store the software instructions in a suitable non-transitory computer-readable storage medium, and one or more processors can be used to execute the instructions in hardware to perform the technology of this disclosure. Any of the foregoing (including hardware, software, combinations of hardware and software, etc.) can be considered as one or more processors.
[0121] The data acquisition device (e.g., an image acquisition device, an audio / video acquisition device) and the encoder 213 can be integrated into a single physical device or set on different physical devices; there is no limitation on this. For example, such as... Figure 2The source device 210 shown includes an image acquisition unit 211 and an encoder 213, indicating that the image acquisition unit 211 and the encoder 213 are integrated into a single physical device. Therefore, the source device 210 can also be referred to as an acquisition device. The source device 210 can be, for example, a mobile phone, tablet computer, computer, laptop computer, camera, wearable device, in-vehicle device, terminal device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, extended reality (XR) device, or other image acquisition devices. If the source device 210 does not include the image acquisition unit 211, it means that the image acquisition unit 211 and the encoder 213 are two different physical devices, and the source device 210 can acquire raw images from other devices (such as image acquisition devices or image storage devices).
[0122] Furthermore, the display 221 and the decoder 223 can be integrated into a single physical device or located on different physical devices; there is no limitation on this. For example, such as... Figure 2 The destination device 220 shown includes a display 221 and a decoder 223, indicating that the display 221 and decoder 223 are integrated into a single physical device. Therefore, the destination device 220 can also be called a playback device. The destination device 220 has the function of decoding and displaying media content in a merged format. The destination device 220 can be, for example, a monitor, television, digital media player, video game console, in-vehicle computer, or other image display device. If the destination device 220 does not include the display 221, it means that the display 221 and decoder 223 are two different physical devices. After decoding the media file, the destination device 220 transmits the merged media content to other display devices (such as a television or digital media player) for display.
[0123] also, Figure 2 The source device 210 and the destination device 220 are shown to be integrated on a single physical device, but they can also be set on different physical devices, without limitation.
[0124] For example, such as Figure 3The diagram in (a) shows a schematic of an encoding / decoding system. Source device 210 is a server, exemplarily a server in a cloud system. Destination device 220 is a display of various possible forms. Source device 210 acquires video of a first scene and generates fused media content based on the video and images within it; alternatively, source device 210 obtains video from a storage device and generates fused media content based on the video and images within it, or generates fused media content based on the video and images contained in various media content. Source device 210 includes an encoding module and an encapsulation module. The encoding module encodes various media content to obtain encoded data, and the encapsulation module encapsulates the encoded data and metadata to obtain a media file, which is then transmitted through a channel.
[0125] The destination device 220 includes a decapsulation module and a decoding module. The decapsulation module uses metadata in the media file to obtain encoded data of various media contents from the media file. The decoding module decodes the encoded data of the various media contents to obtain various media contents. Optionally, the destination device 220 also performs fusion processing on the various media contents according to the metadata to obtain fused media content. For example, the destination device 220 also performs display processing on a first image and / or a first media content according to the metadata. The destination device 220 performs operations such as storing and displaying the fused media content.
[0126] For example, such as Figure 3 The diagram in (b) illustrates an encoding / decoding system where source device 210 and destination device 220 are integrated into the same device, such as a smartphone, tablet, computer, laptop, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, or extended reality (XR) device. This device is capable of encoding and decoding multiple media contents. For example, source device 210 captures video of the user's real-world scene and generates fused media content based on the video and stored images. Destination device 220 displays the real-world scene in a virtual environment.
[0127] In these embodiments, the source device 210 or its corresponding functions and the destination device 220 or its corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. As described, Figure 2 The presence and division of different units or functions in the source device 210 and / or destination device 220 shown may vary depending on the actual device and application, which is obvious to those skilled in the art.
[0128] The structure of the above-described encoding / decoding system is merely illustrative. In some possible implementations, the encoding / decoding system may also include other devices. For example, the encoding / decoding system may also include end-side devices or cloud-side devices. After the source device 210 acquires the original image, it preprocesses the original image to obtain a preprocessed image; and then transmits the preprocessed image to the end-side device or cloud-side device, which performs the encoding / decoding function on the preprocessed image.
[0129] Next, the media data processing process will be explained with reference to the accompanying diagram. Figure 4 Figure (a) shows a flowchart of a media data processing method at an encoding end provided in this application. Figure 2 The following explanation uses the Zhongyuan device 210 as an example to illustrate the process of processing various media content. For example... Figure 4 As shown in (a) of the diagram, the method includes the following steps.
[0130] Step 410: Obtain various media content.
[0131] The multiple media content includes at least images and also other media content. For example, the multiple media content includes a first image and other first media content besides the first image. The first image may include a single image, or multiple images such as a main image (e.g., original image, 3D image, etc.) and secondary images (e.g., thumbnails, etc.). The first media content includes a first video. Optionally, the first media content may also include at least one of a second image, audio, subtitles, or text. Optionally, the second image may be a set of images. Optionally, the second image may include one or more images. Optionally, the second image may include or exclude the first image. Optionally, the content of the first image and the content of the second image may differ.
[0132] This application does not limit the source of various media content. The source device may collect media content itself, obtain media content from memory or other storage, or obtain media content from network resources. As described in the above embodiments, if the source device 210 carries an image acquisition device 211, the source device 210 acquires images through the image acquisition device 211. Optionally, the source device 210 may receive images acquired by other devices; or obtain images from memory or other storage in the source device 210. Images include at least one of real-time acquired real-world images, images stored in the device, and images synthesized from multiple images. This embodiment does not limit the method of image acquisition or the type of images.
[0133] In some embodiments, the source device acquires different media content from different sources. For example, the source device captures images and videos, and acquires at least one of audio, subtitles, or text from memory or network resources. Alternatively, the source device acquires the desired images and videos from a set of images. For example, the set of images includes a set of temporally consecutive images. The source device acquires the desired images and videos from a set of temporally consecutive images. Or, the set of images includes multiple images with different exposure parameters. The source device acquires the desired images and videos from multiple images with different exposure parameters. Or, the set of images includes multiple images with different exposure parameters captured at the same time. The source device acquires the desired images and videos from multiple images with different exposure parameters captured at the same time.
[0134] Optionally, the original media content may also include information about the images or a set of images. This information may include at least one of exposure parameters, timing parameters, or format parameters.
[0135] This application does not limit the methods of acquiring various media content. Examples of methods for acquiring the first image and the first video are given below.
[0136] In some examples, a first image and a first video containing first media content are obtained based on a set of temporally consecutive images. That is, the first image is selected from a set of temporally consecutive images, and images within a certain duration are obtained from it to obtain the first video. Optionally, the first video may or may not contain the first image. Optionally, the first video includes a set of temporally consecutive images.
[0137] In other examples, the first video comprises a subset of images from a set of temporally consecutive images. For instance, the first video is obtained based on a first group of images from a set of temporally consecutive images. The first group of images comprises a subset of images from a set of temporally consecutive images. Alternatively, a second group of images from a set of temporally consecutive images is obtained, and the first video is obtained from the second group of images. For example, the first group of images is obtained by interpolating frames from a set of temporally consecutive images. Or, the second group of images is obtained by interpolating frames from a first group of images from a set of temporally consecutive images.
[0138] In other examples, the first image is derived from a third image group within a set of temporally continuous images. The third image group comprises multiple images with different exposure parameters.
[0139] In other examples, the first image is obtained from a single frame.
[0140] Step 420: Generate media files based on various media content.
[0141] In one possible implementation, a first image is encoded to obtain image data, and an image file is generated based on the image data, the image file including the image data. First media content is encoded to obtain data of the first media content; the data of the first media content is then encoded into the image file to obtain a media file.
[0142] In one embodiment, metadata is encoded into a media file, which is used to indicate data of the first media content.
[0143] For example, reusing fields in an image file to set the data and metadata of the first media content makes the media file also include the data and metadata of the first media content. Alternatively, inserting the data and metadata of the first media content into an image file makes the media file also include the data and metadata of the first media content.
[0144] In another possible implementation, the first image is encoded to obtain image data; the first media content is encoded to obtain data of the first media content; and a media file is generated based on the image data, metadata, and data of the first media content. That is, instead of generating an image file, the image data, the data of the first media content, and the metadata are directly encoded into a media file. In this implementation, relative to the image file, new fields are added to the media file, and the data and metadata of the first media content are located in the newly added fields; or, fields from the image file are reused, and the data and metadata of the first media content are located in the fields of the image file.
[0145] This application does not limit the encoding format of media content such as images, videos, audio, subtitles, or text, nor the format of media files. For example, image codec format types include at least one of the following: Moving Pictures Experts Group (MPEG), Joint Photographic Experts Group (JPEG), Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), or Versatile Video Codec (VVC).
[0146] Audio codec format types include at least one of the following: Pulse Code Modulation (PCM), Differential Pulse Code Modulation (DPCM), Adaptive Differential Pulse Code Modulation (ADPCM), Advanced Audio Coding (AAC), or Free Lossless Audio Codec (FLAC).
[0147] PCM is an uncompressed audio encoding method that converts analog audio signals into digital pulse signals. PCM encoding produces high-quality audio, but also generates a large amount of data.
[0148] DPCM is a lossy audio compression method that reduces the amount of data by predicting the difference between the current sample and the previous sample.
[0149] ADPCM improves upon DPCM by adaptively adjusting the quantization step size to enhance coding efficiency.
[0150] AAC is a lossy audio encoding method that combines the advantages of various audio encoding technologies, resulting in high compression efficiency and sound quality.
[0151] FLAC is a lossless audio encoding method that does not lose any audio information during compression and decompression.
[0152] The subtitle encoding / decoding format type includes at least one of Unicode, UTF-8, or ANSI.
[0153] The text encoding / decoding format type includes at least one of the following: American Standard Code for Information Interchange (ASCII), GBK, Unicode, or UTF-8.
[0154] Media file formats include JPEG (Joint Photographic Experts Group File Interchange Format, JFIF) and High Efficiency Image Format (HEIF). JPEG is a standard for encoding and exchanging digital images. HEIF is an image file format based on HEVC video coding technology.
[0155] For example, the image data obtained by encoding the first image according to the HEIF standard is encoded to obtain an image file in HEIF format. The image data obtained by encoding the first image according to the JPEG or JFIF standard is encoded to obtain an image file in JPEG format.
[0156] As shown in Table 1, the JPEG image file format is illustrated.
[0157] Table 1
[0158]
[0159] The JPEG file format is used to store and distribute a series of JPEG images that are played in sequence to simulate an animation effect.
[0160] The Start of Image (SOI) is used to mark the beginning of JPEG image data. Every JPEG file begins with an SOI marker.
[0161] Application Marker n (APPn), where APP0 stores JFIF information, including version number, thumbnail, etc. JFIF is a common format for JPEG, used to define the structure of JPEG data at the file level.
[0162] The quantization table (DQT) defines the quantization matrix used in the JPEG compression process. Quantization is one of the key steps in converting image data from the spatial domain to the frequency domain, determining the quality of the compressed image and the file size.
[0163] Huffman tables (DHT) are used in JPEG compression with Huffman coding, a variable-length coding technique used to further reduce the size of image data. DHTs optimize based on the statistical properties of image data to improve compression efficiency.
[0164] The Restart Interval (DRI) is used to insert a restart marker during JPEG compression so that decoding can restart from a specific point during decoding, which helps to restore synchronization in case of decoding errors.
[0165] The Start of Frame (SOF) is used to mark the beginning of a JPEG image frame and contains information about the image size, number of color components, sampling factor, etc.
[0166] The Start of Scan (SOS) marker indicates the start of a JPEG image scan and specifies the color components to be scanned and the scanning mode. The actual compressed data follows the SOS marker.
[0167] Compression data is the actual data of a JPEG image, obtained after compression processes such as quantization and Huffman coding. This data is the core part of the image, containing all its information.
[0168] The End of Image (EOI) is used to mark the end of JPEG image data. Every JPEG file ends with an EOI marker.
[0169] As shown in Table 2, the HEIF format image file format is illustrated.
[0170] Table 2
[0171]
[0172] The file type (File-type Box, Box type is ftyp) is used to identify the file's format and version, ensuring that the decoder can correctly parse the file. For example, for an HEVC encoded video file, the ftyp field may indicate that it is an ISO Base Media File Format (isom) file, and version and compatibility flags will also be included.
[0173] Metaboxes (Box type: meta) refer to ISO 14496-12 or ISO 23008-12 and contain various information about media files such as hdlr, pitm, and idat, as shown in Table 2 above. For example, author, title, date, copyright, etc. Optionally, metadata may also include technical information about encoding settings and color spaces.
[0174] The processor or processor type (Handler Box, Box type is hdlr) indicates the type of software or hardware processor that handles the media content. For example, it indicates that the video is processed by a specific decoder or player.
[0175] The Primary Item Box (box type pitm) is used to identify the main item or content in a media file. In files containing multiple media tracks (such as video, audio, subtitles, etc.), the pitm field indicates which track is the primary track or should be played first.
[0176] Item data (Box type idat) is used for contexts that may depend on its implementation. In some cases, it may be used to store additional data related to media items.
[0177] The Item Info Box (box type iinf) contains detailed information about a specific item in a media project. For example, it may contain information such as the resolution, frame rate, and bitrate of a video track.
[0178] Item Info Entry Boxes (Box type: INFE) are used to specifically describe an item. For example, hvc1 indicates that this is an HEVC-encoded video data item.
[0179] The Item Location Box (box type iloc) provides the physical location information of data items in a media file. This helps the decoder quickly locate and read the required media data.
[0180] The Item Properties Box (Box type: IPRP) contains information about the properties of specific items in a media project. For example, it may contain technical details about color spaces, encoding settings, etc.
[0181] The Item Property Container Box (Box type is ipco) contains the actual item property information.
[0182] Color attributes or color information (Color information Box, Box type 'color') provide information about the colors of a media item, such as color space, gamma value, etc. Under Ipco, repetition may indicate different contexts or uses.
[0183] Item reference boxes (box type isiref) provide reference information between different items in a media file. For example, they might be used to indicate which video track a subtitle track is associated with.
[0184] Under `iref`, `dimg` may indicate an image reference or a similar meaning (the specific meaning may depend on the context) referencing image data in a media file.
[0185] A Media Data Box (box type: mdat) contains the actual media data, such as video frames and audio samples. This is the most important data part of a media file.
[0186] Data1 could be a placeholder or example name. In this context, Data1 likely refers to the actual image bitstream data stored in the mdat box. This is the core data the decoder needs to decode and display the image.
[0187] In some embodiments, optionally, the metadata includes tag metadata. Tag metadata is used to indicate that the media file includes first media content.
[0188] The source device determines the tag metadata based on the content of the first media.
[0189] For example, if the source device determines that the acquired media content includes a first media content other than the first image, it determines that the media file contains tag metadata.
[0190] In other words, a media file containing tag metadata indicates that it contains first media content besides the first image. Conversely, a media file without tag metadata indicates that it only contains the first image, meaning it does not contain any other first media content besides the first image. In other words, a media file containing tag metadata indicates that it contains first media content besides the directly decoded image. Conversely, a media file without tag metadata indicates that it only contains the directly decoded image, meaning it does not contain any other first media content besides the directly decoded image.
[0191] For example, different values of the tag metadata contained in a media file indicate whether the media file contains first media content other than the first image. For instance, if the tag metadata contained in a media file has a first value (e.g., a value of 1), it means that the media file contains first media content other than the first image. If the tag metadata contained in a media file has a second value (e.g., a value of 0), it means that the media file contains only the first image, that is, the media file does not contain first media content other than the first image.
[0192] The following describes how to set tag metadata in media files for different media file formats.
[0193] When the media file is in a format that meets the HEIF standard, the inclusion of tag metadata in the media file includes the following methods.
[0194] Method 1: The tag metadata is located in the `ftyp` field of the media file. For example, as shown in Table 3, the tag metadata is located in the data segment corresponding to the `ftyp` field in the media file. That is, in the data segment corresponding to the `ftyp` field of the HEIF format image file shown in Table 2.
[0195] Table 3
[0196] ftyp Tag metadata meta … … … … …
[0197] Method 2: The tag metadata is located in the meta field of the media file. For example, as shown in Table 4, the tag metadata is located in the data segment corresponding to the meta field in the media file. That is, in the data segment corresponding to the meta field of the HEIF format image file shown in Table 2.
[0198] Table 4
[0199] ftyp meta Tag metadata … … … … …
[0200] Method 3: The meta field includes the iinf subfield. The tag metadata is located in the data segment corresponding to the type in the iinf subfield, excluding the image codec format type.
[0201] Among them, the image codec format type includes at least one of MPEG, JPEG, AVC, HEVC or VVC.
[0202] The types mentioned above, excluding image encoding / decoding format types, include at least one of uri, mime, it35, xml, exif, or live.
[0203] For example, the tag metadata is located in the data segment corresponding to the type URI in the iinf subfield, excluding the image codec format type.
[0204] For example, the tag metadata is located in the data segment corresponding to the type MIME in the iinf subfield, excluding the image codec format type.
[0205] For example, the tag metadata is located in the data segment corresponding to type it35 in the iinf subfield, excluding image codec format types.
[0206] For example, the tag metadata is located in the data segment corresponding to the type xml in the iinf subfield, excluding image encoding / decoding format types.
[0207] For example, the tag metadata is located in the data segment corresponding to the type exif in the iinf subfield, excluding image codec format types.
[0208] For example, as shown in Table 5, the tag metadata is located in the data segment corresponding to the first type of the iinf subfield, excluding the image codec format type, such as the data segment corresponding to the live type. That is, the meta field of the HEIF format image file shown in Table 2 contains the data segment corresponding to the first type under the iinf subfield. For example, the tag metadata is located in the idat field corresponding to the item of the type other than the image codec format type in the iinf subfield.
[0209] Table 5
[0210]
[0211] Optionally, the meta field includes an hdlr subfield. Tag metadata resides in the data segment corresponding to the hdlr subfield.
[0212] Method 4: The meta field includes the iloc subfield, iprp subfield, ifef subfield, and idat subfield. Tag metadata resides in at least one of the iloc, iprp, ifef, or idat subfields.
[0213] For example, the tag metadata is located in the data segment corresponding to the iloc subfield.
[0214] For example, the tag metadata is located in the data segment corresponding to the IPRP subfield.
[0215] For example, the tag metadata is located in the data segment corresponding to the ifef subfield.
[0216] For example, the tag metadata is located in the data segment corresponding to the idat subfield.
[0217] For example, as shown in Table 6, the tag metadata is located in the data segment corresponding to the idat subfield. That is, in the data segment corresponding to the idat subfield contained in the meta field of the HEIF format image file shown in Table 2.
[0218] Table 6
[0219]
[0220] Method 5: In the HEIF format image file shown in Table 2, the ftyp, meta, and mdat fields are used as items in the image file. A first field is added to the media file to identify and tag metadata. For example, the media file adds an item of type 1.
[0221] When the media file is in JPEG format, the inclusion of tag metadata in the media file includes the following methods.
[0222] Method 1: The tag metadata is located in the APPn field of the media file. For example, as shown in Table 7, the tag metadata is located in the APPn field of the media file. That is, in the data segment corresponding to the APPn added to the JPEG format image file shown in Table 1.
[0223] Table 7
[0224]
[0225] Method 2: The tag metadata is located in the MPF of the media file. For example, as shown in Table 8, the tag metadata is located in the MPF of the media file, specifically in the data segment corresponding to the APPn added to the JPEG format image file shown in Table 1.
[0226] Table 8
[0227]
[0228]
[0229] For example, the tag metadata is located in at least one of the tags, tags, or payloads of an APPn in the media file. For example, the tag metadata is located in an APPn of the first type in the media file. Or, the tag metadata is located in an APPn in the media file that contains a first preset character.
[0230] Optionally, the metadata may also include at least one of attribute metadata, location metadata, time-related metadata, or transformation-related metadata. Attribute metadata indicates the attributes of the first media content included in the media file. Attribute metadata indicates the types of media content included in the media file. Location metadata indicates the location of the first media content in the media file. Location metadata indicates the location of the media content included in the media file. Time-related metadata indicates the temporal relationship between the first media content and the first image included in the media file. Transformation-related metadata determines the transformation performed on the first media content and / or the first image. Optionally, transformation-related metadata indicates the method of transformation performed on the first media content and / or the first image.
[0231] Optionally, the metadata may also include display metadata. Display metadata indicates display information for various media content.
[0232] In some embodiments, display metadata is used to indicate at least one of the display methods for a cover frame or various media content.
[0233] A cover frame typically refers to one or more frames in a video or image sequence. It serves as the cover image or representative image of a video or image sequence.
[0234] Optionally, a cover frame can also refer to a browse frame or thumbnail. Display metadata also indicates a browse frame or thumbnail. A cover frame is, for example, one or more images displayed in a list of images or videos. A cover frame is, for example, a display image of an image or video in a list of images or videos (e.g., an image displayed after clicking on a video or image in the list).
[0235] The following describes how the metadata indicator cover frame is implemented.
[0236] Method 1: Display metadata indicating that at least one of the multiple images included in various media content is used as the cover frame.
[0237] For example, the metadata indicates that the first image is the cover frame. Similarly, the metadata indicates that the second image in the multi-media content is the cover frame. Again, the metadata indicates that both the first and second images are cover frames. Even if the second image is a set of images, the metadata indicates that one or more frames from that set are the cover frame. And again, the metadata indicates that one or more frames from the first video in the multi-media content are the cover frame.
[0238] Optionally, the cover frame is obtained based on frame selection information indicated by the display metadata. For example, the display metadata includes a frame number (an index of multiple images in the second image or a frame number in the first video, etc.). The frame indicated by the frame number serves as the cover frame.
[0239] Method 2: Display metadata indicating that the cover frame is the result of processing at least one of the multiple images included in the various media content.
[0240] For example, the processing results include the deformation results of at least one of multiple images.
[0241] Optionally, the display metadata includes the frame number of at least one frame and the processing method. The processing method includes information such as the display position, size, scaling, and cropping of the cover frame.
[0242] Optionally, the processing method includes coloring the blank areas. For example, the processing method includes filling the blank areas with a uniform color.
[0243] Optionally, the processing method includes a multi-frame cover presentation. The presentation method can be any of the following: grid, page turning, or custom.
[0244] Optionally, the processing method includes traversing the image / video frames in the file according to the image list information involved, and processing each image / video frame: scaling the image according to the scaling relationship; and drawing the scaled image onto the image to be displayed.
[0245] The processing results include the results of processing at least one item according to the above processing methods.
[0246] For example, the processing result may include the fusion result of at least one of multiple images.
[0247] Optionally, the processing result includes a fusion of an image and the deformed result of that image.
[0248] Optionally, the processing result may include the fusion result of multiple images.
[0249] For example, the processing result includes the fusion result of at least one of multiple images with media content other than images in the first media content.
[0250] Optionally, the display metadata includes the frame number of at least one frame and the media content other than images in the first media content. The media content other than images in the first media content includes at least one of audio, text, or subtitles.
[0251] For example, the processing result includes at least one deformed image and a fusion result with media content other than the image in the first media content.
[0252] Method 3 displays metadata indicating the format of the cover frame. Formats include brightness, contrast, SDR, or HDR, etc.
[0253] In other embodiments, the display method includes display style and / or playback information of multiple frames of images contained in the first image and the first media content.
[0254] Display styles include the presentation format or style of images and / or videos. For example, a display style includes the display style of a processing result obtained by processing at least one image according to the above processing method. Another example is the display style after processing a cover frame. Yet another example is the display style of an image overlaid on video.
[0255] Playback information includes at least one of the following: playback order, starting frame, or image switching method.
[0256] For example, the playback information includes the playback order of the first image and the second image contained in the first media content.
[0257] For example, playback information includes the playback order of the first image and the first video contained in the first media content.
[0258] For example, playback information includes the first image being the starting frame of the video playback.
[0259] For example, playback information may include direct switching operations between two consecutive frames, which may include scaling or other operations.
[0260] For example, playback information includes introducing fade-in and fade-out operations between two consecutive frames.
[0261] For example, playback information includes introducing deformation switching operations between two consecutive frames.
[0262] For example, playback information includes time data for multiple images, and operations such as direct switching, fade-in / fade-out, and fly-out are introduced based on this time data. For instance, playback information includes a timeline data set. This timeline data set contains multiple timeline data points, each associated with one or more operations, and each operation includes one or more other operations. For example, operations such as direct switching, fade-in / fade-out, and fly-out are introduced based on the time data of two consecutive frames. It should be understood that the aforementioned timeline data is used to represent the time (playback) sequence of multiple images.
[0263] For example, playback information includes time data of multiple images, text is added based on the time data of multiple images, and operations such as direct switching, fade-in and fade-out, and fly-out of the screen are introduced.
[0264] Optionally, the display metadata may also include at least one of the following: display order, position, size, or scaling information for the display styles corresponding to various media content.
[0265] The following describes the location of the display metadata in the media file.
[0266] Optionally, the source device sets display metadata in the media file based on the method of setting tag metadata in the media file.
[0267] In some embodiments, the field that identifies the tag metadata is the same as the field that identifies the display metadata.
[0268] For example, as shown in Table 9, both tag metadata and display metadata are located in the ftyp field of the media file.
[0269] Table 9
[0270] ftyp Tag metadata / Display metadata meta … … … … …
[0271] For example, as shown in Table 10, both tag metadata and display metadata are located in the meta field of the media file.
[0272] Table 10
[0273] ftyp meta Tag metadata / Display metadata … … … … …
[0274] For example, both tag metadata and display metadata reside in the same field within the media file, and the data types of the fields identifying tag metadata and display metadata are the same. As shown in Table 11, both tag metadata and display metadata reside in the infe subfield of the iinf subfield contained in the meta field of the media file. For instance, tag metadata resides in the data segment corresponding to the type live of the infe subfield. Display metadata resides in the data segment corresponding to the type live of the infe subfield.
[0275] Table 11
[0276]
[0277] In some embodiments, the field that identifies the tag metadata is different from the field that identifies the display metadata.
[0278] For example, as shown in Table 12, the tag metadata is located in the ftyp field of the media file, and the display metadata is located in the meta field of the media file.
[0279] Table 12
[0280] ftyp Tag metadata meta Display metadata … … … … …
[0281] Optionally, as shown in Table 11, the tag metadata is located in the live type subfield of the iinf field, and the display metadata is located in the idat field.
[0282] For example, as shown in Table 13, the tag metadata is located in the iinf subfield contained in the meta field of the media file. For instance, the tag metadata is located in the data segment corresponding to the live type in the iinf subfield, excluding the image codec format type. The display metadata is located in the iprp subfield contained in the meta field of the media file.
[0283] Table 13
[0284]
[0285] For example, both tag metadata and display metadata reside in the same field within the media file; however, the data types of the field identifying tag metadata and the field identifying display metadata differ. As shown in Table 14, both tag metadata and display metadata reside in the infe subfield of the iinf subfield contained within the meta field of the media file. Tag metadata is located in the data segment corresponding to the live type of the infe subfield. Display metadata is located in the data segment corresponding to the mime type of the infe subfield.
[0286] Table 14
[0287]
[0288] For example, as shown in Table 15, the tag metadata is located in the data segment corresponding to the type `live` of the `infe` subfield. The display metadata is located in the data segment corresponding to the type `xml` of the `infe` subfield.
[0289] Table 15
[0290]
[0291]
[0292] Optionally, the metadata may also include at least one of location metadata, time-related metadata, and transformation-related metadata. The fields identifying tag metadata, location metadata, time-related metadata, transformation-related metadata, and display metadata are the same. As shown in Table 16, tag metadata, location metadata, time-related metadata, transformation-related metadata, and display metadata are all located in the ftyp field of the media file.
[0293] Table 16
[0294]
[0295] Optionally, as shown in Table 17, tag metadata, location metadata, time-related metadata, transformation-related metadata, and display metadata are all located in the meta field of the media file.
[0296] Table 17
[0297]
[0298] Optionally, as shown in Table 18, the tag metadata is located in the ftyp field of the media file. Location metadata, time-related metadata, transition-related metadata, and display metadata are all located in the meta field of the media file.
[0299] Table 18
[0300]
[0301] In some embodiments, metadata is set according to the association between fields that identify tag metadata and fields that identify display metadata.
[0302] For example, a field that identifies tag metadata can be associated with a field that identifies display metadata. Tag metadata and display metadata are placed in the data segment corresponding to the associated field.
[0303] For example, as shown in Table 13, the iinf field, which identifies the metadata of the tag, is associated with the ifef field, which identifies the metadata of the display.
[0304] For example, the type of the identifier tag metadata is associated with the type of the identifier display metadata. For instance, the type refers to the type or tag of box or fullbox. For example, the first type of the identifier tag metadata is associated with the IPRP of the identifier display metadata.
[0305] For example, a field identifying display metadata includes at least one of the IPRP subfields or the `iref` subfield; a field identifying tag metadata is associated with at least one of the IPRP subfields or the `iref` subfield. Alternatively, a field identifying display metadata includes an `iref` subfield, and a field identifying tag metadata is associated with the `iref` subfield.
[0306] In some embodiments, metadata is set through a single item association rule. The type of the metadata is identified as the first type of item in the iinf field, and the iinf field is associated with the iloc field through the item. The field that identifies the displayed metadata is iloc. The iinf field is associated with the iprp field through the item.
[0307] In some embodiments, metadata is set through association rules for multiple items. The types of identifier tag metadata, identifier time-related metadata, and identifier transformation-related metadata are associated through item association rules. For example, as shown in Table 21, image data items use item items of type hvc1, and tag metadata uses item items of type first; the two items are associated through iref.
[0308] In some embodiments, metadata is set through multiple items and a single association rule. The type of the tag metadata and the type of the display metadata are associated through item association rules. For example, as shown in Table 21, image data items use item items of type hvc1, tag metadata uses item items of type 1, and the two items are associated through iref. Alternatively, image data items use item items of type hvc1, video data items use item items of type MIME, URI, or type 1, and tag metadata items use item items of a different type than video (type 1), and the three items are associated through iref. Simultaneously, these image data items, video data items, and tag metadata items are associated with the iloc and iprp fields through a single item association rule.
[0309] In some embodiments, association is achieved through predefined rules for the brand included in ftyp. For example, as shown in Table 18, tag metadata resides in the ftyp field of the media file, while display metadata resides in the meta field of the media file. The ftyp and meta fields belong to two separate items; when these two items are associated, the field identifying the tag metadata is associated with the field identifying the display metadata. The first type of brand needs to include a specific combination of fields from the meta tag.
[0310] In some embodiments, the type of the identifier tag metadata is the first type in the iinf field; the type of the identifier image or video is hvc1, mime, or uri in the iinf field; the iinf field is associated with the iloc field; the type of the identifier location-related metadata is iloc; the iinf field is associated with the IPrp field; the type of the identifier time or attribute-related metadata is IPrp; the iinf field is associated with the idat field; and the identifier transformation-related metadata and display metadata are marked. The first type uses any one of the abbreviations or full names related to continuous shooting and dynamic images, such as Live, motion, movie photo, mpho, Continuous shooting, Live, mpvd, mpic, Live / motion / movie / multiple photo, Continuous shooting, burst mode, etc.
[0311] The first field can use any of the following abbreviations or full names related to continuous shooting and moving images: Live, motion, movie photo, mpho, Continuous shooting, Live, mpvd, mpic, Live / motion / movie / multiple photo, Continuous shooting, burst mode, etc.
[0312] In some embodiments, the first media content includes multiple media contents, and the media file includes display metadata for each type of media content. For example, as shown in Table 19, the tag metadata is located in the iinf subfield contained in the meta field of the media file. For instance, the tag metadata is located in the data segment corresponding to type live in the iinf subfield, excluding the image codec format type. The display metadata for the first type of media content is located in the iprp subfield contained in the meta field of the media file. The display metadata for the second type of media content is located in the idat subfield contained in the meta field of the media file.
[0313] Table 19
[0314]
[0315] Optionally, as shown in Table 20, the tag metadata for the first type of media content is located in the data segment corresponding to the type `live` in the `iinf` subfield, excluding the image codec format type. The display metadata for the first type of media content is located in the `iloc` subfield contained in the `meta` field of the media file. The tag metadata, time-related metadata, transformation-related metadata, and display metadata for the second type of media content are located in the `idat` subfield contained in the `meta` field of the media file.
[0316] Table 20
[0317]
[0318] Add the first media content to the media file. The method for setting the data of the first media content in the media file varies depending on the media file format.
[0319] When the media file is in HEIF format, the first media content can be set in the media file in the following ways.
[0320] Method 1: The data for the first media content is located in the `mdat` field. If the first media content is video, the video stream is located in the `mdat` field.
[0321] Method 2: The data for the first media content is located in the idat subfield of the meta field. If the first media content is video, the video stream is located in the idat subfield of the meta field.
[0322] Method 3: The data of the first media content is located in the moov field. As shown in Table 23, for example, the first media content is video, and the video stream is located in the moov field. In an optional embodiment, the moov packet is obtained from the media content data (e.g., video MP4) and added to the media file. The image data of the first image is added to the mdat field of the media file. The offset metadata of moov is modified, with moov indicating the offset to start from the idat start position.
[0323] Method 4: The data for the primary media content is carried in other fields at the same level as ftyp, meta, mdat, moov, etc. If the primary media content is video, the video stream is carried in the free field. Alternatively, the video stream is carried in any field related to abbreviations or full names of continuous shooting or moving images, such as Live, motion, movie photo, mpho, Continuous shooting, Live, mpvd, mpic, Live / motion / movie / multiple photo, Continuous shooting, burst mode, etc.
[0324] For example, primary media content data includes video streams.
[0325] When the media file is in JPEG format, the data containing the first media content in the media file can include the following methods.
[0326] Method 1: Add the data of the first media content to the end of the image file of the first image. For example, add the data of the first media content between SOS and EOI.
[0327] Method 2: Add the first media content from the JUMBF package. Method 3: Add the first media content to the end of the media file. Method 4: Add the first media content after the independent sub-file format within the media file.
[0328] In JPEG media files, APPX is used to identify and mark metadata.
[0329] JPEG User Marker Blocks for Features (JUMBF) may be an application- or standard-specific term used to describe user-defined marker blocks in a JPEG image. These marker blocks are used to store various additional information or functions related to the image.
[0330] APP11 JUMBF1 may represent the first JUMBF block, used to store certain characteristics or information.
[0331] APP11 JUMBF2 may represent a second JUMBF block used to store another specific feature or information.
[0332] APP11 JUMBF3 may represent the third JUMBF block, and so on.
[0333] APP11 JUMBF4 may represent the fourth JUMBF block.
[0334] For suitability, the above APP11 JUMBF1~4 are data of the primary media content.
[0335] In some embodiments, when the media file is in a JPEG format, the JPEG-formatted media file includes content related to the HEIF file format.
[0336] By encapsulating HEIF-related content within JPEG media files, unified management of media content is achieved. This facilitates the expansion of various media content formats, ensures media file standards are as compatible as possible with existing standards, and enhances the convenience of media content management.
[0337] When the media file format conforms to the JPEG file format, the media file includes a new APPn. The APPn tag indicates HEIF. An example, as shown in Table 21, is a general structure for a media file. The media file includes an image. The media file includes APPn:heif. Optionally, the media file may also include video.
[0338] Table 21
[0339]
[0340]
[0341] Table 22 shows the structure of a multi-layer HDR image-related media file. The media file includes various media content. For example, a media file may include multiple images and videos.
[0342] Table 22
[0343]
[0344] APPn includes APP start, APP length, tag, and payload.
[0345] The APPn marker indicates a data structure related to HEIF, meaning that the APPn includes content related to the HEIF file format. For example, the APPn marker indicates HEIF, HEIC, or MIAF, etc. Optionally, the APPn marker indicates Urn:iso:std:iso:23008:-12, indicating that the APPn includes content related to the HEIF file format.
[0346] The data structures associated with the HEIF file format include the File-type Box, Metabox, and Media Data Box Header.
[0347] As shown in Table 23, the payload of APPn includes the File-type Box, Metabox, and Mdat Box Header.
[0348] Table 23
[0349]
[0350] Optionally, as shown in Table 1, JPEG format image files include SOI and EOI, and between SOI and EOI, DQI, Compression Data, etc., are also included. Therefore, the payload of APPn also includes SOI. For example, as shown in Table 23, the payload field includes SOI.
[0351] Metadata is located in the APPn of the media file. If the APPn is marked as HEIF, the metadata is located in the payload of the APPn in the media file. Optionally, the media file may contain metadata in the manner described in Tables 3 to 6 of the above embodiments.
[0352] The new APPn payload contains structures related to the HEIF file, such as the File-type Box, Metabox, and Mdat Box Header. These structures allow linking to or access to the first image and / or the first media content. For example, after removing the portion before the new APPn payload in the media file, the remaining portion can be processed as a complete HEIF structure (in other words, the new APPn payload combined with subsequent data can be processed as a complete HEIF structure). For example, an SOI can be added after the Mdat Box Header in the new APPn payload. For example, some or all of the data following the new APPn in the current media file can be considered as included in the Mdat data (data[]).
[0353] For example, as shown in Table 24, the SOI in the APPn payload includes data following the APPn in JPEG file format, marked with HEIF. The SOI in the APPn payload also includes other APPn (e.g., APPn:MPF, in this example, APPn:heif before APPn:MPF), DQT, Compressed Data, and EOI, as well as other complete files. For example, the SOI in the APPn payload also includes complete JPEG file format media files and complete MP4 file format media files.
[0354] Table 24
[0355]
[0356]
[0357] For example, as shown in Table 24, the four columns on the left show the payload of APPn in the media file and the subsequent data structure, while the three columns on the right show the result of heif encapsulation / management obtained by combining the heif-related structures in the payload of APPn with the above data structure.
[0358] Metadata related to photo / video and other media content has been added to image files. By improving the structure of media files and setting metadata, various media content can be processed in a unified manner, expanding the forms of fusion of various media content. For example, as shown in Table 25.
[0359] Table 25
[0360] Metadata related to photo / video Metadata related to video or image with otherMedia Other media besides images otherMediaNumber Other media otherMediaFormat Other media types Position video file start position length Video file length
[0361] The storage order and quantity of the metadata in the table above are not restricted.
[0362] The above embodiments illustrate the media data processing method at the encoding end. The media data processing method at the decoding end will be described below.
[0363] Here Figure 2 The following explanation uses the process of performing multiple media content processing on the intermediate target device 220 as an example. Figure 4 As shown in (b) of this application, it is a flowchart illustrating a media data processing method at a decoding end. The method includes the following steps.
[0364] Step 430: Obtain the media file. The media file includes the image data of the first image, metadata, and data of the first media content.
[0365] The destination device receives a media file from the source device. Alternatively, the destination device generates the media file, or retrieves the media file from its own memory or other storage. The description of the media file, the first image, the first media content, and the metadata is as explained in step 410.
[0366] Step 440: Obtain the first image from the media file.
[0367] In some embodiments, a first image is obtained by decoding a media file. The media file is directly decoded as image data to obtain the first image.
[0368] In some embodiments, image data in a media file is decoded to obtain a first image. The process involves acquiring image data from a media file, decoding the image data, and obtaining the first image.
[0369] Optionally, data blocks in the media file are deleted, and the resulting file is decoded to obtain the first image. For example, the unnecessary parts of the media file are determined based on the format of the first image (or the media file format), and then the unnecessary parts are deleted (e.g., deleting content related to the app). The remaining parts are then decoded to obtain the first image. Optionally, since the positions of fields in the remaining content have changed due to the deletion of some content in the media file, the positions of the remaining content in the media file are modified, and decoding is performed based on the modified content to obtain the first image. Optionally, data of the first media content in the media file can also be deleted, and then decoding is performed based on the remaining content to obtain the first image.
[0370] Optionally, the fields required for decoding the first image are obtained from the media file according to the format of the first image (or the media file format), and the data of the first image is decoded based on the obtained fields to obtain the first image.
[0371] For example, when the media file is in the HEIF file format, the first image is obtained from the HEIF file file in one or more of the following ways.
[0372] Method 1: Treat the entire HEIF file as a static image file, decode it according to the HEIF standard, and obtain the first image.
[0373] Method 2: Obtain the primary image and related fields (ftyp, meta, mdat) from the media file as an image file and decode it to obtain the first image.
[0374] Method 3: Obtain other types of standard files (e.g., ISO 21496-1, ISO 22028-5, HEIF tmap, etc.) from the media file as image files and decode them to obtain the first image.
[0375] Method 4: Remove all fields except ftyp, meta, and mdat from the media file and decode it as an image file to obtain the first image.
[0376] Method 5: Remove all fields except ftyp, meta, and mdat from the media file, modify some data in the meta field, and then decode it as an image file to obtain the first image.
[0377] When the media file is in JPEG format, the first image is obtained from the JPEG media file using one or more of the following methods.
[0378] Method 1: Treat the entire JPEG file as a still image file, decode it according to the JPEG standard, and obtain the first image.
[0379] Method 2: Obtain the complete JPEG structure confirmed from the media file, decode it into an image file, and obtain the first image.
[0380] Method 3: Confirm the complete JPEG structure obtained from the media file, remove the APP except for APP0 JFIF, and / or remove the data with the file extension of the complete structure file, and decode the obtained part as an image file to obtain the first image.
[0381] Method 4: Obtain other types of standard files (e.g., ISO 21496-1, ISO 22028-5, HEIF tmap, JPEG standard extension standards, etc.) from the media file as image files and decode them to obtain the first image.
[0382] Step 450: Obtain the first media content from the media file.
[0383] For example, when the media file is in HEIF format, the first media content can be obtained from the media file in the following ways.
[0384] Method 1: The data for the first media content is located in the `mdat` field. If the first media content is video, the video bitstream is obtained from the `mdat` field.
[0385] Method 2: The data for the first media content is located in the `idat` subfield of the `meta` field. If the first media content is video, the video stream is obtained from the `idat` subfield of the `meta` field.
[0386] Method 3: The data for the first media content is located in the moov field. If the first media content is video, the video stream is obtained from the moov field.
[0387] Method 4: The data of the first media content is carried in free or other first type or first field at the same level as meta and mdat.
[0388] For example, when the media file is in JPEG format, obtaining the first media content from the media file includes the following methods.
[0389] Method 1: Obtain the first media content from the end position of the first image file.
[0390] Method 2: Obtain the first media content from the JUMBF package.
[0391] Method 3: Obtain the first media content from the end of the media file.
[0392] Method 4: Add the first media content after the independent sub-file format in the media file.
[0393] Optionally, the metadata includes tag metadata. Tag metadata is used to indicate that the media file includes first media content. The first media content is retrieved from the media file based on the tag metadata.
[0394] In some embodiments, the destination device determines whether a media file contains first media content based on tag metadata. If the media file contains tag metadata, it is determined that the media file contains first media content. If the media file does not contain tag metadata, it is determined that the media file does not contain first media content.
[0395] Optionally, the relevant fields or types carried by the above-mentioned tag metadata conform to the first type or the first field.
[0396] When the media file is in HEIF format, the following methods can be used to obtain tag metadata from the media file.
[0397] Method 1, as shown in Table 3, involves retrieving the tag metadata from the `ftyp` field in the media file.
[0398] Method 2, as shown in Table 4, involves retrieving the tag metadata from the meta field of the media file.
[0399] Method 3: The meta field includes the iinf subfield. As explained in step 420 above, the tag metadata is located in the data segment corresponding to the type other than the image codec format type in the iinf subfield. The tag metadata is retrieved from the data segment corresponding to the type other than the image codec format type in the iinf subfield.
[0400] Method 4: The meta field includes the iloc subfield, iprp subfield, ifef subfield, and idat subfield. As described in step 420 above, the tag metadata is located in at least one of the iloc, iprp, ifef, or idat subfields. The tag metadata is obtained from the data segment corresponding to at least one of the iloc, iprp, ifef, or idat subfields.
[0401] Method 5: Add a new item box to the media file to identify the tag metadata. Retrieve the tag metadata from the new item box.
[0402] Optionally, the relevant fields or types carried by the above-mentioned tag metadata conform to the first type or the first field.
[0403] When the media file is in JPEG format, the following methods can be used to obtain tag metadata from the media file.
[0404] Method 1, as described in step 420 above, involves retrieving the tag metadata from the APPn file within the media file.
[0405] Method 2, as described in step 420 above, involves retrieving the tag metadata from the MPF file within the media file.
[0406] After determining that the media file contains tag metadata, the target device obtains the first media content in the media file based on the tag metadata.
[0407] In some embodiments, when the media file is in JPEG format and the JPEG-formatted media file includes content related to the HEIF file format, the media file is decoded according to the JPEG file format, and the content related to the HEIF file format is decoded according to the HEIF file format.
[0408] For example, if the media file is in JPEG format, it includes a new APPn. The APPn is marked with HEIF, and the metadata is located in the APPn within the media file. The metadata, the first image, and the first media content are then obtained from the media file according to the acquisition method described above for media files in HEIF format. The metadata includes at least one of the following: tag metadata, location metadata, time-related metadata, transformation-related metadata, or display metadata.
[0409] For example, as shown in Table 24, assume that item 1 is a tag metadata field, located in the subfield 'infe' within the 'iinf' subfield. The location information of the data corresponding to item 1 is located in the subfield 'iloc' within the 'meta' field. The data block of item 1 is obtained based on the location information of the data corresponding to item 1.
[0410] Assume item 2 is display metadata, located in the subfield 'infe' within the 'iinf' subfield. The location information for the data corresponding to item 2 is located in the subfield 'iloc' within the 'meta' field. The data block for item 2 is retrieved based on this location information.
[0411] In some embodiments, the media file is used to present media content in a merged form. The merged media content may also be simply referred to as merged media. After the destination device acquires the first image and the first media content, it performs a fusion process on the first image and the first media content. The fusion process includes at least one of editing, displaying, or transcoding. The first media content includes at least one of a second image, a first video, audio, subtitles, or text.
[0412] For example, the first image and the second image are merged to obtain a merged media, which includes the first image and the second image.
[0413] For example, the first image and the first video can be merged to obtain merged media, which includes the first image and the first video.
[0414] For example, the first image and audio can be merged to obtain merged media, which includes the first image and audio.
[0415] Optionally, the metadata may also include display metadata. The target device performs display processing on the first image and the first media content based on the display metadata, making the fused media content presentation richer and more diverse, and the display effect more attractive.
[0416] In some embodiments, metadata is displayed to indicate the cover frame.
[0417] Method 1: The target device uses at least one of the multiple images as the cover frame according to the instructions of the display metadata.
[0418] For example, the first image can be used as the cover frame. Alternatively, a second image from multiple media content can be used as the cover frame. Another example is using both the first image and the second image from multiple media content as the cover frame. Yet another example is where the second image is a set of images, and one or more frames from that set are used as the cover frame. Yet another example is using one or more frames from the first video within multiple media content as the cover frame.
[0419] Optionally, the frame indicated by the frame number contained in the metadata will be displayed as the cover frame.
[0420] Method 2: The target device uses the processing result of at least one of the multiple images as the cover frame according to the instructions of the display metadata.
[0421] For example, deform the first image and use the deformed result as the cover frame.
[0422] For example, at least one frame of the multiple images included in the first media content is deformed, and the deformation result is used as the cover frame.
[0423] For example, at least one frame of the multiple images included in the first media content is deformed along with the first image, and the deformation result is used as the cover frame.
[0424] Optionally, the image / video frames in the media file can be traversed, their sizes can be adjusted, and the adjusted image can be used as the cover frame. For example, size adjustment includes scaling, cropping, etc.
[0425] Optionally, the display position of the cover frame can be set according to the instructions of the display metadata.
[0426] Optionally, fill the blank areas with a uniform color according to the instructions of the displayed metadata.
[0427] Optionally, the presentation format of the cover frame can be set according to the instructions of the display metadata. The presentation format can be any of the following: grid, page turning, or custom.
[0428] Method 3: The target device uses the fusion result of at least one of the multiple images as the cover frame according to the instructions of the display metadata.
[0429] For example, deform the first image, and use the first image and the deformed result as the cover frame.
[0430] Method 4: The target device uses the result of fusing at least one of the multiple images with the media content other than the images in the first media content as the cover frame, according to the instructions of the display metadata.
[0431] For example, at least one includes at least one image obtained according to method one. At least one is merged with media content other than the image in the first media content as a cover frame. The media content other than the image in the first media content includes at least one of audio, text, or subtitles.
[0432] Method 5: The target device merges the deformation result from Method 2 with the media content other than the image in the first media content as a cover frame, according to the instructions of the display metadata.
[0433] Method 6: The target device adjusts the format of the cover frame according to the instructions of the displayed metadata.
[0434] For example, adjust the brightness, contrast, white balance, SDR or HDR of the cover frame, or at least one of these formats.
[0435] Method 7: The target device adjusts the format of the cover frame according to at least one of the display order, position, size, or scaling information in the display style corresponding to various media content indicated by the display metadata.
[0436] For example, such as Figure 5 As shown in (a), the media file contains multiple frames. One frame from these multiple frames is deformed and then used as the cover frame. Figure 5 As shown in (b) above, the multi-frame image with adjusted display position will be used as the cover frame. Figure 5 As shown in (c), the superimposed result of multiple frames of images is used as the cover frame.
[0437] Optionally, the cover frame can also refer to the browsing frame, the default display frame, the thumbnail, etc. The above-mentioned cover frame scheme can also be applied to browsing frames, default display frames, or thumbnails.
[0438] In other embodiments, display metadata is used to indicate how various media content is displayed.
[0439] Optionally, the destination device determines playback information for the first image and the multiple frames contained in the first media content based on display metadata. The playback information includes at least one of the following: playback order, starting frame of playback, or image switching method.
[0440] For example, the playback information includes the playback order of the first image and the second image contained in the first media content.
[0441] For example, playback information includes the playback order of the first image and the first video contained in the first media content.
[0442] For example, playback information includes the first image being the starting frame of the video playback.
[0443] For example, playback information may include direct switching operations between two consecutive frames, which may include scaling or other operations.
[0444] For example, playback information includes introducing fade-in and fade-out operations between two consecutive frames.
[0445] For example, playback information includes introducing deformation switching operations between two consecutive frames.
[0446] For example, playback information might include time data from multiple images, which can be used to initiate operations such as direct switching, fade-in / fade-out, and fly-out from the screen. Alternatively, playback information could include a timeline data set. This timeline data set contains multiple timeline data points, each related to one or more operations, which in turn involve one or more other operations. For instance, operations like direct switching, fade-in / fade-out, and fly-out from the screen could be initiated based on the time data of two consecutive frames.
[0447] For example, playback information includes time data of multiple images, text is added based on the time data of multiple images, and operations such as direct switching, fade-in and fade-out, and fly-out of the screen are introduced.
[0448] The target device displays the first image and the first media content based on the playback information indicated by the display metadata.
[0449] For example, such as Figure 6The image shows the playback order of multiple frames contained in the media file. The multiple frames are displayed according to their playback order. Optionally, the first video contained in the first media content can also be overlaid on the images for playback. Metadata marks the time of each frame or its association with multiple frames.
[0450] Optionally, retrieve the content of display metadata in the media file using one or more of the following methods.
[0451] Information such as file frame count, and / or the number of image and video files contained within, as well as file length and / or offset, and / or playback information. Playback information includes the playlist (which contains the frame number of the image or video file or image frame being played, HEIF item number, JPEG file number (assigned incrementing numbers from front to back to multiple concatenated JPEG files, or to multiple files contained within other payload formats); playback order; playback time; playback processing information, etc.
[0452] This application also provides a method for encapsulating media files. The method includes encapsulating various media contents into a media file, the media file being in a JPEG format. The media file includes an APPn, which contains data structures related to HEIF.
[0453] In some possible implementations, the APPn flag indicates HEIF. Alternatively, the APPn flag may also indicate HEIC or MIAF, etc.
[0454] Optionally, the APPn payload includes a file-type box, a metabox, and a media data box header.
[0455] Optionally, APPn's payload also includes SOI.
[0456] Optionally, the metadata is located in the APPn file of the media file.
[0457] Optionally, the metadata is located in the payload of APPn in the media file.
[0458] For a description of the media files obtained by encapsulating various media content using this encapsulation method, please refer to the above embodiments.
[0459] The media files obtained by this encapsulation method can be applied to the media data processing methods described in the above embodiments.
[0460] For example, an encoder generates a media file from multiple media contents, and a decoder decodes the media file to obtain a first image and / or first media content. The media file is obtained by encapsulating the content according to this encapsulation method.
[0461] For example, the media file obtained by this encapsulation method may include the aforementioned metadata, such as associated metadata, display metadata, etc.
[0462] This application improves the structure of media files and sets metadata to process and display multiple media contents, expanding the forms of media content fusion. The media data processing method provided in this application is a standardized encoding method, which facilitates system implementation, enables unified management of media content, improves the convenience of media content management, reduces application adaptation costs, and enhances the diversity of multi-media content fusion presentation and user cross-platform experience.
[0463] It is understood that, in order to achieve the functions in the above embodiments, the encoder and decoder include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0464] The above text combines Figures 1 to 6 The media data processing method provided according to this embodiment is described in detail below, and will be combined with Figure 7 and Figure 8 This describes the encoding and decoding apparatus provided according to this embodiment.
[0465] Figure 7 This is a schematic diagram of a possible encoding device provided in this embodiment. These encoding devices can be used to implement the encoding function of the media data processing method in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In this embodiment, the encoding device is as follows: Figure 2 The encoder 213 shown can also be a module (such as a chip) applied to a terminal device or server.
[0466] like Figure 7 As shown, the encoding device 700 includes a communication module 710, an encoding module 720, and a storage module 730.
[0467] The encoding device 700 is used to implement the above. Figure 4 The method embodiment shown illustrates the function of the encoder.
[0468] The communication module 710 is used to acquire various media content, including a first image and other media content besides the first image, whereby the first media content includes a first video. For example, the communication module 710 is used to perform... Figure 4 Step 410.
[0469] Encoding module 720 is used to generate a media file based on multiple media contents. The media file includes image data of a first image, data of the first media content, and metadata. The metadata is used to indicate display information for the multiple media contents. For example, encoding module 720 is used to perform... Figure 4 Step 420.
[0470] The communication module 710 is also used to send media files.
[0471] Storage module 730 is used to store media content, metadata, and media files, so that media files can be generated at the encoding end based on various media contents.
[0472] Figure 8 This is a schematic diagram of a possible decoding device provided in this embodiment. These decoding devices can be used to implement the decoding function of the media data processing method in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In this embodiment, the decoding device can be as follows: Figure 2 The decoder 223 shown can also be a module (such as a chip) applied to terminal devices or servers.
[0473] like Figure 8 As shown, the decoding device 800 includes a communication module 810, a decoding module 820, and a storage module 830.
[0474] Decoding device 800 is used to achieve the above. Figure 4 The method embodiment shown illustrates the functionality of the decoder.
[0475] The communication module 810 is also used to acquire a media file, which includes image data of a first image, data of first media content, and metadata. The metadata is used to indicate display information for various media contents, and the first media content includes a first video. For example, the communication module 810 is used to perform... Figure 4 Step 430.
[0476] Decoding module 820 is used to decode the media file to obtain a first image and / or first media content. For example, decoding module 820 is used to perform... Figure 4 Steps 440 and 450.
[0477] Storage module 830 is used to store media content, metadata, and media files, so that the first image and first media content can be obtained from the media files at the decoding end.
[0478] It should be understood that the encoding device 700 and decoding device 800 in the embodiments of this application can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Alternatively, they can be implemented using software. Figure 4 In the method shown, the encoding device 700 and the decoding device 800 and their respective modules can also be software modules.
[0479] For a more detailed description of the communication module, encoding module, decoding module, and storage module mentioned above, please refer to [reference needed]. Figure 4 The relevant descriptions in the method embodiments shown are directly obtained and will not be repeated here.
[0480] Figure 9 This is a schematic diagram of the structure of an encoder 900 provided in this embodiment. Figure 9 As shown, the encoder 900 includes a processor 910, a bus 920, a memory 930, and a communication interface 940.
[0481] It should be understood that in this embodiment, the processor 910 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0482] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0483] The communication interface 940 is used to enable communication between the encoder 900 and external devices or components. In this embodiment, the communication interface 940 is used to acquire various media content.
[0484] Bus 920 may include a pathway for transmitting information between the aforementioned components (such as processor 910 and memory 930). In addition to a data bus, bus 920 may also include a power bus, a control bus, and a status signal bus, etc. However, for clarity, all buses are labeled as bus 920 in the figure.
[0485] As an example, encoder 900 may include multiple processors. A processor may be a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or computing units used to process data (e.g., computer program instructions). Processor 910 generates media files based on various media content.
[0486] It is worth noting that, Figure 9 Taking the encoder 900 as an example, which includes one processor 910 and one memory 930, the processor 910 and the memory 930 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0487] The memory 930 can correspond to the storage medium used to store media content, metadata and media files in the above method embodiments, such as a disk, like a mechanical hard disk or a solid-state hard disk.
[0488] The encoder 900 described above can be a general-purpose device or a special-purpose device. For example, the encoder 900 can be an X108 or ARM-based server, or other special-purpose servers, such as a policy control and charging (PCC) server. This application does not limit the type of encoder 900.
[0489] It should be understood that the encoder 900 in this embodiment can correspond to the encoding device 700 in this embodiment, and can correspond to the execution according to Figure 4 The corresponding subject in any of the methods, and the above and other operations and / or functions of each module in the encoding device 700 are respectively for implementing Figure 4 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.
[0490] Figure 10 This is a schematic diagram of the structure of a decoder 1000 provided in this embodiment. Figure 10 As shown, the decoder 1000 includes a processor 1010, a bus 1020, a memory 1030, and a communication interface 1040.
[0491] It should be understood that in this embodiment, the processor 1010 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0492] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0493] The communication interface 1040 is used to enable communication between the decoder 1000 and external devices or components. In this embodiment, the communication interface 1040 is used to acquire media files.
[0494] Bus 1020 may include a pathway for transmitting information between the aforementioned components (such as processor 1010 and memory 1030). In addition to a data bus, bus 1020 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 1020 in the figure.
[0495] As an example, decoder 1000 may include multiple processors. A processor may be a multi-core (multi-CPU) processor. Here, "processor" can refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions). Processor 1010 is used to decode media files to obtain a first image and / or first media content. Optionally, processor 1010 is used to perform display processing on the first image and / or first media content based on metadata.
[0496] It is worth noting that, Figure 10 Taking the decoder 1000 as an example, which includes one processor 1010 and one memory 1030, the processor 1010 and the memory 1030 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business requirements.
[0497] The memory 1030 can correspond to the storage medium used in the above method embodiments for storing information such as media content, metadata and media files, for example, a disk, such as a mechanical hard disk or a solid-state hard disk.
[0498] The decoder 1000 described above can be a general-purpose device or a dedicated device. For example, the decoder 1000 can be an x86 or ARM-based server, or other dedicated servers, such as a policy control and charging (PCC) server. This application does not limit the type of decoder 1000.
[0499] It should be understood that the decoder 1000 according to this embodiment may correspond to the decoding device 800 in this embodiment, and may correspond to the execution of the decoding device 800 according to this embodiment. Figure 4 The corresponding entities in any of the methods, and the above and other operations and / or functions of each module in the encoding device 700 and the decoding device 800 are respectively implemented for the purpose of... Figure 4 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.
[0500] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a network device or terminal device. Of course, the processor and storage medium can also exist as discrete components in the network device or terminal device.
[0501] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0502] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0503] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
[0504] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. In the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following or similar expressions" refers to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and / or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0505] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0506] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0507] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, throughout the specification, various embodiments do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0508] Some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current underlying solution, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in other scenarios. Correspondingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated upon here.
[0509] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.
Claims
1. A media data processing method, characterized in that, include: Acquire multiple media content, the multiple media content including a first image and other first media content besides the first image, the first media content including at least one of a second image or a first video; A media file is generated based on the various media content. The media file includes image data and metadata of the first image and data of the first media content. The metadata is used to indicate the display information of the various media content.
2. The method according to claim 1, characterized in that, Media files are generated based on the aforementioned multiple media contents, including: The first image is encoded to obtain an image file, the image file including the image data; The first media content is encoded to obtain the data of the first media content; The metadata and the data of the first media content are encoded into the image file to obtain the media file.
3. The method according to claim 1, characterized in that, Media files are generated based on the aforementioned multiple media contents, including: The first image is encoded to obtain the image data; The first media content is encoded to obtain the data of the first media content; The media file is generated based on the image data, the metadata, and the data of the first media content.
4. A media data processing method, characterized in that, include: Obtain a media file, the media file including image data of a first image, metadata and data of a first media content, the metadata being used to indicate display information of multiple media contents, the multiple media contents including the first image and the first media content, the first media content including at least one of a second image or a first video; The media file is decoded to obtain the first image and / or the first media content.
5. The method according to claim 4, characterized in that, Decoding the media file to obtain the first image includes: The image data in the media file is decoded to obtain the first image.
6. The method according to claim 4 or 5, characterized in that, The method further includes: The first image and / or the first media content are displayed based on the metadata.
7. The method according to any one of claims 1-6, characterized in that, The metadata includes display metadata, which indicates at least one of the display methods of the cover frame or the various media content.
8. The method according to claim 7, characterized in that, The display metadata indicates that at least one of the multiple images included in the various media content is used as the cover frame; or, The display metadata indicates that the cover frame is the result of processing at least one of the multiple images included in the multiple media content.
9. The method according to claim 8, characterized in that, The processing result includes any one of the following: a deformation result of at least one of the plurality of images, a fusion result of at least one of the plurality of images, or a fusion result of at least one of the plurality of images with media content other than images in the first media content.
10. The method according to any one of claims 7-9, characterized in that, The display method includes the display style and / or the playback information of multiple frames of images contained in the first image and the first media content.
11. The method according to claim 10, characterized in that, The playback information includes at least one of the following: playback order, starting frame of playback, or image switching method.
12. The method according to any one of claims 7-11, characterized in that, The display metadata also includes at least one of the following: display order, position, size, or scaling information in the display styles corresponding to the various media content.
13. The method according to any one of claims 1-12, characterized in that, The media file is in JPEG format; The metadata is located in APPn of the media file; The APPn contains data structures related to the high-efficiency image file format HEIF.
14. The method according to claim 13, characterized in that, The APPn marker indicates the High Efficiency Image File Format (HEIF).
15. The method according to claim 13 or 14, characterized in that, The payload of the APPn includes the File-type Box, Metabox, and Media Data Box Header.
16. The method according to claim 15, characterized in that, The payload of the APPn also includes the Start of Image (SOI).
17. The method according to any one of claims 13-16, characterized in that, The metadata is located in the APPn of the media file and includes: The metadata is located in the payload of APPn in the media file.
18. An encoding device, characterized in that, The apparatus includes a processing circuit for implementing the method as described in any one of claims 1-3 and 7-17.
19. A decoding device, characterized in that, The apparatus includes a processing circuit for implementing the method as described in any one of claims 4-6 and 7-17.
20. An encoder, characterized in that, The encoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the method as described in any one of claims 1-3 and 7-17.
21. A decoder, characterized in that, The decoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the method as described in any one of claims 4-6, 7-17.
22. A codec system, characterized in that, The encoding / decoding system includes an encoder as described in claim 20 and a decoder as described in claim 21, wherein the encoder is used to perform the operation steps of the method according to any one of claims 1-3 and 7-17, and the decoder is used to perform the method according to any one of claims 4-6 and 7-17.
23. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-17.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions or programs that, when executed on a computer, implement the method as described in any one of claims 1-17.
25. A media file, characterized in that, The media file is obtained by the method according to any one of claims 1-3 and 7-17.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes media files obtained by the method according to any one of claims 1-3 and 7-17.