Method and apparatus for storage of generative artificial intelligence based dynamic content

The method and apparatus enable efficient storage and playback of generative AI entities in multimedia files by organizing them within a media file format structure, addressing the limitations of existing formats and enhancing dynamic content generation.

WO2025219865A1PCT designated stage Publication Date: 2025-10-23NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/053914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing multimedia data storage formats do not effectively accommodate generative artificial intelligence (AI) entities, limiting their integration and utilization in media files.

Method used

A method and apparatus for storing generative AI entities in a media file format structure, linking them within a media file and signaling them in a bitstream, utilizing metaboxes and conventional international standards to organize and manage these entities, enabling their playback and synthesis.

Benefits of technology

Facilitates the efficient storage and playback of generative AI entities, allowing for dynamic content generation and synthesis, enhancing multimedia capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025053914_23102025_PF_FP_ABST
    Figure IB2025053914_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Various embodiments provide methods, apparatuses, and computer program products. An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: storing one or more generative artificial intelligence (AI) entities of a content in a media file format structure, wherein each of the one or more generative AI entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR STORAGE OF GENERATIVE ARTIFICIAL INTELLIGENCE BASED DYNAMIC CONTENTTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to multimedia data storage and, more particularly to, storage of generative artificial intelligence based dynamic content.BACKGROUND

[0002] It is known to provide standardized formats for storage of multimedia data.SUMMARY

[0003] Example 1: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

[0004] Example 2: The apparatus of claim 1, wherein the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images; narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generating prompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for real-time audio and / or video generation; or generated video properties.

[0005] Example 3: The apparatus of claim 2, wherein: the background images or videos, the narrator images, the narrator properties, the narration text, the narration text properties, the neural networks and / or the uniform resource indicators for the neural networks are stored in a metabox.

[0006] Example 4: The apparatus of any of the claims 2 or 3, wherein the generative Al track; generative Al sample entries; generative Al samples are stored in a media track comprising a type which indicates Al usage.

[0007] Example 5: The apparatus of any of the claims 2 to 4, wherein the background video tracks, the background audio tracks, narration timed text tracks, and / or low frame rate video tracks are stored in conventional international standards organization base media file format media tracks.

[0008] Example 6: The apparatus of any of the previous claims, wherein generative Al entities of same type are stored as alternative of each other.

[0009] Example 7: The apparatus of any of the previous claims, wherein tracks that are alternative of each other are stored as alternative entity group.

[0010] Example 8: The apparatus of any of the previous claims, wherein the media file further comprises a handler types or brand to indicate that the file comprises the one or more generative Al entities.

[0011] Example 9: The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: grouping one or more of the following properties and data structures: a narrative text item with associated language properties; a narrator item with associated style properties; an image item as background image which does not have any style related properties; a neural network (NN) item for video generation; an NN item for audio generation; a feature track or a generative Al track; or a baked video track.

[0012] Example 10: The apparatus of any of the claims 2 to 9, wherein the apparatus is further caused to perform: defining a handler type for indicating that the generative Al track is present in the media file.

[0013] Example 11: The apparatus of any of the claims 2 to 10, wherein the generative Al track comprises one or more of the following: one or more supplemental enhancement information (SEI) messages that provide information for Al based video generation; one or more neural-network postfilter characteristics (NNPFC) SEI messages and / or one or more neural-network post-filter activation (NNPFA) SEI messages; generative face video messages; text prompts used as input to the generative Al model; or feature maps used as input tensors to the generative Al model.

[0014] Example 12: The apparatus of any of the previous claims, wherein the media file further comprises a conventional video representing a low frame rate version of the content.

[0015] Example 13: The apparatus of any of the previous claims, wherein the media file comprises the feature track or the generative Al track, and wherein samples of the track are comprised of features.

[0016] Example 14: The apparatus of any of the claims 2 to 13, wherein the generative Al modeluses one or more reconstructed pictures from a video track and one or more samples of the feature track to generate a picture.

[0017] Example 15: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

[0018] Example 16: The apparatus of claim 15, wherein the apparatus is further caused to perform: checking handler type or brands to determine wherein the media file comprises a generative Al related content.

[0019] Example 17: The apparatus of any of claims 15 or 16, wherein the apparatus is further caused to perform: searching for baked video for playing back, when the baked video is present and the apparatus setting or configuration indicates searching for the baked video.

[0020] Example 18: The apparatus of claims 15 or 16, wherein the apparatus is further caused to perform: accessing narrator image, text item, audio track for background audio, NN model, and grouped entities; and initializing Al based audio-visual synthesis process.

[0021] Example 19: The apparatus of claims 15 or 16, wherein the apparatus is further caused to perform: creating a database of alternative narrator images, languages, background images; and prompting the viewer to select one of the languages and narrators.

[0022] Example 20: The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: searching a metabox, comprised in the media file, for an entity group; when the entity group comprises a generative Al track, searching for a track group or the entity group for corresponding timed text track; using the timed text track to define durations of text to be generated; and using NN models comprised in the entity group to generate the audio and / or visual content.

[0023] Example 21: The apparatus of claim 20, wherein the NN models are comprised as uniform resource identifiers (URIs) or uniform resource locators (URLs), and wherein the apparatus is further caused to perform: downloading NN models.

[0024] Example 22: The apparatus of claim 21, wherein the apparatus is further caused to perform: starting audio and / or visual synthesis when a first NN model of the NN models is progressivelydownloaded.

[0025] Example 23: A method comprising: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

[0026] Example 24: The method of claim 23, wherein the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images; narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generating prompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for real-time audio and / or video generation; or generated video properties.

[0027] Example 25: The method of claim 24, wherein: the background images or videos, the narrator images, the narrator properties, the narration text, the narration text properties, the neural networks and / or the uniform resource indicators for the neural networks are stored in a metabox.

[0028] Example 26: The method of any of the claims 24 or 25, wherein the generative Al track; generative Al sample entries; generative Al samples are stored in a media track comprising a type which indicates Al usage.

[0029] Example 27 : The method of any of the claims 24 to 26, wherein the background video tracks, the background audio tracks, narration timed text tracks, and / or low frame rate video tracks are stored in conventional international standards organization base media file format media tracks.

[0030] Example 28: The method of any of the claims 23 to 27, wherein generative Al entities of same type are stored as alternative of each other.

[0031] Example 29: The method of any of the claims 23 to 28, wherein tracks that are alternative of each other are stored as alternative entity group.

[0032] Example 30: The method of any of the claims 23 to 28, wherein the media file further comprises a handler types or brand to indicate that the file comprises the one or more generative Al entities.

[0033] Example 31: The method of any of the claims 23 to 28 further comprising: grouping one or more of the following properties and data structures: a narrative text item with associated language properties; a narrator item with associated style properties; an image item as background image which does not have any style related properties; a neural network (NN) item for video generation; an NN item for audio generation; a feature track or a generative Al track; or a baked video track.

[0034] Example 32: The method of any of the claims 24 to 31 further comprising: defining a handler type for indicating that the generative Al track is present in the media file.

[0035] Example 33: The method of any of the claims 24 to 32, wherein the generative Al track comprises one or more of the following: one or more supplemental enhancement information (SEI) messages that provide information for Al based video generation; one or more neural-network postfilter characteristics (NNPFC) SEI messages and / or one or more neural-network post-filter activation (NNPFA) SEI messages; generative face video messages; text prompts used as input to the generative Al model; or feature maps used as input tensors to the generative Al model.

[0036] Example 34: The method of any of the claims 23 to 33, wherein the media file further comprises a conventional video representing a low frame rate version of the content.

[0037] Example 35: The method of any of the claims 23 to 34, wherein the media file comprises the feature track or the generative Al track, and wherein samples of the track are comprised of features.

[0038] Example 36: The method of any of the claims 24 to 35, wherein the generative Al model uses one or more reconstructed pictures from a video track and one or more samples of the feature track to generate a picture.

[0039] Example 37: A method comprising: receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

[0040] Example 38: The method of claim 37 further comprising: checking handler type or brands to determine wherein the media file comprises a generative Al related content.

[0041] Example 39: The method of any of claims 37 or 38 further comprising: searching for baked video for playing back, when the baked video is present and the method setting or configuration indicates searching for the baked video.

[0042] Example 40: The method of claims 37 or 38 further comprising: accessing narrator image, text item, audio track for background audio, NN model, and grouped entities; and initializing Al based audio- visual synthesis process.

[0043] Example 41: The method of claims 37 or 38 further comprising: creating a database of alternative narrator images, languages, background images; and prompting the viewer to select one of the languages and narrators.

[0044] Example 42: The method of any of the claims 37 to 41 further comprising: searching a metabox, comprised in the media file, for an entity group; when the entity group comprises a generative Al track, searching for a track group or the entity group for corresponding timed text track; using the timed text track to define durations of text to be generated; and using NN models comprised in the entity group to generate the audio and / or visual content.

[0045] Example 43: The method of claim 42, wherein the NN models are comprised as uniform resource identifiers (URIs) or uniform resource locators (URLs), and wherein the method further comprises: downloading NN models.

[0046] Example 44: The method of claim 43 further comprising: starting audio and / or visual synthesis when a first NN model of the NN models is progressively downloaded.

[0047] Example 45: An apparatus comprising: means for storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and means for signaling, in or along bitstream, the media file comprising the media file format.

[0048] Example 46: The apparatus of claim 45, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 24 to 36.

[0049] Example 47: An apparatus comprising: means for receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and means for generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

[0050] Example 48: The apparatus of claim 47, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 38 to 44.

[0051] Example 49: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

[0052] Example 50: The apparatus of claim 49, wherein the apparatus is further caused to perform methods as claimed in any of the claims 24 to 36.

[0053] Example 51: The computer readable medium of any of the claims 49 or 50, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0054] Example 52: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

[0055] Example 53: The apparatus of claim 52, wherein the apparatus is further caused to perform methods as claimed in any of the claims 24 to 36.

[0056] Example 54: The computer readable medium of any of the claims 52 or 53, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0058] FIG. 1 shows schematically an apparatus employing embodiments of the examples described herein.

[0059] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.

[0060] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.

[0061] FIG. 4 is a block diagram illustrating a system in accordance with an example.

[0062] FIG. 5 illustrates an example of how entities are encapsulated into an ISOBMFF-compliant format, in accordance with an embodiment.

[0063] FIG. 6 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.

[0064] FIG. 7 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0065] FIG. 8 is an example method performed with an encoder, based on the examples described herein.

[0066] FIG. 9 is an example method performed with a decoder, based on the examples described herein.DETAIEED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0067] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU coding unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl or Fl-C interface between CU and DU control interface gNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsREC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingSI interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LIE networkXn interface between two NG-RAN nodes

[0068] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments may be shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.

[0069] Described herein is a method and apparatus for storage of generative artificial intelligence based dynamic content.

[0070] The following describes in detail a suitable apparatus and possible method for storage of generative artificial intelligence based dynamic content according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an electronic device or apparatus 100. The apparatus 100 may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.

[0071] The apparatus 100 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.

[0072] The apparatus 100 may comprise a housing 101 for incorporating and protecting the device. The apparatus 100 further may comprise a display 102 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 100 may further comprise a keypad 104. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0073] The apparatus may comprise a microphone 106 or any suitable audio input which may be a digital or analog signal input. The apparatus 100 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 108, speaker, or an analog audio or digital audio output connection. The apparatus 100 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus 100 may further comprise a camera 109 capable of recording or capturing images and / or video. The apparatus 100 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 100 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0074] The apparatus 100 may comprise a controller 110, processor or processor circuitry forcontrolling the apparatus 100. The controller 110 may be connected to memory 112 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 110. The controller 110 may further be connected to codec circuitry 114 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0075] The apparatus 100 may further comprise a card reader 118 and a smart card 116, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.

[0076] The apparatus 100 may comprise radio interface circuitry 120 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 100 may further comprise an antenna 122 connected to the radio interface circuitry 120 for transmitting radio frequency signals generated at the radio interface circuitry 120 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0077] The apparatus 100 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec circuitry 114 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 100 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 100 described above represent examples of means for performing a corresponding function.

[0078] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 300 comprises multiple communication devices which can communicate through one or more networks. The system 300 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.

[0079] The system 300 may include both wired and wireless communication devices and / or apparatus 100 suitable for implementing embodiments of the examples described herein.

[0080] For example, the system shown in FIG. 3 shows a mobile telephone network 301 and a representation of the internet 302. Connectivity to the internet 302 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connectionsincluding, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0081] The example communication devices shown in the system 300 may include, but are not limited to, an electronic device or apparatus 100, a combination of a personal digital assistant (PDA) and a mobile telephone 304, a PDA 306, an integrated messaging device (IMD) 308, a desktop computer 310, a notebook computer 312, or a head-mounted apparatus. The head-mounted apparatus may be a head-mounted display (HMD), or glasses having a device such as a camera configured to encode and / or decode images and / or video. The apparatus 100 may be stationary or mobile when carried by an individual who is moving. The apparatus 100 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.

[0082] The embodiments may also be implemented in a set-top box; e.g., a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0083] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 314 to a base station 316. The base station 316 may be connected to a network server 318 that allows communication between the mobile telephone network 301 and the internet 302. The system may include additional communication devices and communication devices of various types.

[0084] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0085] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, forexample a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0086] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP -based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).

[0087] FIG. 4 is a block diagram illustrating a system or apparatus 400 in accordance with several examples. In an example, the encoder 402 is used to encode an image or video from the scene 404, and the encoder 402 is implemented in a transmitting apparatus 406. The encoder 402 produces a bitstream 408 comprising signaling that is received by the receiving apparatus 410, which implements a decoder 412. The encoder 402 sends the bitstream 408 that comprises the herein described signaling. The decoder 412 forms the image or video for the scene 404-1, and the receiving apparatus 410 would present this to the user, e.g., via a smartphone, television, or projector among many other options.

[0088] In some examples, the transmitting apparatus 406 and the receiving apparatus 410 are at least partially within a common apparatus, and for example, are located within a common housing 414. In other examples the transmitting apparatus 406 and the receiving apparatus 410 are at least partially not within a common apparatus and have at least partially different housings. Therefore in some examples, the encoder 402 and the decoder 412 are at least partially within a common apparatus, and for example are located within a common housing 414. For example, the common apparatus comprising the encoder 402 and decoder 412 implements a codec. In other examples, the encoder 402 and the decoder 412 are at least partially not within a common apparatus and have at least partially different housings, but when together still implement a codec.

[0089] In some examples, 3D media from the capture (e.g., volumetric capture) at a viewpoint 416 of the scene 404, which includes a person 418) is converted via projection to a series of 2D representations with occupancy, geometry, attributes and / or displacements. Additional atlas information is also included in the bitstream to enable inverse reconstruction. For decoding, the received bitstream 408 is separated into its components with atlas information; occupancy, geometry, displacement, and attribute 2D representations. A 3D reconstruction is performed to reconstruct the scene 404-1 created looking at the viewpoint 416-1 with a “reconstructed” person 418-1. The “-1” areused to indicate that these are reconstructions of the original. As indicated at 420, the decoder 412 performs an operation(s) or action(s) based on the received signaling.

[0090] Encoding 422 performs encoding of constituent rectangles based on the examples described herein. Decoding 424 performs decoding of constituent rectangles, based on the examples described herein.

[0091] Having thus introduced a suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described in detail.

[0092] Features as described herein may generally relate, for example, to the ISO base media file format (ISOBMFF).

[0093] Generative artificial intelligence

[0094] The term generative artificial intelligence (Al), or generative modeling, or generative machine learning (and other similar terms), are commonly used to indicate a class of models learned from data that are capable of generating new data. State-of-the-art generative models are based on neural networks. The basic components or layers of a generative neural network (NN) are usually not different from the components or layers of a non-generative NN. Example of such components are convolution layers, non-linear layers, fully-connected layers, normalization layers, attention layers, and the like.

[0095] One typical example of neural network architecture that allows for generating text data is a Transformer-based “decoder”, where “decoder” may not refer to a decoder that is part of a codec performing compression of input data into a small bitstream. Instead, the decoder is a neural network that gets a set of input words or parts of words or tokens extracted from input words, and outputs a set of output words or parts of words or tokens. At inference time, such a NN is run in auto-regressive mode, where the generated word(s) or token(s) is provided as part of the input word(s) or token(s). In order for such a NN to generate data, it is trained to predict the next word(s) (or an estimate of a probability distribution over the next words) given a set of input words. The NN may be based on the Transformer architecture, which comprises the use of the self-attention mechanism, where an attention score is assigned to each input token or word based on all other input tokens or words, including the previously generated words or tokens. During training of a decoder-style Transformer architecture, the future data items (words or tokens) are masked so not to leak information from the future. In some cases, decoder-style Transformer architectures are referred to as “uni-directional”, because they use or process information from left-to-right, as opposed to some encoder-style Transformer architectures that are referred to as “bi-directional” (because they use or process information from left-to-right and from right- to-left).

[0096] Another example of generative modeling is visual temporal extrapolation, where a picture is generated by a NN based on one or more previously decoded or generated pictures and on one or more other data items. The one or more previously decoded or generated pictures may be pictures decoded by a process that does not involve generative modeling, such as a traditional codec, e.g., a versatile video coding (VVC)-compliant codec. The one or more data items may include parameters or features that describe the differences between the one or more previously decoded pictures and the current picture to be temporally extrapolated. Examples of such parameters are facial parameters (such as facial keypoints or facial landmarks and their positions or differential positions with respect to the facial landmarks of a previous picture), or parameters of other objects. The one or more data items may be signaled from encoder to decoder.

[0097] Prompt generative artificial intelligence is capable of generating text, images or other data in response to user prompts. An example of prompt-driven generative Al based audio- visual content are services such as Synthesia, available from [http: / / www.synthesia.io, (last accessed on April 8, 2024)], D-ID, available from Qjt^^wwwxLid om / , (last accessed on April 8, 2024)], or Sora, available from UlSgs^o^iai^om^ora, (last accessed on April 8, 2024)]. Those type of services provide possibility to generate a video content based on given text, narration, narrator image, selected voice, and / or language.

[0098] ISO base media file format

[0099] Available media file format standards include International Standards Organization (ISO) base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF), Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496- 14, also known as the MP4 format), file format for NAL (Network Abstraction Layer) unit structured video (ISO / IEC 14496-15) and High Efficiency Video Coding standard (HEVC or H.265 / HEVC).

[0100] Some concepts, structures, and specifications of ISOBMFF are described below as an example of a container file format, based on which some embodiments may be implemented. The features of the embodiments described herein are not limited to ISOBMFF, but rather the description is given for one possible basis on top of which at least some embodiments may be partly or fully realized.

[0101] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. Box type is typically identified by an unsigned 32-bit integer, interpreted as a four character code (4CC). A box may enclose other boxes, and the ISO file format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may beallowable to have more than one box present in a file. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.

[0102] In files conforming to the ISO base media file format, the media data may be provided in one or more instances of MediaDataBox (‘mdat‘) and the MovieBox (‘moov’) may be used to enclose the metadata for timed media. In some cases, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The ‘moov’ box may include one or more tracks, and each track may reside in one corresponding TrackBox (‘trak’). Each track is associated with a handler, identified by a four-character code, specifying the track type. Video, audio, and image sequence tracks can be collectively called media tracks, and they include an elementary media stream. Other track types comprise hint tracks and timed metadata tracks.

[0103] Tracks comprise samples, such as audio or video frames. For video tracks, a media sample may correspond to a coded picture or an access unit.

[0104] A media track refers to samples (which may also be referred to as media samples) formatted according to a media compression format (and its encapsulation to the ISO base media file format). A hint track refers to hint samples, including cookbook instructions for constructing packets for transmission over an indicated communication protocol. A timed metadata track may refer to samples describing referred media and / or hint samples.

[0105] The 'trak' box includes in its hierarchy of boxes the SampleDescriptionBox, which gives detailed information about the coding type used, and any initialization information needed for that coding. The SampleDescriptionBox includes an entry-count and as many sample entries as the entrycount indicates. The format of sample entries is track-type specific but derived from generic classes (e.g. VisualSampleEntry, AudioSampleEntry). Which type of sample entry form is used for derivation of the track-type specific sample entry format is determined by the media handler of the track.

[0106] The track reference mechanism can be used to associate tracks with each other. The TrackReferenceBox includes box(es), each of which provides a reference from the including track to a set of other tracks. These references are labeled through the box type (e.g., the four-character code of the box) of the contained box(es).

[0107] The ISO Base Media File Format includes three mechanisms for timed metadata that can be associated with particular samples: sample groups, timed metadata tracks, and sample auxiliary information. A derived specification may provide similar functionality with one or more of these three mechanisms.

[0108] A sample grouping in the ISO base media file format and its derivatives, such as the advancedvideo coding (AVC) file format and the scalable video coding (SVC) file format, may be defined as an assignment of each sample in a track to be a member of one sample group, based on a grouping criterion. A sample group in a sample grouping is not limited to being contiguous samples and may include non- adjacent samples. As there may be more than one sample grouping for the samples in a track, each sample grouping may have a type field to indicate the type of grouping. Sample groupings may be represented by two linked data structures: (1) a SampleToGroupBox (sbgp box) represents the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox (sgpd box) includes a sample group entry for each sample group describing the properties of the group. There may be multiple instances of the SampleToGroupBox and SampleGroupDescriptionBox based on different grouping criteria. These may be distinguished by a type field used to indicate the type of grouping. SampleToGroupBox may comprise a grouping_type_parameter field that can be used e.g. to indicate a sub-type of the grouping.

[0109] In ISOMBFF, an edit list provides a mapping between the presentation timeline and the media timeline. Among other things, an edit list provides for the linear offset of the presentation of samples in a track, provides for the indication of empty times and provides for a particular sample to be dwelled on for a certain period of time. The presentation timeline may be accordingly modified to provide for looping, such as for the looping videos of the various regions of the scene. One example of the box that includes the edit list, the EditListBox, is provided below:

[0110] aligned(8) class EditListBox extends FullBox(‘elst’, version, flags) { unsigned int(32) entry _count; for (i= 1; i <= entry _count; i++) { if (version==l) { unsigned int(64) segment_duration; int(64) media lime;} else { / / version==0 unsigned int(32) segment_duration; int(32) media lime;}int(16) media_rate_integer; int( 16) media_rate_fraction = 0;}}

[0111] In ISOBMFF, an EditListBox may be included in EditBox, which is included in TrackBox ('trak').

[0112] In this example of the edit list box, flags specifies the repetition of the edit list. By way of example, setting a specific bit within the box flags (the least significant bit, i.e., flags & 1 in ANSI-C notation, where & indicates a bit-wise AND operation) equal to 0 specifies that the edit list is not repeated, while setting the specific bit (i.e., flags & 1 in ANSI-C notation) equal to 1 specifies that the edit list is repeated. The values of box flags greater than 1 may be defined to be reserved for future extensions. As such, when the edit list box indicates the playback of zero or one samples, (flags & 1) shall be equal to zero. When the edit list is repeated, the media at time 0 resulting from the edit list follows immediately the media having the largest time resulting from the edit list such that the edit list is repeated seamlessly.

[0113] In ISOBMFF, a Track group enables grouping of tracks based on certain characteristics or the tracks within a group have a particular relationship. Track grouping, however, does not allow any image items in the group.

[0114] The syntax of TrackGroupBox in ISOBMFF is as follows:

[0115] aligned(8) class TrackGroupBox extends Box('trgr'){} aligned(8) class TrackGroupTypeBox(unsigned int(32) track_group_type) extends FullBox(track_group_type, version = 0, flags = 0) { unsigned int(32) track_group_id; / / the remaining data may be specified for a particular track_group_type}

[0116] track_group_type indicates the grouping_type and shall be set to one of the following values,or a value registered, or a value from a derived specification or registration:

[0117] 'msrc' indicates that this track belongs to a multi-source presentation. The tracks that have the same value of track_group_id within a TrackGroupTypeBox of track_group_type 'msrc' are mapped as being originated from the same source. For example, a recording of a video telephony call may have both audio and video for both participants, and the value of track_group_id associated with the audio track and the video track of one participant differs from value of track_group_id associated with the tracks of the other participant.

[0118] The pair of track_group_id and track_group_type identifies a track group within the file. The tracks that include a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type belong to the same track group.

[0119] TrackGroupTypeBox with track_group_type equal to 'ster' indicates that this track is either the left or right view of a stereo pair suitable for playback on a stereoscopic display. The tracks that have the same value of track_group_id within StereoVideoGroupBox form a stereo pair, and there shall be no more than two of such tracks for the same value of track_group_id. Usually there are two tracks indicated to be a stereo pair with the StereoVideoGroupBox having the same value of track_group_id. However, only one track can be associated with a stereo pair in specific cases. For example, the file can be edited in a manner that one of the tracks forming a stereo pair gets removed. In another example, only one of the tracks of a stereo pair is selected for transmission, e.g. using DASH.

[0120] Files conforming to the ISOBMFF may include any non-timed objects, referred to as items, meta items, or metadata items, in a meta box (four-character code: ‘meta’ ). While the name of the meta box refers to metadata, items can generally include metadata or media data. The meta box may reside at the top level of the file, within a movie box (four-character code: ‘moov’), and within a track box (four-character code: ‘trak’), but at most one meta box may occur at each of the file level, movie level, or track level. The meta box may be required to include a ‘hdlr’ box indicating the structure or format of the ‘meta’ box contents. The meta box may list and characterize any number of items that can be referred and each one of them can be associated with a file name and are uniquely identified with the file by item identifier (item_id) which is an integer value. The metadata items may be for example stored in the 'idat' box of the meta box or in an 'mdaf box or reside in a separate file. When the metadata is located external to the file then its location may be declared by the DatalnformationBox (four- character code: ‘dinf’). In the specific case that the metadata is formatted using extensible Markup Language (XML) syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (four-character code: ‘xml ‘) or the BinaryXMLBox (four- character code: ‘bxml’). An item may be stored as a contiguous byte range, or it may be stored in severalextents, each being a contiguous byte range. In other words, items may be stored fragmented into extents, e.g. to enable interleaving. An extent is a contiguous subset of the bytes of the resource. The resource can be formed by concatenating the extents.

[0121] A common base structure is used to include general untimed metadata. This structure is called the MetaBox as it was originally designed to carry metadata, i.e. data that is annotating other data. However, it is now used for a variety of purposes including the carriage of data that is not annotating other data, especially when present at ‘file level’ .

[0122] The MetaBox is required to include a HandlerBox indicating the structure or format of the MetaBox contents.

[0123] All other included boxes are specific to the format specified by the HandlerBox.

[0124] The other boxes defined here may be defined as optional or mandatory for a given format. When they are used, then they shall take the form specified here. These optional boxes include a DatalnformationBox, which documents other files in which metadata values (e.g. pictures) are placed, and an ItemLocationBox, which documents where in those files each item is located (e.g. in the common case of multiple pictures stored in the same file).

[0125] At most one MetaBox may occur at each of the file level, segment, movie level, or track level.

[0126] If an ItemProtectionBox occurs, then some or all of the metadata, including possibly the primary resource, may have been protected and be un-readable unless the protection system is taken into account.

[0127] The MetaBox is unusual in that it is a container box yet extends FullBox, but not Box.

[0128] Metadata items are identified by item_ID. Within a given MetaBox, a given item_ID shall uniquely refer to a single item. When an item is updated in movie fragments, the item_ID refers to the latest received version.

[0129] Derived specifications may further restrict the criteria for uniqueness: unique among the item_IDs in both file and movie-level boxes, or unique within that set extended with the track_ID of the tracks in a movie box. The item_ID value of 0 should not be used, and shall not be used when the set is extended to include track_IDs.

[0130] There are three scopes for item_IDs: file and segments; MovieBox and MovieFragmentBox; and TrackBox and TrackFragmentBox. In other words, there shall be only one item with a given item_ID within a given scope (e.g. in the TrackBox and all TrackFragmentBox with the same track_ID).

[0131] aligned(8) class MetaBox (handler_type) extends FullBoxfmeta', version = 0, 0) {HandlerBox(handler_type) theHandler;Primary ItemB ox primary _resource; / / optionalDatalnformationBox file_locations; / / optionalItemLocationB ox item_locations; / / optionalItemProtectionBox protections; / / optionalItemlnfoBox item_infos; / / optionalIPMPControlB ox IPMP_control; / / optionalItemReferenceB ox item_refs; / / optionalItemDataBox item_data; / / optionalBox other_boxes [] ; / / optional}

[0132] The structure or format of the metadata is declared by the handler. In the case that the primary data is identified by a primary item, and that primary item has an item information entry with an item_type, the handler type may be the same as the item_type.

[0133] The ItemPropertiesBox enables the association of any item with an ordered set of item properties. Item properties may be regarded as small data records. The ItemPropertiesBox includes two parts: ItemPropertyContainerBox that includes an implicitly indexed list of item properties, and one or more ItemProperty AssociationBox(es) that associate items with item properties.

[0134] High Efficiency Image File Format (HEIF)

[0135] High Efficiency Image File Format (HEIF) is a standard developed by the Moving Picture Experts Group (MPEG) for storage of images and image sequences. Among other things, the standard facilitates file encapsulation of data coded according to the High Efficiency Video Coding (HEVC) standard. HEIF includes features building on top of the used ISO Base Media File Format (ISOBMFF).

[0136] The ISOBMFF structures and features are used to a large extent in the design of HEIF. The basic design for HEIF comprises still images that are stored as items and image sequences that are stored as tracks. An item in HEIF is defined as the data that does not require timed processing, as opposed to sample data, and is described by the boxes included in a MetaBox

[0137] In the context of HEIF, the following boxes may be included within the root-level 'meta' box and may be used as described in the following. In HEIF, the handler value of the Handler box of the 'meta' box is 'pict'. The resource (whether within the same file, or in an external file identified by a uniform resource identifier) including the coded media data is resolved through the Data Information ('dinf ) box, whereas the Item Location filoc') box stores the position and sizes of every item within the referenced file. The Item Reference firef) box documents relationships between items using typed referencing. When there is an item among a collection of items that is in some way to be considered the most important compared to others then this item is signaled by the Primary Item ('pitm') box. Apart from the boxes mentioned here, the 'meta' box is also flexible to include other boxes that may be necessary to describe items.

[0138] Any number of image items can be included in the same file. Given a collection of images stored by using the 'meta' box approach, it sometimes is essential to qualify certain relationships between images. Examples of such relationships include indicating a cover image for a collection, providing thumbnail images for some or all of the images in the collection, and associating some or all of the images in a collection with an auxiliary image such as an alpha plane. A cover image among the collection of images is indicated using the 'pitm' box. A thumbnail image or an auxiliary image is linked to the primary image item using an item reference of type 'thmb' or 'auxl', respectively.

[0139] The services such as Synthesia available from [http: / / www.synthesia.io, (last accessed on April 8, 2024)], D-ID, available from [htqjs: / ^(last accessed on April 8, 2024)], orSora, available from [hU^^opgt iLcom / s^i, (last accessed on April 8, 2024)] allow to produce an AI- generated video with the narrator reading the narration text in the given language with natural humanlike gestures and facial expressions. However, the final video is baked, and it is not possible to modify it further after distribution. Baked content does not allow interactivity, hence real-time synthesis of the video and voice is currently not possible. Moreover, there is no way to stream such content unless it is baked into an MP4 file.

[0140] In the future, Al generated content could be generated on the fly by the target device in real-time, without the need for baking the video beforehand. There is no well-defined and interoperable format to share such dynamic Al generated content.

[0141] Various embodiments provide a method to store such generative Al content for real-timerendering based on the prompts and narration as well as configuration inputs.

[0142] Following are some example objectives of various embodiments:

[0143] storage in an ISOBMFF-compliant format structure where each of the entities of the, to- be, Al-generated content are stored and linked to each other.ISOBMFF compliancy brings default compatibility with video and audio formats and enable segmentation as well as non-timed asset storage (e.g. NN model / filter or narrator input image) for such content.

[0144] generation of the final audio-visual content on-the-fly from the entities that are stored in the ISOBMFF-compliant file and / or playback of the generated content to the viewer. generation on-the-fly enables modification capability, as well as interactive capabilities where the media content is generated based on, for example, inputs or preferences of the viewer.

[0145] desired language, narrator image, voice type or image background could be modified and changed by the viewer preferences during playback, on-the-fly.

[0146] the same format could be used for re-editing the baked visual content, when desired. Baked audio and video could also be stored in the same file and then be modified or re-edited by, for example, modifying the text or narrator person image.

[0147] such a format enables interoperability between different products in the Al-generated content workflow.

[0148] Following are one or more examples of entities that may be stored in the file format:Background images or videosOverlays (e.g. logos or banners in different languages)Narrator portrait images (for selecting the narrator)Narrator related properties such as emotion and style or sentimentA generated video (baked video)A base video which could be used as input for Al-based video generationA timed text track or narration texts, that may be utilized a prompt information for the Al generation model.Audio track (when voice to be the same as the audio track voice or background audio, such as a low volume music or alike)A feature track, which may comprise, e.g., 3D keypoints or 3D landmarks of the narrator face, and may be used to represent, e.g., head movements and / or facial gesturesA generative Al track that provide information on how to combines these entities.NNs or NN uniform resource identifiers (URIs) that perform the inference for real-time audio / video generationGenerated video properties such as language

[0149] FIG. 5 illustrates an example of how entities are encapsulated into an ISOBMFF- compliant format, in accordance with an embodiment. In this example, background images, NNs for generative Al, NN properties or URIs, narrator images, narrator properties, narration text, and narration text properties are stored in a Metabox 502; generative Al track, generative Al sample entry, and generative Al sample are stored as a new track type 504; and background video tracks, background audio tracks, narration timed-text-track, and low-frame-rate video are stored legacy tracks 506, for example, conventional international standards organization base media file format media tracks.

[0150] Entities of the same type may be stored as alternatives of each other by utilizing entity grouping, such as ‘altr’ entity grouping. For example, same type means having the same 4CC code or identifier that map to the same entity type as defined in ISOBMFF file format. Tracks that are alternatives of each other may be stored as ‘altr’ entity groups and / or indicated to be alternatives using the alternate_group flag in the TrackHeaderBox. Alternative tracks may also be grouped as alternate track groups with type ‘altr’ .

[0151] Narrator items may be stored as image items, with properties including name. Narrator emotions and style could be stored as an item property and then associated with the narrator item. In an embodiment, this may be stored as an XML or HTML or JSON data structure.

[0152] Neural Network (NN) models may be stored as NN items with type ‘ainn’. NN model data may be stored in or along / outside the file.

[0153] In the file: the NN model data may be stored in Box of type ‘mdat’ or type ‘idat’.

[0154] In another embodiment, NN models may be stored and compressed by utilizing the mechanisms defined in ISO / IEC 15938-17 “Compression of neural networks for multimedia content description and analysis” standard. For example, an NN model may be compressed into a bitstream conforming to ISO / IEC 15938-17 and stored as an 'nnrl' item as specified in ISO / IEC 14496-15 Edition 6, Draft Amendment 3.

[0155] Along / outside the file the: the information how to fetch the NN model data could be stored in the file, for example by providing URI / URLs that point to the NN model data location.

[0156] NN items may include associated properties that indicate that they are for generating video or audio.

[0157] When narration text is not timed, then there is not a duration constraint on the video synthesis, and it may take as long as the NN model generated video duration. In this case, narration text may be text items of type ‘nrtx’ .

[0158] The file may include specific handler types or brands that indicate that the file includes generative Al entities. Generative Al representation may be signaled using an Entity Grouping of type ‘aigv’. It may group one or more of the following properties and data structures:A narrative text item with associated language properties;A narrator item with associated style properties;A image item as background image which will not have any style related properties;A NN item for video generation;A NN item for audio generation. Alternatively, an audio track may be part of the grouping. In such a case, the duration of the audio track may define the duration of the video generation as well;A feature track or a generative Al track where the samples are the features; and / orOptionally, a baked video track.

[0159] In an embodiment, the video track is intended to be used for extracting features, such as facial landmarks, for generative Al or directly used as one of the inputs to generative Al.

[0160] In an embodiment, a derived track of a new handler type that indicates generative Al track is present in a file.

[0161] The generative Al track may be grouped with a time text track which includes samples that have text of a particular language and duration.

[0162] The generative Al track may be grouped with a narrator item, background image, and / or NN items for audio and video generation. The grouping may be done via an entity grouping of, for example, type ‘aigv’

[0163] Each alternative language, track, item, NN model may use the ‘altr’ track grouping.

[0164] NN models may have language specific properties associated with them, in order to indicate the language that they synthesize.

[0165] Video generation NN models may also have language assigned to them, which may indicate “nation-specific gestures” (e.g. Italian style vs. Finnish style narration)

[0166] In another embodiment, the timed text track and the generative Al track may be listed in the ‘aigv’ entity group and not use track grouping.

[0167] In another embodiment, certain prompts that may be input to generative Al model may be stored as timed metadata or text tracks.

[0168] Additional embodiments

[0169] In an embodiment, a generative Al track may comprise one or more SEI messages that provide information for video generation, for example, Al based video generation. It is to be understood that SEI messages may interchangeably refer to any syntax structure(s), which may carry information for decoder-side post-processing.

[0170] In an embodiment, a generative Al track may comprise one or more neural-network postfilter characteristics (NNPFC) SEI messages and / or one or more neural-network post-filter activation (NNPFA) SEI messages as specified in the Versatile Supplemental Enhancement Information (VSEI) standard (ISO / IEC 23007-2 I ITU-T H.274). The NNPFC SEI message(s) may be included in samples of the generative Al track or may be represented by an NNPFC sample group as specified in ISO / IEC 14496-15 Edition 6, Draft Amendment 3.

[0171] In an embodiment, a generative Al track may comprise generative face video SEI messages, which may be similar to what has been described in JVET-AG2032. A generative face video (GFV) SEI message provides information for visual temporal extrapolation of faces in videos. One of the functions of the GFV SEI message is to signal facial parameters, which can be used to derive input data to the generative NN that performs the temporal extrapolation.

[0172] In an embodiment, a generative Al track may comprise text prompts used as input to the generative Al model.

[0173] In an embodiment, a generative Al track may comprise feature maps used as input tensors to the generative Al model. A feature map may represent a tensor having particular values.

[0174] In an embodiment, a file comprises a “conventional” video track, which may represent a low frame rate version of the content. Additionally, the file comprises a feature track or a generative Al track where the samples are the features. The generative Al model uses one or more reconstructed pictures from the video track and one or more samples of the feature track to generate a picture. Thus, by processing the video track and the feature track, a high frame rate version of the content can be reconstructed.

[0175] The samples of the feature track may be associated to the samples of the video track through decoding time or composition time. For example, the generative Al model may use as input a feature track sample and the video track sample that coincides with the feature track sample in composition time or has the last composition time preceding the composition time of the feature track sample. The picture generated by the generative Al model may be inferred to have the composition time of the feature track sample that was used as input to the generative Al model.

[0176] In an embodiment, a generative Al sample auxiliary information is defined. The generative Al sample auxiliary information carries the generative face video SEI messages, which may be similar to what has been described in JVET-AG2032.

[0177] In an embodiment, a generative Al sample auxiliary information may comprise text prompts used as input to the generative Al model.

[0178] In an embodiment, a generative Al sample auxiliary information may comprise feature maps used as input tensors to the generative Al model.

[0179] In an embodiment, a file comprises a “conventional” video track, which may represent a low frame rate version of the content. Additionally, the file comprises a feature sample auxiliary information or a generative Al sample auxiliary information where the sample auxiliary information are the features. The generative Al model uses one or more reconstructed pictures from the video track and the corresponding sample auxiliary information of the feature to generate a picture. Thus, by processing the video track and the feature sample auxiliary information, a high frame rate version of the content can be reconstructed.

[0180] For example, the generative Al model may use as input a feature sample auxiliary information and the corresponding video track sample that coincides with the feature sample auxiliary information.

[0181] CLIENT

[0182] When a client receives a file, the client may process the file as follows:The client checks handler type or brands to understand that this is a generative Al related content;The client may search for the metabox and then find the ‘aigv’ entity group;The client may search for the baked video to play it back, when present and when indicated by its settings / configuration to do so;Via the entity group, the client may access the narrator image, text item, audio track for background audio, NN model and other grouped entities and initialize the Al based audio-visual synthesis process;The client may also build a database of alternative narrator images, languages and background images and prompt the viewer to select one of the languages and narrators;Alternatively, when the ’aigv’ entity group includes a generative Al track (with the proper handler type), it may search for a track group to find the corresponding timed text track or search the entity group to see when a timed text track is present. When present, client uses the timed text track to define the durations of the text to be generated and use the NN models in the entity group to generate the audio-visual content;In an embodiment, the NN models may be only present as URIs or URLs and the client would first download them before the audio-visual synthesis; and / orIn another embodiment, generative Al may start audio visual synthesis when the first NN model is progressively downloaded instead of waiting for all NN models to be downloaded, (e.g. start audio synthesis while downloading the NN model for video).

[0183] FIG. 6 is an example apparatus 600, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 600 comprises at least one processor 602 (e.g., an FPGA and / or CPU), at least one memory 604 including computer program code 605, the computer program code 605 having instructions to carry out the methods described herein, wherein the at leastone memory 604 and the computer program code 605 are configured to, with the at least one processor 602, cause the apparatus 600 to implement circuitry, a process, component, module, or function (implemented with control module 606) to implement the examples described herein, including storage of generative artificial intelligence based dynamic content. Optionally included encoder 608 of the control module 606 implements encoding based on the examples described herein, and optionally included decoder 610 implements decoding based on the examples described herein. The at least one memory 604 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).

[0184] The apparatus 600 includes a display and / or I / O interface 612, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 600 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 614. The communication I / F(s) 614 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 616. The communication I / F(s) 614 may comprise one or more transmitters or one or more receivers.

[0185] The transceiver 618 comprises one or more transmitters 620 and one or more receivers 622. The transceiver 618 and / or communication I / F(s) 614 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 624 used for communication over wireless link 626.

[0186] The control module 606 of the apparatus 600 comprises one of or both parts 606-1 and / or 606-2, which may be implemented in a number of ways. The control module 606 may be implemented in hardware as control module 606-1, such as being implemented as part of the at least one processor 602. The control module 606-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 606 may be implemented as control module 606-2, which is implemented as computer program code (having corresponding instructions) 605 and is executed by the at least one processor 602. For instance, the at least one memory 604 store instructions that, when executed by the at least one processor 602, cause the apparatus 600 to perform one or more of the operations as described herein. Furthermore, the at least one processor 602, the at least one memory 604, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.

[0187] The apparatus 600 to implement the functionality of control module 606 may correspond toany of the apparatuses depicted herein. Alternatively, apparatus 600 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 600 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.

[0188] The apparatus 600 may also be distributed throughout the network including within and between apparatus 600 and any network element (such as a base station and / or terminal device and / or user equipment).

[0189] Interface 628 enables data communication and signaling between the various items of apparatus 600, as shown in FIG. 6. For example, the interface 628 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 605, including control module 606 may comprise object-oriented software configured to pass data or messages between objects within computer program code 605. The apparatus 600 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 600 may at least partially reside in a housing 630, or a subset of the various components of apparatus 600 may at least partially be located in different housings, which different housings may include housing 630.

[0190] FIG. 7 shows a schematic representation of non-volatile memory media 700a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 700b (e.g. universal serial bus (USB) memory stick) and 700c (e.g. cloud storage for downloading instructions and / or parameters 702 or receiving emailed instructions and / or parameters 702) storing instructions and / or parameters 702 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein. Instructions and / or parameters 702 may represent or correspond to a non- transitory computer readable medium.

[0191] FIG. 8 is an example method 800 performed with an encoder, based on the example embodiments described herein. At 802, the method 800 includes storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file. At 804, the method 800 signaling, in or along bitstream, the media file comprising the media file format.

[0192] In an example embodiment, the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images; narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generatingprompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for realtime audio and / or video generation; or generated video properties.

[0193] The method 800 may be performed with an encoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, for example, the transmitting apparatus 406 with the encoder 402, or the apparatus 400 with the encoder 402.

[0194] FIG. 9 is an example method 900 performed with a decoder, based on the example embodiments described herein. At 902, the method 900 includes receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other. At 904, the method 900 includes generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

[0195] In an example embodiment, the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images; narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generating prompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for realtime audio and / or video generation; or generated video properties.

[0196] The method 900 may be performed with an decoding apparatus, such as the apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, the receiving apparatus 410 with a decoder 412, or the apparatus 400 with the decoder 412.

[0197] As described above, FIGs. 8 and 9 include flowcharts of an apparatus (e.g. 100, 600, or any other apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody theprocedures described above may be stored by a memory (e.g. 58, 125, or 604) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g., 56, 120, or 602) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.

[0198] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer -readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 8 and 9. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.

[0199] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.

[0200] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.

[0201] Some embodiments have been described in relation to one or more neural networks performing visual temporal extrapolation. It is to be understood that embodiments can be realized with any generative modelling neural networks.

[0202] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0203] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0204] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0205] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.

[0206] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not onlycomputers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.

[0207] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

[0208] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0209] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

2. The apparatus of claim 1, wherein the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images; narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generating prompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for real-time audio and / or video generation; or generated video properties.

3. The apparatus of claim 2, wherein: the background images or videos, the narrator images, the narrator properties, the narration text, the narration text properties, the neural networksand / or the uniform resource indicators for the neural networks are stored in a metabox.

4. The apparatus of any of the claims 2 or 3, wherein the generative Al track; generative Al sample entries; generative Al samples are stored in a media track comprising a type which indicates Al usage.

5. The apparatus of any of the claims 2 to 4, wherein the background video tracks, the background audio tracks, narration timed text tracks, and / or low frame rate video tracks are stored in conventional international standards organization base media file format media tracks.

6. The apparatus of any of the previous claims, wherein generative Al entities of same type are stored as alternative of each other.

7. The apparatus of any of the previous claims, wherein tracks that are alternative of each other are stored as alternative entity group.

8. The apparatus of any of the previous claims, wherein the media file further comprises a handler types or brand to indicate that the file comprises the one or more generative Al entities.

9. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: grouping one or more of the following properties and data structures: a narrative text item with associated language properties; a narrator item with associated style properties; an image item as background image which does not have any style related properties; a neural network (NN) item for video generation; an NN item for audio generation; a feature track or a generative Al track; or a baked video track.

10. The apparatus of any of the claims 2 to 9, wherein the apparatus is further caused to perform: defining a handler type for indicating that the generative Al track is present in the media file.

11. The apparatus of any of the claims 2 to 10, wherein the generative Al track comprises one or more of the following: one or more supplemental enhancement information (SEI) messages that provideinformation for Al based video generation; one or more neural-network post-filter characteristics (NNPFC) SEI messages and / or one or more neural-network post-filter activation (NNPFA) SEI messages; generative face video messages; text prompts used as input to the generative Al model; or feature maps used as input tensors to the generative Al model.

12. The apparatus of any of the previous claims, wherein the media file further comprises a conventional video representing a low frame rate version of the content.

13. The apparatus of any of the previous claims, wherein the media file comprises the feature track or the generative Al track, and wherein samples of the track are comprised of features.

14. The apparatus of any of the claims 2 to 13, wherein the generative Al model uses one or more reconstructed pictures from a video track and one or more samples of the feature track to generate a picture.

15. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

16. The apparatus of claim 15, wherein the apparatus is further caused to perform: checking handler type or brands to determine wherein the media file comprises a generative Al related content.

17. The apparatus of any of claims 15 or 16, wherein the apparatus is further caused to perform: searching for baked video for playing back, when the baked video is present and the apparatus setting or configuration indicates searching for the baked video.

18. The apparatus of claims 15 or 16, wherein the apparatus is further caused to perform:accessing narrator image, text item, audio track for background audio, NN model, and grouped entities; and initializing Al based audio-visual synthesis process.

19. The apparatus of claims 15 or 16, wherein the apparatus is further caused to perform: creating a database of alternative narrator images, languages, background images; and prompting the viewer to select one of the languages and narrators.

20. The apparatus of any of the previous claims, wherein the apparatus is further caused to perform: searching a metabox, comprised in the media file, for an entity group; when the entity group comprises a generative Al track, searching for a track group or the entity group for corresponding timed text track; using the timed text track to define durations of text to be generated; and using NN models comprised in the entity group to generate the audio and / or visual content.

21. The apparatus of claim 20, wherein the NN models are comprised as uniform resource identifiers (URIs) or uniform resource locators (URLs), and wherein the apparatus is further caused to perform: downloading NN models.

22. The apparatus of claim 21, wherein the apparatus is further caused to perform: starting audio and / or visual synthesis when a first NN model of the NN models is progressively downloaded.

23. A method comprising: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

24. The method of claim 23, wherein the one or more generative Al entities comprises one or more of the following: background images or videos; overlays; narrator images;narrator properties; narration text; narration text properties; a video generated based on Al (Al generated video); a base video used as an input for generating the Al generated video; a timed text track or narration text used for generating prompt information for a generative Al model; an audio track; a feature track comprising three dimensional key-points or landmark information; a generative Al track comprising information on how to combine the one or more generative Al entities; generative Al sample entries; generative Al samples; neural networks or uniform resource indicators for the neural networks that perform inference for real-time audio and / or video generation; or generated video properties.

25. The method of claim 24, wherein: the background images or videos, the narrator images, the narrator properties, the narration text, the narration text properties, the neural networks and / or the uniform resource indicators for the neural networks are stored in a metabox.

26. The method of any of the claims 24 or 25, wherein the generative Al track; generative Al sample entries; generative Al samples are stored in a media track comprising a type which indicates Al usage.

27. The method of any of the claims 24 to 26, wherein the background video tracks, the background audio tracks, narration timed text tracks, and / or low frame rate video tracks are stored in conventional international standards organization base media file format media tracks.

28. The method of any of the claims 23 to 27, wherein generative Al entities of same type are stored as alternative of each other.

29. The method of any of the claims 23 to 28, wherein tracks that are alternative of each other are stored as alternative entity group.

30. The method of any of the claims 23 to 28, wherein the media file further comprises ahandler types or brand to indicate that the file comprises the one or more generative Al entities.

31. The method of any of the claims 23 to 28 further comprising: grouping one or more of the following properties and data structures: a narrative text item with associated language properties; a narrator item with associated style properties; an image item as background image which does not have any style related properties; a neural network (NN) item for video generation; an NN item for audio generation; a feature track or a generative Al track; or a baked video track.

32. The method of any of the claims 24 to 31 further comprising: defining a handler type for indicating that the generative Al track is present in the media file.

33. The method of any of the claims 24 to 32, wherein the generative Al track comprises one or more of the following: one or more supplemental enhancement information (SEI) messages that provide information for Al based video generation; one or more neural-network post-filter characteristics (NNPFC) SEI messages and / or one or more neural-network post-filter activation (NNPFA) SEI messages; generative face video messages; text prompts used as input to the generative Al model; or feature maps used as input tensors to the generative Al model.

34. The method of any of the claims 23 to 33, wherein the media file further comprises a conventional video representing a low frame rate version of the content.

35. The method of any of the claims 23 to 34, wherein the media file comprises the feature track or the generative Al track, and wherein samples of the track are comprised of features.

36. The method of any of the claims 24 to 35, wherein the generative Al model uses one or more reconstructed pictures from a video track and one or more samples of the feature track to generate a picture.

37. A method comprising:receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

38. The method of claim 37 further comprising: checking handler type or brands to determine wherein the media file comprises a generative Al related content.

39. The method of any of claims 37 or 38 further comprising: searching for baked video for playing back, when the baked video is present and the method setting or configuration indicates searching for the baked video.

40. The method of claims 37 or 38 further comprising: accessing narrator image, text item, audio track for background audio, NN model, and grouped entities; and initializing Al based audio-visual synthesis process.

41. The method of claims 37 or 38 further comprising: creating a database of alternative narrator images, languages, background images; and prompting the viewer to select one of the languages and narrators.

42. The method of any of the claims 37 to 41 further comprising: searching a metabox, comprised in the media file, for an entity group; when the entity group comprises a generative Al track, searching for a track group or the entity group for corresponding timed text track; using the timed text track to define durations of text to be generated; and using NN models comprised in the entity group to generate the audio and / or visual content.

43. The method of claim 42, wherein the NN models are comprised as uniform resource identifiers (URIs) or uniform resource locators (URLs), and wherein the method further comprises: downloading NN models.

44. The method of claim 43 further comprising: starting audio and / or visual synthesis when a first NN model of the NN models is progressively downloaded.

45. An apparatus comprising: means for storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and means for signaling, in or along bitstream, the media file comprising the media file format.

46. The apparatus of claim 45, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 24 to 36.

47. An apparatus comprising: means for receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each of the one or more generative Al entities are linked to each other; and means for generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

48. The apparatus of claim 47, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 38 to 44.

49. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: storing one or more generative artificial intelligence (Al) entities of a content in a media file format structure, wherein each of the one or more generative Al entities are linked to each other, and wherein the media file format structure is comprised in a media file; and signaling, in or along bitstream, the media file comprising the media file format.

50. The apparatus of claim 49, wherein the apparatus is further caused to perform methods as claimed in any of the claims 24 to 36.

51. The computer readable medium of any of the claims 49 or 50, wherein the computer readable medium comprises a non-transitory computer readable medium.

52. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: receiving, in or from a bitstream, a media file comprising a media file format structure storing one or more generative artificial intelligence (Al) entities of a content, wherein each ofthe one or more generative Al entities are linked to each other; and generating an audio and / or visual content from the one or more entities that are stored in the media file format and / or playing back the generated content to a viewer.

53. The apparatus of claim 52, wherein the apparatus is further caused to perform methods as claimed in any of the claims 24 to 36.

54. The computer readable medium of any of the claims 52 or 53, wherein the computer readable medium comprises a non-transitory computer readable medium.

Citation Information

Patent Citations

  • Method and apparatus for encapsulating images or sequences of images with proprietary information in a file

    GB2575288A

  • An apparatus and a method for artificial intelligence

    US20210349943A1