Systems for decoding multimedia files, multimedia archives and systems for coding demultimedia files
Patent Information
- Application Number
- BR122018005939
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Publication Date
- 2026-09-15
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure 00000118_0000 
Figure 00000119_0000 
Figure 00000120_0000
Description
Descriptive Report of the Invention Patent for SYSTEMS FOR DECODING MULTIMEDIA FILES, MULTIMEDIA FILES AND SYSTEMS FOR ENCODING MULTIMEDIA FILES.
[0001] Divided from PI0416738-4, filed on December 8, 2004. Background of the Invention
[0002] The present invention generally relates to the encoding, transmission and decoding of multimedia files. More specifically, the invention relates to the encoding, transmission and decoding of multimedia files that include tracks in addition to a single audio track and a single video track.
[0003] The development of the Internet has led to the development of file formats for multimedia information to allow for standardized generation, distribution, and display of files. Typically, a single multimedia file includes a single video track and a single audio track. When multimedia is written to a high-volume, physically transportable medium, such as a Digital Video Disc (DVD), multiple files can be used to provide a number of video tracks, audio tracks, and subtitle tracks. Additional files may be provided containing information that can be used to generate an interactive menu. Summary of the Invention
[0004] The embodiments of the present invention include multimedia files and systems for generating, distributing, and decoding multimedia files. In one aspect of the invention, the multimedia files include a plurality of encoded video tracks. In another aspect of the invention, the multimedia files include a plurality of encoded audio tracks. In another aspect of the invention, the multimedia files include at least one subtitle track. Petition 870200018202, dated 07 / 02 / 2020, page 13 / 20 2 / 111 In another aspect of the invention, the multimedia files include encoded 'metadata'. In another aspect of the invention, the multimedia files include encoded menu information.
[0005] A multimedia file according to one embodiment of the invention includes a plurality of encoded video tracks. In further embodiments of the invention, the multimedia file comprises a plurality of concatenated 'RIFF' blocks and each encoded video track is contained in a separate 'RIFF' block. In addition, the video is encoded using psychovisual enhancements and each video track has at least one audio track associated with it.
[0006] In another embodiment, each video track is encoded as a series of 'video' blocks within a 'RIFF' block, and the audio track accompanying each video track is encoded as a series of 'audio' blocks interleaved within the 'RIFF' block that contains the 'video' blocks of the associated video track. Furthermore, each 'video' block may contain the information that can be used to generate a single video frame from a video track, and each 'audio' block contains the audio information of the portion of the audio track that accompanies the frame generated using a 'video' block. Additionally, the 'audio' block may be interleaved before the corresponding 'video' block within the 'RIFF' block.
[0007] A system for encoding multimedia files according to one embodiment of the present invention includes a processor configured to encode a plurality of video tracks, concatenate the encoded video tracks, and write the concatenated encoded video tracks into a single file. In another embodiment, the processor is configured to encode the video tracks so that each video track is contained within a separate 'RIFF' block, and the processor is configured to encode the video. Petition 870180148460, dated 06 / 11 / 2018, page 8 / 125 3 / 111 using psychovisual enhancements. In addition, each video track can have at least one associated audio track.
[0008] In another embodiment, the processor is configured to encode each video track as a series of 'video' blocks within a 'RIFF' block and to encode at least one audio track that accompanies each video track as a series of interleaved 'audio' blocks within the 'RIFF' block that contains the 'video' blocks of the associated video track.
[0009] In another additional embodiment, the processor is configured to encode the video tracks so that each 'video' block contains the information that can be used to generate a single video frame from a video track and to encode the audio tracks associated with a video track so that each 'audio' block contains the audio information from the portion of the audio track that accompanies the frame generated using a 'video' block generated from the video track. Furthermore, the processor may be configured to interleave each 'audio' block before the corresponding 'video' block within the 'RIFF' block.
[00010] A system for decoding a multimedia file containing a plurality of encoded video tracks according to an embodiment of the present invention includes a processor configured to extract information from the multimedia file. The processor is configured to extract information regarding the number of encoded video tracks contained in the multimedia file.
[00011] In a further embodiment, the processor is configured to locate an encoded video track within a 'RIFF' block. Furthermore, a first encoded video track may be contained in a first 'RIFF' block that has a standard 4cc code, a second video track may be contained in a second Petition 870180148460, dated 06 / 11 / 2018, page 9 / 125 4 / 111 'RIFF' block that has a specialized 4cc code, and the specialized 4cc code may have as its last two characters the first two characters of a standard 4cc code.
[00012] In an additional embodiment, each encoded video track is contained in a separate 'RIFF' block.
[00013] In another additional embodiment, the decoded video track is similar to the original video track that was encoded in the creation of the multimedia file, and at least some of the differences between the decoded video track and the original video track are located in dark portions of frames in the video track. Furthermore, some of the differences between the decoded video track and the original video track may be located in high-motion scenes in the video track.
[00014] In an additional mode again, each video track has at least one audio track associated with it.
[00015] In an additional embodiment, the processor is configured to display video from a video track by decoding a series of 'video' blocks within a 'RIFF' block and to generate audio from an audio track that accompanies the video track by decoding a series of interleaved 'audio' blocks within the 'RIFF' block that contains the 'video' blocks of the associated video track.
[00016] In yet another additional embodiment, the processor is configured to use the information extracted from each 'video' block to generate a single frame of the video track and to use the information extracted from each 'audio' block to generate the portion of the audio track that accompanies the frame generated using a 'video' block. Furthermore, the processor may be configured to locate the 'audio' block before the 'video' block with which it is associated in the 'RIFF' block. Petition 870180148460, dated 06 / 11 / 2018, page 10 / 125 5 / 111
[00017] A multimedia file according to an embodiment of the present invention includes a series of encoded video frames and encoded audio interspersed between the encoded video frames. The encoded audio includes two or more tracks of audio information.
[00018] In an additional embodiment, at least one of the audio information tracks includes a plurality of audio channels.
[00019] Another modality includes header information that identifies the number of audio tracks contained in the multimedia file and description information about at least one of the audio information tracks.
[00020] In an additional embodiment again, each encoded video frame is preceded by encoded audio information, and the encoded audio information preceding the video frame includes the audio information for the portion of each audio track that accompanies the encoded video frame.
[00021] In another embodiment again, video information is stored as blocks within the multimedia file. Furthermore, each block of video information can include a single video frame. Moreover, audio information can be stored as blocks within the multimedia file, and audio information from two separate audio tracks is not contained within a single audio information block.
[00022] In yet another embodiment, the 'video' blocks are separated by at least one 'audio' block from each of the audio tracks, and the 'audio' blocks that separate the 'video' blocks contain the audio information for the portions of the audio tracks that accompany the video information contained within the 'video' block following the 'audio' block.
[00023] A system for encoding multimedia files of Petition 870180148460, dated 06 / 11 / 2018, page 11 / 125 6 / 111 according to one embodiment of the invention includes a processor configured to encode a video track, encode a plurality of audio tracks, interleave the video track information with the information from the plurality of audio tracks, and write the interleaved video and audio information into a single file.
[00024] In an additional embodiment, at least one of the audio tracks includes a plurality of audio channels.
[00025] In another mode, the processor is also configured to encode the header information that identifies the number of the encoded audio tracks and write the header information to the single file.
[00026] In an additional embodiment again, the processor is further configured to encode the header information that identifies the description information on at least one of the encoded audio tracks and write the header information to the single file.
[00027] In another embodiment again, the processor encodes the video track as blocks of video information. In addition, the processor can encode each audio track as a series of blocks of audio information. Furthermore, each block of audio information can contain the audio information of a single audio track, and the processor can be configured to interleave the blocks of audio information between the blocks of video information.
[00028] In yet another embodiment, the processor is configured to encode the portion of each audio track that accompanies the video information in a 'video' block into an 'audio' block, and the processor is configured to interleave the 'video' blocks with the 'audio' blocks so that each 'video' block is preceded by 'audio' blocks containing the audio information. Petition 870180148460, dated 06 / 11 / 2018, page 12 / 125 7 / 111 of each of the audio tracks that accompany the video information contained in the 'video' block. Additionally, the processor may be configured to encode the video track so that a single video frame is contained within each 'video' block.
[00029] In yet another form again, the processor is a general-purpose processor.
[00030] In yet another embodiment, the processor is a dedicated circuit.
[00031] A system for decoding a multimedia file containing a plurality of audio tracks according to the present invention includes a processor configured to extract information from the multimedia file. The processor is configured to extract information regarding the number of audio tracks contained in the multimedia file.
[00032] In an additional embodiment, the processor is configured to select a single audio track from a plurality of audio tracks and the processor is configured to decode the audio information of the selected audio track.
[00033] In another embodiment, at least one of the audio tracks includes a plurality of audio channels.
[00034] In yet another mode, the processor is configured to extract header information from the multimedia file that includes descriptive information about at least one of the audio tracks.
[00035] A system for communicating multimedia information according to an embodiment of the invention includes a network, a storage device containing a multimedia file and connected to the network via a server, and a client connected to the network. The client can request the transfer of the multimedia file from the server, and the multimedia file includes at least one track of Petition 870180148460, dated 06 / 11 / 2018, page 13 / 125 8 / 111 video and a plurality of audio tracks that accompany the video track.
[00036] A multimedia file according to the present invention includes a series of encoded video frames and at least one encoded subtitle track interspersed between the encoded video frames.
[00037] In a further embodiment, at least one encoded caption track comprises a plurality of encoded caption tracks.
[00038] Another option includes header information that identifies the number of encoded subtitle tracks contained in the multimedia file.
[00039] An additional mode also includes header information which includes description information about at least one of the encoded legend tracks.
[00040] In yet another embodiment, each legend track includes a series of bitmaps and each legend track may include a series of compressed bitmaps. Furthermore, each bitmap is compressed using run extension encoding.
[00041] In yet another embodiment, the series of encoded video frames are encoded as a series of video blocks, and each encoded caption track is encoded as a series of caption blocks. Each caption block includes information capable of being represented as text on a display. Furthermore, each caption block may contain information pertaining to a single caption. Moreover, each caption block may include information pertaining to the portion of the video sequence over which the caption is to be superimposed.
[00042] In yet another modality, each caption block includes information relating to the portion of a display in which the caption appears. Petition 870180148460, dated 06 / 11 / 2018, page 14 / 125 9 / 111 needs to be located.
[00043] In an additional embodiment, each legend block includes information regarding the legend color, and the information regarding the color may include a color palette. Furthermore, the legend blocks may comprise a first legend block that includes information regarding a first color palette and a second legend block that includes information regarding a second color palette that replaces the information regarding the first color palette.
[00044] A system for encoding multimedia files according to an embodiment of the invention may include a processor configured to encode a video track, encode at least one subtitle track, interleave the video track information with the information from at least one subtitle track, and write the interleaved video and subtitle information into a single file.
[00045] In an additional embodiment, at least one caption track includes a plurality of caption tracks.
[00046] In another mode, the processor is also configured to encode and write to the single file the header information that identifies the number of subtitle tracks contained in the multimedia file.
[00047] In an additional mode again, the processor is further configured to encode and write to the single file the description information about at least one of the legend tracks.
[00048] In another additional embodiment, the video track is encoded as video blocks and each of the at least one subtitle track is encoded as subtitle blocks. Furthermore, each of the subtitle blocks contains a single subtitle that accompanies a portion of the video track and the interleaver can be configured Petition 870180148460, dated 06 / 11 / 2018, page 15 / 125 10 / 111 rado to interleave each caption block before the video blocks that contain the portion of the video track that the caption within the caption block accompanies.
[00049] In yet another mode, the processor is configured to generate a caption block by encoding the caption as a bitmap.
[00050] In yet another mode, the subtitle is encoded as a compressed bitmap. Furthermore, the bitmap can be compressed using run extension encoding. Moreover, the processor can include in each subtitle block the information relating to the portion of the video sequence over which the subtitle should be superimposed.
[00051] In yet another additional mode, the processor includes in each caption block the information relating to the portion of a display in which the caption should be located.
[00052] In yet another mode, the processor includes in each caption block the information regarding the caption color.
[00053] In yet another additional embodiment, the color information includes a color palette. Furthermore, the legend blocks may include a first legend block that includes information relating to a first color palette and a second legend block that includes information relating to a second color palette that replaces the information relating to the first color palette.
[00054] A system for decoding multimedia files according to an embodiment of the present invention includes a processor configured to extract information from the multimedia file. The processor is configured to inspect the multimedia file to determine if there is at least one subtitle track. Furthermore, the presence of at least one subtitle track may comprise Petition 870180148460, dated 06 / 11 / 2018, page 16 / 125 11 / 111 gives a plurality of subtitle tracks and the processor may be configured to determine the number of subtitle tracks in the multimedia file.
[00055] In an additional mode, the processor is also configured to extract the header information that identifies the number of subtitle tracks in the multimedia file.
[00056] In yet another mode, the processor is also configured to extract descriptive information about at least one of the subtitle tracks from the multimedia file.
[00057] In a further embodiment again, the multimedia file includes at least one video track encoded as video blocks and the multimedia file includes at least one subtitle track encoded as subtitle blocks.
[00058] In another mode again, each caption block includes information relating to a single caption.
[00059] In yet another embodiment, each subtitle is encoded in the subtitle blocks as a bitmap, the processor is configured to decode the video track, and the processor is configured to construct a video frame for display by overlaying the bitmap onto a portion of the video sequence. Furthermore, the subtitle may be encoded as a compressed bitmap, and the processor is configured to decompress the bitmap. Still more, the processor may be configured to decompress a run-extension encoded bitmap.
[00060] In yet another mode, each caption block includes information regarding the portion of the video track over which the caption should be superimposed, and the processor is configured to generate a sequence of video frames for display by overlaying the caption bitmap over each video frame indicated by the information in the caption block. Petition 870180148460, dated 06 / 11 / 2018, page 17 / 125 12 / 111
[00061] In yet another additional mode, each caption block includes information regarding the position within a frame where the caption should be located, and the processor is configured to overlay the caption at the position within each video frame indicated by the information within the caption block.
[00062] In another additional embodiment, each legend block includes information regarding the legend color, and the processor is configured to overlay the legend with the color or colors indicated by the color information within the legend block. Furthermore, the color information within a legend block may include a color palette, and the processor is configured to overlay the legend using the color palette to obtain the color information used in the legend bitmap.Furthermore, the legend blocks may comprise a first legend block that includes information relating to a first color palette and a second legend block that includes information relating to a second color palette, and the processor may be configured to overlay the legend using the first color palette to obtain information relating to the colors used in the legend's bitmap after the first block is processed, and the processor may be configured to overlay the legend using the second color palette to obtain information relating to the colors used in the legend's bitmap after the second block is processed.
[00063] A system for communicating multimedia information according to one embodiment of the invention includes a network, a storage device containing a multimedia file and connected to the network via a server, and a client connected to the network. The client requests the transfer of the multimedia file from the server, and the multimedia file includes at least one video track and at least one subtitle track that accompanies the video track. Petition 870180148460, dated 06 / 11 / 2018, page 18 / 125 13 / 111
[00064] A multimedia file according to one embodiment of the invention includes a series of encoded video frames and encoded menu information. In addition, the encoded menu information can be stored as a block.
[00065] An additional embodiment also includes at least two separate 'menu' blocks of menu information and at least two separate 'menu' blocks may be contained within at least two separate 'RIFF' blocks.
[00066] In another embodiment, the first 'RIFF' block containing a 'menu' block includes a standard 4cc code and the second 'RIFF' block containing a 'menu' block includes a specialized 4cc code where the first two characters of a standard 4cc code appear as the last two characters of the specialized 4cc code.
[00067] In an additional mode again, at least two separate 'menu' blocks are contained within a single 'RIFF' block.
[00068] In another embodiment again, the 'menu' block includes blocks that describe a series of menus and an 'MRIF' block that contains media associated with the menu series. In addition, the 'MRIF' block may contain media information including video tracks, audio tracks, and overlay tracks.
[00069] In yet another embodiment, the blocks describing a series of menus may include a block describing the total menu system, at least one block grouping the menus by language, at least one block describing an individual menu display and accompanying background audio, and at least one block describing a button in a menu, at least one block describing the button's location on the screen, and at least one block describing various actions associated with a button. Petition 870180148460, dated 06 / 11 / 2018, p. 19 / 125 14 / 111
[00070] Yet another mode also includes a connection to a second file. The encoded menu information is contained in the second file.
[00071] A system for encoding multimedia files according to an embodiment of the invention includes a processor configured to encode menu information. The processor is also configured to generate a multimedia file that includes an encoded video track and the encoded menu information. Furthermore, the processor may be configured to generate an object model of the menus, convert the object model into a configuration file, parse the configuration file into blocks, generate AVI files containing the media information, interleave the media in the AVI files into an 'MRIF' block, and concatenate the parsed blocks with the 'MRIF' block to create a 'menu' block. Moreover, the processor may also be configured to use the object model to generate a second, smaller 'menu' block.
[00072] In an additional embodiment, the processor is configured to encode a second menu and the processor can insert the first encoded menu into a first 'RIFF' block and insert the second encoded menu into a second 'RIFF' block.
[00073] In another mode, the processor includes the first and second menus encoded in a single 'RIFF' block.
[00074] In an additional embodiment again, the processor is configured to insert into the multimedia file a reference to a menu encoded in a second file.
[00075] A system for decoding multimedia files according to the present invention includes a processor configured to extract information from the multimedia file. The processor is configured to inspect the multimedia file to determine if it contains the encoded menu information. In addition, Petition 870180148460, dated 06 / 11 / 2018, page 20 / 125 15 / 111 so, the processor can be configured to extract menu information from a 'menu' block within a 'RIFF' block, and the processor can be configured to construct menu displays using the video information stored in the 'menu' block.
[00076] In an additional embodiment, the processor is configured to generate background audio that accompanies a menu display using the audio information stored in the 'menu' block.
[00077] In another embodiment, the processor is configured to generate a menu display by overlaying a 'menu' block overlay on the video information of the 'menu' block.
[00078] A system for communicating multimedia information according to the present invention includes a network, a storage device containing a multimedia file and connected to the network via a server, and a client connected to the network. The client can request the transfer of the multimedia file from the server, and the multimedia file includes encoded menu information.
[00079] A multimedia file that includes a series of encoded video frames and encoded metadata about the multimedia file. The encoded metadata includes at least one sentence comprising a subject, a predicate, an object, and an authority. In addition, the subject may contain information that identifies a file, an item, a person, or an organization that is described by the metadata; the predicate may contain information indicative of a characteristic of the subject; the object contains descriptive information about the characteristic of the subject identified by the predicate; and the authority may contain information pertaining to the source of the sentence.
[00080] In an additional embodiment, the subject is a block that includes a type and a value, where the value contains the information and the type indicates whether the block is a resource or an anonymous node. Petition 870180148460, dated 06 / 11 / 2018, page 21 / 125 16 / 111
[00081] In another embodiment, the predicate is a block that includes a type and a value, where the value contains the information and the type indicates whether the value information is a predicate URI or an ordinary list entry.
[00082] In a further embodiment again, the object is a block that includes a type, a language, a data type, and a value, where the value contains the information, the type indicates whether the value information is a UTF-8 literal, an integer literal, or XML data literal, the data type indicates the type of the value information, and the language contains the information that identifies a specific language.
[00083] In another modality again, the authority is a block that includes a type and a value, where the value contains the information and the type indicates whether the value information is the authority of the statement.
[00084] In yet another embodiment, at least a portion of the encoded data is represented as binary data.
[00085] In yet another embodiment, at least a portion of the encoded data is represented as 64-bit ASCII data.
[00086] In yet another embodiment, at least a first portion of the encoded data is represented as binary data and at least a second portion of the encoded data is represented as additional blocks containing the data represented in a second format. Furthermore, the additional blocks may each contain a single metadata segment.
[00087] A system for encoding multimedia files according to an embodiment of the present invention includes a processor configured to encode a video track. The processor is also configured to encode metadata relating to the multimedia file, and the encoded metadata includes at least one sentence comprising a subject, a predicate, an object, and Petition 870180148460, dated 06 / 11 / 2018, page 22 / 125 17 / 111 an authority. In addition, the subject may contain information that identifies a file, an item, a person, or an organization that is described by the metadata; the predicate may contain information indicative of a characteristic of the subject; the object may contain descriptive information about the characteristic of the subject identified by the predicate; and the authority may contain information regarding the source of the sentence.
[00088] In an additional embodiment, the processor is configured to encode the subject as a block that includes a type and a value, where the value contains the information and the type indicates whether the block is a resource or an anonymous node.
[00089] In another embodiment, the processor is configured to encode the predicate as a block that includes a type and a value, where the value contains the information and the type indicates whether the value information is a predicate URI or an ordinary list entry.
[00090] In an additional embodiment again, the processor is configured to encode the object as a block that includes a type, a language, a data type, and a value, where the value contains the information, the type indicates whether the value information is a UTF-8 literal, an integer literal, or XML data literal, and the data type indicates the type of the value information and the language contains the information that identifies a specific language.
[00091] In another mode again, the processor is configured to encode the authority as a block that includes a type and a value, where the value contains the information and the type indicates whether the value information is the authority of the statement.
[00092] In yet another additional mode, the processor is also configured to encode at least a portion of the metadata relating to the multimedia file as binary data.
[00093] In yet another mode, the processor is still con Petition 870180148460, dated 06 / 11 / 2018, page 23 / 125 18 / 111 figuratively used to encode at least a portion of the metadata relating to the multimedia file as 64-bit ASCII data.
[00094] In an additional embodiment, the processor is further configured to encode at least a first portion of the metadata relating to the multimedia file as binary data and to encode at least a second portion of the metadata relating to the multimedia file as additional blocks containing the data represented in a second format. In addition, the processor may also be configured to encode the additional blocks with a single metadata segment.
[00095] A system for decoding multimedia files according to the invention includes a processor configured to extract information from the multimedia file. The processor is configured to extract metadata information pertaining to the multimedia file, and the metadata information includes at least one sentence comprising a subject, a predicate, an object, and an authority. Furthermore, the processor may be configured to extract, from the subject, information identifying a file, item, person, or organization described by the metadata. Moreover, the processor may be configured to extract information indicative of a characteristic of the subject of the predicate, the processor may be configured to extract descriptive information of the characteristic of the subject identified by the predicate of the object, and the processor may be configured to extract information pertaining to the source of the sentence of the authority.
[00096] In an additional embodiment, the subject is a block that includes a type and a value, and the processor is configured to identify that the block contains the subject information by inspecting the type, and the processor is configured to extract the value information. Petition 870180148460, dated 06 / 11 / 2018, page 24 / 125 19 / 111
[00097] In yet another embodiment, the predicate is a block that includes a type and a value, and the processor is configured to identify that the block contains the predicate information by inspecting the type, and the processor is configured to extract the value information.
[00098] In an additional embodiment again, the object is a block that includes a type, a language, a data type, and a value; the processor is configured to identify that the block contains object information by inspecting the type; the processor is configured to inspect the data type to determine the data type of the information contained in the value; the processor is configured to extract information of a type indicated by the data type of the value; and the processor is configured to extract information that identifies a specific language from the language.
[00099] In another mode again, the authority is a block that includes a type and a value, and the processor is configured to identify that the block contains the authority information by inspecting the type, and the processor is configured to extract the information from the value. [000100] In yet another mode, the processor is configured to extract the metadata sentence information and display at least a portion of the information. [000101] In yet another mode, the processor is configured to construct data structures indicative of an identified graph directed into memory using metadata. [000102] In yet another mode, the processor is configured to search through metadata for information by inspecting at least one of the subject, predicate, object, and authority for a plurality of sentences. [000103] In yet another mode, the processor is configured Petition 870180148460, dated 06 / 11 / 2018, page 25 / 125 20 / 111 to display search results as part of a graphical user interface. Additionally, the processor can be configured to perform a search in response to a request from an external device. [000104] In yet another embodiment, at least a portion of the metadata information relating to the multimedia file is represented as binary data. [000105] In another additional embodiment, at least a portion of the metadata information relating to the multimedia file is represented as 64-bit ASCII data. [000106] In another additional embodiment, at least a first portion of the metadata information relating to the multimedia file is represented as binary data and at least a second portion of the metadata information relating to the multimedia file is represented as additional blocks containing the data represented in a second format. Furthermore, the additional blocks may contain a single metadata segment. [000107] A system for communicating multimedia information according to the present invention includes a network, a storage device containing a multimedia file that is connected to the network via a server, and a client connected to the network. The client can request the transfer of the multimedia file from the server, and the multimedia file includes metadata relating to the multimedia file, and the metadata includes at least one sentence comprising a subject, a predicate, an object, and an authority. [000108] A multimedia file according to the present invention includes at least one encoded video track, at least one encoded audio track, and a plurality of encoded text sequences. The encoded text sequences describe the characters Petition 870180148460, dated 06 / 11 / 2018, page 26 / 125 21 / 111 characteristics of at least one video track and at least one audio track. [000109] In an additional embodiment, a plurality of text sequences describes the same feature of a video track or an audio track using different languages. [000110] Another embodiment also includes at least one encoded caption track. The plurality of encoded text sequences includes sequences that describe the characteristics of the caption track. [000111] A system for creating a multimedia file according to the present invention includes a processor configured to encode at least one video track, encode at least one audio track, interleave at least one of the encoded audio tracks with a video track, and insert text sequences that describe each of a number of features of the at least one video track and the at least one audio track in a plurality of languages. [000112] A system for displaying a multimedia file that includes audio, video and encoded text sequences according to the present invention, which includes a processor configured to extract the encoded text sequences from the file and generate a drop-down menu display using the text sequences. Brief Description of the Drawings [000113] Figure 1 is a diagram of a system according to an embodiment of the present invention for encoding, distributing and decoding files. [000114] Figure 2.0 is a diagram of a multimedia file structure according to an embodiment of the present invention. [000115] Figure 2.0.1 is a diagram of a multimedia file structure according to another embodiment of the present invention. Petition 870180148460, dated 06 / 11 / 2018, page 27 / 125 22 / 111 action. [000116] Figure 2.1 is a conceptual diagram of an 'hdrl' list block according to an embodiment of the invention. [000117] Figure 2.2 is a conceptual diagram of a 'strl' block according to an embodiment of the invention. [000118] Figure 2.3 is a conceptual diagram of the memory allocated to store a 'DXDT' block of a multimedia file according to an embodiment of the invention. [000119] Figure 2.3.1 is a conceptual block diagram of 'metadata' that can be included in a 'DXDT' block of a multimedia file according to an embodiment of the invention. [000120] Figure 2.4 is a conceptual diagram of the 'DMNU' block according to one embodiment of the invention. [000121] Figure 2.5 is a conceptual diagram of menu blocks contained within a 'WowMenuManager' block according to an embodiment of the invention. [000122] Figure 2.6 is a conceptual diagram of menu blocks contained within a 'WowMenuManager' block according to another embodiment of the invention. [000123] Figure 2.6.1 is a conceptual diagram that illustrates the relationships between the various blocks contained in a 'DMNU' block. [000124] Figure 2.7 is a conceptual diagram of the 'movi' list block of a multimedia file according to an embodiment of the invention. [000125] Figure 2.8 is a conceptual diagram of the 'movi' list block of a multimedia file according to an embodiment of the invention that includes a DRM. [000126] Figure 2.9 is a conceptual diagram of the 'DRM' block according to an embodiment of the invention. [000127] Figure 3.0 is a block diagram of a system for Petition 870180148460, dated 06 / 11 / 2018, page 28 / 125 23 / 111 generate a multimedia file according to an embodiment of the invention. [000128] Figure 3.1 is a block diagram of a system for generating a 'DXDT' block according to an embodiment of the invention. [000129] Figure 3.2 is a block diagram of a system for generating a 'DMNU' block according to an embodiment of the invention. [000130] Figure 3.3 is a conceptual diagram of a media model according to an embodiment of the invention. [000131] Figure 3.3.1 is a conceptual diagram of objects of a media model that can be used to automatically generate a small menu according to an embodiment of the invention. [000132] Figure 3.4 is a flowchart of a process that can be used to re-block audio according to an embodiment of the invention. [000133] Figure 3.5 is a block diagram of a video encoder according to an embodiment of the invention. [000134] Figure 3.6 is a flowchart of a method for performing a psychovisual smoothing enhancement on an I-frame according to embodiments of the invention. [000135] Figure 3.7 is a flowchart of a process for performing a macroblock SAD psychovisual enhancement according to an embodiment of the invention. [000136] Figure 3.8 is a flowchart of a process for controlling the rate of a pass according to an embodiment of the invention. [000137] Figure 3.9 is a flowchart of a process for performing the VBV rate control of Na pass according to an embodiment of the invention. [000138] Figure 4.0 is a flowchart for a process to locate the required multimedia information from a multimedia file and display the multimedia information according to a modality. Petition 870180148460, dated 06 / 11 / 2018, page 29 / 125 24 / 111 validity of the invention. [000139] Figure 4.1 is a block diagram of a decoder according to an embodiment of the invention. [000140] Figure 4.2 is an example of a menu displayed according to an embodiment of the invention. [000141] Figure 4.3 is a conceptual diagram showing the information sources used to generate the display shown in Figure 4.2 according to one embodiment of the invention. Detailed Description of the Invention [000142] Referring to the drawings, embodiments of the present invention are capable of encoding, transmitting, and decoding multimedia files. Multimedia files according to embodiments of the present invention may contain multiple video tracks, multiple audio tracks, multiple subtitle tracks, data that can be used to generate a menu interface to access the file content, and 'metadata' relating to the file. Multimedia files according to various embodiments of the present invention also include references to video tracks, audio tracks, subtitle tracks, and metadata external to the file. System Description [000143] Looking now at Figure 1, a system according to an embodiment of the present invention for encoding, distributing and decoding files is shown. The system 10 includes a computer 12, which is connected to a variety of other computing devices via a network 14. Devices that may be connected to the network include a server 16, a laptop computer 18 and a personal digital assistant (PDA) 20. In various embodiments, the connections between the devices and the network may be either wired or wireless and implemented using any of a variety of network protocols. Petition 870180148460, dated 06 / 11 / 2018, page 30 / 125 25 / 111 [000144] In operation, computer 12 can be used to encode multimedia files according to an embodiment of the present invention. Computer 12 can also be used to decode multimedia files according to embodiments of the present invention and distribute multimedia files according to embodiments of the present invention. The computer can distribute the files using any of a variety of transfer protocols, including over a peer-to-peer network. Furthermore, computer 12 can transfer the multimedia files according to embodiments of the present invention to a server 18, where the files can be accessed by other devices. The other devices may include any variety of computing devices or even a dedicated decoder device. In the illustrated embodiment, a laptop computer and a PDA are shown.In other configurations, frequency converters, desktop computers, game consoles, consumer electronics devices, and other devices can be connected to the network, load multimedia files, and decode them. [000145] In one embodiment, devices access multimedia files from the server over the network. In other embodiments, devices access multimedia files from a number of computers over a peer-to-peer network. In several embodiments, multimedia files can be written to a portable storage device, such as a disk drive, a CD-ROM, or a DVD. In many embodiments, electronic devices can access multimedia files written to portable storage devices. Description of the File Structure [000146] Multimedia files according to embodiments of the present invention can be structured to be in con Petition 870180148460, dated 06 / 11 / 2018, p. 31 / 125 26 / 111 conformity with the Resource Interchange File Format ('RIFF file format'), defined by Microsoft Corporation of Redmond, WA and International Business Machines Corporation of Armonk, NY. RIFF is a format for storing multimedia data and associated information. A RIFF file typically has an 8-byte RIFF header, which identifies the file and provides the residual file length after the header (i.e., file_length 8). The entire remainder of the RIFF file comprises blocks and lists. Each block has an 8-byte block header that identifies the block type and provides the length in bytes of the data following the block header. Each list has an 8-byte header that identifies the list type and provides the length in bytes of the data following the list header. The data in a list comprises blocks and / or other lists (which in turn may comprise blocks and / or lists).RIFF lists are also sometimes referred to as blocks of lists. [000147] An AVI file is a special form of RIFF file that follows the format of a RIFF file, but includes multiple blocks and lists with defined identifiers that contain multimedia data in specific formats. The AVI format was developed and defined by Microsoft Corporation. AVI files are typically created using an encoder that can output multimedia data in AVI format. AVI files are typically decoded by any one of a group of software programs known collectively as AVI decoders. [000148] The RIFF and AVI formats are flexible because they only define the blocks and lists that are part of the defined file format, but allow files to also include lists and / or blocks that are outside the RIFF and / or AVI file format definitions without making the file unreadable by a decoder. Petition 870180148460, dated 06 / 11 / 2018, p. 32 / 125 27 / 111 RIFF and / or AVI decoders. In practice, AVI (and similarly RIFF) decoders are implemented in such a way that they simply ignore lists and blocks containing header information not found in the AVI file format definition. The AVI decoder must still read through these non-AVI blocks and lists, and thus the operation of the AVI decoder may be slowed down, but otherwise, these generally have no effect and are ignored by an AVI decoder. [000149] A multimedia file according to an embodiment of the present invention is illustrated in Figure 2.0. The multimedia file 30 includes a character set block ('CSET' block) 32, an information list block ('INFO' list block) 34, a file header block ('hdrl' list block) 36, a metadata block ('DXDT' block) 38, a menu block ('DMNU' block) 40, a trash block ('trash' block) 41, a movie list block ('movi' list block) 42, an optional index block ('idx1' block) 44, and a second menu block ('DMNU' block) 46. Some of these blocks and portions of others are defined in the AVI file format while others are not contained in the AVI file format. In many, but not all, cases, the discussion below identifies the blocks or portions of blocks that are defined as part of the AVI file format. [000150] Another multimedia file according to an embodiment of the present invention is shown in Figure 2.0.1. The multimedia file 30' is similar to that shown in Figure 2.0 except that the file includes multiple concatenated 'RIFF' blocks. The 'RIFF' blocks may contain a 'RIFF' block similar to that shown in Figure 2.0 which may exclude the second 'DMNU' block 46 or may contain the menu information in the form of a 'DMNU' block 46'. [000151] In the illustrated version, the multimedia includes multiple blocks Petition 870180148460, dated 06 / 11 / 2018, page 33 / 125 28 / 111 of concatenated 'RIFF', where the first 'RIFF' block 50 includes a character set block ('CSET' block) 32', an information list block ('INFO' list block) 34', a file header block ('hdrl' list block) 36', a metadata block ('DXDT' block) 38', a menu block ('DMNU' block) 40', a junk block ('junk' block) 41', a movie list block ('movi' list block) 42', and an optional index block ('idx1' block) 44'. The second 'RIFF' block 52 contains a second menu block ('DMNU' block) 46'. Additional 'RIFF' blocks containing additional titles can be included after the 'RIFF' menu block 52. The additional 'RIFF' blocks can contain independent media conforming to the AVI file format.In one embodiment, the second menu block 46' and the additional 'RIFF' blocks have specialized 4-character codes (defined in the AVI format and discussed below) such that the first two characters of the 4-character code appear as the second two characters and the second two characters of the 4-character code appear as the first two characters. 2.1 The 'CSET' Block [000152] The 'CSET' block 32 is a block defined in the Audio Video Interleaved format (AVI file format), created by Microsoft Corporation. The 'CSET' block defines the character set and language information of the multimedia file. The inclusion of a 'CSET' block according to the embodiments of the present invention is optional. [000153] A multimedia file according to an embodiment of the present invention does not use the 'CSET' block and uses UTF-8, which is defined by the Unicode Consortium, for the character set as a standard combined with the RFC 3066 Language Specification, which is defined by the Internet Engineering Task Force. Petition 870180148460, dated 06 / 11 / 2018, page 34 / 125 29 / 111 for information language. 2.2 The 'Info' List Block [000154] The 'INFO' list block 34 can store information that helps identify the content of the multimedia file. The 'INFO' list is defined in AVI file format and its inclusion in a multimedia file according to the embodiments of the present invention is optional. Many embodiments that include a 'DXDT' block do not include an 'INFO' list block. [000155] 2.3 The 'hdrl' List Block [000156] The 'hdrl' list block 38 is defined in the AVI file format and provides information regarding the data format in the multimedia file. The inclusion of an 'INFO' list block or a block containing similar descriptive information is generally required. The 'hdrl' list block or a block containing similar descriptive information is generally required. The 'hdrl' list block includes one block for each video track, each audio track, and each subtitle track. [000157] A conceptual diagram of an 'hdrl' list block 38 according to an embodiment of the invention includes a single video track 62, two audio tracks 64, and an external audio track 66, two subtitle tracks 68 and an external subtitle track 70 is illustrated in Figure 2.1. The 'hdrl' list 60 includes an 'avih' block. The 'avih' block 60 contains global information for the entire file, such as the number of streams within the file and the width and height of the video contained in the multimedia file. The 'avih' block can be implemented according to the AVI file format. [000158] In addition to the 'avih' block, the 'hdrl' list includes a stream descriptor list for each audio, video, and subtitle track. In one embodiment, the stream descriptor list is implemented using 'strl' blocks. A 'strl' block according to an embodiment Petition 870180148460, dated 06 / 11 / 2018, page 35 / 125 30 / 111 of the present invention is illustrated in Figure 2.2. Each 'strl' block serves to describe each track in the multimedia file. The 'strl' blocks for the video and subtitle audio tracks within the multimedia file include a 'strl' block referencing a 'strh' block 92, a 'strf' block 94, a 'strd' block 96, and a 'strn' block 98. All these blocks can be implemented according to the AVI file format. Of particular interest is the 'strh' block 92, which specifies the media track type, and the 'strd' block 96, which can be modified to indicate whether the video is protected by digital rights management. A discussion of various digital rights management implementations according to embodiments of the present invention is provided below. [000159] Multimedia files according to embodiments of the present invention may contain references to external files containing multimedia information, such as an additional audio track or an additional subtitle track. References to these tracks may be contained either in the 'hdrl' block or in the 'garbage' block 41. In either case, the reference may be contained in the 'strh' block 92 of a 'strl' block 90, with references either to a local file or a remotely stored file. The referenced file may be a standard AVI file or a multimedia file according to an embodiment of the present invention containing the additional track. [000160] In additional modalities, the referenced file may contain any of the blocks that may be present in the reference file, including 'DMNU' blocks, 'DXDT' blocks, and blocks associated with audio, video, and / or subtitle tracks for a multimedia presentation. For example, a first multimedia file could include a 'DMNU' block (discussed in more detail below) that references the first multimedia presentation. Petition 870180148460, dated 06 / 11 / 2018, page 36 / 125 31 / 111 is located within the 'movi' list block of the first multimedia file, and a second multimedia presentation is located within the 'movi' list block of the second multimedia file. Alternatively, both 'movi' list blocks may be included in the same multimedia file, which does not need to be the same file as the file in which the 'DMNU' block is located. 2.4. The 'dxdt' Block [000161] The 'DXDT' block 38 contains so-called 'metadata'. 'Metadata' is a term used to describe data that provides information about the content of a file, document, or transmission. The 'metadata' stored within the 'DXDT' block of multimedia files according to the embodiments of the present invention can be used to store such specific content information as the title, author, copyright holder, and cast. In addition, technical details about the codec used to encode the multimedia file can be provided, such as the CLI options used and the quantifier distribution after each pass. [000162] In one embodiment, metadata is represented within the 'DXDT' block as a series of sentences, where each sentence includes a subject, a predicate, an object, and an authority. The subject is a reference to what is being described. The subject can refer to a file, an item, a person, or an organization. The subject can refer to anything that has characteristics capable of being described. The predicate identifies a characteristic of the subject being described. The object is a description of the identified characteristic of the subject, and the authority identifies the source of the information. [000163] The following is a table showing an example of how various 'metadata' segments can be represented as an object, a predicate, a subject, and an authority: Petition 870180148460, dated 06 / 11 / 2018, page 37 / 125 Sujeito Predicado Objeto Autoridade _:file281 http: / / purl.org / dc / elements / 1.1 / title The Matrix _:auth42 _:file281 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#Person _:cast871 _:auth42 _:file281 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#Person _:cast872 _:auth42 _:file281 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#Person _:cast873 _:auth42 _:cast871 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#name Keanu Reeves _:auth42 _:cast871 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#role Actor _:auth42 _:cast871 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#character Neo _:auth42 _:cast282 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#name Andy Wachowski _:auth42 _:cast282 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#role Director _:auth42 _:cast283 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#name Larry Wachowski _:auth42 _:cast283 http: / / xmlns.divxnetworks.com / 2004 / 11 / cast#role Director _:auth42 _:file281 http: / / purl.org / dc / elements / 1.1 / rights Copyright 1998 Warner Brothers. All Rights Reserved._:auth42 _:file281 Series _:file321 _:auth42 _:file321 Episode 2 _:auth42 _:file321 http: / / purl.org / dc / elements / 1.1 / title The Matrix Reloaded _:auth42 _:file321 Series _:file122 _:auth42 _:file122 Episode 3 _:auth42 _:auth42 http: / / xmlns.com / foaf / 0.1 / Organization _:foaf92 _:auth42 _:foaf92 http: / / xmlns.com / foaf / 0.1 / name Warner Brothers _:auth42 _:file281 http: / / xmllns.divxnetworks.com / 2004 / 11 / track#track _:track#dc00 _:auth42 _:track#dc00 http: / / xmlns.divxnetworks.com / 2004 / 11 / track#resolution 1024x768 _:auth42 _:file281 http: / / xmlns.divxnetworks.com / 2004 / 11 / content#certificationLevel HT _:auth42 _:track#dc00 http: / / xmlns.divxnetworks.com / 2004 / 11 / track#frameTypeDist 32,1,3,5 _:auth42 _:track#dc00 http: / / xmlns.divxnetworks.com / 2004 / 11 / track#codecSettings bv1 276 -psy 0 -key 300 -b 1 -sc 50 -pq 5 -vbv 6951200,3145728,2359296 -profile 3 -nf _:auth42. 32 / 111 Tabela 1 - Representação conceituai de 'metadados' Petition 870180148460, dated 06 / 11 / 2018, page 38 / 125 33 / 111 [000164] In one embodiment, the expression of the subject, predicate, object, and authority is implemented using binary representations of the data, which can be considered to form Directed Identified Graphs (DLGs). A DLG consists of nodes that are either resources or literals. Resources are identifiers, which may conform to a naming convention, such as a Universal Resource Identifier (URI) as defined in RFC 2396 by the Internet Engineering Task Force (http: / / www.ietf.org / rfc / rfc2396.txt), or refer to data specific to the system itself. Literals are representations of a real value, rather than a reference. [000165] One advantage of DLGs is that they allow the inclusion of a flexible number of data items that are of the same type, such as the cast members of a film. In the example shown in Table 1, three cast members are included. However, any number of cast members can be included. DLGs also allow relational connections with other data types. In Table 1, there is a 'metadata' item that has a subject _:file281, a predicate Series, and an object _:file321. The subject _:file281 indicates that the 'metadata' refers to the content of the file referenced as _:file321 (in this case, the film The Matrix). The predicate is Series, which indicates that the object will have information about other films in the series to which The Matrix belongs. However, _:file321 is not the title or any other specific information about the series that includes The Matrix, but rather a reference to another entry that provides more information about _:file321.The next 'metadata' entry, with the subject _:file321, however includes data about _:fi1e321, namely that the Title as specified by the Dublin Core Vocabulary as indicated by http: / / pur1.org / dc / elements / 1.1 / title of this sequence is The Matrix Reloaded. Petition 870180148460, dated 06 / 11 / 2018, p. 39 / 125 34 / 111 [000166] The additional metadata sentences in Table 1 specify that Keanu Reeves was a cast member playing the role of Neo and that both Larry and Andy Wachowski were the directors. Technical information is also expressed in the metadata. The metadata sentences identify that _ :file281 includes the track _ :track#dcOO. The metadata provides information including the video track resolution, the video track certification level, and codec settings. Although not shown in Table 1, the metadata may also include a unique identifier assigned to a track at the time of encoding. When unique identifiers are used, encoding the same content multiple times will result in a different identifier for each version of the content. However, a copy of an encoded video track would retain the identifier of the track from which it was copied. [000167] The entries shown in Table 1 can be replaced by other vocabularies, such as the UPnP vocabulary, which is defined by the UPnP forum (see http: / / www.upnpforum.org). Another alternative would be the Digital Item Declaration Language (DIDL) or DIDL-Lite vocabularies developed by the International Standards Organization as part of the work towards the MPEG-21 standard. The following are examples of predicates within the UPnP vocabulary: [000168] urn:schemas-upnp-org:metadata-1-0 / upnp / artist [000169] urn:schemas-upnp-org:metadata-1-0 / upnp / actor [000170] urn:schemas-upnp-org:metadata-1-0 / upnp / author [000171] urn:schemas-upnp-org:metadata-1-0 / upnp / producer [000172] urn:schemas-upnp-org:metadata-1-0 / upnp / director [000173] urn:schemas-upnp-org:metadata-1-0 / upnp / genre [000174] urn:schemas-upnp-org:metadata-1-0 / upnp / album [000175] urn:schemas-upnp-org:metadata-1-0 / upnp / playlist [000176] urn:schemas-upnp-org:metadata-1Petition 870180148460, dated 11 / 06 / 2018, p.40 / 125. 35 / 111 0 / upnp / originalT rackNumber [000177] urn:schemas-upnp-org:metadata-1-0 / upnp / userAnnotation [000178] The authority for all 'metadata' is '_:auth42.' The 'metadata' statements show that '_:auth42' is 'Warner Brothers.' The authority allows for the evaluation of both the quality of the file and the 'metadata' statements associated with the file. [000179] The nodes in a graph are connected via named feature nodes. A 'metadata' sentence consists of a subject node, a predicate node, and an object node. Optionally, an authority node may be connected in the DLG as part of the 'metadata' sentence. [000180] For each node, there are certain characteristics that further explain the node's functionality. The possible types can be represented as follows using the ANSI C programming language: / ** Invalid Type * / #define RDF_IDENTIFIER_TYPE_UNKNOWN 0x00 / ** Resource URI rdf:about * / #define RDF_IDENTIFIER_TYPE_RESOURCE0x01 / ** rdf:NodeId, _:file or generated N-Triples * / #define RDF_IDENTIFIER_TYPE_ANONYMOUS0x02 / ** Predicate URI * / #define RDF_IDENTIFIER_TYPE_PREDICATE0x03 / ** rdf:li, rdf:_ <n>* / #define RDF_IDENTIFIER_TYPE_ORDINAL0x04 / ** Authority URI * / #define RDF_IDENTIFIER_TYPE_AUTHORITY0x05 / ** UTF-8 formatted literal * / #define RDF_IDENTIFIER_TYPE_LITERAL0x06 / ** Integer Literal * / Petition 870180148460, dated 06 / 11 / 2018, page 41 / 125 36 / 111 #define RDF_IDENTIFIER_TYPE_INT / ** Literal XML data * / 0x07 #define RDF_IDENTIFIER_TYPE_XML_LITERAL 0x08 [000181] An example of a data structure (represented by the ANSI C programming language) that represents the 'metadata' blocks contained within the 'DXDT' block is as follows: typedef struct RDFDataStruct { RDFHeader uint32 t RDFStatement Header; numOfStatements; statements(RDF_MAX_STATEMENTS]; RDFData [000182] The 'RDFData' block includes a block referred to as an 'RDFHeader' block, a 'numOfStatements' value, and a list of 'RDFStatement' blocks. [000183] The 'RDFHeader' block contains information about the mode in which the 'metadata' is formatted in the block. In one embodiment, the data in the 'RDFHeader' block can be represented as follows (represented in ANSI C): typedef struct RDFHeaderStruct { uint16t uint16t uint16t uint16t RDFSchema versionMajor; versionMinor; versionFix; numOfSchemas; schemas(RDF_MAX_SCHEMAS]; RDFHeader; [000184] The 'RDFHeader' block includes a 'version' number that indicates the version of the feature description format to allow for forward compatibility. The header includes a second number. Petition 870180148460, dated 06 / 11 / 2018, p. 42 / 125 37 / 111 'numOfSchemas' represents the number of 'RDFSchema' blocks in the 'schemas' list, which is also part of the 'RDFHeader' block. In various ways, 'RDFSchema' blocks are used to allow complex features to be represented more efficiently. In one way, the data contained in an 'RDFSchema' block can be represented as follows (represented in ANSI C): typedef struct RDFSchemaStruct { wchar_t* prefix; wchar_t* uri; RDFSchema; [000185] The 'RDFSchema' block includes a first text string such as 'dc' identified as 'prefix' and a second text string, such as 'http: / / purl.org / dc / elements / 1-1 / ' identified as 'uri'. The 'prefix' defines a term that can be used in 'metadata' in place of the 'uri'. The 'uri' is a Universal Resource Identifier, which can conform to a specific standardized vocabulary or be a vocabulary specific to a particular system. [000186] Returning to the discussion of the 'RDFData' block. In addition to an 'RDFHeader' block, the 'RDFData' block also includes a 'numOfStatements' value and a list of 'RDFStatement' blocks. The 'numOfStatements' value indicates the actual number of 'RDFStatement' blocks in the 'RDFStatement' list that contain the information. In one embodiment, the data contained in the 'RDFStatement' block can be represented as follows (represented in ANSI C): typedef struct RDFStatementStruct { RDFSubject subject; RDFPredicate predicate; Petition 870180148460, dated 06 / 11 / 2018, p. 43 / 125 38 / 111 RDFObject object; RDF Authority; RDFStatement; [000187] Each 'RDFStatement' block contains a snippet of 'metadata' relating to the multimedia file. The 'subject', 'predicate', 'object', and 'authority' blocks are used to contain the various 'metadata' components described above. [000188] The 'subject' is a block of 'RDFSubject', which represents the subject portion of the 'metadata' described above. In one embodiment, the data contained in the 'RDFSubject' block can be represented as follows (represented in ANSI C): typedef struct RDFSubjectStruct { uint16_t type; wchar_t* value; RDFSubject; [000189] The 'RDFSubject' block shown above includes a 'type' value indicating that the data is either a Resource or an anonymous node of a 'metadata' snippet, and a unicode text string 'value', which contains the data representing the subject of the 'metadata' snippet. In modes where the 'RDFSchema' block has been defined, the value may be a defined term rather than a direct reference to a resource. [000190] The 'predicate' in an 'RDFStatement' block is an 'RDFPredicate' block, which represents the predicate portion of a 'metadata' section. In one embodiment, the data contained in an 'RDFPredicate' block can be represented as follows (represented in ANSI C): typedef struct RDFPredicateStruct { Petition 870180148460, dated 06 / 11 / 2018, p. 44 / 125 39 / 111 uint16_t type; wchar_t* value; RDFPredicate; [000191] The 'RDFPredicate' block shown above includes a 'type' value indicating that the data is the URI predicate or an ordinal list entry from a 'metadata' snippet, and a 'value' text string which contains the data representing the predicate in a 'metadata' snippet. In modes where an 'RDFSchema' block has been defined, the value can be a defined term rather than a direct reference to a resource. [000192] The 'object' in an 'RDFStatement' block is an 'RDFObject' block which represents the object portion of a 'metadata' segment. In one embodiment, the data contained in the 'RDFObject' block can be represented as follows (represented in ANSI C): typedef struct RDFObjectStruct { uint16_t type; wchar_t* language; wchar_t* dataTypeURI; wchar_t* value; RDFObject; [000193] The 'RDFObject' block shown above includes a 'type' value indicating that the data snippet is a UTF-8 literal string, an integer literal, or XML literal data from a 'metadata' snippet. The block also includes three values. The first value, 'language', is used to represent the language in which the 'metadata' snippet is expressed (for example, a movie title may vary in different languages). In various embodiments, a standard representation may be used to identify the language (such as RFC 3066 - Identification of Languages). Petition 870180148460, dated 06 / 11 / 2018, p. 45 / 125 40 / 111 (for Language Identification specified by the Internet Engineering Task Force, see http: / / www.ietf.org / rfc / rfc3066.txt). The second value of 'dataTypeURI' is used to indicate the type of data contained in the 'value' field if this cannot be explicitly indicated by the 'type' field. The URI specified by dataTypeURI points to the general RDF URI Vocabulary used to describe the specific type of data used. The different formats in which the URI can be expressed are described at http: / / www.w3.org / TR / rdfconcepts / #section-Datatypes. In one embodiment, the 'value' is a 'wide character.' In other embodiments, the 'value' can be any of a variety of data types from a single bit to an image or a video sequence. The 'value' contains the object snippet of the 'metadata.' [000194] An 'authority' in an 'RDFStatement' block is an 'RDFAuthority' block, which represents the authority portion of a 'metadata' segment. In one embodiment, the data contained in the 'RDFAuthority' block can be represented as follows (represented in ANSI C): typedef struct RDFAuthorityStruct { uint16_t type; wchar_t* value; RDF Authority; [000195] The 'RDFAuthority' data structure shown above includes a 'type' value that indicates whether the data is a Resource or an anonymous node of a 'metadata' snippet. The 'value' contains the data that represents the authority for the 'metadata'. In modes where an 'RDFSchema' block has been defined, the value may be a defined term rather than a direct reference to a resource. [000196] A conceptual representation of the storage of a Petition 870180148460, dated 06 / 11 / 2018, p. 46 / 125 41 / 111 block of 'DXDT' of a multimedia file according to an embodiment of the present invention is shown in Figure 2.3. Block 'DXDT' 38 includes an 'RDFHeader' block 110, a 'numOfStatements' value 112, and a list of RDFStatement blocks 114. Block 110 includes a 'version' value 116, a 'numOfSchemas' value 118, and a list of 'Schema' blocks 120. Each 'RDFStatement' block 114 includes an 'RDFSubject' block 122, an 'RDFPredicate' block 124, an 'RDFObject' block 126, and an 'RDFAuthority' block 128. The 'RDFSubject' block includes a 'type' value 130 and a 'value' value 132. Block 124 also includes a 'type' value 134 and a value of 'value' 136. The 'RDFObject' block 126 includes a 'type' value 138, a 'language' value 140 (shown in the figure as 'lang'), a 'dataTypeURI' value 142 (shown in the figure as 'dataT'), and a 'value' value 144.The 'RDFAuthority' block 128 includes a 'type' value 146 and a 'value' value 148. Although the illustrated 'DXDT' block is shown as including a single 'Schema' block and a single 'RDFStatement' block, a person skilled in the art will readily appreciate that different numbers of 'Schema' blocks and 'RDFStatement' blocks can be used in a block describing the 'metadata.' [000197] As discussed below, multimedia files according to embodiments of the present invention can be continuously modified and updated. Determining in advance the 'metadata' to associate with the file itself and the 'metadata' to access remotely (e.g., via the Internet) can be difficult. Typically, sufficient 'metadata' is contained in a multimedia file according to an embodiment of the present invention to describe the file's content.Additional information can be obtained if the device reviewing the file is able to access other devices containing the same files over a network. Petition 870180148460, dated 06 / 11 / 2018, page 47 / 125 42 / 111 'metadata' referenced from within the file. [000198] The methods for representing the 'metadata' described above can be extensible and can provide the ability to add or remove different 'metadata' fields stored within the file as the need for them changes over time. Furthermore, the representation of 'metadata' can be forward-compatible between revisions. [000199] The structured way in which 'metadata' is represented according to the embodiments of the present invention allows devices to query the multimedia file to better determine its content. The query could then be used to update the content of the multimedia file, to obtain additional 'metadata' relating to the multimedia file, to generate a menu relating to the file's content, or to perform any other function involving the automatic processing of data represented in a standard format. Furthermore, defining the length of each analyzable element of the 'metadata' can increase the ease with which devices with limited amounts of memory, such as consumer electronic devices, can access the 'metadata'. [000200] In other embodiments, the 'metadata' is represented using individual blocks for each 'metadata' segment. Several 'DXDT' blocks according to the present invention include a binary block containing the 'metadata' encoded as described above and additional blocks containing individual 'metadata' segments formatted either as described above or in another format. In embodiments where the binary 'metadata' is included in the 'DXDT' block, the binary 'metadata' may be represented using 64-bit encoded ASCII. In other embodiments, other binary representations may be used. Petition 870180148460, dated 06 / 11 / 2018, page 48 / 125 43 / 111 [000201] Examples of individual blocks that can be included in the 'DXDT' block according to the present invention are illustrated in Figure 2.3.1. The 'metadata' includes a 'MetaData' block 150 which may contain a 'PixelAspectRatioMetaData' block 152a, an 'EncoderURIMetaData' block 152b, a 'CodecSettingsMetaData' block 152c, a 'FrameTypeMetaData' block 152d, a 'VideoResolutionMetaData' block 152e, a 'PublisherMetaData' block 152f, a 'CreatorMetaData' block 152g, a 'GenreMetaData' block 152h, a 'CreatorToolMetaData' block 152i, a 'RightsMetaData' block 152j, a 'RunTimeMetaData' block 152k, a 'QuantizerMetaData' block 152l, and a 'CodecInfoMetaData' block. 152m, a block of 'EncoderNameMetaData' 152n, a block of 'FrameRateMetaData' 152o, a block of 'InputSourceMetaData' 152p, a block of 'FileIDMetaData' 152q, a block of 'TypeMetaData' 152r, a block of 'TitleMetaData' 152s and / or a block of 'CertLevelMetaData' 152t. [000202] The 'PixelAspectRatioMetaData' block 152a includes information regarding the pixel aspect ratio of the encoded video. The 'EncoderURIMetaData' block 152b includes information regarding the encoder. The 'CodecSettingsMetaData' block 152c includes information regarding the codec settings used to encode the video. The 'FrameTypeMetaData' block 152d includes information regarding the video frames. The 'VideoResolutionMetaData' block 152e includes information regarding the video resolution of the encoded video. The 'PublisherMetaData' block 152f includes information regarding the person or organization that published the media. The 'CreatorMetaData' block 152g includes information regarding the creator of the content. The 'GenreMetaData' block 152h includes information regarding the genre of the media. The 'CreatorToolMetaData' block 152i includes information about the tool used to... Petition 870180148460, dated 06 / 11 / 2018, p. 49 / 125 44 / 111 create the file. The 'RightsMetaData' block 152j includes DRM information. The 'RunTimeMetaData' block 152k includes media duration information. The 'QuantizerMetaData' block 152l includes information about the quantizer used to encode the video. The 'CodecInfoMetaData' block 152m includes codec information. The 'EncoderNameMetaData' block 152n includes encoder name information. The 'FrameRateMetaData' block 152o includes media frame rate information. The 'InputSourceMetaData' block 152p includes input source information. The 'FileIDMetaData' block 152q includes a unique identifier for the file. The 'TypeMetaData' block 152r includes information regarding the multimedia file type.The 'TitleMetaData' block 152s includes information regarding the media title, and the 'CertLevelMetaData' block 152t includes information regarding the media certification level. In other embodiments, additional blocks containing additional metadata may be included. In several embodiments, a block containing metadata in a binary format as described above may be included in the 'MetaData' block. In one embodiment, the binary metadata block is encoded as 64-bit ASCII. 2.5. The 'DMNU' Blocks [000203] Referring to Figures 2.0 and 2.0.1, a first 'DMNU' block 40 (40') and a second 'DMNU' block 46 (46') are shown. In Figure 2.0, the second 'DMNU' block 46 is part of multimedia file 30. In the embodiment illustrated in Figure 2.0.1, the 'DMNU' block 46' is contained in a separate RIFF block. In both cases, the first and second 'DMNU' blocks contain data that can be used to display navigable menus. In one embodiment, the first 'DMNU' block 40 (40') contains the data that can be used to create a simple menu that does not include characters. Petition 870180148460, dated 06 / 11 / 2018, p. 50 / 125 45 / 111 advanced features, such as extended background animations. In addition, the second block of 'DMNU' 46 (46') includes data that can be used to create a more complex menu that includes such advanced features as an extended animated background. [000204] In several embodiments, the provision of a simple and a complex menu may allow a device to choose which menu it wishes to display. Placing the smaller of the two menus before the 'movi' list block 42 allows devices according to embodiments of the present invention that cannot display menus to quickly advance over information that cannot be displayed. [000205] In other embodiments, the data required to create a single menu is split between the first and second 'DMNU' blocks. Alternatively, the 'DMNU' block may be a single block before the 'movi' block containing the data for a single menu set or multiple menu sets. In other embodiments, the 'DMNU' block may be a single block or multiple blocks located elsewhere throughout the multimedia file. [000206] In several multimedia files according to the present invention, the first block of 'DMNU' 40 (40') can be automatically generated based on a 'richer' menu in the second block of 'DMNU' 46 (46'). The automatic menu generation is discussed in more detail below. [000207] The structure of a 'DMNU' block according to an embodiment of the present invention is shown in Figure 2.4. The 'DMNU' block 158 is a list block containing a menu block 160 and an 'MRIF' block 162. The menu block contains the information necessary to build and navigate through the menus. The 'MRIF' block contains the media information that can be used to provide subtitles, background video and background audio for Petition 870180148460, dated 06 / 11 / 2018, page 51 / 125 46 / 111 menus. In various modes, the 'DMNU' block contains menu information that allows menus to be displayed in several different languages. [000208] In one embodiment, the 'WowMenu' block 160 contains the hierarchy of menu block objects that are conceptually illustrated in Figure 2.5. At the top of the hierarchy is the 'WowMenuManager' block 170. The 'WowMenuManager' block may contain one or more 'LanguageMenus' blocks 172 and one 'Media' block 174. [000209] The use of 'Language Menus' blocks 172 allows the 'DMNU' block 158 to contain menu information in different languages. Each 'LanguageMenus' block 172 contains the information used to generate a complete set of menus in a specified language. Therefore, the 'LanguageMenus' block includes an identifier that identifies the language of the information associated with the 'LanguageMenus' block. The 'LanguageMenus' block also includes a list of 'WowMenu' blocks 175. [000210] Each 'WowMenu' 175 block contains all the information to be displayed on the screen for a specific menu. This information may include background video and audio. The information may also include data relating to the actions of the buttons that can be used to access other menus or to exit the menu and begin displaying a portion of the multimedia file. In one embodiment, the 'WowMenu' 175 block includes a list of media references. These references refer to information contained in the 'Media' 174 block, which will be discussed further below. The media references may define the background video and background audio for a menu. The 'WowMenu' 175 block also defines an overlay that can be used to highlight a specific button when a menu is first accessed. [000211] In addition, each block of 'WowMenu' 175 includes a number Petition 870180148460, dated 06 / 11 / 2018, page 52 / 125 47 / 111 of 'ButtonMenu' blocks 176. Each 'ButtonMenu' block 176 defines the properties of a button on the screen. The 'ButtonMenu' block can describe things such as the overlay to use when the button is highlighted by the user, the button's name, and what to do in response to various actions performed by a user navigating through the menu. Responses to actions are defined by referencing an 'Action' block 178. A single action, for example, selecting a button, can result in several 'Action' blocks being accessed. In modes where the user is able to interact with the menu using a device, such as a mouse that allows a pointer on the screen to move around the display in an unrestricted way, the on-screen location of the buttons can be defined using a 'MenuRectangle' block 180.Knowing the button's location on the screen allows the system to determine if a user is selecting a button when using a freely accessible input device. [000212] Each 'Action' block identifies one or more of a number of different types of action-related blocks, which may include a 'PlayAction' block 182, a 'MenuTransitionAction' block 184, a 'ReturnToPlayAction' block 186, an 'AudioSelectAction' block 188, a 'SubtitleSelectAction' block 190, and a 'ButtonTransitionAction' block 191. A 'PlayAction' block 182 identifies a portion of each of the video, audio, and subtitle tracks within a multimedia file. The 'PlayAction' block references a portion of the video track using a reference to a 'MediaTrack' block (see discussion below). The 'PlayAction' block identifies the audio and subtitle tracks using the 'SubtitleTrack' 192 and 'AudioTrack' 194 blocks. The 'SubtitleTrack' and 'AudioTrack' blocks both contain references to a 'MediaTrack' 198 block.When a 'PlayAction' block forms the basis of an action according to the embodiments of the present invention, the audio tracks. Petition 870180148460, dated 06 / 11 / 2018, page 53 / 125 48 / 111 and the legend that are selected are determined by the values of variables initially set as default and then potentially modified by user interactions with the menu. [000213] Each 'MenuTransitionAction' block 184 contains a reference to a 'WowMenu' block 175. This reference can be used to obtain the information for transitioning to and displaying another menu. [000214] Each 'ReturnToPlayAction' block 186 contains the information that allows a player to return to a portion of the multimedia file that was being accessed before the user brought up a menu. [000215] Each 'AudioSelectAction' block 188 contains the information that can be used to select a specific audio track. In one embodiment, the audio track is selected from audio tracks contained in a multimedia file according to an embodiment of the present invention. In other embodiments, the audio track may be located in an externally referenced file. [000216] Each 'SubtitleSelectAction' block 190 contains the information that can be used to select a specific subtitle track. In one embodiment, the subtitle track is selected from a subtitle contained in a multimedia file according to an embodiment of the present invention. In other embodiments, the subtitle track may be located in an externally referenced file. [000217] Each 'ButtonTransitionAction' block 191 contains the information that can be used to transition to another button in the same menu. This is executed after other actions associated with a button have been performed. [000218] Block 174 of 'Media' includes a number of 'Me' blocks. Petition 870180148460, dated 06 / 11 / 2018, page 54 / 125 49 / 111 diaSource' 166 and 'MediaTrack' blocks 198. The 'Media' block defines all multimedia tracks (e.g., audio, video, subtitle) used by the movie and the menu system. Each 'MediaSource' block 196 identifies a RIFF block within the multimedia file according to an embodiment of the present invention, which, in turn, may include multiple RIFF blocks. Each 'MediaTrack' block 198 identifies a portion of a multimedia track within a RIFF block specified by a 'MediaSource' block. [000219] The 'MRIF' block 162 is essentially its own small multimedia file conforming to the RIFF format. The 'MRIF' block contains audio, video, and subtitle tracks that can be used to provide background audio and video and overlays for menus. The 'MRIF' block can also contain a video to be used as overlays to indicate highlighted menu buttons. In modes where less menu data is required, the background video can be a still frame (a variation of the AVI format) or a short sequence of identical frames. In other modes, more elaborate video sequences can be used to provide the background video. [000220] As discussed above, the various blocks that make up a 'WowMenu' block 175 and the 'WowMenu' block itself contain references to actual media tracks. Each of these references is typically to a media track defined in the 'hdrl' LIST block of a RIFF block. [000221] Other blocks that can be used to create a 'DMNU' block according to the present invention are shown in Figure 2.6. The 'DMNU' block includes a 'WowMenuManager' block 170. The 'WowMenuManager' block 170 may contain at least one 'LanguageMenus' block 172, at least one 'Media' block 174. Petition 870180148460, dated 06 / 11 / 2018, p. 55 / 125 50 / 111 and at least one 'TranslationTable' block 200. [000222] The content of the 'LanguageMenus' block 172 is largely similar to that of the 'LanguageMenus' block 172 illustrated in Figure 2.5. The main difference is that the 'PlayAction' block 182 does not contain the 'SubtitleTrack' blocks 192 and the 'AudioTrack' blocks 194. [000223] The 'Media' block 174 is significantly different from the 'Media' block 174 shown in Figure 2.5. The 'Media' block 174 contains at least one 'Title' block 202 and at least one 'MenuTracks' block 204. The 'Title' block refers to a title within the multimedia file. As discussed above, multimedia files according to embodiments of the present invention may include more than one title (for example, multiple episodes in a television series, a relative series of full-length films, or simply a selection of different films). The 'MenuTracks' block 204 contains information relating to the media information used to create a menu display and the audio track and subtitles that accompany the display. [000224] The 'Title' block may contain at least one 'Chapter' block 206. The 'Chapter' block 206 references a scene within a specific title. The 'Chapter' block 206 contains references to the portions of the video track, each audio track, and each subtitle track that correspond to the scene indicated by the 'Chapter' block. In one embodiment, the references are implemented using 'MediaSource' blocks 196 and 'MediaTrack' blocks 198 similar to those described above in relation to Figure 2.5. In several embodiments, a 'MediaTrack' block references the appropriate portion of the video track, and a number of additional 'MediaTrack' blocks each reference one of the audio tracks or subtitle tracks. In one embodiment, all audio tracks and subtitle tracks that correspond Petition 870180148460, dated 06 / 11 / 2018, p. 56 / 125 51 / 111 responses to a specific video track are referenced using separate 'MediaTrack' blocks. [000225] As described above, the 'MenuTracks' blocks 204 contain references to the media used to generate the audio, video, and overlay media for the menus. In one embodiment, the media information references are made using the 'MediaSource' blocks 196 and the 'MediaTrack' blocks 198 contained within the 'MenuTracks' block. In one embodiment, the 'MediaSource' blocks 196 and the 'MediaTrack' blocks 198 are implemented in the manner described above in relation to Figure 2.5. [000226] The 'TranslationTable' block 200 can be used to contain text strings that describe each title and chapter in a variety of languages. In one embodiment, the 'TranslationTable' block 200 includes a 'TranslationLookup' block 208. Each 'TranslationLookup' block 208 is associated with a 'Title' block 202, a 'Chapter' block 206, or a 'MediaTrack' block 196 and contains a number of 'Translation' blocks 210. Each of the 'Translation' blocks in a 'TranslationLookup' block contains a text string that describes the block associated with the 'TranslationLookup' block in a language indicated by the 'Translation' block. [000227] A diagram conceptually illustrating the relationships between the various blocks contained within a 'DMNU' block is shown in Figure 2.6.1. The figure shows the containment of one block by another block using a solid arrow. The direction in which the arrow points indicates the block contained by the block from which the arrow originates. References from one block to another block are indicated by a dashed line, where the referenced block is indicated by the dashed arrow. 2.6. The 'junk' block [000228] The 'junk' block 41 is an optional block that can be included in multimedia files according to the pre-programming guidelines. Petition 870180148460, dated 06 / 11 / 2018, p. 57 / 125 52 / 111 feels invention. The nature of the 'junk' block is specified in the AVI file format. 2.7. The 'movi' List Block [000229] The 'movi' list block 42 contains a number of 'data' blocks. Examples of information that 'data' blocks may contain are audio, video, or subtitle data. In one embodiment, the 'movi' list block includes data for at least one video track, multiple audio tracks, and multiple subtitle tracks. [000230] The interleaving of 'data' blocks in the 'movi' list block 42 of a multimedia file containing one video track, three audio tracks, and three subtitle tracks is illustrated in Figure 2.7. For convenience, a 'data' block containing video will be described as a 'video' block, a 'data' block containing audio will be referred to as an 'audio' block, and a 'data' block containing subtitles will be referred to as a 'subtitle' block. In the illustrated 'movi' list block 42, each 'video' block 262 is separated from the next 'video' block by 'audio' blocks 264 from each of the audio tracks.In several formats, the 'audio' blocks contain the portion of the audio track that corresponds to the portion of video contained in the 'video' block following the 'audio' block. [000231] Adjacent 'video' blocks may also be separated by one or more 266 'subtitle' blocks from one of the subtitle tracks. In one embodiment, the 266 'subtitle' block includes a subtitle and a start time and a stop time. In several embodiments, the 'subtitle' block is interleaved within the 'movi' list block so that the 'video' block following the 'subtitle' block includes the portion of video that occurs at the subtitle start time. In other embodiments, the start time of all 'subtitle' and 'audio' blocks is ahead of the equivalent video start time. In one embodiment, the 'audio' and 'subtitle' blocks may Petition 870180148460, dated 06 / 11 / 2018, p. 58 / 125 53 / 111 being placed within 5 seconds of the corresponding 'video' block, and in other modes the 'audio' and 'subtitle' blocks may be placed within a time relative to the amount of video that can be stored by a device capable of displaying audio and video within the file. [000232] In one embodiment, the 'data' blocks include a 'FOURCC' code to identify the stream to which the 'data' block belongs. The 'FOURCC' code consists of a two-digit stream number followed by a two-character code that defines the type of information in the block. (Raider) An alternative 'FOURCC' code consists of a two-character code that defines the type of information in the block followed by the two-digit stream number. Examples of the two-character code are shown in the following table: Two-character code Description db Uncompressed video frame dc Compressed video frame dd Key information for the video frame pc Palette change wb Audio data st Subtitle (text mode) sb Subtitle (bitmap mode) ch Chapter Table 2 - Selected two-character codes used in FOURCC codes [000233] In one embodiment, the structure of the 'video' blocks 262 and the 'audio' blocks 264 conforms to the AVI file format. Other formats for the blocks may be used that specify the nature of the media and contain the encoded media. [000234] In several modalities the data contained in a 'subtitle' block 266 can be represented as follows: typedef struct _subtitlebloco { Petition 870180148460, dated 06 / 11 / 2018, page 59 / 125 54 / 111 FOURCC fcc; DWORD cb; STR duration; STR subtitle; } SUBTITLEBLOCO; [000235] The value of 'fcc' is the FOURCC code that indicates the legend track and the nature of the legend track (text mode or bitmap). The value of 'cb' specifies the size of the structure. The value of 'duration' specifies the time at the start and end points of the legend. In one embodiment, this may be in the form hh:mm:ss.xxxhh:mm:ss.xxx. hh represents hours, mm minutes, ss seconds, and xxx milliseconds. The value of 'subtitle' contains either the Unicode text of the legend in text mode or a bitmap image of the legend in bitmap mode. Several embodiments of the present invention use compressed bitmap images to represent legend information. In one embodiment, the 'subtitle' field contains information regarding the width, height, and screen position of the legend. In addition, the 'subtitle' field may also contain color information and the actual pixels of the bitmap.In several applications, an operation-length encoding is used to reduce the amount of pixel information required to represent the bitmap. [000236] Multimedia files according to embodiments of the present invention may include digital rights management. This information can be used in video-on-demand applications. Multimedia files that are protected by digital rights management can only be played correctly on a player to which the specific reproduction right has been granted. In one embodiment, the fact that a track is protected by digital rights management may be indicative Petition 870180148460, dated 06 / 11 / 2018, page 60 / 125 55 / 111 cado in the track information in the 'hdrl' list block (see description above). A multimedia file according to an embodiment of the present invention that includes a track protected by digital rights management may also contain the digital rights management information in the 'movi' list block. [000237] A 'movi' list block of a multimedia file according to an embodiment of the present invention that includes a video track, multiple audio tracks, and at least one subtitle track and information enabling digital rights management is illustrated in Figure 2.8. The 'movi' list block 42' is similar to the 'movi' list block shown in Figure 2.7 with the addition of a 'DRM' block 270 before each video block 262'. The 'DRM' blocks 270 are 'data' blocks containing digital rights management information, which can be identified by a 'nndd' code FOURCC. The first two characters 'nn' refer to the track number and the second two characters are 'dd' to signify that the block contains digital rights management information. In one embodiment, the 'DRM' block 270 provides the digital rights management information for the 'video' block 262 following the 'DRM' block.A device attempting to play a video track protected by digital rights management uses the information in the 'DRM' block to decode the video information in the 'video' block. Typically, the absence of a 'DRM' block before a 'video' block is interpreted as meaning that the 'video' block is unprotected. [000238] In a cryptographic system according to an embodiment of the present invention, video blocks are only partially encrypted. Where partial encryption is used, the 'DRM' blocks contain a reference to the portion of a 'video' block that is encrypted and a reference to the key that can be used. Petition 870180148460, dated 06 / 11 / 2018, page 61 / 125 56 / 111 to decrypt the encrypted portion. The decryption keys may be located in a 'DRM' header, which is part of the 'strd' block (see description above). The decryption keys are mixed and encrypted with a master key. The 'DRM' header also contains information that identifies the master key. [000239] A conceptual representation of the information in a 'DRM' block is shown in Figure 2.9. The 'DRM' block can include a 'frame' value of 280, a 'status' value of 282, an 'offset' value of 284, a 'number' value of 286, and a 'key' value of 288. The 'frame' value can be used to reference the encrypted video frame. The 'status' value can be used to indicate whether the frame is encrypted, the 'offset' value of 284 points to the beginning of the encrypted block within the frame, and the 'number' value of 286 indicates the number of encrypted bytes in the block. The 'key' value of 288 references the decryption key that can be used to decrypt the block. 2.8. The 'idx1' Block [000240] The 'idx1' block 44 is an optional block that can be used to index the 'data' blocks in the 'movi' list block 42. In one embodiment, the 'idx1' block can be implemented as specified in the AVI format. In other embodiments, the 'idx1' block can be implemented using data structures that reference the location within the file of each of the 'data' blocks in the 'movi' list block. In several embodiments, the 'idx1' block identifies each 'data' block by the number of data tracks and the data type. The FOURCC codes mentioned above can be used for this purpose. Encoding a Media File [000241] The embodiments of the present invention can be used to generate multimedia files in a number of ways. In one case, the systems according to the embodiments of the present invention Petition 870180148460, dated 06 / 11 / 2018, p. 62 / 125 57 / 111 tion can generate multimedia files containing separate video tracks, audio tracks, and subtitle tracks. In such cases, other information, such as menu information and metadata, may be authorized and inserted into the file. [000242] Other systems according to embodiments of the present invention can be used to extract information from various files on a Digital Video Disc (DVD) and authorize a single multimedia file according to an embodiment of the present invention. Where a DVD is the initial source of information, systems according to embodiments of the present invention can use a codec to achieve greater compression and can re-block the audio so that the audio blocks correspond to the video blocks in the newly created multimedia file. Furthermore, menu information in the DVD menu system can be analyzed and used to generate the menu information included in the multimedia file. [000243] Other embodiments may generate a new multimedia file by adding additional content to an existing multimedia file according to an embodiment of the present invention. An example of adding additional content would be adding an additional audio track to the file, such as an audio track containing commentary (e.g., director's commentary, the subsequently created narration of a vacation video). The additional audio track information interspersed in the multimedia file could also be accompanied by a modification of the menu information in the multimedia file to allow playback of the new audio track. 3.1. Generation Using Stored Data Tracks [000244] A system according to an embodiment of the present invention for generating a multimedia file is illustrated in Figure Petition 870180148460, dated 06 / 11 / 2018, page 63 / 125 58 / 111 3.0. [000245] The main component of the system 350 is the interleaver 352. The interleaver receives the information blocks and interleaves them to create a multimedia file according to an embodiment of the present invention in the format described above. The interleaver also receives metadata information from a metadata manager 354. The interleaver outputs a multimedia file according to the embodiments of the present invention to a storage device 356. [000246] Typically, the blocks provided to the interleaver are stored on a storage device. In several embodiments, all blocks are stored on the same storage device. In other embodiments, blocks can be provided to the interleaver from a variety of storage devices or generated and provided to the interleaver in real time. [000247] In the embodiment illustrated in Figure 3.0, the 'DMNU' block 358 and the 'DXDT' block 360 have already been generated and are stored on storage devices. The video source 362 is stored on a storage device and is decoded using a video decoder 364 and then encoded using a video encoder 366 to generate a 'video' block. The audio sources 368 are also stored on a storage device. The audio blocks are generated by decoding the audio source using an audio decoder 370 and then encoding the decoded audio using an audio encoder 372. The 'Subtitle' blocks are generated from text captions 374 stored on a storage device. The captions are provided to a first transcoder 376 which converts any of a number of caption formats into a raw bitmap format. In one embodiment, the stored caption format can be a Petition 870180148460, dated 06 / 11 / 2018, p. 64 / 125 59 / 111 format, such as SRT, SUB, or SSA. Additionally, the bitmap format can be that of a four-bit bitmap that includes a color palette lookup table. The color palette lookup table includes a 24-bit color depth identifier for each of the sixteen possible four-bit color codes. A single multimedia file can include more than one color palette lookup table (see the FOURCC pc palette code in Table 2 above). The four-bit bitmap thus allows each menu to have 16 different simultaneous colors taken from a palette of 16 million colors. In alternative embodiments, different numbers of bits per pixel and different color depths are used. The output of the first 376 transcoder is provided to a second 378 transcoder, which compresses the bitmap. In one embodiment, operation length encoding is used to compress the bitmap.In other applications, other suitable compression formats are used. [000248] In one embodiment, the interfaces between the various encoders, decoders, and transcoders conform to the Direct Show standards specified by Microsoft Corporation. In other embodiments, the software used to perform encoding, decoding, and transcoding does not need to conform to such standards. [000249] In the embodiment shown, separate processing components are shown for each media source. In other embodiments, resources may be shared. For example, a single audio decoder and audio encoder could be used to generate the audio blocks from all sources. Typically, the entire system can be implemented on a computer using software and connected to a storage device, such as a hard disk drive. Petition 870180148460, dated 06 / 11 / 2018, page 65 / 125 60 / 111 [000250] In order to use the interleaver in the mode described above, the 'DMNU' block, the 'DXDT' block, the 'video' blocks, the 'audio' blocks and the 'subtitle' blocks according to the embodiments of the present invention must be generated and provided to the interleaver. The process for generating each of the various blocks in a multimedia file according to an embodiment of the present invention is discussed below in greater detail. 3.2. Generating a 'DXDT' Block [000251] The 'DXDT' block can be generated in one of a number of ways. In one mode, metadata are inserted into data structures through a graphical user interface and then parsed into a 'DXDT' block. In another mode, metadata are expressed as subject, predicate, object, and authority series sentences. In another mode, metadata sentences are expressed in any of a variety of formats. In several modes, each metadata sentence is parsed in a separate block. In other modes, several metadata sentences in a first format (such as subject, predicate, object, authority expressions) are parsed in a first block, and other metadata sentences in other formats are parsed in separate blocks. In one mode, metadata sentences are written in an XML configuration file, and the XML configuration file is parsed to create the blocks within a 'DXDT' block. [000252] One embodiment of a system for generating a 'DXDT' block from a series of 'metadata' sentences contained in an XML configuration file is shown in FIG. 3.1. The system 380 includes an XML configuration file 382, which can be provided to a parser 384. The XML configuration file includes the 'metadata' encoded as XML. The parser parses the XML and Petition 870180148460, dated 06 / 11 / 2018, page 66 / 125 61 / 111 generates a 386-bit 'DXDT' block by converting the 'metadata' sentence into blocks that are written to the 'DXDT' block according to any of the 'metadata' block formats described above. 3.3. Generating a 'DMNU' Block [000253] A system that can be used to generate a 'DMNU' block according to an embodiment of the present invention is illustrated in Figure 3.2. The menu block generation system 420 requires as inputs a media model 422 and media information. The media information can take the form of a video source 424, an audio source 426 and an overlay source 428. [000254] Generating a 'DMNU' block using the entries for the menu block generation system involves creating a number of intermediate files. Media model 422 is used to create an XML configuration file 430, and the media information is used to create a number of AVI files 432. The XML configuration file is created by a model 434 transcoder. The AVI files 432 are created by interleaving the video, audio, and overlay information using an interleaver 436. The video information is obtained by using a video decoder 438 and a video encoder 440 to decode the video source 424 and re-encode it in the manner discussed below. The audio information is obtained by using an audio decoder 442 and an audio encoder 444 to decode the audio and re-encode it in the manner described below.The overlay information is generated using a first 446 transcoder and a second 448 transcoder. The first 446 transcoder converts the overlay into a graphical representation, such as a standard bitmap, and the second transcoder takes the graphical information and formats it as required for inclusion in the multimedia file. Once the XML file and the AVI files containing the information are complete... Petition 870180148460, dated 06 / 11 / 2018, page 67 / 125 62 / 111 of the information required to build the menus has been generated; menu generator 450 can use the information to generate a 'DMNU' block 358. 3.3.1. The Menu Model [000255] In one embodiment, the media model is an object-oriented model that represents all menus and their subcomponents. The media model organizes the menus in a hierarchical structure, which allows the menus to be organized by language selection. A media model according to one embodiment of the present invention is illustrated in Figure 3.3. The media model 460 includes a top-level 'MediaManager' object 462, which is associated with a number of 'LanguageMenus' objects 463, a 'Media' object 464, and a 'TranslationTable' object 465. The 'Menu Manager' also contains the default menu language. In one embodiment, the default language can be indicated by the two-letter ISO 639 language code. [000256] The 'LanguageMenus' objects organize information for various menus by language selection. All 'Menu' objects 466 for a given language are associated with the 'LanguageMenus' object 463 for that language. Each 'Menu' object is associated with a number of 'Button' objects 468 and references a number of 'MediaTrack' objects 488. The referenced 'MediaTrack' objects 488 indicate the background video and background audio for the 'Menu' object 466. [000257] Each 'Button' object 468 is associated with an 'Action' object 470 and a 'Rectangle' object 484. The 'Button' object 468 also contains a reference to a 'MediaTrack' object 488 that indicates the overlay to be used when the button is highlighted on a display. Each 'Action' object 470 is associated with a number of objects that may include a 'MenuTransition' object 472, Petition 870180148460, dated 06 / 11 / 2018, page 68 / 125 63 / 111 a 'ButtonTransition' object 474, a 'ReturnToPlay' object 476, a 'Subtitle Selection' object 478, an 'AudioSelection' object 480, and a 'PlayAction' object 482. Each of these objects defines the menu system's response to various user inputs. The 'MenuTransition' object contains a reference to a 'Menu' object indicating a menu that should transition in response to an action. The 'ButtonTransition' object indicates a button that should be highlighted in response to an action. The 'ReturnToPlay' object can cause a player to continue playing a movie. The 'SubtitleSelection' and 'AudioSelection' objects contain references to 'Title' objects 487 (discussed below). The 'PlayAction' object contains a reference to a 'Chapter' object 492 (discussed below). The 'Rectangle' object 484 indicates the portion of the screen occupied by the button. [000258] The 'Media' object 464 indicates the media information referenced by the menu system. The 'Media' object has a 'MenuTracks' object 486 and a number of 'Title' objects 487 associated with it. The 'MenuTracks' object 486 references the 'MediaTrack' objects 488 which are indicative of the media used to construct the menu (i.e., background audio, background video, and overlays). [000259] The 'Title' objects 487 are indicative of a multimedia presentation and have a number of 'Chapter' objects 492 and 'MediaSource' objects 490 associated with them. 'Title' objects also contain a reference to a 'TranslationLookup' object 494. 'Chapter' objects indicate a certain point in a multimedia presentation and have a number of 'MediaTrack' objects 488 associated with them. 'Chapter' objects also contain a reference to a 'TranslationLookup' object 494.Each 'MediaTrack' object associated with a 'Chapter' object indicates a point in any audio, video, or caption track of the multimedia presentation and references it. Petition 870180148460, dated 06 / 11 / 2018, page 69 / 125 64 / 111 'MediaSource' object 490 and a 'TranslationLookup' object 494 (discussed below). [000260] The 'TranslationTable' object 465 groups a number of text strings that describe the various parts of multimedia presentations indicated by the 'Title' objects, the 'Chapter' objects, and the 'MediaTrack' objects. The 'TranslationTable' object 465 has a number of 'TranslationLookup' objects 494 associated with it. Each 'TranslationLookup' object is indicative of a specific object and has a number of 'Translation' objects 496 associated with it. The 'Translation' objects are each indicative of a text string that describes the object indicated by the 'TranslationLookup' object in a specific language. [000261] A media object model can be constructed using software configured to generate the various objects described above and establish the required associations and references between the objects. 3.3.2. Generating an XML File [000262] An XML configuration file is generated from the menu model, which represents all menus and their subcomponents. The XML configuration file also identifies all media files used by the menus. The XML can be generated by implementing an appropriate parser application that parses the object model into XML code. [000263] In other modes, a video editing application may provide a user with a user interface that allows the direct generation of an XML configuration file without creating a menu template. [000264] In modes where another menu system is the basis of the menu model, such as a DVD menu, the menus can be trimmed by the user to eliminate menu options relating to a Petition 870180148460, dated 06 / 11 / 2018, page 70 / 125 65 / 111 content not included in the multimedia file generated according to the practice of the present invention. In one embodiment, this can be done by providing a graphical user interface that allows the elimination of objects from the menu template. In another embodiment, menu pruning can be achieved by providing a graphical user interface or a text interface that can edit the XML configuration file. 3.3.3. Media Information [000265] When the 'DMNU' block is generated, the media information provided to the 450 menu generator includes the data required to provide the background video, background audio, and foreground overlays for the buttons specified in the menu template (see description above). In one embodiment, a video editing application, such as VideoWave distributed by Roxio, Inc. of Santa Clara, CA, is used to provide the source media tracks that represent the video, audio, and button selection overlays for each individual menu. 3.3.4. Generating Intermediate AVI Files [000266] As discussed above, the media tracks used in background video, background audio, and foreground button overlays are stored in a single AVI file for one or more menus. The blocks containing the media tracks in a menu AVI file can be created using software designed to interleave video, audio, and button overlay tracks. The 'audio', 'video', and 'overlay' blocks (i.e., the 'subtitle' blocks containing the overlay information) are interleaved into an AVI-compliant file using an interleaver. [000267] As mentioned above, a separate AVI file can be created for each menu. In other modes, other file formats or a single file could be used to contain the Petition 870180148460, dated 06 / 11 / 2018, page 71 / 125 66 / 111 media information used to provide background audio, background video, and foreground overlay information. 3.3.5. Combining the XML Configuration File and AVI Files [000268] In one embodiment, a computer is configured to parse the information in the XML configuration file to create a 'WowMenu' block (described above). Additionally, the computer can create an 'MRIF' block (described above) using the AVI files containing the media for each menu. The computer can then complete the generation of the 'DMNU' block by creating the necessary references between the 'WowMenu' block and the media blocks in the 'MRIF' block. In several embodiments, the menu information can be encrypted. Encryption can be achieved by encrypting the media information contained in the 'MRIF' block in a manner similar to that described below for the 'video' blocks. In other embodiments, various alternative encryption techniques are used. 3.3.6. Automatic Generation of Object Model Menus [000269] Referring back to Figure 3.3, a menu containing less content than the full menu can be automatically generated from the menu template simply by examining the 'Title' objects 487 associated with the 'Media' object 464. The objects used to automatically generate a menu according to an embodiment of the invention are shown in Figure 3.3.1. Software can generate an XML configuration file for a simple menu that simply allows the selection of a specific section of a multimedia presentation and the selection of the audio and subtitle tracks to be used. 3.3.7. Generating 'DXDT' and 'DMNU' Blocks Using a Single Configuration File [000270] Systems according to various embodiments of the present invention are capable of generating a single configuration file. Petition 870180148460, dated 06 / 11 / 2018, page 72 / 125 67 / 111 XML files that contain both metadata and menu information, and which use the XML file to generate DXDT and DMNU blocks. These systems derive the XML configuration file using the metadata information and the menu object model. In other modes, the configuration file does not need to be in XML. 3.4. Generating 'audio' blocks [000271] The 'audio' blocks in the 'movi' list block of multimedia files according to the embodiments of the present invention can be generated by decoding an audio source and then encoding the source into 'audio' blocks according to the practice of the present invention. In one embodiment, the 'audio' blocks can be encoded using an mp3 codec. 3.4.1. Audio Re-block [000272] When the audio source is provided in blocks that do not contain the audio information corresponding to the content of a corresponding 'video' block, then embodiments of the present invention can perform audio re-blocking. A process that can be used to perform audio re-blocking is illustrated in Figure 3.4. The process 480 involves identifying (482) a 'video' block, identifying (484) the audio information that accompanies the 'video' block and extracting (486) the audio information from the existing audio blocks to create (488) a new 'audio' block. The process is repeated until the decision (490) is made that the entire audio source has been re-blocked. At which point, the audio re-blocking execution is complete (492). 3.5. Generating the 'video' blocks [000273] As described above, the process of creating video blocks may involve decoding the video source and encoding the decoded video into 'video' blocks. In one embodiment, each 'video' block contains the information for a single video frame. The pro Petition 870180148460, dated 06 / 11 / 2018, page 73 / 125 68 / 111 The decoding process simply involves taking the video in a specific format and decoding the video from that format into a standard video format, which may be uncompressed. The encoding process involves taking the standard video, encoding the video, and generating 'video' blocks using the encoded video. [000274] A video encoder according to an embodiment of the present invention is conceptually illustrated in Figure 3.5. The video encoder 500 preprocesses 502 the standard video information 504. A motion estimate 506 is then performed on the preprocessed video to provide a motion compensation 508 for the preprocessed video. A discrete cosine transform (DCT transform) 510 is performed on the motion-compensated video. After the DCT transform, the video is quantized 512 and a prediction 514 is performed. A compressed bitstream 516 is then generated by combining a texture-encoded version 518 of the video with a motion encoding 520 generated using the results of the motion estimate. The compressed bitstream is then used to generate the 'video' blocks. [000275] In order to perform motion estimation 506, the system must know how the previously processed video frame will be decoded by a decoding device (for example, when the video is compressed and decompressed to be viewed by a spectator). This information can be obtained by inverse quantization 522 of the output of the quantizer 512. An inverse DCT 524 can then be performed on the output of the inverse quantizer and the result placed in a frame store 526 for access during the motion estimation process. [000276] The multimedia files according to the embodiments of the present invention may also include a number of psychovisual enhancements 528. The psychovisual enhancements may be Petition 870180148460, dated 06 / 11 / 2018, page 74 / 125 69 / 111 video compression methods based on human perceptions of vision. These techniques are further discussed below and generally involve modifying the number of bits used by the quantizer to represent various aspects of video. Other aspects of the encoding process may also include psychovisual enhancements. [000277] In one embodiment, the entire 500 coding system can be implemented using a computer configured to perform the various functions described above. Examples of detailed implementations of these functions are provided below. 3.5.1. Pre-processing [000278] The 502 preprocessing operations that are optionally performed by an encoder 500 according to an embodiment of the present invention may utilize a number of signal processing techniques to improve the quality of the encoded video. In one embodiment, the 502 preprocessing may involve one or all of deinterlacing, temporal / spatial noise reduction, and resizing. In embodiments where all three of these preprocessing techniques are used, deinterlacing is typically performed first followed by temporal / spatial noise reduction and resizing. 3.5.2. Motion Estimation and Compensation [000279] A video encoder according to an embodiment of the present invention can reduce the number of pixels required to represent a video track by searching for pixels that are repeated in multiple frames. Essentially, each frame in a video typically contains many of the same pixels as the one before it. The encoder can conduct various types of searches for pixel matches between each frame (such as macroblocks, pixels, half-pixels, and quarter-pixels) and eliminates these redundancies without Petition 870180148460, dated 06 / 11 / 2018, page 75 / 125 70 / 111 as much as possible without reducing image quality. Using motion estimation, the encoder can represent most of the image simply by recording the changes that will occur from the last frame instead of storing the entire image for each frame. During motion estimation, the encoder divides the frame it is analyzing into a grid of uniform blocks, often referred to as 'macroblocks'. For each 'macroblock' in the frame, the encoder can try to find a matching block in the previous frame. The process of trying to find matching blocks is called a 'motion search'. The motion of the 'macroblock' can be represented as a two-dimensional vector, i.e., an (x, y) representation. The motion search algorithm can be run with varying degrees of accuracy.A full image element search is one where the coder will attempt to locate matching blocks by walking through the reference frame in any dimension one pixel at a time. In a half-pixel search, the coder searches for a matching block by walking through the reference frame in any dimension by half a pixel at a time. The coder may use a quarter pixel, or other pixel fractions, or searches that involve a granularity greater than one pixel. [000280] The encoder embodiment illustrated in Figure 3.5 performs motion estimation according to an embodiment of the present invention. During motion estimation, the encoder has access to the pre-processed video 502 of the previous frame, which is stored in a frame store 526. The previous frame is generated by taking the output of the quantizer, performing an inverse quantization 522 and an inverse DCT transformation 524. The reason for performing the inverse functions is such that the frame in the frame store is as it will appear when decoded by Petition 870180148460, dated 06 / 11 / 2018, page 76 / 125 71 / 111 a reproducer according to an embodiment of the present invention. [000281] Motion compensation is performed by taking the blocks and vectors generated as a result of motion estimation. The result is an approximation of the encoded image that can be matched to the real image by providing additional texture information. 3.5.3. Discrete Cosine Transform [000282] The DCT and inverse DCT performed by the encoder illustrated in Figure 3.5 are in accordance with the standard specified in ISO / IEC 14496-2:2001(E), Annex A.1 (encoding transforms). 3.5.3.1. Transform Description [000283] DCT is a method for transforming a set of spatial domain data points into a frequency domain representation. In the case of video compression, a two-dimensional DCT converts image blocks into a form where redundancies are more readily exploitable. A frequency domain block can be a sparse matrix that is easily compressed by entropy coding. 3.5.3.2. Psychovisual Enhancements in the Transformed [000284] DCT coefficients can be modified to improve quantized image quality by reducing quantization noise in areas where it is readily apparent to a human observer. Additionally, file size can be reduced by increasing quantization noise in portions of the image where it is not readily discernible by a human observer. [000285] Encoders according to an embodiment of the present invention can perform what is referred to as 'slow' psychovisual enhancement. 'Slow' psychovisual enhancement analyzes video image blocks and decides whether allowing some noise there can save some bits without degrading the appearance of the video. The process Petition 870180148460, dated 06 / 11 / 2018, p. 77 / 125 72 / 111 only uses a metric per block. The process is referred to as a 'slow' process because it performs a considerable amount of computation to avoid blocking or ring artifacts. [000286] Other embodiments of encoders according to the embodiments of the present invention implement a 'fast' psychovisual enhancement. The 'fast' psychovisual enhancement is able to control where noise appears within a block and can model the quantization noise. [000287] Both 'slow' and 'fast' psychovisual enhancements are discussed in greater detail below. Other psychovisual enhancements may be implemented according to embodiments of the present invention, including enhancements that control noise and image edges and that seek to concentrate higher levels of quantization noise in areas of the image where it is not readily apparent to human vision. 3.5.3.3. Slow Psychovisual Improvement [000288] Slow psychovisual enhancement analyzes video image blocks and determines whether allowing some noise can save bits without degrading the video's appearance. In one embodiment, the algorithm includes two stages. The first involves generating a differentiated image for the input luminance pixels. The differentiated image is generated in the mode described below. The second stage involves modifying the DCT coefficients before quantization. 3.5.3.3.1. Differentiated Image Generation [000289] Each pixel p xy of the differentiated image is computed from the uncompressed pixel source, pxy, according to the following: p'xy = max(px +1 y - pxy |,|Px-1 y — pxy |,|pxy +1 — pxy |,|pxy-1 — pxy |) where p xy will be within the range of 0 to 255 (assuming an 8-bit video). Petition 870180148460, dated 06 / 11 / 2018, page 78 / 125 73 / 111 3.5.3.3.2. Modification of DCT Coefficients [000290] Modifying DCT coefficients may involve computing a block ring factor, computing block energy, and actually modifying the coefficient values. 3.5.3.3.3. Block Ring Factor Computation [000291] For each image block, a ring factor is calculated based on the differentiated local image region. In modes where the block is defined as an 8 x 8 block, the ring factor can be determined using the following method. [000292] Initially, a threshold is determined based on the maximum and minimum luminance pixel values within the 8 x 8 block: threshold,, . = floor ((maxw, - min,, .) / 8)+ 2 block d W block block / / [000293] The differentiated image and boundary are used to generate a map of smooth pixels in the block's neighborhood. The potential for each block to have a different boundary prevents the creation of a smooth pixel map for the entire frame. The map is generated as follows: flatxy = 1When P'xy <thresholdblock flatxy = 0 otherwise [000294] The smooth pixel map is filtered according to a simple logical operation: flat'xy= 1 when flat = 1 and flatx,y= 1 and flat , = 1 and flatx,y, = 1 flat'xyotherwise [000295] The smooth pixels in the filtered map are then counted over the 9 x 9 region that covers the 8 x 8 block. flatcountblock= Σ flat'xyfor 0 = x = 8 and 0 = y = 8 [000296] The risk of visible ring artifacts can be assessed using the following expression: ringingbriski,iock=((flatccuntblockk-10 )x 256 + 20) / 40 [000297] The ring factor of 8 x 8 blocks can be derived using Petition 870180148460, dated 06 / 11 / 2018, p. 79 / 125 74 / 111 using the following expression: Ringingfactor = 0 when ringingrisk > 255 = 255 when ringingrisk < 0 = 255 - ringingrisk otherwise 3.5.3.3.4. Block Energy Computation [000298] The energy for the image blocks can be calculated using the following procedure. In several modalities, 8 x 8 blocks of the image are used. [000299] A direct DCT is performed on the source image: T = fDCT(S) where S are the 64 luminance values of the source image of the 8 x 8 block in question and T is the transformed version of the same portion of the source image. [000300] The energy at a specific coefficient position is defined as the square of that coefficient value: ek= tk2for 0 = k = 63 where tk is the ko coefficient of the transformed block T. 3.5.3.3.5. Coefficient Modification [000301] The modification of DCT coefficients can be performed according to the following process. In several modalities, the process is performed for each non-zero AC DCT coefficient before quantification. The magnitude of each coefficient is changed by a small delta, the value of delta being determined according to psychovisual techniques; [000302] The DCT coefficient modification of each non-zero AC coefficient ck is performed by calculating an energy based on local and block energies using the following formula: energy = max (ak x ek, 0.12 x total energy) where ak is a constant whose value depends on the position of the coefficient as described in the following table: Petition 870180148460, dated 06 / 11 / 2018, page 80 / 125 75 / 111 0.0 1.0 1.5 2.0 2.0 2.0 2.0 2.0 1.0 1.5 2.0 2.0 2.0 2.0 2.0 2.0 1.5 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 2.0 Table 3 - Coefficient table [000303] The energy can be modified according to the block ring factor using the following relationship: energy'k= ringingfactor x energyk [000304] The resulting value is shifted and clipped before being used as an entry for a lookup table (LUT). ek= min (1023.4 x energy'k) dk= LUTi where i = ek [000305] The lookup table is computed as follows: LUTi = min (floor(kte3turex((i+0-5) / 4) 12+kflatx offset]2 xQp) [000306] The value 'displacement' depends on the quantifier, Qp, as described in the table. Qp Displacement Qp Displacement 1 -0.5 16 8.5 2 1.5 17 7.5 3 1.0 18 9.5 4 2.5 19 8.5 5 1.5 20 10.5 6 3.5 21 9.5 7 2.5 22 11.5 8 4.5 23 10.5 Petition 870180148460, dated 06 / 11 / 2018, page 81 / 125 76 / 111 Qp Displacement Qp Displacement 9 3.5 24 12.5 10 5.5 25 11.5 11 4.5 26 13.5 12 6.5 27 12.5 13 5.5 28 14.5 14 7.5 29 13.5 15 6.5 30 15.5 31 14.5 Table 4 - Displacement as a function of Qp values [000307] The variables ktexture and kflat control the intensity of the psychovisual effect in smooth and textured regions, respectively. In a modality, these take values in the range of 0 to 1, with 0 meaning no effect and 1 meaning full effect. In a modality, the values for ktexture and kflat are set as follows: [000308] Luminance: k =10 ^texture1·νkflcit11.0 [000309] Chrominance: k = 10 texture = .kflat = 0.0 [000310] The lookup table output (dk) is used to modify the magnitude of the DCT coefficient by an additive process: c' k = ck -min (dk, \ck |)x sgn(ck) [000311] Finally, the DCT coefficient ck is replaced by the modified coefficient c'k and passed on for quantification. 3.5.3.4. 'Rapid' Psychovisual Improvement [000312] A 'quick' psychovisual enhancement can be performed on DCT coefficients by computing an 'importance' map for the input luminance pixels and then modifying the Petition 870180148460, dated 06 / 11 / 2018, page 82 / 125 77 / 111 DCT coefficients. 3.5.3.4.1. Computing an 'importance' Map [000313] An 'importance' map can be generated by calculating an 'importance' value for each pixel in place of the luminance of the input video frame. In many ways, the 'importance' value approximates the sensitivity of the human eye to any distortion located at that specific pixel. The 'importance' map is a network of pixel 'importance' values. [000314] The 'importance' of a pixel can be determined first by calculating the dynamic range of a block of pixels surrounding the pixel (dxy). In several embodiments the dynamic range of a 3 x 3 block of pixels centered on the pixel location (x, y) is computed by subtracting the value of the darkest pixel in the area from the value of the lightest pixel in the area. [000315] The 'importance' of a pixel (mxy) can be derived from the pixel's dynamic range as follows: m^ = 0.08 / max(d,3)+ 0.001 xy x xy 3.5.3.4.2. Modifying dct Coefficients [000316] In one embodiment, the modification of DCT coefficients involves the generation of basis function energy matrices and delta lookup tables. 3.5.3.4.3. Basis Function Energy Matrix Generation [000317] A set of basis function energy matrices can be used in modifying DCT coefficients. These matrices contain constant values that can be computed before encoding. An 8 x 8 matrix is used for each of the 64 DCT basis functions. Each matrix describes how each pixel in an 8 x 8 block will be impacted by modifying its corresponding coefficient. The basis function energy matrix is derived by taking an 8 x 8 matrix Ak with the corresponding coefficient. Petition 870180148460, dated 06 / 11 / 2018, page 83 / 125 78 / 111 adjusted to 100 and the other coefficients adjusted to 0. akn = 100 if n = k = 0 otherwise where n represents the position of the coefficient within the 8 x 8 matrix; 0 = n = 63. [000318] An inverse DCT is performed on the matrix to generate an additional 8 x 8 matrix A'k. The matrix elements (a'kn) represent the ka basis function of DCT. A\. = iDCT (Ak) [000319] Each value in the transformed matrix is then squared: bn =a'kn2for0=n= 63 [000320] The process is executed 64 times to produce the basis function energy matrices Bk, 0 = k = 63, each comprising 64 natural values. Each matrix value is a measure of how much a pixel at that position in the 8 x 8 block will be impacted by any error or modification of the coefficient k. 3.5.3.4.4. Generating a Delta Lookup Table [000321] A lookup table (LUT) can be used to accelerate the computation of the coefficient modification delta. The table content can be generated in a mode that is dependent on the desired strength of the 'fast' psychovisual enhancement and the quantification parameter (Qp). [000322] The values in the lookup table can be generated according to the following relationship: LUTt= min(floor(l28XXstrength / (i + 0.5)+ f Xoffset + 0.5),2XQp) where i is the position within the table, 0 = i = 1023. [000323] Intensity and displacement depend on quantification Petition 870180148460, dated 06 / 11 / 2018, page 84 / 125 79 / 111 pain, Qp, as described in the following table: Qp intensity displacement Qp intensity displacement 1 0.2 -0.5 16 2.0 8.5 2 0.6 1.5 17 2.0 7.5 3 1.0 1.0 18 2.0 9.5 4 1.2 2.5 19 2.0 8.5 5 1.3 1.5 20 2.0 10.5 6 1.4 3.5 21 2.0 9.5 7 1.6 2.5 22 2.0 11.5 8 1.8 4.5 23 2.0 10.5 9 2.0 3.5 24 2.0 12.5 10 2.0 5.5 25 2.0 11.5 11 2.0 4.5 26 2.0 13.5 12 2.0 6.5 27 2.0 12.5 13 2.0 5.5 28 2.0 14.5 14 2.0 7.5 29 2.0 13.5 15 2.0 6.5 30 2.0 15.5 31 2.0 14.5 Table 5 - Relationship between intensity and displacement values and the Qp value [000324] ktexture and kflat control the intensity of the psychovisual effect in smooth and textured regions respectively. In a modality, these take values in the range of 0 to 1, with 0 meaning no effect and 1 meaning full effect. In a modality, the values for ktexture and kflat are set as follows: [000325] Luminance: k =10 ^texture1·νkflcit11.0 [000326] Chrominance: k =10 texture= .f = 0.0 3.5.3.4.5. Modification of DCT Coefficients [000327] DCT coefficients can be modified using Petition 870180148460, dated 06 / 11 / 2018, page 85 / 125 80 / 111 the values calculated above. In one modality, each non-zero AC DCT coefficient is modified according to the following procedure before quantification. [000328] Initially, an 'energy' (ek) value is computed by taking the dot product of the basis function energy matrix and the appropriate 8 x 8 block of the importance map. This 'energy' is a measure of how quantization errors in the specific coefficient would be perceived by the human observer. This is the sum of the pixel importance product and the pixel basis function energy: ek = M · Bk where M contains the importance map values of the 8 x 8 block; and Bk is the basis function energy kamatriz. [000329] The resulting 'energy' value is shifted and clipped before being used as an index (dk) in the delta lookup table. e'k= min [1023, floor(ek / 32768)] dk= LUTi where i = e'k [000330] The output of the delta lookup table is used to modify the magnitude of the DCT coefficient by an additive process: c'k = ck -min (dk,\ckI )χ sign(Ck) [000331] The DCT coefficient ck is replaced by the modified c'k and passed on for quantification. 3.5.4. Quantification [000332] Encoders according to the embodiments of the present invention may use a standard quantifier, such as the quantifier defined by the International Telecommunication Union as Video Coding for Low Bitrate Communication, ITU-T Recommendation H.263, Petition 870180148460, dated 06 / 11 / 2018, page 86 / 125 81 / 111 1996. 3.5.4.1. Psychovisual Enhancements for Quantification [000333] Some encoders according to embodiments of the present invention use a psychovisual enhancement that exploits the psychovisual effects of human vision to achieve more efficient compression. The psychovisual effect can be applied at a frame level and a macroblock level. 3.5.4.2. Psychovisual Improvements at the Framing Level [000334] When applied at a frame level, enhancement is part of the rate control algorithm and its goal is to adjust the encoding so that a given bitrate is best utilized to ensure maximum visual quality as perceived by the human eye. Psychovisual frame rate enhancement is motivated by the theory that human vision tends to ignore details when the action is loud and human vision tends to notice details when the image is static. In one modality, the amount of motion is determined by observing the sum of the absolute differences (SAD) for a frame. In one modality, the SAD value is determined by the sum of the absolute differences of the luminance pixels placed from two blocks. In several modalities, the absolute differences of 16 x 16 pixel blocks are used.In modes that deal with fractional pixel shifts, interpolation is performed as specified in the MPEG-4 standard (an ISO / IEC standard developed by the ISO / IEC Moving Picture Experts Group) before the sum of absolute differences is calculated. [000335] Frame-level psychovisual enhancement applies only to P-frames in the video track and is based on the frame's SAD value. During encoding, the psychovisual module keeps a record of the average SAD (i.e., SAD) of all P-frames in the video track and the average SAD distance of each frame from its SAD. Petition 870180148460, dated 06 / 11 / 2018, p. 87 / 125 82 / 111 total (this is dsad). The average can be calculated using an exponential moving average algorithm. In one embodiment, the one-pass rate control algorithm described above can be used here as the averaging period (see description above). [000336] For each frame P of the encoded video track, the frame quantifier Q (obtained from the rate control module) will have a psychovisual correction applied to it. In one modality, the process involves calculating a ratio R using the following formula: S SAD - SADT R — —--1 DSAD where I is a constant and currently set to 0.5. OR is clipped within the bound of (-1, 1]. [000337] The quantifier is then adjusted according to the ratio R, through the calculation shown below: Qad, — Q Le.(1 + RS,.., )J where Sframe is a constant intensity for psychovisual enhancements at the frame level. [000338] The Sframe constant determines how strong an adjustment can be for the psychovisual at the frame level. In one codec mode, the option to adjust Sframe to 0.2, 0.3, or 0.4 is available. 3.5.4.3. Psychovisual Improvements at the Macroblock Level [000339] Encoders according to embodiments of the present invention that utilize a macroblock-level psychovisual enhancement attempt to identify macroblocks that are prominent for the visual quality of the video for a human observer and attempt to encode these macroblocks with higher quality. The effect of macroblock-level psychovisual enhancements is to remove bits from the less important parts of a frame and apply them to the other parts. Petition 870180148460, dated 06 / 11 / 2018, page 88 / 125 83 / 111 most important aspects of the framework. In several modalities, enhancements are achieved using three technologies, which are based on smoothness, brightness, and macroblock SAD. In other modalities, any of the techniques alone or in combination with another technique, or another technique entirely, may be used. [000340] In one modality, all three macroblock-level psychovisual enhancements described above share a common parameter, SMB, which controls the intensity of the macroblock-level psychovisual enhancement. The maximum and minimum quantifiers for the macroblocks are then derived from the intensity parameter and the Qframe quantifier through the calculations shown below: QMBMaxQframe(1-SMB ) andQMBMin Qr. fajrame '(1 - SMB ) where QMBMax is the maximum quantifier. QMBMin is the minimum quantifier. [000341] The QMBMax and QMBMin values define the upper and lower limits for macroblock quantifiers for the entire frame. In one mode, the option to adjust the SMB value to any of the values 0.2, 0.3, and 0.4 is provided. In other modes, other values for SMB may be used. 3.5.4.3.1. Brightness Enhancement [000342] In modes where psychovisual enhancement is performed based on macroblock brightness, the encoder attempts to encode the brightest macroblocks with higher quality. The theoretical basis of this enhancement is that the relatively dark parts of the frame are more or less ignored by human observers. This macroblock psychovisual enhancement is applied to I-frames and P-frames of the video track. For each frame, the encoder observes Petition 870180148460, dated 06 / 11 / 2018, p. 89 / 125 84 / 111 scan the entire frame first. The average brightness (BR) is calculated and the average brightness difference from the average (dbr) is also calculated. These values are then used to develop two limits (TBRLower, TBRUpper), which can be used as indicators for whether psychovisual enhancement should be applied: T . = BR - DBR TBRUpper = BR +(BR -RBRLower ) [000343] Brightness enhancement is then applied based on the two limits using the conditions presented below to generate a desired quantifier (QMB) for the macroblock: QMB =QMBMin whenBR> TBR. UpperQMB =Qframe whenTBRLower ^BR^TBRUpper, and QMB =QMBMaxwhen BR <TBRLower onde [000344] BR is the brightness value for that specific macroblock. [000345] In modes where the encoder conforms to the MPEG-4 standard, the macroblock-level psychovisual brightness enhancement technique cannot change the quantifier by more than ±2 from one macroblock to the next. Therefore, the calculated Qmb may require modification based on the quantifier used in the previous macroblock. 3.5.4.3.2. Smoothness Enhancement [000346] Encoders according to embodiments of the present invention that include a psychovisual smoothness enhancement modify the quantizer based on the spatial variation of the image being encoded. The use of a psychovisual smoothness enhancement may be motivated by the theory that human vision has increased sensitivity to quantization artifacts in the smooth parts of an image. The psychovisual smoothness enhancement may therefore involve increasing the number of bits to Petition 870180148460, dated 06 / 11 / 2018, pages 90 / 125 85 / 111 represents the smoother portions of the image and reduces the number of bits where there is a high degree of spatial variation in the image. [000347] In one embodiment, the smoothness of a portion of an image is measured as the average difference in pixel luminance in a macroblock to the macroblock brightness (dr). A method for performing psychovisual smoothness enhancement on an I-frame according to embodiments of the present invention is shown in Figure 3.6. The process 540 involves examining the entire frame to calculate (542) dr. The threshold for applying smoothness enhancement, Tdr, can then be derived (544) using the following calculation: t = DrR DR [000348] The following smoothness enhancement is performed (546) based on the threshold. QmB = Qframe when DR > TDR, and QmB = QMBMin whenDR <Tdr onde Qmb is the desired quantifier for the macroblock. DR is the deviation value for the macroblock (that is, average luminance - average brightness). [000349] The methods that encode files according to the MPEG-4 standard are limited, as described above, by the fact that the macroblock-level quantifier shift can be at most ±2 from one macroblock to the next. 3.5.4.3.3. Macroblock Sad Improvement [000350] Encoders according to embodiments of the present invention may use a macroblock SAD psychovisual enhancement. A macroblock SAD psychovisual enhancement may be used to increase detail for static macroblocks and allow for reduced detail in portions of a frame that are used. Petition 870180148460, dated 06 / 11 / 2018, p. 91 / 125 86 / 111 shots in a high-action scene. [000351] A process for performing a macroblock SAD psychovisual enhancement according to an embodiment of the present invention is illustrated in Figure 3.7. The process 570 includes inspecting (572) an entire frame I to determine the average SAD (i.e., MBSAD) for all macroblocks in the entire frame and the average difference of a macroblock SAD from the average (i.e., DMBSAD) is also obtained. In one embodiment, both of these macroblocks are averaged over the inter-frame encoded macroblocks (i.e., macroblocks encoded using motion compensation or other dependencies on previous encoded video frames). Two limits for applying the macroblock SAD enhancement are then derived (574) from these averages using the following formulas: T . = MBSAD - DMBSAD, and T m MBSADUpper = MBSAD + DMBSAD where Tbsadlower is the lower limit. TBSADUpper is the upper limit, which can be limited by 1024 if necessary. [000352] The macroblock SAD enhancement is then applied (576) based on these two limits according to the following conditions: QmB = QmBMox whenMBSAD>TMBSADUpper, QmB =Qframewhen TMADLower -MBSAD^TMBSADUpper QMB = QMBMin when MBSAD <TMBSADLower onde Qb is the desired quantifier for the macroblock. MBSAD is the SAD value for that specific macroblock. Petition 870180148460, dated 06 / 11 / 2018, p. 92 / 125 87 / 111 [000353] The methods that encode files according to the MPEG-4 specification are limited, as described above, by the fact that the macroblock-level quantifier shift can be at most ±2 from one macroblock to the next. 3.5.5. Rate Control [000354] The rate control technique used by an encoder according to an embodiment of the present invention can determine how the encoder uses the allocated bitrate to encode a video sequence. An encoder will typically seek to encode at a predetermined bitrate, and the rate control technique is responsible for matching the bitrate generated by the encoder as closely as possible to the predetermined bitrate. The rate control technique may also seek to allocate the bitrate in a way that will ensure the highest visual quality of the video sequence when it is decoded. Much of the rate control is performed by adjusting the quantizer. The quantizer determines how finely the encoder encodes the video sequence. A smaller quantizer will result in higher quality and higher bit consumption.Therefore, the rate control algorithm seeks to modify the quantizer in a way that balances the competing interests of video quality and bit consumption. [000355] Encoders according to embodiments of the present invention may utilize any of a variety of different rate control techniques. In one embodiment, a single-pass rate control technique is used. In other embodiments, a double (or multiple) pass rate control technique is used. In addition, a 'video storage-verified' rate control may be implemented as required. Specific examples of these techniques are discussed below. However, any rate control technique may be used in an encoder according to the present invention. Petition 870180148460, dated 06 / 11 / 2018, page 93 / 125 88 / 111 with the practice of the present invention. 3.5.5.1. Single-Pass Rate Control [000356] One embodiment of a one-pass rate control technique according to an embodiment of the present invention seeks to enable high bit rate peaks for high-motion scenes. In several embodiments, the one-pass rate control technique seeks to increase the bit rate slowly in response to an increase in the amount of motion in a scene and rapidly decrease the bit rate in response to a reduction in the motion of a scene. [000357] In one embodiment, the one-pass rate control algorithm uses two averaging periods to track the bit rate. A long-term average to ensure total bit rate convergence and a short-term average to allow response to variations in the amount of action in a scene. [000358] A one-pass rate control technique according to an embodiment of the present invention is illustrated in Figure 3.8. The one-pass rate control technique 580 begins (582) by initializing (584) the encoder with a desired bit rate, the video frame rate, and a variety of other parameters (further discussed below). A floating-point variable is stored, which is indicative of the quantizer. If a frame requires quantization (586), then the floating-point variable is retrieved (588) and the quantizer is obtained by rounding the floating-point variable to the nearest integer. The frame is then encoded (590). Observations are made during the encoding of the frame that allow the determination (592) of a new quantizer value. The process decides (594) to repeat unless there are no more frames. At which point, the encoding is complete (596). [000359] As discussed above, the encoder is initialized (584) Petition 870180148460, dated 06 / 11 / 2018, pp. 94 / 125 89 / 111 with a variety of parameters. These parameters are 'bitrate', 'frame rate', 'Max keyframe interval', 'Maximum quantifier', 'Minimum quantifier', 'average period', 'reaction period', and 'fall / rise ratio'. The following is a discussion of each of these parameters. 3.5.5.1.1. The 'Bit Rate' [000360] The 'bit rate' parameter determines the target bit rate for encoding. 3.5.5.1.2. The 'Frame Rate' [000361] The 'frame rate' defines the time between video frames. 3.5.5.1.3. The 'Max Keyframe Range' [000362] The 'Max Keyframe Interval' specifies the maximum interval between keyframes. Keyframes are normally inserted automatically into the encoded video when the codec detects a scene change. In circumstances where a scene continues for a long interval without a single cut, keyframes may be inserted to ensure that the interval between keyframes is always less than or equal to the 'Max Keyframe Interval'. In one mode, the 'Max Keyframe Interval' parameter can be set to a value of 300 frames. In other modes, other values may be used. 3.5.5.1.4. The 'Maximum Quantifier' and the 'Minimum Quantifier' [000363] The 'Maximum Quantifier' and 'Minimum Quantifier' parameters adjust the upper and lower limits of the quantifier used in the encoding. In one mode, the quantifier limits are set to values between 1 and 31. 3.5.5.1.5. The 'Averaging Period' [000364] The 'media period' parameter controls the amount of video that is considered when modifying the quantifier. One Petition 870180148460, dated 06 / 11 / 2018, pp. 95 / 125 A longer 90 / 111 media period will typically result in the encoded video having a more accurate total bitrate. In one mode, an 'average period' of 2000 is used. Although in other modes other values may be used. 3.5.5.1.6. The 'Reaction Period' [000365] The 'reaction period' parameter determines how quickly the encoder adapts to changes in motion in recent scenes. A longer 'reaction period' value can result in higher quality high-motion scenes and lower quality low-motion scenes. In one mode, a 'reaction period' of 10 is used. Although in other modes other values may be used. 3.5.5.1.7. The 'Rate of Descent / Rise' [000366] The 'fall / rise ratio' parameter controls the relative sensitivity for quantizer adjustment in reaction to high or low motion scenes. A higher value typically results in higher quality high-motion scenes and increased bit consumption. In one mode, a 'fall / rise ratio' of 20 is used. Although in other modes, other values may be used. 3.5.5.1.8. Calculating the Quantifier Value [000367] As discussed above, the one-pass rate control technique involves calculating a quantifier value after encoding each frame. The following is a description of a technique according to an embodiment of the present invention that can be used to update the quantifier value. [000368] The encoder maintains two exponential moving averages that have periods equal to the 'average period' (Paverage) and the 'reaction period' (Preaction) as a moving average of the bit rate. The two exponential moving averages can be calculated according to the relation Petition 870180148460, dated 06 / 11 / 2018, p. 96 / 125 91 / 111 action: P - TT At= A.. PT- + B.' t t-1p p where At is the average at instance t; At-1 is the average in the tT instance (usually the average in the previous frame); T represents the interval period (usually the frame time); and P is the averaging period, which can be either P average or Preaction. [000369] The moving average calculated above is then adjusted for bitrate by dividing by the time interval between the current instance and the last instance in the video, using the following calculation: Rt = AtT where Rt is the bitrate; At is any of the moving averages; and T is the time interval between the current instance and the last instance (it is usually the inverse of the frame rate). [000370] The encoder can calculate the target bit rate (Rtarget) of the next frame as follows: R = R ,, + (R ,,-R ) / arg et ^overall \ overall ^average / where Roverall is the total bitrate determined for the entire video; and Raverage is the average bit rate using the long averaging period. [000371] In several modes, the target bit rate is limited below 75% of the total bit rate. If the target bit rate falls Petition 870180148460, dated 06 / 11 / 2018, p. 97 / 125 92 / 111 below that limit, so it will be forced upwards to the limit to ensure video quality. [000372] The encoder then updates the internal quantifier with ba if the difference between Rtarget and Rreaction is equal. If Rreaction is less than Rtarget, then there is a chance that the previous frame was of relatively low complexity. Therefore, the quantifier can be decreased by performing the following calculation: Qint ernal Qint ernal reaction [000373] When Rreaction is greater than Rtarget, there is a significant probability that the previous frame had a relatively high level of complexity. Therefore, the quantifier can be increased by performing the following calculation: Q 'int ernalQint ernal sp .. reaction where S is the 'rate of rise / rate of fall'. 3.5.5.1.9. B-VOP Coding [000374] The algorithm described above can also be applied to B-VOP encoding. When B-VOP is enabled in the encoding, the quantifier for B-VOP (Qb) is chosen based on the quantifier for P-VOP (Qp) after B-VOP. The value can be obtained according to the following relationships: Qb = 2Qp for Qp < 4 3 ~ „ Qb = 5 + -QP for 4 < Qp < 20 Qb = Qp for Qp > 20 3.5.5.2. Two-Pass Rate Control [000375] Encoders according to an embodiment of the present invention that use a two-pass (or multiple-pass) rate control technique can determine the properties of a Petition 870180148460, dated 06 / 11 / 2018, pp. 98 / 125 93 / 111 video sequence in a first pass and then encode the video sequence with knowledge of the properties of the entire sequence. Therefore, the encoder can adjust the quantization level for each frame based on its relative complexity compared to other frames in the video sequence. [000376] In a two-pass rate control technique according to an embodiment of the present invention, the encoder performs a first pass in which the video is encoded according to the one-pass rate control technique described above and the complexity of each frame is recorded (any of a variety of different metrics for averaging complexity can be used). The average complexity and therefore the average quantizer (Qref) can be determined based on the first pass. In the second pass, the bitstream is encoded with the quantizers determined based on the complexity values calculated during the first pass. 3.5.5.2.1. Quantifiers for I-VOPs [000377] The Q quantifier for I-VOPs is set to 0.75 x Qref provided the next frame is not an I-VOP. If the next frame is also an I-VOP, the Q (for the current frame) is set to 1.25 x Qref. 3.5.5.2.2. Quantifiers for P-VOPs [000378] The quantifier for P-VOPs can be determined using the following expression. C, .. / complexity / C / complexity where c, ... ..... Complexity is the complexity of the framework; C ~ . .............. .. Complexity is the average complexity of the video sequence; Petition 870180148460, dated 06 / 11 / 2018, page 99 / 125 94 / 111 F(x) is a function that provides the number by which the frame complexity must be multiplied to give the number of bits required to encode the frame using a quantifier with a quantification value x; F-1(x) is the inverse function of F(x); ek is the intensity parameter. [000379] The following table defines a form of a function F(Q) that can be used to generate the factor by which the complexity of a frame must be multiplied in order to determine the number of bits required to encode a frame using an encoder with a Q quantifier. QF(Q) QF(Q) 1 1 9 0.013 2 0.4 10 0.01 3 0.15 11 0.008 4 0.08 12 0.0065 5 0.05 13 0.005 6 0.032 14 0.0038 7 0.022 15 0.0028 8 0.017 16 0.002 Table 6 - Values of F(Q) with respect to Q. [000380] If the intensity parameter k is chosen to be 0, then the result is a constant quantifier. When the intensity parameter is chosen to be 1, then the quantifier is proportional to ac, .... . . . ...... complexity. Several encoders according to the modalities of the present invention have an intensity parameter k equal to 0.5. 3.5.5.2.3. Quantifiers for B-VOPs [000381] The Q quantifier for B-VOPs can be chosen using the same technique as for choosing the quantifier for B-VOPs in the one-pass technique described above. Petition 870180148460, dated 06 / 11 / 2018, pages 100 / 125 95 / 111 3.5.5.3. Rate Control Verified in Temporary Video Storage [000382] The number of bits required to represent a frame can vary depending on the characteristics of the video sequence. Most communication systems operate at a constant bit rate. One problem that can be encountered with variable bit rate communications is allocating sufficient resources to handle peak resource utilization. Several encoders according to embodiments of the present invention encode video with a view to preventing the overflow of a temporary video storage in the decoder when the bit rate of the variable bit rate communication is at its peak. [000383] The objectives of video buffer rate control (VBV) may include generating video that will not exceed the buffer of a decoder when transmitted. Additionally, it may be desirable for the encoded video to match a target bitrate and for the rate control to produce high-quality video. [000384] The encoders according to various embodiments of the present invention provide a choice of at least two VBV rate control techniques. One of the VBV rate control techniques is referred to as a casual rate control and the other technique is referred to as the Na-pass rate control. 3.5.5.3.1. Casualty Rate Control [000385] Casual VBV rate control can be used in conjunction with the one-pass rate control technique and generates outputs simply based on the current and previous quantifier values. [000386] An encoder according to an embodiment of the present invention includes a random rate control that involves adjusting the Petition 870180148460, dated 06 / 11 / 2018, pages 101 / 125 96 / 111 quantifier for frame n (i.e., Qn) according to the following relation. Q Qn-1 bitrate1velocity1size Qn Qn drift where Q'n is the quantifier estimated by the single-pass rate control; Xbitrate is calculated by determining a target bit rate based on the deviation from the desired bit rate; Xvelocity is calculated based on the estimated time until the temporary VBV storage experiences an overflow or a negative overflow; Xsize is applied to the P-VOPs result only and is calculated based on the rate at which the size of the compressed P-VOPs is changing over time; Xdrift is the deviation from the desired bit rate. [000387] In several modes, casual VBV rate control can be forced to drop frames and insert padding to respect the VBV model. If a compressed frame unexpectedly contains too many or too few bits, then it can be dropped or padded. 3.5.5.3.2. Pass Rate Control [000388] The Na-pass VBV rate control can be used in conjunction with a multi-pass rate control technique and utilizes information gathered during the previous analysis of the video sequence. Encoders according to various embodiments of the present invention perform Na-pass VBV rate control according to the process illustrated in Figure 3.9. The process begins with a first pass, during which a Petition 870180148460, dated 06 / 11 / 2018, pages 102 / 125 97 / 111 analysis is executed (602). A map generation is executed (604) and a strategy is generated (606). The Na-pass rate control is then executed (608). 3.5.5.3.3. Analysis [000389] In one mode, the first pass uses some form of random rate control and data is recorded for each frame relating to things such as frame duration, frame encoding type, quantifier used, motion bits produced and texture bits produced. In addition, global information such as timescale, resolution and codec settings may also be recorded. 3.5.5.3.4. Map Generation [000390] The analysis information is used to generate a map of the video sequence. The map can specify the encoding type used for each frame (I / B / P) and can include data for each frame regarding frame duration, motion complexity, and texture complexity. In other modalities, the map may also contain information that allows for better prediction of the influence of quantized and other parameters on compressed frame size and perceptual distortion. In several modalities, map generation is performed after N-1a is completed. 3.5.5.3.5. Strategy Generation [000391] The map can be used to plan a strategy for how the past rate control will operate. The ideal level of temporary VBV storage after each frame is encoded can be planned. In another embodiment, strategy generation results in information for each frame including the desired compressed frame size, an estimated frame quantifier. In several embodiments, strategy generation is executed after the Petition 870180148460, dated 06 / 11 / 2018, pages 103 / 125 98 / 111 map generation and before the last one. [000392] In one embodiment, the strategy generation process involves using an iterative process to simulate the encoder and determine the desired quantizer values for each frame, attempting to keep the quantizer as close as possible to the average quantizer value. A binary search can be used to generate a base quantizer for the entire video sequence. The base quantizer is the constant value that causes the simulator to achieve the desired target bitrate. Once the base quantizer is found, the strategy generation process involves consideration of VBV constraints. In one embodiment, a constant quantizer is used if this does not modify the VBV constraints. In other embodiments, the quantizer is modulated based on the motion complexity in the video frames. This can be further extended to incorporate scene change masking and other temporal effects. 3.5.5.3.6. Loop Pass Rate Control [000393] In one embodiment, the looped Na rate control utilizes the strategy and map to make the best possible prediction of the influence of the quantifier and other parameters on the compressed frame size and perceptual distortion. There may be a limited criterion for deviating from the strategy to take a short-term corrective strategy. Typically, following the strategy will prevent a violation of the VBV model. In one embodiment, the looped Na rate control utilizes a PID control loop. The return in the control loop is the accumulated deviation from the optimal bitrate. [000394] Although strategy generation does not involve dropping frames, rate control in looped past may require video padding to be inserted to prevent a VBV overflow. Petition 870180148460, dated 06 / 11 / 2018, pages 104 / 125 99 / 111 3.5.6. Predictions [000395] In one embodiment, an AD / DC prediction is performed in a mode that conforms to the standard referred to as ISO / IEC 14496-2:2001(E), section 7.4.3. (DC and AC prediction) and 7.7.1. (field DC and AC prediction). 3.5.7. Texture Coding [000396] An encoder according to an embodiment of the present invention can perform texture encoding in a mode that conforms to the standard referred to as ISO / IEC 14496-2:2001(E), Annex B (variable length codes) and 7.4.1 (variable length decoding). 3.5.8. Motion Coding [000397] An encoder according to an embodiment of the present invention can perform texture encoding in a mode that conforms to the standard referred to as ISO / IEC 14496-2:2001(E), Annex B (variable length codes) and 7.6.3 (motion vector decoding). 3.5.9. Generating 'video' blocks [000398] The video track can be considered as a sequence of frames 1 to N. Systems according to the embodiments of the present invention are capable of encoding the sequence to generate a compressed bitstream. The bitstream is formatted by segmenting it into blocks 1 to N. Each video frame n has a corresponding block n. [000399] The blocks are generated by appending bits from the bitstream to block n until this, together with blocks 1 to n-1, contains sufficient information for a decoder according to an embodiment of the present invention to decode video frame n. In cases where sufficient information is contained in blocks 1 to n-1 to generate video frame n, an encoder according to Petition 870180148460, dated 06 / 11 / 2018, pages 105 / 125 100 / 111 embodiments of the present invention may include a marker block. In one embodiment, the marker block is an uncoded P-frame with timing information identical to that of the previous frame. 3.6. Generating 'subtitle' blocks [000400] An encoder according to an embodiment of the present invention can take legends in one of a number of standard formats and then convert the legends into bitmaps. The information in the bitmaps is then compressed using op-length encoding. The op-length encoded bitmaps are then formatted into a block, which also includes start-time and stop-time information for the specific legend contained in the block. In several embodiments, information regarding the color, size, and position of the legend on the screen may also be included in the block. The blocks may be included in the legend track that determines the palette for the legends and that indicate that the palette has changed. Any application capable of generating a legend in a standard legend format can be used to convert user-entered text directly into legend information. 3.7. Intercalation [000401] Once the interleaver has received all the blocks described above, the interleaver constructs a multimedia file. The construction of the multimedia file may involve creating a 'CSET' block, an 'INFO' list block, an 'hdrl' block, a 'movi' list block, and an idx1 block. The methods according to the embodiments of the present invention for creating these blocks and for generating the multimedia files are described below. 3.7.1. Generating a 'cset' block [000402] As described above, the 'CSET' block is optional and can be generated by the interleaver according to the AVI Container Format. Petition 870180148460, dated 06 / 11 / 2018, pages 106 / 125 101 / 111 Specification. 3.7.2. Generating an 'INFO' list block [000403] As described above, the 'INFO' list block is optional and can be generated by the interleaver according to the AVI Container Format Specification. 3.7.3. Generating the 'hdrl' List Block [000404] The 'hdrl' list block is generated by the merger based on information in the various blocks provided to the merger. The 'hdrl' list block references the location within the file of the referenced blocks. In one embodiment, the 'hdrl' list block uses file offsets to establish the references. 3.7.4. Generating the 'movi' List Block [000405] As described above, the 'movi' list block is created by encoding the audio, video, and subtitle tracks to create 'audio', 'video', and 'subtitle' blocks, and then interleaving these blocks. In several forms, the 'movi' list block may also include digital rights management information. 3.7.4.1. Interleaving video / audio / subtitles [000406] A variety of rules can be used to interleave audio, video, and subtitle blocks. Typically, the interleaver establishes a number of queues for each of the video and audio tracks. The interleaver determines which queue should be written to the output file. Queue selection may be based on the interleaving period, writing from the queue with the fewest interleaving periods written. The interleaver may need to wait for an entire interleaving period to be present in the queue before the block can be written to the file. [000407] In one embodiment, the generated 'audio,' 'video,' and 'subtitle' blocks are interleaved so that the 'audio' and Petition 870180148460, dated 06 / 11 / 2018, pages 107 / 125 In versions 102 / 111, 'subtitle' blocks are located within the file before the 'video' blocks containing information about the video frames to which they correspond. In other versions, 'audio' and 'subtitle' blocks may be located after the 'video' blocks to which they correspond. The timing differences between the location of 'audio,' 'video,' and 'subtitle' blocks largely depend on the temporary storage capabilities of the players used to play the devices. In versions where temporary storage is limited or unknown, the interleaver interleaves 'audio,' 'video,' and 'subtitle' blocks so that the 'audio' and 'subtitle' blocks are located between the 'video' blocks, where the 'video' block immediately following the 'audio' and 'subtitle' blocks contains the first video frame corresponding to the audio or subtitle. 3.7.4.2. Generating DRM information [000408] In modes where DRM is used to protect the video content of a multimedia file, DRM information can be generated concurrently with the encoding of the video blocks. As each block is generated, the block can be encrypted and a DRM block generated containing the information relating to the encryption of the video block. 3.7.4.3. Interleaving DRM Information [000409] An interleaver according to one embodiment of the present invention interleaves a DRM block containing the encryption information of a video block before the video block. In one embodiment, the DRM block for video block n is located between video block n-1 and video block n. In other embodiments, the spacing of the DRM before and after video block n is dependent on the amount of temporary storage provided within the device that decodes the media file. Petition 870180148460, dated 06 / 11 / 2018, pages 108 / 125 103 / 111 3.7.5. Generating the 'idxl' block [000410] Once the 'movi' list block has been generated, generating the 'idx1' block is a simple process. The 'idx1' block is created by reading the location within the 'movi' list block of each 'data' block. This information is combined with information read from the 'data' block regarding the track to which the 'data' block belongs and the content of the 'data' block. All this information is then inserted into the 'idx1' block in a manner appropriate to whichever of the formats described above is being used to represent the information. 4. Multimedia File Transmission and Distribution [000411] Once a multimedia file is generated, the file can be distributed over any of a variety of networks. The fact that in many modes the elements required to generate a multimedia presentation and menus, among other things, are contained in a single file simplifies the transfer of information. In several modes, the multimedia file can be distributed separately from the information required to decipher the content of the multimedia file. [000412] In one embodiment, multimedia content is provided to a first server and encoded to create a multimedia file according to the present invention. The multimedia file may then be located either on the first server or on a second server. In other embodiments, the DRM information may be located on the first server, the second server, or a third server. In one embodiment, the first server may be consulted to ensure the location of the encoded multimedia file and / or to ensure the location of the DRM information. 5. Decoding the Multimedia File [000413] The information of a multimedia file according to Petition 870180148460, dated 06 / 11 / 2018, pages 109 / 125 104 / 111 One embodiment of the present invention can be accessed by a computer configured using appropriate software, a dedicated player that is fixed to the multimedia file access information, or any other device capable of parsing an AVI file. In several embodiments, the devices can access all the information in the multimedia file. In other embodiments, the device may be unable to access all the information in the multimedia file according to an embodiment of the present invention. In one specific embodiment, a device is unable to access any of the above-described information that is in blocks not specified in the AVI file format. In embodiments where not all information can be accessed, the device will typically discard those blocks not recognized by the device. [000414] Typically, a device capable of accessing the information contained in a multimedia file according to an embodiment of the present invention is capable of performing a number of functions. The device may display a multimedia presentation involving the display of video on a visual display, generate audio from one of potentially a number of audio tracks in an audio system, and display subtitles from potentially one of a number of subtitle tracks. Several embodiments may also display menus on a visual display while playing the accompanying audio and / or video. These display menus are interactive, with features such as selectable buttons, and pop-up menus and submenus. In some embodiments, the menu items may point to audio / video content outside the multimedia file currently being accessed.External content can be located either locally to the device accessing the multimedia file, or the file itself can be located remotely, such as over a local area. Petition 870180148460, dated 06 / 11 / 2018, pages 110 / 125 105 / 111 wide area or a public network. Many modalities can also search one or more multimedia files according to the 'metadata' included in the multimedia file(s) or the 'metadata' referenced by one or more of the multimedia files. 5.1. Multimedia Presentation Display [000415] Given the ability of multimedia files according to the embodiments of the present invention to support multiple audio tracks, multiple video tracks, and multiple subtitle tracks, the display of a multimedia presentation using such a multimedia file that combines video, audio, and / or subtitles may require the selection of a specific audio track, video track, and / or subtitle track either through a visual menu system or an instant menu system (the operation of which is discussed below) or through the standard settings of the device used to generate the multimedia presentation. Once an audio track, a video track, and potentially a subtitle track are selected, the display of the multimedia presentation can proceed. [000416] A process for locating the required multimedia information from a multimedia file that includes DRM and displaying the multimedia information according to an embodiment of the present invention is illustrated in Figure 4.0. The process 620 includes obtaining the encryption key required to decrypt the DRM header (622). The encryption key is then used to decrypt (624) the DRM header and the first DRM block is located (626) within the list and 'movi' block of the multimedia file. The encryption key required to decrypt the 'DRM' block is obtained (628) from the table in the 'DRM' header and the encryption key is used to decrypt an encrypted video block. The required audio block and any required subtitle block accompanying the video block are then decoded (630) and the audio, video information Petition 870180148460, dated 06 / 11 / 2018, pages 111 / 125 106 / 111 deo and any legend are presented (632) through the display and sound system. [000417] In several modes, the chosen audio track may include multiple channels to provide stereo or surround sound audio. When a subtitle track is chosen to be displayed, a determination may be made as to whether the previous video frame includes a subtitle (this determination may be made in any of a variety of ways that achieve the result of identifying a previous 'subtitle' block that contained the subtitle information that should be displayed over the currently decoded video frame). If the previous subtitle included a subtitle and the timing information for the subtitle indicates that the subtitle should be displayed with the current frame, then the subtitle is overlaid on the decoded video frame.If the previous frame did not include a caption, or the timing information for the caption in the previous frame indicates that the caption should not be displayed in conjunction with the currently decoded frame, then a 'subtitle' block for the selected caption track is searched for. If a 'subtitle' block is found, then the caption is overlaid onto the decoded video. The video (including all overlaid captions) is then displayed with the accompanying audio. [000418] Returning to the discussion of Figure 4.0, the process determines (634) whether there are any additional DRM blocks. If there are, then the next DRM block is located (626) and the process continues until no additional DRM blocks remain. At which point, the presentation of the audio, video and / or subtitle tracks is complete (636). [000419] In several modes, a device can search for a specific portion of multimedia information (for example, a specific scene from a movie with an accompanying audio track) Petition 870180148460, dated 06 / 11 / 2018, pages 112 / 125 107 / 111 specific track and optionally a specific accompanying caption track) using the information contained in the 'hdrl' block of a multimedia file according to the present invention. In many embodiments, the decoding of the 'video' block, the 'audio' block and / or the 'subtitle' block can be performed in parallel with other tasks. [000420] An example of a device capable of accessing multimedia file information and displaying a video together with a specific audio track and / or a specific subtitle track is a computer configured in the manner described above using software. Another example is a DVD player equipped with a codec that includes these capabilities. In other embodiments, any device configured to locate or select (either intentionally or arbitrarily) the 'data' blocks that correspond to specific media tracks and decode these tracks for presentation is capable of generating a multimedia presentation using a multimedia file according to the practice of the present invention. [000421] In several embodiments, a device can reproduce multimedia information from a multimedia file in combination with multimedia information from an external file. Typically, such a device would do this by sourced an audio track or a subtitle track from a local file referenced in a multimedia file of the type described above. If the referenced file is not stored locally and the device is networked with the location where the device is stored, then the device can obtain a local copy of the file. The device would then access both files, establishing a chain of video, audio, and subtitle (if required) in which the various multimedia tracks are fed from the different file sources. Petition 870180148460, dated 06 / 11 / 2018, pages 113 / 125 108 / 111 5.2. Menu Generation [000422] A decoder according to an embodiment of the present invention is illustrated in Figure 4.1. The decoder 650 processes a multimedia file 652 according to an embodiment of the present invention by providing the file to a demultiplexer 654. The demultiplexer extracts the 'DMNU' block from the multimedia file and extracts all 'LanguageMenus' blocks from the 'DMNU' block and provides them to a menu analyzer 656. The demultiplexer also extracts all 'Media' blocks from the 'DMNU' block and provides them to a media renderer 658. The menu analyzer 656 analyzes the information from the 'LanguageMenu' blocks to construct a state machine that represents the menu structure defined in the 'LanguageMenu' block. The state machine representing the menu structure can be used to provide displays for the user and respond to user commands. The state machine is provided for a 660 menu state controller.The menu state controller monitors the current state of the menu state machine and receives commands from the user. User commands can cause a state transition. The initial display provided to a user and any display updates that accompany a menu state transition can be controlled using a 662 menu player interface. The 662 menu player interface can be connected to the menu state controller and the media renderer. The menu player interface instructs the media renderer which media should be extracted from the media blocks and provided to the user through the 664 player connected to the media renderer. The user can provide the player with instructions using an input device, such as a keyboard, mouse, or remote control. Generally, the multimedia file dictates the menu initially displayed to the user and the user's instructions. Petition 870180148460, dated 06 / 11 / 2018, pages 114 / 125 109 / 111 dictate the audio and video displayed after the initial menu is generated. The system illustrated in Figure 4.1 can be implemented using a computer and software. In other embodiments, the system can be implemented using function-specific integrated circuits or a combination of firmware software. [000423] An example of a menu according to an embodiment of the present invention is illustrated in Figure 4.2. The menu display 670 includes four button areas 672, a background video 674, which includes a title 676, and a pointer 678. The menu also includes background audio (not shown). The visual effect created by the display can be misleading. The visual appearance of the buttons is typically part of the background video, and the buttons themselves are simply defined regions of the background video that have specific actions associated with them when the region is activated by the pointer. A pointer is typically an overlay. [000424] Figure 4.3 conceptually illustrates the source of all information in the display shown in Figure 4.2. The background video 674 may include a menu title, the visual appearance of the buttons, and the display background. All these elements and additional elements may appear static or animated. The background video is extracted using information contained in a 'MediaTrack' block 700 that indicates the location of the background video within a video track 702. The background audio 706 that accompanies the menu can be located using a 'MediaTrack' block 708 that indicates the location of the background audio within an audio track 710. As discussed above, the pointer 678 is part of an overlay 713. The overlay 713 may also include a graphic that appears to highlight the portion of the background video that appears as a button.In one embodiment, the 713 overlay is obtained using a 'MediaTrack' Block 712 which indicates the location of the overlay within a sound track. Petition 870180148460, dated 06 / 11 / 2018, pages 115 / 125 110 / 111 overlay 714. The way in which the menu interacts with a user is defined by the 'Action' blocks (not shown) associated with each of the buttons. In the illustrated mode, a 'PlayAction' block 716 is shown. The 'PlayAction' block indirectly references (the other blocks referenced by the 'PlayAction' block are not shown) a scene within a multimedia presentation contained in the multimedia file (i.e., an audio track, a video track, and possibly a subtitle track). The 'PlayAction' block 716 finally references the scene using a 'MediaTrack' block 718, which indicates the scene within the movie track. A point in a selected or default audio track and potentially a subtitle track are also referenced. [000425] As the user enters commands using the input device, the display may be updated not only in response to the selection of button areas but also simply because the pointer is located within the button area. As discussed above, typically all media information used to generate the menus is located within the multimedia files and more specifically within a 'DMNU' block. Although in other modes, the information may be located anywhere within the file and / or in other files. 5.3. Access to Metadata [000426] Metadata is a standardized method of representing information. The standardized nature of metadata allows data to be accessed and understood by automated processes. In one modality, metadata is extracted and provided to a user for observation. Several modalities allow multimedia files on a server to be inspected to provide information regarding users, viewing habits, and viewing preferences; such information could be used by Petition 870180148460, dated 06 / 11 / 2018, pages 116 / 125 111 / 111 software applications recommend other multimedia files that a user might enjoy viewing. In one embodiment, the recommendations may be based on multimedia files contained on other users' servers. In other embodiments, a user may request a multimedia file, and the file may be located by a search engine and / or intelligent agents that inspect the 'metadata' of multimedia files in a variety of locations. Furthermore, the user may choose from several multimedia files containing a specific multimedia presentation based on the 'metadata' relating to the mode in which each of the different versions of the presentation was encoded. [000427] In various embodiments, the 'metadata' of multimedia files according to the embodiments of the present invention can be accessed for cataloging purposes or to create a simple menu to access the file content. [000428] Although the above description contains many specific embodiments of the invention, these should not be considered as imitations of the scope of the invention, but rather as an example of one embodiment thereof. For example, a multimedia file according to an embodiment of the present invention may include a single multimedia presentation or multiple multimedia presentations. Furthermore, such a file may include one or more menus and any variety of different types of 'metadata'. Consequently, the scope of the invention should be determined not only by the embodiments illustrated, but by the appended claims and their equivalents.< / n>
Claims
1. System for decoding multimedia files (30, 30'), wherein the decoding system comprises: a processor configured and instructed to extract information from a multimedia file (30, 30'), wherein: the multimedia file (30, 30') includes a plurality of blocks to carry information, said blocks including a plurality of video blocks (262) and a plurality of audio blocks (264);characterized in that the video blocks (262) are parts of at least one video track, said at least one video track (702) comprising a series of encoded video frames including at least 1 to N encoded video frames, said video blocks (262) including at least 1 to N encoded video blocks, wherein each of said 1 to N video frames corresponds to a block of 1 to N video blocks, wherein the video blocks (262) include a marker video block N where sufficient information is contained in blocks 1 to N-1 to generate the N video frame, wherein the N marker video block includes an unencoded P frame with timing information identical to that of a previous encoded video frame in the series of encoded video frames; and the audio blocks (264) are parts of at least one audio track;The processor is further configured and instructed to access multimedia file information by: decoding frame N using the N-marker video block along with video blocks 1 to N-1.
2. System, according to claim 1, characterized in that said audio blocks and video blocks (264, 262) are interleaved in such a way that the audio blocks (264) are located within the file Petition 870200018202, dated 07 / 02 / 2020, page 14 / 20 2 / 6 before the video blocks (262) containing information regarding the video frames to which they correspond.
3. System according to claim 1 or 2, characterized in that the encoded audio track is provided in blocks that do not contain audio information corresponding to the content of a corresponding video block (262), and encoding of at least one encoded audio track as audio blocks (264) involves identifying the audio information in the encoded audio track accompanying the video block (262) and extracting the audio information from the existing audio blocks to create a new audio block (264).
4. System, according to any one of claims 1 to 3, characterized in that the audio blocks (264) are interleaved between the video blocks (262) such that each particular audio block (264) is positioned relative to a corresponding video block (262) based on an amount of video capable of being loaded (buffered) by a device capable of displaying the multimedia file (30, 30').
5. Multimedia file (30, 30') including: a plurality of blocks for carrying information, said blocks including a plurality of video blocks (262) and a plurality of audio blocks (264); said audio blocks (264) being parts of at least one audio track; and characterized in that said video blocks (262) being parts of at least one video track, said video track comprising a series of encoded video frames including at least 1 to N encoded video frames, said video blocks (262) including at least 1 to N video blocks, wherein each of said video frames 1 to N Petition 870200018202, dated 07 / 02 / 2020, p.15 / 20 3 / 6 corresponds to a block of video blocks 1 to N, wherein video blocks (262) include a marker video block N where sufficient information is contained in blocks 1 to N-1 to decode video frame N, wherein the marker video block N includes an unencoded frame P with timing information identical to that of a previous encoded video frame in the series of encoded video frames.
6. System for encoding multimedia files, comprising a processor configured to: encode at least one video track as a plurality of video blocks (262), wherein: the at least one video track comprises a series of encoded video frames including at least video frames 1 to N; characterized in that the video blocks (262) are parts of the at least one video track and the video blocks (262) include at least video blocks 1 to N, wherein the video blocks (262) include a marker video block N where sufficient information is contained in blocks 1 to N-1 to generate the video frame N, wherein the marker video block N includes an unencoded frame P with timing information identical to that of a previous encoded video frame in the series of encoded video frames;and frame N can be decoded using the video marker block N along with video blocks 1 to N-1, where each of the aforementioned encoded video frames 1 to N corresponds to a block of video blocks 1 to N; encode at least one audio track as audio blocks (264); and write the video blocks (262) and audio blocks (264) to a single file.
7. System for decoding multimedia files characterized by the fact that it comprises: Petition 870200018202, dated 07 / 02 / 2020, page.16 / 20 4 / 6 a processor configured to extract information from a multimedia file, wherein the multimedia file comprises at least one video track encoded as a plurality of video blocks, a set of DRM digital rights management blocks, and an index block; wherein each DRM block comprises DRM information for decoding a video block into a plurality of video blocks; wherein the index block includes information regarding the location of data blocks within the multimedia file that includes the video block locations of at least one video track along with the DRM blocks of the DRM block set; and wherein the processor is configured to decode the video block and identify a corresponding DRM block that contains the DRM information for the video block decoded using the index block; and wherein the processor is configured to construct a video frame for display.
8. System according to claim 7, characterized in that each DRM block is identified by a FOURCC code that identifies a track number associated with the DRM block.
9. A system according to claim 7, characterized in that at least one video block is partially encrypted such that only a portion of the at least one video block is encrypted, wherein a DRM block associated with the at least one video block comprises a reference to the portion of the video block that is encrypted.
10. System, according to claim 9, characterized in that the DRM block additionally comprises a reference to a DRM key that can be used to decrypt the encrypted part. Petition 870200018202, dated 07 / 02 / 2020, page 17 / 20 5 / 6 11. System, according to claim 7, characterized in that at least one unencrypted video block out of the plurality of video blocks does not have a corresponding DRM block in the DRM block set.
12. System according to claim 7, characterized in that the location of a particular DRM block for a particular video block is dependent on an amount of buffering provided within the system.
13. A system for encoding multimedia files, characterized in that it comprises: a memory storing an encoding application; a network interface; a processing unit, wherein, upon execution, the encoding application directs the processing unit to: obtain source media using the network interface, wherein the source media comprises DRM digital rights management information for the video; encode at least one video track as a plurality of video blocks, the video blocks being parts of at least one video track, the video track comprising a series of encoded video frames, wherein at least one part of the plurality of video blocks is encoded with DRM information; encode the DRM information as a set of DRM blocks, wherein each DRM block comprises DRM information to decode a video block from the plurality of video blocks;Interleave the video blocks and the DRM blocks so that a DRM block for encoding a particular video block is located before a particular video block; and encode at least one index block that includes information regarding the locations of video blocks and the locations of DRM blocks within a multimedia file; write the interleaved blocks to a single file; and transmit the single file using the network interface.