Frame conversion for adaptive stream alignment

By refactoring the metadata of audio and video frames to align their frame types with timestamps, the problem of aligning audio I-frames and video I-frames in adaptive streaming is solved, enabling seamless splicing and switching of audio and video content. This is applicable to MPEG-DASH, HLS, MMT and MPEG-2 transport stream formats.

CN115802046BActive Publication Date: 2026-05-26DOLBY LABORATORIES LICENSING CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2019-06-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In adaptive streaming, audio I-frames and video I-frames are difficult to align, making it difficult to achieve seamless splicing and switching of audio and video content at segment boundaries.

Method used

By re-creating the metadata of audio and video frames, an AV bitstream is generated, aligning its frame type with the timestamp, thereby ensuring I-frame synchronization of audio and video content and ensuring that each segment begins with an I-frame and contains subsequent P-frames.

Benefits of technology

It achieves seamless splicing and switching of audio and video content at segment boundaries, meets the requirements of MPEG-DASH, HLS, MMT and MPEG-2 transport stream formats, and ensures audio and video synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115802046B_ABST
    Figure CN115802046B_ABST
Patent Text Reader

Abstract

This application relates to frame conversion for adaptive streaming alignment. A method for generating an AV bitstream (e.g., an MPEG-2 transport stream or bitstream segment with an adaptive streaming format) such that the AV bitstream includes at least one video I-frame synchronized with at least one audio I-frame includes, for example, re-encoding at least one video or audio frame (as a re-encoded I-frame or a re-encoded P-frame). Typically, a content segment of the AV bitstream containing the re-encoded frame begins with an I-frame and includes at least one subsequent P-frame. Other aspects are methods for adapting such AV bitstreams, audio / video processing units configured to perform any embodiment of the methods of the present invention, and audio / video processing units including buffer memory storing at least one segment of an AV bitstream generated according to any embodiment of the methods of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information related to divisional application

[0002] This case is a divisional application. The parent application of this divisional application is the invention patent application filed on June 27, 2019, with application number 201980043163.9 and invention title "Frame Conversion for Adaptive Stream Alignment".

[0003] Cross-reference to related applications

[0004] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 690,998 and European Patent Application No. 18180389.1, both filed on June 28, 2018, which are hereby incorporated by reference. Technical Field

[0005] This document relates to audio and video signal processing, and more specifically, to the generation and adaptation of bitstreams (e.g., bitstreams, bitstream segments, or transport streams used in adaptive streaming formats or methods / protocols) comprising video frames (encoded video data and optionally metadata) and audio frames (encoded audio data and optionally metadata). Some embodiments of the invention generate or adapt (e.g., align) bitstream segments (e.g., having an MPEG-2 transport stream format, or a format according to MMT or MPEG-DASH or another streaming method / protocol, or another adaptive streaming format, or another standards-compliant format) comprising encoded audio data (e.g., encoded audio data in a format conforming to or compatible with AC-4, or MPEG-D USAC, or MPEG-H audio standards). Background Technology

[0006] In adaptive streaming, data having (or being used for) an adaptive streaming format may include separate bitstreams (or bitstream segments) for each video and audio representation. Therefore, data may not include a single bitstream (e.g., a single transport stream), and instead may include two or more separate bitstreams. Hereinafter, the expression “AV bitstream” (defined below) is used to indicate a signal (or multiple signals) that indicates a bitstream or bitstream segment, or two or more bitstream segments (e.g., a transport stream or bitstream segments used in an adaptive streaming format), which contains video data and / or encoded audio data, and typically also includes metadata.

[0007] An AV bitstream (e.g., a transport stream or a bitstream segment used in an adaptive streaming format or streaming method or protocol) may indicate at least one audio / video (audio and / or video) program (“program”) and may contain (for each program thus indicated) video data frames (which define at least one video elementary stream) and encoded audio data frames corresponding to the video data (which define at least one audio elementary stream). The video data frames may contain or include video data I-frames (video I-frames) and video data P-frames (video P-frames), and the encoded audio data frames may contain or include I-frames or encoded audio data (audio I-frames) and encoded audio data P-frames (audio P-frames).

[0008] In this document, as included in the claims, "I-frame" refers to an independently decodable frame that can be decoded using information from itself alone. In this document, as included in the claims, a non-I-frame (e.g., a predicted encoded frame) is referred to as a "P-frame". In AV bitstreams, P-frames typically require information from the preceding I-frame in order to be decoded. I-frames (or P-frames) may contain video data (and often also metadata), and such frames are sometimes referred to herein as video frames or video data frames. I-frames (or P-frames) may contain encoded audio data (and often also metadata), and such frames are sometimes referred to herein as audio frames or audio data frames.

[0009] Many modern audio codecs (e.g., AC-4 Audio, MPEG-H Audio, and MPEG-D USAC Audio codecs) and video codecs utilize the concepts of independently decorable frames (as defined above as “I-frames”) and other frames (e.g., non-independently decorable frames, as defined above as “P-frames”), making bitstreams containing audio and / or video content encoded by such codecs typically contain both I-frames and P-frames. Many packaged media delivery formats or protocols (e.g., MPEG-DASH (Dynamic Adaptive Streaming over HTTP, published under ISO / IEC 23009-1:2012), HLS (Apple HTTP Real-Time Streaming), MMT (MPEG Media Transport), and MPEG-2 Transport Streaming Format) require segments of audio (or video) content to begin with an I-frame to enable seamless splicing (or switching) at segment boundaries and can benefit from audio and video segment alignment. Because audio encoders and video encoders typically operate independently, and both are allowed to decide when to create I-frames without knowing the other, alignment between audio I-frames and video I-frames is often difficult. Summary of the Invention

[0010] Some embodiments of the present invention generate AV bitstreams that include video data I-frames (video I-frames) synchronized with encoded audio data I-frames (audio I-frames) to resolve alignment issues between segments of the AV bitstream's underlying streams (e.g., a video underlying stream and one or more corresponding audio underlying streams). In a typical embodiment, the AV bitstream indicates at least one audio / video program ("program") and includes (for each program thus indicated) video data frames (which define at least one video underlying stream) and corresponding encoded audio data frames (which define at least one audio underlying stream), the video data frames including video data I-frames (video I-frames) and video data P-frames (video P-frames), and the encoded audio data frames including encoded audio data I-frames (audio I-frames) and encoded audio data P-frames (audio P-frames). At least one frame of the AV bitstream (at least one video frame among the video frames and / or at least one audio frame among the audio frames) has been remade (as a remade I-frame or a remade P-frame) such that a content segment of the stream containing the remade frame begins with an I-frame and contains at least one subsequent P-frame, and typically also such that the I-frame of the content segment is aligned (time-aligned, e.g., because the I-frame of the content segment has the same timestamp value as the I-frame of the corresponding content segment of the stream). In some embodiments, at least one audio frame (but no video frame) of the AV bitstream has been remade (as a remade audio I-frame or a remade audio P-frame) such that an audio content segment of the stream containing the remade frame begins with an audio I-frame (aligned with the video I-frame of the corresponding video content segment) and contains at least one subsequent audio P-frame. In a typical embodiment, the AV bitstream has a packaged media delivery format that requires each segment of the AV bitstream's content (audio or video content) to begin with an I-frame to achieve seamless adaptation (e.g., splicing or switching) at segment boundaries (e.g., at the beginning of a video content segment in a video elementary stream, the beginning being aligned with the beginning of a corresponding audio content segment in at least one audio elementary stream). In some (but not all) embodiments, the packaged media delivery format conforms to MPEG-DASH, HLS, or MMT methods / protocols, or the MPEG-2 transport stream format, or is typically based on the ISO Basic Media Format (MPEG-414496-12).

[0011] In this document, "remake" refers to a frame containing content (audio or video content) (or a "remake" frame) that means replacing the metadata of the frame with different metadata without modifying the content of the frame (e.g., without modifying the content by decoding such content and then re-encoding the decoded content), or replacing the metadata of the frame with different metadata (and modifying the content of the frame without decoding such content and then re-encoding the decoded content). For example, in some embodiments of the invention, an audio P-frame is remade into a remade audio I-frame by replacing all or some of the metadata (which may be in the header of the audio P-frame) with a different set of metadata (e.g., different sets of metadata consisting of or containing metadata, which are copied from or generated by modifying the metadata obtained from the previous audio I-frame). In another instance, an audio P-frame is remade into a remade audio I-frame, comprising replacing the encoded audio content of the P-frame (which was encoded using incremental time (inter-frame) coding) with encoded audio content already encoded using intra-frame coding (i.e., the same audio content, but encoded using incremental frequency or absolute coding), without decoding the original audio content of the P-frame. In yet another instance, in some embodiments of the invention, an audio P-frame is remade into a remade audio I-frame by copying at least one previous audio P-frame into the P-frame (containing the encoded audio content of each previous P-frame), such that the resulting remade I-frame contains a different set of metadata than the P-frame (i.e., containing the metadata of each previous P-frame as well as the original metadata of the P-frame), and also contains additional encoded audio content (the encoded audio content of each previous P-frame). In another instance, in some embodiments of the invention, an audio (or video) I-frame is remade into a remade audio (video) P-frame by replacing all or some of the metadata (which may be in the header of the I-frame) with a different set of metadata (e.g., different sets of metadata consisting of or containing metadata, which are copied from previous frames or generated by modifying metadata obtained from previous frames). This is done without modifying the content of the I-frame (audio or video content) (e.g., without modifying the encoded content of the I-frame by decoding such content and then re-encoding the decoded content).

[0012] In this document, a frame of an AV bitstream, or a frame that may be contained in an AV bitstream (referred to as the “first” frame for clarity, although it may be any video or audio frame) having the expression “first decoding type” indicates that the frame is an I-frame (a frame that has been encoded using information only from itself and is independently decodeable, such that its decoding type is the decoding type of an I-frame) or a P-frame (such that its decoding type is the decoding type of a P-frame), and another frame of the AV bitstream (“second frame”) or another frame that may be contained in the AV bitstream having the expression “second decoding type” (different from the first decoding type) indicates that the second frame is a P-frame (if the first frame is an I-frame) or an I-frame (if the first frame is a P-frame).

[0013] Some embodiments of the method of the present invention include the following steps:

[0014] (a) Providing an input AV bitstream, the input AV bitstream comprising: frames indicating audio and video content (e.g., transmitted, delivered, or otherwise provided to an encoder or NBMP entity, or an input transport stream or other input AV bitstream containing the frames), the frames comprising frames of a first decoding type; and optionally metadata associated with each frame in the frames, wherein each frame of the first decoding type comprises a P-frame or an I-frame and indicates audio content (e.g., encoded audio content) or video content (e.g., encoded video content);

[0015] (b) Modify some of the metadata associated with at least one of the frames of the first decoding type to different metadata to generate at least one remade frame of a second decoding type different from the first decoding type, and

[0016] (c) Generating an output AV bitstream in response to the input AV bitstream, comprising remaking at least one frame of the first decoding type into a remade frame of a second decoding type different from the first decoding type (e.g., if the frame of the first decoding type is a P-frame, then the remade frame is a remade I-frame, or if the frame of the first decoding type is an I-frame, then the remade frame is a remade P-frame), such that the AV bitstream includes a segment containing the content of the remade frame, and the segment of the content begins with an I-frame and includes at least one P-frame following the I-frame, for aligning the I-frames of the video content with the I-frames of the audio content. For example, step (b) may include remaking at least one audio P-frame such that the remade frame is an audio I-frame, or remaking at least one audio I-frame such that the remade frame is an audio P-frame, or remaking at least one video P-frame such that the remade frame is a video I-frame, or remaking at least one video I-frame such that the remade frame is a video P-frame.

[0017] In some such embodiments, step (a) includes the following steps: in a first system (e.g., a production unit or encoder), generating an input AV bitstream (e.g., an input transport stream) containing the frame indicating the content; and delivering (e.g., transmitting) the input AV bitstream to a second system (e.g., an NBMP entity); and steps (b) and (c) are performed in the second system.

[0018] In some such embodiments, at least one audio P-frame (containing metadata and encoded audio data, e.g., AC-4 encoded audio data) is remade into an audio I-frame (containing the same encoded audio data and different metadata). In some other embodiments, at least one audio I-frame (containing metadata and encoded audio data, e.g., AC-4 encoded audio data) is remade into an audio P-frame (containing the same encoded audio data and different metadata). In some embodiments, at least one audio P-frame (containing metadata and encoded audio content already encoded using incremental time (inter-frame) coding) is remade into a remade audio I-frame by replacing the original encoded audio content of the P-frame with encoded audio content already encoded using intra-frame coding (i.e., the same audio content, but encoded using incremental frequency or absolute coding), without decoding the original audio content of the P-frame.

[0019] In one embodiment, the different metadata may be metadata associated with a frame of the first or second decoding type prior to the remade frame.

[0020] In some embodiments, the frame being remade is an audio P-frame. The remade frame is an audio I-frame. The metadata of the audio P-frame is modified by replacing some of the metadata of such an audio P-frame with different metadata copied from a previous audio I-frame, such that the remade frame contains the different metadata.

[0021] In some other embodiments, the frame being remade is an audio P-frame. The remade frame is an audio I-frame. The metadata of the audio P-frame is modified by modifying the metadata from the previous audio I-frame and replacing at least some of the metadata in the audio P-frame with the modified metadata, such that the remade frame (I-frame) contains the modified metadata.

[0022] In some other embodiments, the frame being remade is an audio P-frame. The remade frame is an audio I-frame. The metadata of the audio P-frame is modified by copying at least one previous P-frame into the audio P-frame, such that the resulting remade I-frame contains a different set of metadata than the P-frame (i.e., it contains the metadata of each previous P-frame as well as the original metadata of the P-frame), and also contains additional encoded audio content (the encoded audio content of each previous P-frame).

[0023] In some embodiments, generating the output AV bitstream includes determining segments of the audio and video content of the output AV bitstream as segments of the video content (e.g., the input video content of the input AV bitstream) that begin with an I-frame. In other words, the segments of the audio and video content are defined by the (e.g., input) video content. The segments of the audio and video content begin with an I-frame of the (e.g., input) video content. The audio frames are remade such that the I-frames of the audio content are aligned with the I-frames of the video content. The segments of the audio and video content also begin with an I-frame of the audio content.

[0024] In some embodiments, generating the output AV bitstream includes transmitting audio and video content as well as metadata of segments of the content from the input AV bitstream, which has not yet been modified for the output AV bitstream. For example, a buffer may be used to buffer (e.g., store) segments of the content from the input AV bitstream. A subsystem may be configured to recreate one of these buffered segment frames of the content and transmit them as recreated, unmodified frames. Any media content and metadata unmodified by the subsystem are transmitted from the input AV bitstream to the output AV bitstream. The generated output AV bitstream includes the unmodified frames from the input AV bitstream and the recreated frames.

[0025] In a first type of embodiment (sometimes referred to herein as an embodiment implementing "Method 1"), the method of the present invention includes: (a) generating audio I-frames and audio P-frames indicating content (e.g., in a conventional manner); and (b) generating an AV bitstream, comprising re-encoding at least one of the audio P-frames into a re-encoded audio I-frame, such that the AV bitstream includes a segment containing content of the re-encoded audio I-frame, and the segment of content begins with the re-encoded audio I-frame. Typically, the segment of content also includes at least one audio P-frame following the re-encoded audio I-frame. Steps (a) and (b) can be performed in an audio encoder (e.g., a production unit that includes or implements an audio encoder), comprising performing step (a) by operating the audio encoder.

[0026] In some embodiments of implementing method 1, step (a) is performed in an audio encoder (e.g., an AC-4 encoder) that generates the audio I-frame and the audio P-frame, step (b) includes remaking at least one of the audio P-frames (corresponding to the time when the audio I-frame is needed) into the remade audio I-frame, and step (b) further includes including the remade audio I-frame instead of the one audio P-frame in the audio P-frame in the AV bitstream.

[0027] Some embodiments of method 1 include: in a first system (e.g., a production unit or encoder), generating an AV bitstream (e.g., an input transport stream) containing the audio I-frames and audio P-frames generated in step (a); delivering (e.g., transmitting) the input AV bitstream to a second system (e.g., an NBMP entity); and performing step (b) in the second system.

[0028] In a second type of embodiment (sometimes referred to herein as an embodiment implementing "Method 2"), the method of the present invention includes: (a) generating an audio I-frame indicating content (e.g., in a conventional manner); and (b) generating an AV bitstream, comprising re-encoding at least one of the audio I-frames into a re-encoded audio P-frame, such that the AV bitstream includes a segment of the content comprising the re-encoded audio P-frame, and the segment of the content begins with one of the audio I-frames generated in step (a). Steps (a) and (b) can be performed in an audio encoder (e.g., a production unit that includes or implements an audio encoder), comprising performing step (a) by operating the audio encoder.

[0029] In some embodiments of implementing method 2, step (a) is performed in an audio encoder (e.g., an AC-4 encoder) that generates the audio I-frame, and step (b) includes remaking each audio I-frame corresponding to a time other than a segment boundary (and therefore not appearing at a segment boundary), thereby determining at least one remade audio P-frame, and the step of generating the AV bitstream includes the step of including the at least one remade audio P-frame in the AV bitstream.

[0030] Some embodiments of method 2 include: in a first system (e.g., a production unit or encoder), generating an AV bitstream (e.g., an input transport stream) containing the audio I-frame generated in step (a); delivering (e.g., transmitting) the input AV bitstream to a second system (e.g., an NBMP entity); and performing step (b) in the second system.

[0031] Some embodiments of the present invention's method for generating AV bitstreams (e.g., involving remaking at least one frame of the input AV bitstream) satisfy at least one current potential network constraint or other constraint (e.g., the generation of the AV bitstream involves remaking at least one frame of the input bitstream and is performed to ensure that the AV bitstream satisfies at least one current potential network constraint on the input bitstream). For example, when the AV bitstream is generated by an NBMP entity (e.g., as a CDN server or an MPEG NBMP entity included in a CDN server), the NBMP entity can be implemented to insert or remove I-frames from the AV bitstream in a manner dependent on network and other constraints. Examples of such constraints include, but are not limited to, available bitrate, the time required to load into a program, and / or the segment duration of a potential MPEG-DASH or MMT AV bitstream.

[0032] In some embodiments, a method for generating an AV bitstream includes:

[0033] Provide frames, wherein at least one of the frames is a hybrid frame comprising P-frames and I-frames, wherein the I-frames indicate an encoded version of the content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and

[0034] Generating the AV bitstream includes selecting at least one of the P-frames of the mixed frame and including each selected P-frame in the AV bitstream such that the AV bitstream includes a segment that begins with an I-frame and includes at least one of the selected P-frames following the I-frame.

[0035] In some embodiments, the method of the present invention includes the step of adapting (e.g., switching or splicing) an AV bitstream that indicates at least one audio / video program (“program”) and has been generated according to any embodiment of the method of the present invention for generating AV bitstreams (e.g., having MPEG-2 format, or MPEG-DASH or MMT format) to generate an adapted (e.g., spliced) AV bitstream. In a typical embodiment, the adapted AV bitstream is generated without modifying any encoded audio elementary stream of the AV bitstream. In some embodiments that include the step of adapting (e.g., switching or splicing) the AV bitstream, the adaptation is performed at an adaptation point (e.g., time in the program indicated by the AV bitstream) of the AV bitstream corresponding to a video I-frame (e.g., a re-encoded video I-frame) and at least one corresponding re-encoded audio I-frame of the AV bitstream. In a typical embodiment where the AV bitstream contains (for each program thus indicated) video data frames (which define at least one video elementary stream) and corresponding encoded audio data frames (which define at least one audio elementary stream), the adaptation (e.g., splicing) is performed in a manner that ensures audio / video synchronization (“A / V” synchronization) is maintained without requiring an adapter (e.g., splicer) to modify any encoded audio elementary stream of the AV bitstream.

[0036] The AV bitstream generated by any typical embodiment of the method according to the invention has (i.e., satisfies) the property of I-frame synchronization (i.e., video and audio are encoded in synchronization such that for each program indicated by the AV bitstream, for each video I-frame in the video elementary stream of the program, there is at least one matching audio I-frame in the audio elementary stream of the program (i.e., at least one audio I-frame synchronized with the video I-frame)). In this sense, for each program indicated by the AV bitstream, the data of the AV bitstream indicating the program has the property of I-frame synchronization.

[0037] Another aspect of this invention is an audio / video processing unit (AVPU) configured to perform any embodiment of the method of the invention (e.g., for generating and / or adapting AV bitstreams). For example, the AVPU may be an NBMP entity (e.g., an MPEG NBMP entity, which may be a CDN server or may be included in a CDN server) programmed or otherwise configured to perform any embodiment of the method of the invention. For another instance, the AVPU may be an adapter (e.g., a splicer) configured to perform any embodiment of the AV bitstream adaptation (e.g., splicing) method of the invention. In another type of embodiment of the invention, the AVPU includes a buffer memory (buffer) that stores at least one segment of an AV bitstream generated by any embodiment of the method of the invention (e.g., in a non-transitory manner). Instances of AVPU include, but are not limited to, encoders (e.g., code converters), NBMP entities (e.g., NBMP entities configured to generate AV bitstreams and / or perform adaptations thereto), decoders (e.g., decoders configured to decode the contents of AV bitstreams and / or perform adaptations (e.g., splicing) of AV bitstreams to generate adapted (e.g., spliced) AV bitstreams and decode the contents of said adapted AV bitstreams), codecs, AV bitstream adapters (e.g., splicers), preprocessing systems (preprocessors), postprocessing systems (postprocessors), AV bitstream processing systems, and combinations of such elements.

[0038] Some embodiments of the present invention include systems or apparatus configured (e.g., programmed) to perform any embodiment of the methods of the present invention, and computer-readable media (e.g., disks) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the methods of the present invention or steps thereof. For example, the system of the present invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data (including embodiments of the methods of the present invention or steps thereof). Such a general-purpose processor may be or include a computer system comprising input devices, memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the methods of the present invention (or steps thereof) in response to data asserted thereto. Attached Figure Description

[0039] Figure 1 This is a block diagram of a content delivery network, wherein NBMP entity 12 (and optionally at least one other element) is configured according to an embodiment of the invention.

[0040] Figure 2 This is a diagram of an example of an MPEG-2 transport stream.

[0041] Figure 3 This is a block diagram of an embodiment of the system, wherein one or more elements of the system can be configured according to an embodiment of the invention.

[0042] Figure 4 It includes the configuration according to an embodiment of the present invention. Figure 3 Production unit 3 and Figure 3 A block diagram of the implementation scheme of the capture unit 1.

[0043] Figure 5 It is a diagram of a segment of an AV bitstream in MPEG-DASH format.

[0044] Figure 6 This is a diagram of a hybrid frame generated according to an embodiment of the present invention.

[0045] Figure 6A This is a diagram of a hybrid frame generated according to an embodiment of the present invention.

[0046] Figure 7 This is a block diagram of a system for generating AV bitstreams in MPEG-DASH format according to an embodiment of the present invention.

[0047] Symbols and terms

[0048] Throughout this disclosure, the expression “to” an operation on a signal or data (e.g., to filter, scale, transform, or apply gain to the signal or data) is used broadly to mean performing an operation directly on the signal or data, or on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0049] Throughout this disclosure, and as included in the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem implementing encoding can be referred to as an encoder system, and a system containing such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs, and the other XM inputs are received from an external source) can also be referred to as an encoder system.

[0050] Throughout this disclosure, the expression "AV bitstream," as included in the claims, refers to a signal (or multiple signals) indicating a bitstream, or a bitstream segment, or two or more bitstream segments (e.g., a transport stream or bitstream segments used in an adaptive streaming format), which contains video data and / or encoded audio data, and typically also contains metadata. The expression "AV data" is sometimes used herein to refer to such video data and / or such encoded audio data. Typically, a transport stream is a signal indicating a serial bitstream that contains a sequence of segments of encoded audio data (e.g., data packets), segments of video data (e.g., data packets), and segments of metadata (e.g., metadata containing splicing support) (e.g., headers or other segments). An AV bitstream (e.g., a transport stream) may indicate multiple programs, and each program may contain multiple elementary streams (e.g., a video elementary stream and two or more audio elementary streams). Typically, each elementary stream of a transport stream has an associated descriptor containing information related to the elementary stream.

[0051] Throughout this disclosure, and as included in the claims, the term "processor" is used broadly to refer to a system or apparatus that is programmable or otherwise configured (e.g., with software or firmware) to perform operations on data (e.g., audio and / or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0052] Throughout this disclosure, the terms "audio / video processing unit" (or "AV processing unit" or "AVPU") and "AV processor" are used interchangeably and are broadly used to refer to a system configured to process AV bitstreams (or video data and / or encoded audio data of AV bitstreams). Examples of AV processing units include, but are not limited to, encoders (e.g., codecs), decoders, codecs, NBMP entities, splicers, preprocessing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). In one example, the AV processing unit is a splicer configured to determine the outgoing point (i.e., time) of a first transport stream and the incoming point (i.e., another time) of a second transport stream (which may be the first transport stream or a different transport stream that is not the first transport stream), and to generate a spliced ​​transport stream (e.g., a spliced ​​transport stream containing data from the first bitstream that occurs before the outgoing point and data from the second bitstream that occurs after the incoming point).

[0053] Throughout this disclosure, and as included in the claims, the expression "metadata" refers to data that is separate from and distinct from the corresponding audio and / or video data (audio and / or video content of a bitstream that also includes metadata). Metadata is associated with audio and / or video data ("AV" data) and indicates at least one characteristic or feature of the AV data (e.g., what type of processing has been performed or should be performed on the AV data or the trajectory of an object indicated by the audio data of the AV data). The association between metadata and AV data is time-synchronized. Therefore, current (most recently received or updated) metadata can indicate that the corresponding AV data simultaneously possesses the indicated characteristics and / or includes the result of audio and / or video data processing of the indicated type.

[0054] Throughout this disclosure, and as included in the claims, the terms "coupled" or "coupled" are used to mean a direct or indirect connection. Thus, if a first device is coupled to a second device, the connection can be achieved either through a direct connection or through an indirect connection via other devices and connections. Detailed Implementation

[0055] Some embodiments of the present invention are methods and systems for generating AV bitstreams, the AV bitstreams comprising video data I-frames (video I-frames) synchronized with encoded audio data I-frames (audio I-frames), wherein the AV bitstreams provide a solution to the alignment problem between segments of the AV bitstream's elementary streams (e.g., a video elementary stream and one or more corresponding audio elementary streams). For example, Figure 1 NBMP Entity 12 of Content Delivery Network Figure 7 The system and Figure 4 Each of the production units 3 is a system configured to generate such AV bitstreams according to an embodiment of the invention.

[0056] Independently decodable frames generated according to some video and audio codec formats and defined as I-frames as above may not be formally defined (in the formal specification of the relevant codec format) or are commonly referred to as "I-frames" (and may instead be referred to by other names, such as IDR frames, IPF frames, or IF frames). In some codecs, each video frame and / or each audio frame is independently decodable, and each such frame is an "I-frame".

[0057] In some cases, the encoded audio data of an audio P-frame (e.g., an audio P-frame whose audio content is AC-4 encoded audio data) can be decoded using information from within the P-frame alone. However, in other cases, the decoded audio may not sound exactly as intended during playback. In some situations, to ensure that the decoded version of the audio content of the P-frame sounds as intended during playback, metadata from the preceding I-frame (e.g., dynamic range control metadata, downmixing metadata, loudness metadata, spectral spread metadata, or coupling metadata) typically needs to be available (and usually needs to be used) during decoding.

[0058] In some cases, P-frames can be correctly decoded by referencing data transmitted in previous frames. Video P-frames use previous data along with transmitted data to generate a predicted version of the video frame's image, and only transmit the differences to the image. Audio P-frames can use differential coding of parameters that can be predicted from previous frames, such as differential coding of the spectral envelope used in shifted data like ASPX or ACPL in AC-4.

[0059] In this document, the terms “adaptation” (AV bitstream, e.g., transport stream or another stream) and “adaptation” (to a bitstream) are used broadly to refer to performing any operation on a bitstream that includes accessing content indicated by the bitstream (e.g., content of one or more elementary streams of a program indicated by the bitstream) where the content appears (corresponding to) a specific time (e.g., the time of an I-frame contained in the bitstream according to an embodiment of the invention). Examples of bitstream adaptation include (but are not limited to): adaptive bitrate switching (e.g., switching between different bitrate versions of content indicated by the bitstream at a specific time in the bitstream); alignment of the elementary streams (audio elementary stream or video and audio elementary streams) of the program indicated by the bitstream; splicing the program indicated by the bitstream with another program (e.g., indicated by another stream) at a specific time; connecting the program indicated by the bitstream (or a segment thereof) with another program at a specific time; or initiating playback of the program indicated by the bitstream (or a segment thereof) at a specific time. The expression "adaptation point" (of an AV bitstream, such as a transport stream or another stream) is used in a broad sense to indicate the time (e.g., as indicated by a timestamp value) when an I-frame (or a video I-frame and at least one corresponding audio I-frame) appears in the bitstream, such that adaptation (e.g., switching or splicing) can be performed on the bitstream at the adaptation point.

[0060] An AV bitstream (e.g., a transport stream) can indicate multiple programs. Each program can contain multiple elementary streams (e.g., a video elementary stream and one or more encoded audio elementary streams). An elementary stream can be a video elementary stream (containing video data) or an audio elementary stream (e.g., corresponding to a video elementary stream) that contains encoded audio data (e.g., encoded audio data output from an audio encoder).

[0061] While the methods and systems according to embodiments of the invention are not limited to the generation and / or splicing (or other adaptation) of transport streams having an MPEG-2 transport stream format (or having an AV bitstream having a format conforming to MPEG-DASH or MMT or any other specific format) or containing encoded audio data of any specific format, some embodiments are methods and systems for generating and / or adapting (e.g., splicing) MPEG-2 transport streams (or AV bitstreams having a format conforming to MPEG-DASH or MMT), wherein the audio frames of the MPEG-2 transport stream contain encoded audio data in a format known as AC-4 (or MPEG-DUSAC or MPEG-H audio format). If each such AV bitstream (e.g., bitstream segment) contains P-frames and I-frames of video data and / or P-frames and I-frames of encoded audio data, then according to other embodiments of the invention, transport streams having other formats (or AV bitstreams that are not transport streams, e.g., having an adaptive streaming format or bitstream segments used in streaming methods or protocols) can be generated and / or adapted (e.g., switched or spliced). Frames (e.g., audio frames) of AV bitstreams (e.g., transport streams) generated and / or adapted (e.g., spliced) according to a typical embodiment of the present invention typically also include metadata.

[0062] The adaptive streaming protocol known as MPEG-DASH allows media content to be streamed over the Internet using HTTP web servers. According to MPEG-DASH, the media content of a media presentation (which also includes metadata) is divided into sequences of segments. These segments are organized into representations (each representation contains a sequence of segments, for example, containing...). Figure 5 The media presentation consists of two sequences of "MPEG-DASH segments" as indicated in the document, an adapter (each adapter contains a representation sequence), and a period (each period contains an adapter sequence). A media presentation can contain many periods.

[0063] While media presentation segments can contain any media data, the MPEG-DASH specification provides guidelines and formats for use with two types of media data containers: ISO basic media file formats (e.g., MPEG-4 file format) or MPEG-2 transport stream formats. Therefore, in some cases, media content (and metadata) in MPEG-DASH format can be delivered as a sequence of MPEG-4 files rather than as a single bitstream.

[0064] MPEG-DASH is agnostic to audio / video codecs. Therefore, the media content of an MPEG-DASH media presentation may include encoded audio content (e.g., AC-4 encoded audio content) encoded according to one audio encoding format and additional encoded audio content encoded according to another audio encoding format, as well as video content encoded according to one video encoding format and additional video content encoded according to another video encoding format.

[0065] Data in MPEG-DASH format can contain one or more representations of a multimedia file (e.g., versions with different resolutions or bit rates). Selection among such representations can be implemented based on network conditions, device capabilities, and user preferences, thereby enabling adaptive bit rate streaming.

[0066] Figure 5 This is a diagram illustrating an instance of a segment of an AV bitstream (as defined herein) with an MPEG-DASH adaptive streaming format. Although AV bitstreams (or Figure 5 The segments shown may not be delivered as a single bitstream (e.g., with a single bit rate), but AV bitstreams (and) Figure 5 The fragment shown is an instance of an "AV bitstream" as defined herein. Figure 5 The media content of the shown segment includes encoded audio data encoded as an AC-4 audio frame sequence according to the AC-4 encoding standard, and corresponding video content encoded as a video frame sequence. The AC-4 audio frames include audio I-frames (e.g., AC-4 frame 1 and AC-4 frame 7) and audio P-frames (e.g., AC-4 frame 2). The video frames include video I-frames (e.g., video frame 1 and video frame 5) and video P-frames. Figure 5 In this context, each video frame marked with "P" (indicating "predictive" coding) or "B" (indicating "bidirectional predictive" coding) is a video "P-frame" as defined herein. Figure 5 The video P-frame contains video frame 2 and video frame 3.

[0067] Figure 5 Each video I-frame is aligned with its corresponding audio I-frame (e.g., AC-4 frame 7, which is an audio I-frame, is aligned with video frame 5, which is a video I-frame). Therefore, GOPs are labeled GOP 1 (where GOP stands for "group of pictures"), GOP 2, GOP 3, and GOP 4. Figure 5Each segment in the AV bitstream begins at a time of an "adaptation point" (as defined herein) in the AV bitstream. Therefore, adaptation (e.g., switching) can be performed seamlessly on the AV bitstream at any of these adaptation points. Segments GOP 1 and GOP 2 together correspond to a longer segment of the AV bitstream (labeled "MPEG-DASH Segment 1"), and segments GOP 3 and GOP 4 together correspond to a longer segment of the AV bitstream (labeled "MPEG-DASH Segment 2"), and each of these longer segments also begins at an adaptation point.

[0068] The MPEG-2 transport stream format is a standard format for transmitting and storing video data, encoded audio data, and related data. The MPEG-2 transport stream format is defined in the standard known as MPEG-2 Part 1 – Systems (ISO / IEC Standard 13818-1 or ITU-T REC.H.222.0). MPEG-2 transport streams have a defined container format for encapsulating packet elementary streams.

[0069] MPEG-2 transport streams are typically used to broadcast audio and video content, for example, in the form of DVB (Digital Video Broadcasting) or ATSC (Advanced Television Systems Committee) TV broadcasts. It is generally expected that splicing will be performed between two MPEG-2 transport streams.

[0070] MPEG-2 transport streams carry elementary streams (i.e., data containing indicators of the elementary streams) within data packets (e.g., an elementary stream of video data output from a video encoder, and at least one corresponding elementary stream of encoded audio data output from an audio encoder). Each elementary stream is encapsulated by encapsulating consecutive data bytes from the elementary streams within packet elementary stream (“PES”) data packets with a PES header. Typically, elementary stream data (output from the video and audio encoders) is encapsulated as PES packets, which are then encapsulated within transport stream (TS) packets, and the TS packets are then multiplexed to form the transport stream. Typically, each PES packet is encapsulated as a sequence of TS packets. PES packets may indicate audio or video frames (e.g., audio frames including metadata and encoded audio data, or video frames including metadata and video data).

[0071] At least one (or more) TS packets in a (transport stream) TS packet sequence indicating an audio frame may contain metadata necessary for decoding the audio content of that frame. At least one (or more) TS packets in a TS packet sequence indicating an audio I-frame contain metadata sufficient to independently decode the audio content of that frame. Although at least one (or more) TS packets in a TS packet sequence indicating an audio P-frame contain metadata necessary for decoding the audio content of that frame, additional metadata (from the preceding I-frame) is typically required to decode the audio content of that frame.

[0072] An MPEG-2 transport stream can indicate one or more audio / video programs. Each individual program is described by a Program Map (PMT), which has a unique identifier value (PID), and one or more elementary streams associated with the program have PIDs listed in the PMT. For example, a transport stream could indicate three television programs, each corresponding to a different television channel. In an example, each program (channel) could consist of one video elementary stream and multiple (e.g., one or two) encoded audio elementary streams, along with any necessary metadata. A receiver wishing to decode a particular program (channel) decodes the payload of each elementary stream associated with the program using the PID.

[0073] An MPEG-2 transport stream contains Program Specific Information (PSI), which typically includes data indicating four PSI tables: the Program Association Table (PAT), the Program Map Table (PMT) for each program, the Conditional Access Table (CAT), and the Network Information Table (NIT). The PAT lists all programs indicated by the transport stream (contained within it), and each program within a program has an associated PID value for its PMT. The PMT for each program lists each primary stream of the program and contains data indicating other information about the program.

[0074] An MPEG-2 transport stream includes a Presentation Timestamp (“PTS”) value, which is used to synchronize the individual elementary streams (e.g., video stream and encoded audio stream) of the program in the transport stream. The PTS value is given in units related to the overall clock reference of the program, which is also transmitted in the transport stream. All TS packets, including audio or video frames (indicated by PES packets), have the same PTS (timestamp) value.

[0075] AV bitstreams (e.g., MPEG-2 transport streams or AV bitstreams in MPEG-DASH format) may contain encoded audio data (typically compressed audio data indicating one or more audio content channels), video data, and metadata indicating at least one feature of the encoded audio (or encoded audio and video) content. While embodiments of the invention are not limited to the generation of AV bitstreams (e.g., transport streams) whose audio content is audio data encoded according to AC-4 format (“AC-4 encoded audio data”), typical embodiments are methods and systems for generating and / or adapting AV bitstreams containing AC-4 encoded audio data (e.g., bitstream segments used in adaptive streaming formats).

[0076] The AC-4 format, used for encoding audio data, is well-known and was released in April 2014 in a document entitled "ETSITS 103 190V1.1.1(2014-04), Digital Audio Compression (AC-4) Standard".

[0077] MPEG-2 transport streams are commonly used to broadcast audio and video content, for example, in the form of DVB (Digital Video Broadcasting) or ATSC (Advanced Television Systems Committee) TV broadcasts. Sometimes it is desirable to perform splicing between two MPEG-2 (or other) transport streams. For example, it might be desirable for a transmitter to perform splicing within a first transport stream to insert an advertisement (e.g., a segment of another stream) between two segments of the first transport stream. Conventional systems known as transport stream splicers are used to perform such splicing. Conventional splicers vary in complexity, and conventional transport streams are typically generated under the assumption that the splicer knows and can understand all the codecs contained within (i.e., can parse its video and encoded audio content as well as metadata) in order to perform splicing on it. This leaves a large margin of error in the implementation of splicing and leads to numerous interoperability problems between the multiplexer (which performs multiplexing to generate transport streams or other AV bitstreams) and the splicer.

[0078] Typical embodiments of the present invention may be contained within network (e.g., content delivery network or “CDN”) elements configured to deliver media (including audio and video) content to end users. Network-based media processing (NBMP) is a framework that allows service providers and end users to describe media processing operations to be performed by such a network. An example of NBMP is the MPEG-I Part 8 Network-based Media Processing Framework under development. Some embodiments of the present invention’s devices are considered to be implemented as NBMP entities programmed or otherwise configured according to embodiments of the invention (e.g., MPEG NBMP entities, which may be CDN servers or may be included in CDN servers) (e.g., optionally inserting or removing I-frames from AV bitstreams in a manner dependent on network and / or other constraints). NBMP describes the composition of network-based media processing services from a set of network-based media processing functions and makes these network-based media processing services accessible via an application programming interface (API). NBMP media processing entities (NBMP entities) perform media processing tasks on the media data (and associated metadata) input therein. NBMP also provides control functions for writing and configuring media processing.

[0079] Figure 1 This is a block diagram of a content delivery network, wherein NBMP entity 12, according to an embodiment of the invention, is configured to generate an AV bitstream containing at least one remade frame (and deliver it to playback device 16). In a typical embodiment, the AV bitstream indicates at least one audio / video program (“program”) and contains (for each program thus indicated) video data frames (which define at least one video elementary stream) and corresponding encoded audio data frames (which define at least one audio elementary stream). The video data frames contain video data I-frames (video I-frames) and video data P-frames (video P-frames), and the encoded audio data frames contain encoded audio data I-frames (audio I-frames) and encoded audio data P-frames (audio P-frames). At least one frame of the AV bitstream (at least one video frame in the video frames and / or at least one audio frame in the audio frames) has been remade by entity 12 (as a remade I-frame or a remade P-frame), such that a content segment containing the remade frame begins with an I-frame and contains at least one subsequent P-frame. In some embodiments, at least one audio frame (but no video frame) of the AV bitstream has been remade by entity 12 (as a remade audio I-frame or a remade audio P-frame), such that an audio content segment containing the remade frame begins with an audio I-frame (aligned with the video I-frame of the corresponding video content segment) and contains at least one subsequent audio P-frame.

[0080] exist Figure 1In this embodiment, bitstream source 10 is configured to generate an input AV bitstream (e.g., a standards-compliant transport stream, such as a transport stream compatible with a version of an MPEG standard (e.g., MPEG-2 standard), another transport stream, or an input AV bitstream that is not a transport stream), said input AV bitstream comprising video frames (video data and corresponding metadata) and audio frames (e.g., encoded audio data compatible with an audio format such as AC-4 and metadata). In some embodiments of source 10, the input AV bitstream is generated in a conventional manner. In other embodiments, source 10 generates the input AV bitstream according to an embodiment of the AV bitstream generation method of the present invention (which includes remaking at least one frame).

[0081] It should be understood that this has been taken into consideration. Figure 1 Many implementations (and variations) of the system. For example, in some implementations, elements 10, 12, 14, and 16 are implemented in different devices or systems coupled to the network (at different locations). In one such implementation: source 10 (or a device or subsystem coupled thereto) is configured to package the content generated by source 10 (e.g., encoded audio data, or both video data and encoded audio data) and metadata (e.g., as an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol) for delivery over a network (e.g., delivery to NBMP entity 12); and NBMP entity 12 (or a device or subsystem coupled thereto) is configured to package the content generated by entity 12 (containing at least one remade frame) (e.g., encoded audio data, or both video data and encoded audio data) and metadata (e.g., as an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol) for delivery over a network (e.g., delivery to playback device 16).

[0082] In another instance, in some implementations, elements 10 and 12, and optionally 14, are implemented as a single device or system coupled to the network (e.g., performing...). Figure 3(The production unit 3 is a production unit with the function of production unit 3). In one such implementation: source 10 is configured to provide NBMP entity 12 with frames containing content (e.g., encoded audio data, or both video data and encoded audio data) and metadata; NBMP entity 12 is configured (according to an embodiment of the invention) to re-encode at least one frame of the frames, thereby generating a modified set of frames; and NBMP entity 12 (or a subsystem coupled thereto) is configured to package the modified set of frames (e.g., as an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol) for delivery over a network (e.g., to playback device 16).

[0083] In another instance, elements 10, 12, and 14 are configured to generate and process only audio content (and corresponding metadata), comprising reproducing at least one frame containing encoded audio data by means of an embodiment of the invention (in element 12). Optionally, a video processing subsystem is employed. Figure 1 (Not shown) to generate and process only video content (and corresponding metadata), for example, including at least one frame containing video data remade by means of an embodiment of the invention. An apparatus or subsystem may be coupled to element 12 (and coupled to the video processing subsystem) and configured to package (e.g., as an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol), for example, for delivery over a network (e.g., to playback apparatus 16).

[0084] In some implementations, NBMP entity 14 is omitted (and the described functions implemented by entity 14 are instead performed by NBMP entity 12).

[0085] In a first example embodiment, NBMP entity 12 may be implemented as a single device configured to perform both audio processing (including audio frame re-creation) and video processing (including video frame re-creation), and a second NBMP entity coupled to this device may package the audio output (e.g., raw AC-4 bitstream or other data stream) and video output of entity 12 into an AV bitstream (e.g., an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol). In a second example embodiment, NBMP entity 12 may be implemented as a single device configured to perform the functions of both entity 12 and the second NBMP entity in the first example embodiment. In the third example embodiment, NBMP entity 12 is implemented as a first device configured to perform audio processing (including audio frame re-creation), another NBMP entity is implemented as a second device configured to perform video processing (including video frame re-creation), and a third NBMP entity coupled to both devices can package the audio output of the first device (e.g., a raw AC-4 bitstream or other data stream) and the video output of the second device into an AV bitstream (e.g., an AV bitstream having a format conforming to MPEG-DASH, or HLS, or MMT, or another adaptive streaming format or method / protocol).

[0086] In some embodiments, source 10 and NBMP entity 12 (or source 10, NBMP entity 12, and a packing system configured to pack the output of entity 12 into an AV bitstream) are implemented at an encoding / packing facility, and entity 12 (or the packing system coupled thereto) is configured to output an AV bitstream, and playback device 16 is located remotely from this facility. In some other embodiments, source 10 (and a packing system coupled and configured to pack the output of source 10 into an AV bitstream) are implemented at an encoding / packing facility, and NBMP entity 12 and playback device 16 are located together at a remote playback location, such that NBMP 12 is coupled and configured to generate an AV bitstream containing at least one remade frame in response to an AV bitstream delivered from the encoding / packing facility to the playback location. In any of these embodiments, NBMP entity 14 may be implemented at a location remote from NBMP entity 12, or at the same location as NBMP 12, or entity 14 may be omitted (and entity 12 may be implemented to perform the functions described for entity 14). Figure 1 Other variations of the system can be used to implement other embodiments of the present invention.

[0087] Next, for Figure 1 The example implementation of the system will be described in more detail.

[0088] Typically, the input AV bitstream generated by source 10 indicates at least one audio / video program (“program”) and contains video data frames (which define at least one video elementary stream) and encoded audio data frames corresponding to the video data (which define at least one audio elementary stream). The input AV bitstream may include a sequence of I-frames (video I-frames) containing video data, I-frames (audio I-frames) containing encoded audio data, P-frames (video P-frames) containing video data, and P-frames (audio P-frames) containing encoded audio data, or it may contain a sequence of video I-frames and video P-frames, and a sequence of audio I-frames. The audio I-frames of the input AV bitstream may appear at a first rate (e.g., once every time interval X, such as...). Figure 1 (As indicated).

[0089] The input AV bitstream may have a packaged media delivery format (e.g., a standardized format, such as, but not limited to, a format conforming to MPEG-DASH, HLS, or MMT methods / protocols, or an MPEG-2 transport stream format), which requires that each segment of the AV bitstream's content (audio or video content) begins with an I-frame to achieve seamless adaptation (e.g., splicing or switching) at segment boundaries (e.g., at the beginning of a segment of video content in a video elementary stream (indicated by a video I-frame), the beginning being aligned with the start time of a segment of corresponding audio content in at least one audio elementary stream (indicated by an audio I-frame)). In some cases, the audio content of the input AV bitstream is time-aligned with the corresponding video content of the input AV bitstream. In some other cases, the audio content of the input AV bitstream may not be time-aligned with the corresponding video content of the input AV bitstream. In this sense, a video I-frame may appear at some time in the stream (e.g., as indicated by a timestamp value) but no audio I-frame may appear simultaneously in the stream, or an audio I-frame may appear at some time in the stream (e.g., as indicated by a timestamp value) but no video I-frame may appear simultaneously in the stream. For example, the audio content of the stream may consist only of audio I-frames (e.g., when NBMP entity 12 is configured to generate a transport stream or other AV bitstream in response to an input AV bitstream and in accordance with an embodiment of the second type of embodiment of the AV bitstream generation method described below).

[0090] The input AV bitstream output from source 10 passes through Figure 1The network delivery is sent to a Network-Based Media Processing (NBMP) entity 12, which may be implemented as (or included in) a CDN server. Considered in some embodiments, the NBMP entity 12 is implemented according to a standardized format. In one instance, such a standard may be an MPEG-compliant standard, and the NBMP entity 12 may be implemented as an MPEG NBMP entity, for example, according to the MPEG-I Part 8 Network-Based Media Processing framework or any one or more other future versions of this framework. The NBMP entity 12 has an input 12A coupled to receive an input AV bitstream, a buffer 12B coupled to the input 12A, a processing subsystem 12C coupled to the buffer 12B, and a packetizing subsystem 12D coupled to the subsystem 12C. Fragments of the AV bitstream delivered to the input 12A are buffered (stored in a non-transitory manner) in the buffer 12B. The buffered fragments are asserted from the buffer 12B to the processing subsystem 12C (processor). Subsystem 12C, according to embodiments of the invention, is programmed (or otherwise configured) to perform re-encapsulation of at least one frame of the input AV bitstream in a manner dependent on network and / or other constraints (thereby generating at least one re-encapsulated frame), and to pass any unmodified media content and metadata of the input AV bitstream to subsystem 12D. The media content and metadata output from subsystem 12C (e.g., the original frame of the input AV bitstream and each re-encapsulated frame generated in subsystem 12C) are provided to packaging subsystem 12D for packaging into an AV bitstream. Subsystem 12D is programmed (or otherwise configured) to generate an AV bitstream (to be delivered to playback device 16) to include each of the re-encapsulated frames (e.g., including by inserting at least one re-encapsulated I-frame into the AV bitstream, and / or replacing at least one I-frame of the input AV bitstream with a re-encapsulated P-frame), and to perform any other necessary packaging (including any necessary metadata generation) to generate an AV bitstream having the desired format (e.g., MPEG-DASH format). Network constraints (and / or other constraints) are indicated by control bits, which are transmitted via... Figure 1 The network originates from another NBMP entity (e.g., Figure 1 The NBMP entity 14) delivers the control bit to the NBMP entity 12. Examples of such constraints (e.g., constraints indicating the available bit rate for delivering a bit stream over the network) will be described below. The NBMP entity 12 is configured to perform control functions in a manner subject to control bit constraints delivered from the entity 14 to write and configure media processing operations (including frame re-creation) required to perform the implemented embodiments of the method of the present invention.

[0091] Figure 1Each of the source 10, NBMP entity 12, NBMP entity 14, and playback device 16 is an audio / video processing unit (“AVPU”) as defined herein. Each of the source 10, NBMP entity 12, and playback device 16 can be configured to perform adaptation (e.g., switching or splicing) on ​​the AV bitstream thus generated or provided to it, thereby generating an adapted AV bitstream. Embodiments of each of the source 10, NBMP entity 12, NBMP entity 14, and playback device 16 include a buffer memory and at least one audio / video processing subsystem coupled to the buffer memory, wherein the buffer memory stores at least one segment of the AV bitstream in a non-transitory manner, the AV bitstream having been generated by embodiments of the inventive method for generating the AV bitstream.

[0092] NBMP entity 12 is coupled and configured to generate an AV bitstream (in response to an input AV bitstream delivered from source 10) according to an embodiment of the AV bitstream generation method of the present invention, and to deliver the generated AV bitstream (e.g., asserted to a network for delivery over the network) to playback device 16. Playback device 16 is configured to perform playback of the audio and / or video content of the AV bitstream. Playback device 16 includes input 16A, buffer 16B coupled to input 16A, and processing subsystem 16C coupled to buffer 16B. Segments of the AV bitstream delivered to input 16A are buffered (stored in a non-transitory manner) in buffer 16B. These buffered segments are asserted from buffer 16B to processing subsystem 16C (processor). Subsystem 16C is configured to parse the AV bitstream and perform any necessary decoding on the encoded audio and / or encoded video content of the AV bitstream, and playback device 16 may include a display for displaying the parsed (or parsed and decoded) video content of the AV bitstream. Playback device 16 (e.g., subsystem 16C, or Figure 1 Another subsystem of device 16 (not specifically shown in the figure) may also be configured to present parsed (and decoded) audio content of the AV bitstream to generate at least one speaker feed, and optionally include at least one speaker for emitting sound in response to each such speaker feed.

[0093] As described above, some implementations of the playback device 16 are configured to decode the audio and / or video content of an AV bitstream generated according to embodiments of the present invention. Therefore, these implementations of the device 16 are examples of decoders configured to decode the audio and / or video content of an AV bitstream generated according to embodiments of the present invention.

[0094] Typically, the AV bitstream output from NBMP entity 12 has a packed media delivery format (e.g., a format conforming to MPEG-DASH, HLS, or MMT methods / protocols, or an MPEG-2 transport stream format), which requires that each segment of the AV bitstream's content (audio or video content) begins with a I-frame. Typically, each segment of the video content of the AV bitstream's video elementary stream begins with a video I-frame (which may be a re-enhanced video I-frame generated according to embodiments of the invention), the video I-frame being time-aligned with the start of at least one segment of the corresponding audio content (indicated by an audio I-frame, which may be a re-enhanced audio I-frame generated according to embodiments of the invention) of at least one segment of the audio content in at least one audio elementary stream of the AV bitstream.

[0095] For example, in some embodiments, entity 12 is configured to reconstruct audio P-frames of the input AV bitstream as audio I-frames as needed. For example, entity 12 can identify video I-frames in the input AV bitstream and determine which audio I-frames need to be time-aligned with the video I-frames. In this case, entity 12 can reconstruct the audio P-frames of the input AV bitstream (time-aligned with the video I-frames) as audio I-frames and insert the reconstructed audio I-frames into the AV bitstream to be delivered to playback device 16 in place of the audio P-frames.

[0096] For another instance, entity 12 can be configured to recognize that audio I-frames (or audio I-frames and video I-frames) of the input AV bitstream appear at a first rate (e.g., once every time interval X, such as...). Figure 1 As indicated), however, the adapter points (e.g., splice points) to be delivered to the AV bitstream of playback device 16 need to appear at a higher rate (e.g., once every time interval Y, such as...). Figure 1 As indicated, where Y is greater than X). In this case, entity 12 can recreate the audio P-frames (or audio P-frames and video P-frames time-aligned with the audio P-frames) of the input AV bitstream as audio I-frames (or audio I-frames and video I-frames time-aligned with the audio I-frames), and insert the recreated audio I-frames (or each set of time-aligned recreated audio I-frames and recreated video I-frames) into the AV bitstream to be delivered to playback device 16 instead of the original P-frames, thereby providing an adaptation point in the AV bitstream at the desired rate (Y).

[0097] Other embodiments of the method of the present invention (NBMP entity 12 may be configured to perform these embodiments) will be described below.

[0098] Figure 2 This is a diagram of an example of an MPEG-2 transport stream. Figure 1 Some embodiments of NBMP entity 12 (and Figure 3 Some embodiments of the production unit 3 described below are configured to generate an MPEG-2 transport stream, comprising reproducing at least one video or audio frame.

[0099] As described above, according to embodiments of the present invention, generation (e.g., through...) Figure 1 NBMP entity 12 or Figure 3 The AV bitstream (e.g., transport stream) of production unit 3) may indicate at least one audio / video program (“program”) and contains (for each program thus indicated) video data frames (which define at least one video elementary stream) and corresponding encoded audio data frames (which define at least one audio elementary stream). Video data frames (e.g., Figure 2 Video frames #1, #2, #3, #4, #5, and #6 can contain video data I-frames (e.g., Figure 2 The video I-frames #1 and #5), and the encoded audio data frames (e.g., identified as Figure 2 The audio frames of AC-4 frames #1, #2, #3, #4, #5, and #6 can contain encoded audio data I-frames (e.g., Figure 2 The audio I-frames AC-4 frames #1, AC-4 frames #3, and AC-4 frames #5.

[0100] In order to generate Figure 2 The MPEG-2 transport stream, audio frames are packaged into PES packets (in... Figure 2 (Seen in a magnified version in the top row), and the video frames are packaged into PES packets (in... Figure 2 (Seen in enlarged version in the bottom row). Each PES packet in the PES packet indicating an audio frame has one of a different PTS value of 0, 3600, 7200, 10800, 14400, and 18000, and each PES packet in the PES packet indicating a video frame has one of a different PTS value of 0, 3600, 7200, 10800, 14400, and 18000.

[0101] Each PES packet is packaged into a set of Transport Stream (TS) packets, and the MPEG-2 transport stream includes an indicator sequence of TS packets (in... Figure 2 (As shown in the middle row). The transport stream splicer that processes the transport stream may splice at the positions marked with bars S1 and S2, each of which occurs exactly before the video I-frame and therefore does not interfere with the audio. To simplify the example, Figure 2 All frames indicated in the text are PES packet aligned (even between I frames where PES packet alignment is not required).

[0102] Figure 2 The transport stream has (i.e., satisfies) the property of I-frame synchronization (i.e., video and audio encoding are synchronized, such that for each program indicated by the transport stream, for each video I-frame in the video elementary stream of the program, there exists at least one matching audio I-frame in the audio elementary stream of the program (i.e., at least one audio I-frame synchronized with the video I-frame)). In this sense, for each program indicated by the transport stream, the data of the transport stream indicating the program has the property of I-frame synchronization. A transport stream having this property (e.g., a transport stream generated according to an embodiment of the invention) can (e.g., through...) Figure 3 7) The splicer can seamlessly splice or otherwise adapt any audio base stream of any program without modifying the transport stream.

[0103] Figure 3 This is a block diagram of an example of an audio processing chain (audio data processing system), wherein one or more of the system's elements can be configured according to embodiments of the invention. The system includes the following elements coupled together as shown: a capture unit 1, a production unit 3 (including an encoding subsystem), a delivery subsystem 5, and a splicing unit (splicer) 7, and optionally a capture unit 1', a production unit 3' (including an encoding subsystem), and a delivery subsystem 5'. In variations of the illustrated system, one or more of the elements are omitted, or additional processing units are included (e.g., subsystem 5' is omitted, and the output of unit 3' is delivered by subsystem 5 to unit 7).

[0104] Capture unit 1 is typically configured to generate PCM (temporal) samples and video data samples that include audio content, and output PCM audio samples and video data samples. For example, PCM samples can indicate multiple audio streams captured by a microphone. Production unit 3, typically operated by a broadcaster, is configured to accept PCM audio and video samples as input and generate and output AV bitstreams indicating the audio and video content. Figure 3 In some implementations of the system, production unit 3 is configured to output an MPEG-2 transport stream (e.g., its audio content is encoded according to the AC-4 standard such that each audio primary stream of the MPEG-2 transport stream includes an MPEG-2 transport stream with compressed audio data in AC-4 format).

[0105] The encoding performed on the audio content of the AV bitstream (e.g., MPEG-2 transport stream) generated according to any of the various embodiments of the present invention may be AC-4 encoding, or it may be any other audio encoding aligned with the video (i.e., such that each frame of the video corresponds to an integer (i.e., non-fractional) frame of encoded audio (AC-4 encoding may be performed to have this latter property).

[0106] The AV bitstream output from Unit 3 may include an encoded (e.g., compressed) audio bitstream (sometimes referred to herein as the “master mix”) indicating at least some of the audio content in the audio content and a video bitstream indicating the video content, and optionally at least one additional bitstream or file indicating some of the audio content in the audio content (sometimes referred to herein as the “sub-mix”). The data of the AV bitstream indicating the audio content (and the data of each generated sub-mix, if generated) is sometimes referred to herein as “audio data”.

[0107] The audio data of an AV bitstream (e.g., its master mix) can indicate one or more speaker channels, and / or the audio sample stream can indicate object channels.

[0108] like Figure 4 As shown, Figure 3 An implementation of production unit 3 includes an encoding subsystem 3B coupled to receive video and audio data from unit 1. Subsystem 3B is configured to perform necessary encoding on the audio data (and optionally on the video data) to generate encoded audio data frames and optionally encoded video data frames. Figure 4 The re-creation and multiplexing subsystem 3C of unit 3 has: an input 3D, which is coupled to receive frames (e.g., audio frames, or audio and video frames) of data output from subsystem 3B; and a subsystem (processor) coupled to the input and configured to re-create one or more frames output from subsystem 3B as needed according to embodiments of the invention, and to package (including by packetization and multiplexing) the output of subsystem 3B (or the frames output by subsystem 3B, and each re-created frame generated therefrom for replacing one of the frames output from subsystem 3B) into an AV bitstream according to embodiments of the invention. For example, the AV bitstream may be an MPEG-2 transport stream, the audio content of which may be encoded according to the AC-4 standard, such that each audio elementary stream of the MPEG-2 transport stream includes compressed audio data having AC-4 format. A buffer 3A is coupled to the output of subsystem 3C. Fragments of the AV bitstream generated in subsystem 3C are buffered (stored in a non-transitory manner) in buffer 3A. Because the output of subsystem 3C of unit 3 is an AV bitstream containing the encoded audio content generated in unit 3 (and typically also video content packaged with it), unit 3 is an instance of an audio encoder. The AV bitstream generated by unit 3 is typically (after being buffered in buffer 3A) asserted to a delivery system (e.g., Figure 3 Delivery subsystem 5).

[0109] Figure 3The delivery subsystem 5 is configured to store and / or transmit (e.g., broadcast or deliver via network) the transmission bit stream generated by unit 3 (e.g., including each of its submixes if any submixes are generated).

[0110] The capture unit 1', production unit 3' (including buffer 3A'), and delivery subsystem 5' respectively perform the functions of capture unit 1, production unit 3, and delivery subsystem 5 (and are generally the same as them). It can operate to generate (and deliver to input 8B of splicer 7) a second AV bitstream (e.g., a transport stream or other AV bitstream generated according to an embodiment of the invention), which is spliced ​​with a first AV bitstream (e.g., a first transport stream or other AV bitstream) via splicer 7. The first AV bitstream is generated in production unit 3 (e.g., according to an embodiment of the invention) and delivered to input 8A of splicer 7.

[0111] Figure 3 The splicer 7 includes inputs 8A and 8B. Input 8A is coupled to receive (e.g., read) at least one AV bit stream delivered to the splicer 7 by the delivery subsystem 5, and input 8B is coupled to receive (e.g., read) at least one AV bit stream delivered to the splicer 7 by the delivery subsystem 5'. The splicer 7 also includes, as Figure 3 The splicer 7 is shown to be coupled with buffer memories (buffers) 7A, 7D, parsing subsystem 7E, parsing subsystem 7B, and splicing subsystem 7C. Optionally, the splicer 7 includes a memory 9, which is coupled (as shown) and configured to store the AV bitstreams to be spliced. During typical operation of the splicer 7, fragments of at least one selected AV bitstream received at inputs 8A and / or 8B (e.g., a sequence of fragments of a selected sequence of AV bitstreams received at inputs 8A and 8B) are buffered (stored in a non-transitory manner) in buffers 7A and / or 7D. The buffered fragments are asserted from buffer 7A to parsing subsystem 7B for parsing, and from buffer 7D to parsing subsystem 7E. Optionally, a fragment of at least one AV bitstream stored in memory 9 is asserted to parsing subsystem 7B for parsing (or a fragment of a selected sequence of AV bitstreams stored in memory 9 and / or received at input 8A is asserted from buffer 7A and / or memory 9 to parsing subsystem 7B for parsing). Typically, each AV bitstream to be parsed (in subsystem 7B or 7E) and spliced ​​(in splicing subsystem 7C) has been generated according to embodiments of the invention.

[0112] The splicer 7 (e.g., its subsystems 7B and 7E and / or subsystem 7C) is also coupled and configured to determine splicing points in each AV bitstream to be spliced ​​(e.g., a first transport stream delivered to the splicer 7 by the delivery subsystem 5 and / or a second transport stream delivered to the splicer 7 by the delivery subsystem 5', or a first transport stream stored in the memory 9 and / or a second transport stream delivered to the splicer 7 by the delivery subsystem 5 or 5'), and subsystem 7C is configured to splice one or more streams to generate at least one spliced ​​AV bitstream. Figure 3 (The "spliced ​​output"). In some cases, splicing omits segments of a single AV bitstream, and splicer 7 is configured to determine the out point (i.e., time) and in point (later time) of the AV bitstream, and generate a spliced ​​AV bitstream by concatenating stream segments that appear before the out point with stream segments that appear after the in point. In other cases, a second AV bitstream is inserted between segments of the first AV bitstream (or between segments of the first and third AV bitstreams), and splicer 7 is configured to determine the exit point (i.e., time) of the first AV bitstream, the entry point (later time) of the first (or third) AV bitstream, the entry point (i.e., time) of the second AV bitstream, and the exit point (later time) of the second AV bitstream, and generate a spliced ​​AV bitstream containing data of the first AV bitstream that appears before the exit point of the stream, data of the second AV bitstream that appears between the entry and exit points of the stream, and data of the first (or third) AV bitstream that appears after the entry point of the first (or third) AV bitstream.

[0113] In some implementations, splicer 7 is configured to splice one or more AV bitstreams according to embodiments of the splicing method of the present invention, wherein at least one of the one or more AV bitstreams has been generated according to embodiments of the present invention to generate at least one spliced ​​AV bitstream. Figure 3 The "stitched output" of the invention. In such embodiments of the stitching method of the invention, each stitching point (e.g., an in point or an out point) appears at an audio I-frame of an audio content segment (which may be a re-created frame generated according to an embodiment of the invention), which is aligned with a video I-frame of a corresponding video content segment (which may be a re-created frame generated according to an embodiment of the invention). Typically, each such audio content segment contains at least one audio P-frame following the audio I-frame, and each such video content segment contains at least one video P-frame following the video I-frame.

[0114] Some embodiments of the present invention relate to an adaptation (e.g., splicing or switching) of an AV bitstream (e.g., an MPEG-2 transport stream) generated according to any embodiment of the present invention's method for generating AV bitstreams, thereby generating an adapted (e.g., spliced) AV bitstream (e.g., configured to perform such splicing methods). Figure 3 The method of outputting the splicer 7 in the embodiment. In a typical embodiment, adaptation is performed without modifying any encoded audio base stream of the AV bitstream (but in some cases, the adapter may need to perform AV bitstream re-multiplexing or other modifications in a manner that does not include data that modifies any encoded audio base stream of the AV bitstream).

[0115] Typically, the playback system decodes and presents the spliced ​​AV bitstream output from splicer 7. A playback system typically includes a subsystem for parsing the audio and video content of the AV bitstream, a subsystem configured to decode and present the audio content, and another subsystem configured to decode and present the video content.

[0116] Figure 7 This is a block diagram of a system for generating AV bitstreams in MPEG-DASH format according to an embodiment of the present invention. Figure 7 In the system, the audio encoder 20 is coupled to receive audio data (e.g., from...) Figure 4 (An embodiment of unit 1 or another capture unit), and video encoder 21 is coupled to receive video data (e.g., from...). Figure 4 (An embodiment of unit 1 or another capture unit). Encoder 20 is coupled and configured to perform audio encoding (e.g., AC-4 encoding) on ​​audio data to generate encoded audio data frames and assert the frames to DASH packer 22. Encoder 21 is coupled and configured to perform video encoding (e.g., H.265 video encoding) on ​​video data to generate encoded video data frames and assert the frames to DASH packer 22.

[0117] Packer 22 is an audio / video processing unit that includes an I-frame conversion subsystem 24, a segmenter 28, an audio analyzer 26, an MPEG-4 multiplexer 30, and an MPD generation subsystem 32, which are coupled as shown.

[0118] Segmenter 28 is a video processing subsystem (processor) programmed (or otherwise configured) to determine segments of video content (video frames) provided from encoder 21 to segmenter 28 and to provide the segments to MPEG-4 multiplexer 30. Typically, each segment in a segment begins with a video I-frame.

[0119] I-frame conversion subsystem 24 is an audio processing subsystem (processor) that, according to embodiments of the invention, is programmed (or otherwise configured) to rework at least one audio frame provided to subsystem 24 from encoder 20 (thereby generating at least one reworked frame) and pass any frames from its unmodified frames (output from encoder 20) to audio analyzer 26. Typically, subsystem 24 performs rework to ensure the existence of an audio I-frame (e.g., a reworked I-frame) aligned with each video I-frame identified by segmenter 28. Audio content and metadata output from subsystem 24 (i.e., the original audio frame output from encoder 20 and each reworked audio frame generated in subsystem 24) are provided to audio analyzer 26.

[0120] Audio frames output from subsystem 24 (e.g., the original audio frames from encoder 20 and each reconstructed audio frame generated in subsystem 24) are provided to audio analyzer 26. Typically, the audio frames contain AC-4 encoded audio data. Analyzer 26 is configured to analyze the frame's metadata (e.g., AC-4 metadata) and use the results of the analysis to generate any new metadata it determines needs to be packaged with the audio frames (in MPEG-4 multiplexer 30). Audio frames output from subsystem 24 and any new metadata generated in analyzer 26 are provided to MPEG-4 multiplexer 30.

[0121] The MPEG-4 multiplexer 30 is configured to multiplex audio frames (and any new metadata generated in the analyzer 26) output from the analyzer 26 and video frames output from the segmenter 28 to generate multiplexed audio and video content (and metadata) in MPEG-4 file format for inclusion in an AV bitstream in MPEG-DASH format.

[0122] MPD generation subsystem 32 is coupled to receive multiplexed content output from multiplexer 30 and is configured to generate an AV bitstream in MPEG-DASH format (and instructing MPEG-DASH media presentation) containing the multiplexed content and metadata.

[0123] In some embodiments used to generate AV bitstreams (e.g., AV bitstreams in MPEG-H or MPEG-D USAC format), the step of remaking an audio P-frame into a remade audio I-frame is performed as follows: At least one previous audio P-frame is copied into the P-frame (e.g., in the extended payload of the P-frame). Thus, the resulting remade I-frame contains the encoded audio content of each previous P-frame, and the resulting remade I-frame contains a different set of metadata than the P-frame (i.e., it contains the metadata of each previous P-frame as well as the original metadata of the P-frame). In one instance, this copying of at least one previous P-frame into a P-frame can be performed using the AudioPreRoll() syntax element described in Section 5.5.6 (“Audio Pre-Roll”) of the MPEG-H standard. Considering this, Figure 7 Some implementations of subsystem 24 perform P-frame re-creation (as re-creation of I-frames) in this manner.

[0124] In some embodiments, the method of the present invention includes: (a) providing a frame (to an encoder, for example, Figure 4 (An embodiment of subsystem 3C of production unit 3) or NBMP entity (e.g., Figure 1 (An embodiment of NBMP entity 12) for example, transmitting, delivering, or otherwise providing a frame or an input transport stream or other input AV bit stream containing a frame, or generating a frame (e.g., in...) Figure 4 In the embodiment of production unit 3), each frame in the frame indicates audio content or video content, and the frame contains a frame of a first decoding type; and (b) generating an AV bitstream (e.g., via...). Figure 4 The embodiment of unit 3 operates on the frame thus generated, or through Figure 1 An embodiment of NBMP entity 12 (operation on input AV bitstream frames provided thereto) includes remaking at least one frame of a first decoding type into a remade frame of a second decoding type different from the first decoding type (e.g., if one of the frames of the first decoding type is a P-frame, then the remade frame is an I-frame, or if one of the frames of the first decoding type is an I-frame, then the remade frame is a P-frame), such that the AV bitstream includes a content segment containing the remade frame, and the content segment begins with an I-frame and includes at least one P-frame following the I-frame. For example, step (b) may include remaking at least one audio P-frame such that the remade frame is an audio I-frame, or remaking at least one audio I-frame such that the remade frame is an audio P-frame, or remaking at least one video P-frame such that the remade frame is a video I-frame, or remaking at least one video I-frame such that the remade frame is a video P-frame. In some such embodiments, step (a) includes the following steps: in a first system (e.g., Figure 4 Production unit 3 or Figure 1 In the embodiment of source 10, an input AV bitstream containing a frame containing an indication content is generated; and a second system (e.g., Figure 1 (In an embodiment of NBMP entity 12) the input AV bit stream is delivered (e.g., transmitted); and step (b) is performed in the second system.

[0125] In some such embodiments, at least one audio P-frame (containing metadata and encoded audio data) is remade into an audio I-frame (containing the same encoded audio data and different metadata). The encoded audio data may be AC-4 encoded audio data. For example, in some embodiments of the invention, the audio P-frame is remade into a remade audio I-frame by replacing all or some of the metadata (which may be in the header of the audio P-frame) with a different set of metadata (e.g., different sets of metadata consisting of or containing metadata copied from or generated by modifying metadata obtained from a previous audio I-frame), without modifying the encoded audio content of the frame (e.g., without modifying the encoded audio content by decoding such content and then re-encoding the decoded content). In some other embodiments, at least one audio I-frame (containing metadata and encoded audio data) is remade into an audio P-frame (containing the same encoded audio data and different metadata). This encoded audio data may be AC-4 encoded audio data.

[0126] The AC-4 encoded audio data of an audio P-frame (whose audio content is AC-4 encoded audio data) can be decoded using information solely from within the P-frame. However, in some cases (i.e., when the encoding assumes that spectral extension metadata and / or coupling metadata from the preceding I-frame are available for decoding the encoded audio data), the decoded audio may not sound exactly as originally intended during encoding. In these cases, to ensure that the decoded version of the P-frame's encoded audio content sounds as expected during playback, the spectral extension metadata (and / or coupling metadata) from the preceding I-frame typically needs to be available (and usually needs to be used) during decoding. Instances of such spectral extension metadata are ASPX metadata, and instances of such coupling metadata are ACPL metadata.

[0127] Therefore, in some embodiments of the invention, a re-encoding of an audio P-frame (e.g., an audio P-frame whose audio content is AC-4 encoded audio data) is performed by copying metadata from a previous audio I-frame and appropriately inserting the copied metadata to replace all or some of the original metadata in the original metadata of the P-frame, in order to generate a re-encoded audio I-frame (without modifying the audio content of the P-frame). This type of re-encoding is typically performed when the original encoding of the P-frame does not require both specific metadata from the previous I-frame (e.g., spectral spread metadata and / or coupling metadata) and corresponding metadata from the P-frame itself to be available for decoding the encoded audio data of the P-frame (and thus the re-encoded I-frame) in a manner that does not result in an unacceptable alteration of the original intended sound during playback of the decoded audio. However, in cases where the original encoding of a P-frame (e.g., an audio P-frame whose audio content is AC-4 encoded audio data) truly requires both specific metadata from the previous I-frame (e.g., spectrum spread metadata and / or coupling metadata) and corresponding metadata from the P-frame itself to be available for decoding the encoded audio data of the P-frame (in a way that does not unacceptably alter the original intended sound when playing back the decoded audio), remaking a P-frame (as a remade audio I-frame) according to some embodiments of the invention comprises the following steps: saving specific metadata from the previous audio I-frame (e.g., spectrum spread metadata and / or coupling metadata); typically modifying the saved metadata using at least some of the original metadata from the original metadata of the P-frame (thereby generating modified metadata, which, when included in the remade I-frame, is sufficient to enable decoding of the encoded audio data of the P-frame (and therefore the I-frame) using only information from within the remade I-frame); and inserting the modified metadata in place of all or some of the original metadata from the original metadata of the P-frame. Similarly, upon consideration, re-creating a video P-frame (as a re-created video I-frame) according to some embodiments of the invention comprises the following steps: saving specific metadata from the previous video I-frame; modifying the saved metadata, for example, using at least some of the original metadata from the original metadata of the P-frame (thereby generating modified metadata, which, when included in the re-created I-frame, is sufficient to enable decoding of the video content of the P-frame (and thus the I-frame) using only information from within the re-created I-frame); and inserting the modified metadata to replace all or some of the original metadata from the original metadata of the P-frame.

[0128] Alternatively, the encoding of each P-frame (which can be remade according to embodiments of the invention) is performed in such a way that metadata from another frame (e.g., a previous I-frame) can be simply copied from the other frame (without modifying the copied metadata) into the P-frame to replace the original metadata of the P-frame (thus remaking the P-frame into an I-frame), making it possible to decode the content (audio or video content) of the P-frame (and thus the remade I-frame) in a way that would indeed result in unacceptable perceptual differences (as intended at the time of encoding) when playing back the decoded content. For example, AC-4 encoding of the audio data (of the frame that can be remade according to embodiments of the invention) can be performed without including spectral spread metadata and / or coupling metadata in the resulting audio frame. This allows the audio P-frame (as already generated in the example) to be remade by copying metadata from a previous audio I-frame and appropriately inserting the copied metadata to replace all or some of the original metadata in the original metadata of the P-frame, thereby generating a remade audio I-frame (without modifying the audio content of the P-frame). In this example, the need to modify (re-encode) metadata from previous I-frames in the encoder (in order to recreate P-frames) is avoided by not utilizing inter-channel dependencies at the cost of a slightly higher bit rate.

[0129] In a first type of embodiment (sometimes referred to herein as an embodiment implementing "method 1"), the method of the present invention includes the steps of: (a) generating (e.g., in a conventional manner) audio I-frames and audio P-frames (e.g., in...) indicating content. Figure 4 (a) in subsystem 3B of the embodiment of production unit 3; and (b) generating AV bitstream (e.g., in Figure 4 In a subsystem 3C of an embodiment of production unit 3, the method includes remaking at least one of the audio P-frames into a remade audio I-frame, such that the AV bitstream includes a segment containing the content of the remade audio I-frame, and the segment of content begins with the remade audio I-frame. Typically, the segment of content also includes at least one audio P-frame following the remade audio I-frame. Steps (a) and (b) can be performed in an audio encoder (e.g., a production unit that includes or implements an audio encoder), including performing step (a) by operating the audio encoder.

[0130] In some embodiments of method 1, step (a) is performed in the audio encoder that generates the audio I-frame and the audio P-frame (e.g., in an implementation of an AC-4 encoder). Figure 4The step (b) is performed in the subsystem 3B of the production unit 3, and includes remaking at least one audio P-frame (corresponding to the time when an audio I-frame is required) into the remade audio I-frame, and step (b) also includes including the remade audio I-frame instead of the one audio P-frame in the audio P-frame in the AV bitstream.

[0131] Some embodiments of method 1 include the following steps: in a first system (e.g., a production unit or encoder, for example, Figure 4 An embodiment of production unit 3 or Figure 7 In the system), an input AV bitstream containing the audio I-frames and audio P-frames generated in step (a) is generated; and the second system (e.g., Figure 1 The embodiment of NBMP entity 12 delivers (e.g., transmits) an input AV bitstream; and performs step (b) in the second system. Typically, the step of generating the input AV bitstream involves packaging encoded audio content (indicated by the audio I-frames and audio P-frames generated in step (a)) and video content together to generate the input AV bitstream, and the AV bitstream generated in step (b) contains the encoded audio content packaged together with the video content. In some embodiments of implementing method 1, at least one audio frame of the AV bitstream is a re-created audio I-frame (but the video frames of the AV bitstream are not re-created frames), and the audio content segment of the AV bitstream begins with a re-created audio I-frame (aligned with the video I-frame of the corresponding video content segment of the AV bitstream) and contains at least one subsequent audio P-frame.

[0132] In a second type of embodiment (sometimes referred to herein as an embodiment implementing "method 2"), the method of the present invention includes the steps of: (a) generating (e.g., in a conventional manner) an audio I-frame indicating the content (e.g., in... Figure 4 (a) in subsystem 3B of the embodiment of production unit 3; and (b) generating (e.g., in Figure 4 In a subsystem 3C of an embodiment of production unit 3, an AV bitstream is generated by remaking at least one of the audio I-frames into a remade audio P-frame, such that the AV bitstream includes a segment containing the content of the remade audio P-frame, and the segment of content begins with one of the audio I-frames generated in step (a). Steps (a) and (b) can be performed in an audio encoder (e.g., a production unit that includes or implements an audio encoder), including performing step (a) by operating the audio encoder.

[0133] In some embodiments of method 2, step (a) involves the audio encoder that generates the audio I-frame (e.g., implemented as an AC-4 encoder). Figure 4The process is performed in subsystem 3B of production unit 3, and step (b) includes remaking each audio I-frame corresponding to a time other than the segment boundary (and therefore not appearing at the segment boundary), thereby determining at least one remade audio P-frame, and the step of generating the AV bitstream includes the step of including the at least one remade audio P-frame in the AV bitstream.

[0134] Some embodiments of method 2 include the following steps: in a first system (e.g., a production unit or encoder, for example, Figure 4 An embodiment of production unit 3 or Figure 7 In the system), an input AV bitstream containing the audio I-frame generated in step (a) is generated; and the second system (e.g., Figure 1 The embodiment of NBMP entity 12 delivers (e.g., transmits) an input AV bitstream; and performs step (b) in the second system. Typically, the step of generating the input AV bitstream involves packaging encoded audio content (indicated by the audio I-frame generated in step (a)) and video content together to generate the input AV bitstream, and the AV bitstream generated in step (b) contains the encoded audio content packaged together with the video content. In some embodiments, at least one audio frame (but no video frame) of the AV bitstream has been remade (as a remade audio P-frame) such that an audio content segment containing the remade audio P-frame begins with an audio I-frame (aligned with the video I-frame of the corresponding video content segment) and contains at least one remade audio P-frame following the audio I-frame.

[0135] In a third type of embodiment (sometimes referred to herein as an embodiment implementing "Method 3"), the AV bitstream generation method of the present invention includes generating (or providing) a blended frame that allows P-frames to be included in the AV bitstream, containing data "blocks" selected (or otherwise used). As used herein, a data "block" may be data containing one of the blended frames, or data indicating at least one sequence of frames (e.g., two substreams) of frames that are blended, or a predefined portion of said at least one sequence (e.g., in the case where the frame contains AC-4 encoded audio data, the entire substream may include the entire data block). In one instance, Method 3 includes generating a blended frame to ensure that adjacent frames of the AV bitstream match. In a typical embodiment, the encoder generates the blended frame (e.g., the stream indicating the blended frame) without double processing, as many processes only need to be performed once. In some embodiments, the blended frame contains an instance of common data for P-frames and I-frames. In some embodiments, the packer may synthesize an I-frame or P-frame from at least one data block of the blended frame (or a sequence of frames indicating multiple blended frames containing the blended frame), for example, where the block does not include the entire I-frame or P-frame.

[0136] For example, Figure 6 It is a set of ten mixed frames, and Figure 6A This is another set of ten such mixed frames. In a typical embodiment of method 3, the mixed frames (e.g., Figure 6 Mixed frames, or Figure 6A The mixed frames) are encoded by the encoder (e.g., Figure 7 The implementation of the audio encoder 20) generates the packager (e.g., Figure 7 The implementation of the packer 22 generates an AV bitstream, which includes at least one P-frame by selecting at least one of the mixed frames in the mixed frames.

[0137] refer to Figure 6 and 6A In some embodiments, each hybrid frame may include:

[0138] Two I-frames. For example, Figure 6 Each of the mixed frames H1, H5, and H9 contains two I-frames (both labeled "I"); or

[0139] An I-frame. For example, Figure 6A Each of the hybrid frames H'1, H'5, and H'9 consists of an I-frame (labeled "I"); or

[0140] An I-frame (containing encoded audio data and metadata) and a P-frame (containing the same encoded audio data but different metadata). For example, frames other than H1, H5, or H9. Figure 6 Each mixed frame and all frames except H'1, H'5, and H'9 Figure 6A Each hybrid frame contains such I-frames (labeled "I") and such P-frames (labeled "P").

[0141] For mixed frames (e.g., Figure 6 Instance or Figure 6A As shown in the example, when the packer (or other AV bitstream generation system, which may be a subsystem of another system) determines that an I-frame should be included at a certain time (in the generated AV bitstream), the packer can select an I-frame from a blended frame corresponding to the relevant time (which has been previously generated by the encoder and is available for selection). When the system determines that a P-frame should be included at a certain time (in the generated AV bitstream), the system can: select a P-frame from a blended frame corresponding to the relevant time (if the blended frame contains a P-frame); or remake the I-frame of such a blended frame into a remade P-frame. In some embodiments, remaking may involve copying metadata from another blended frame or modifying metadata obtained from a previous frame.

[0142] In some embodiments of an encoder (or other system or apparatus) that generates mixed frames containing AC-4 encoded audio (e.g., mixed frames of AC-4 encoded audio substreams) according to method 3, the encoder (or other system or apparatus) may include the following types of metadata in each mixed frame (or at least some mixed frames):

[0143] ASF (Audio Spectrum Front End) metadata. If a constant bit rate is not required, the entire `sf_info` and `sf_data` portions are identical in both the I-frame and P-frame of the mixed frame. If a constant bit rate is required, the `sf_data` portion of the I-frame can compensate for the I-frame size overhead and can be made smaller so that the overall size of the I-frame is the same as the corresponding P-frame. In both cases, the `sf_info` portion is identical to ensure a perfect window shape match;

[0144] ASPX (Spectrum Extension) metadata. The I-frame of a hybrid frame contains an aspx_config that matches the aspx_config of the P-frame of the hybrid frame. The aspx_data of the I-frame uses only intra-frame coding, while the aspx_data of the P-frame can use either intra-frame or inter-frame coding. This typically does not impose additional overhead on the encoder, as the encoder usually performs both methods to select the one with the highest bit rate efficiency; and

[0145] ACPL (coupled) metadata. The I-frame of the hybrid frame contains an acpl_config that matches the acpl_config of the P-frame of the hybrid frame. The acpl_framing_data is the same in both P-frames and I-frames. All instances of the acpl_ec_data of the I-frame are restricted to diff_type = DIFF_FREQ (intra-frame coding only).

[0146] In some embodiments, an encoder (or other system or apparatus) that generates a mixed frame containing AC-4 encoded audio (e.g., a mixed frame of AC-4 encoded audio substreams) according to method 3 can produce two frame types (e.g., an I-frame stream and a corresponding P-frame stream) in a single process. For parametric encoding tools like ASPX or ACPL, inter-frame coding and intra-frame coding are generated from the same set of data. Some components (e.g., general components) are configured to generate two sets of data (a set of I-frames and a set of corresponding P-frames), and later decide which uses fewer bits. Intra-frame encoded audio is always contained within a mixed frame containing I-frames.

[0147] In the encoder (or other system or device) that generates hybrid frames according to method 3, other analysis tools (e.g., the frame generator in ASPX or the block switching decision-maker in ASF) can be run only once. For perfect matching, the results and decisions are used in both the I-frame and the corresponding P-frame.

[0148] Bit rate and buffer control can also be run only once (when generating P-frame streams that allow inter-frame coding). The results can then be used for all I-frame streams generated in the same way. To prevent audio quality degradation, the overhead of I-frames can be considered when determining the target bits for ASF coding used for P-frames.

[0149] The hybrid frame sequence generated for implementing method 3 (e.g., determined by an I-frame sequence and a corresponding P-frame sequence) may have an AC-4 metadata TOC (Table of Contents) format or some other wrapper that combines I-frames and corresponding P-frames in a single stream.

[0150] In another embodiment of the invention, each frame in the hybrid frame sequence contains a header, which may be independent of the underlying media format and contains a description of how to generate I-frames or P-frames from the data provided in such hybrid frames. The description may contain one or more commands for synthesizing I-frames by copying data ranges or by deleting data ranges from the hybrid frame. The description may also contain one or more commands for synthesizing P-frames by copying data ranges or by deleting data ranges from the hybrid frame. The packer can then synthesize I-frames or P-frames without knowing the underlying media format by following the instructions in the header of such hybrid frames. For example, a hint track from ISOBMFF can be used.

[0151] In one frame selection implementation, the encoder simultaneously generates a P-frame stream and its corresponding I-frame stream (such that the two streams together determine the mixed frame sequence). If buffer model requirements need to be met, the frames in each pair of corresponding frames (each I-frame and its corresponding P-frame) have equal frame sizes. Whenever an I-frame is needed, the multiplexer selects a frame from all I-frame streams (from the encoder output). If the multiplexing does not require an I-frame, a P-frame can be selected (e.g., from the full P-stream from the encoder). The advantage of the frame replacement implementation is that the multiplexer has minimal complexity. The disadvantage is that the link bandwidth from the encoder to the multiplexer is doubled.

[0152] In another frame selection (block replacement) implementation, the hybrid frame available to the multiplexer comprises an I-frame, a corresponding P-frame (“I-frame replacement” data block), and instructions for I-frame replacement (selecting one or more P-frames to replace one or more corresponding I-frames). This method requires selecting an I-frame replacement data block that is not byte-aligned and may require shifting a large portion of the existing P-frame data during the replacement process.

[0153] In some embodiments, a method for generating an AV bitstream (implementation 3 thereof) includes the following steps:

[0154] (a) Providing frames, wherein at least one of the frames is a hybrid frame comprising P-frames and I-frames, wherein the I-frames indicate an encoded version of the content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and

[0155] (b) Generate an AV bitstream comprising the P-frames by selecting at least one of the mixed frames, and including each selected P-frame in the AV bitstream such that the AV bitstream comprises a segment that begins with an I-frame and includes at least the selected P-frames following the I-frame.

[0156] Step (a) may include generating (e.g., in) Figure 7 In an embodiment of encoder 20, a first I-frame sequence (e.g., a full I-frame stream) indicates an encoded version of the content and a second P-frame sequence (e.g., a full P-frame stream) indicates a different encoded version of the content, and at least one of the mixed frames contains one of the I-frames of the first sequence and one of the P-frames of the second sequence.

[0157] In some embodiments, a method for generating an AV bitstream (implementation 3 thereof) includes the following steps:

[0158] (a) Providing frames, wherein at least one of the frames is a hybrid frame containing at least one data block that can be used to determine P-frames and I-frames, wherein the I-frames indicate an encoded version of content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and

[0159] (b) Generating an AV bitstream, comprising synthesizing at least one I-frame or P-frame using at least one of the data blocks of at least one of the mixed frames (the at least one data block may or may not include an entire I-frame or P-frame), thereby generating at least one composite frame, and including each of the composite frames in the AV bitstream such that the AV bitstream includes a segment that begins with an I-frame and includes at least one composite P-frame following the I-frame, or begins with a composite I-frame and includes at least one P-frame following the composite I-frame. An example of step (b) includes synthesizing at least one P-frame (or I-frame) from at least one data block of a mixed frame, or from at least one data block indicating a frame sequence of at least two mixed frames. In some embodiments, at least one of the mixed frames contains at least one instance of a common data block (for P-frames and I-frames).

[0160] Some embodiments of the present invention's method for generating AV bitstreams (e.g., including re-encapsulating at least one frame of the input AV bitstream) ensure that the AV bitstream satisfies at least one current potential network constraint or other constraint (e.g., the generation of the AV bitstream includes re-encapsulating at least one frame of the input bitstream and is performed such that the AV bitstream satisfies at least one current potential network constraint on the input bitstream). For example, when the AV bitstream is generated by an NBMP entity (e.g., implemented as an MPEG NBMP entity...). Figure 1 When an NBMP entity 12 (e.g., as a CDN server or an entity contained within a CDN server) is generated, the NBMP entity can be implemented to insert or remove I-frames from the AV bitstream in a manner dependent on network and / or other constraints. For example, network constraints (and / or other constraints) can be indicated by control bits derived from another NBMP entity (e.g., Figure 1 NBMP entity 14) delivery (e.g., via Figure 1 The network) to the NBMP entity (e.g., Figure 1 (NBMP entity 12). Examples of such constraints include, but are not limited to, available bit rate, the time required to load into a program, and / or the segment duration of a potential MPEG-DASH or MMT AV bitstream.

[0161] In one exemplary embodiment of this type, the input AV bitstream (e.g., by...) Figure 1 The input transport stream generated by source 10 has a bit rate R (e.g., R = 96 kilobits / second) and contains adaptation points (e.g., splicing points). For example, adaptation splicing points are determined by the occurrence of video I-frames and corresponding audio I-frames of the input AV bitstream every 2 seconds or at some other rate. Embodiments of the method of the present invention (e.g., by operating...) Figure 1 The NBMP entity 12) generates an AV bitstream in response to an input AV bitstream. If the available bit rate for delivering the generated AV bitstream is a bit rate R (e.g., from...), then... Figure 1 The NBMP entity 14 delivers the data to the NBMP entity 12 as indicated by the bit. Furthermore, if the generation of the AV bitstream involves remaking one or more P-frames of the input AV bitstream to insert new adaptation points into the generated AV bitstream (e.g., by remaking the P-frames of the input stream into I-frames with more metadata bits than the P-frames, and including the I-frames in the generated stream instead of the P-frames), this insertion of new adaptation points will undesirably increase the bit rate required to deliver the generated AV bitstream unless compensation is taken. Therefore, in the exemplary embodiment, the generation of the AV bitstream (e.g., by operating...) Figure 1The NBMP entity 12) further includes the following steps: re-encapsulating at least one I-frame of the input AV bitstream as a P-frame (with fewer metadata bits than the I-frame), and including each re-encapsulated P-frame in the generated AV bitstream in place of each I-frame, thereby reducing the bit rate required to deliver the generated AV bitstream so that it does not exceed the available bit rate R.

[0162] In another exemplary embodiment, the input AV bitstream (e.g., by...) Figure 1 The input AV bitstream generated by source 10 indicates the audio / video program and includes adaptation points that appear at a first rate (e.g., determined by the occurrence of a video I-frame and a corresponding audio I-frame of the input AV bitstream every 100 milliseconds). Because the adaptation point is the available time to start playing back the program, it corresponds to a consumer selectable (e.g., by operating...). Figure 1 The playback device 16) starts the playback of the program at the "fetch" time (rewind / fast forward point). If the generation of the AV bitstream is constrained (e.g., from...), Figure 1 The NBMP 14 or playback device 16 is delivered to Figure 1 If the generated AV bitstream includes adapter points appearing at a predetermined rate (e.g., once every 50 milliseconds) greater than a first rate (i.e., such that the generated AV bitstream has more adapter points than the input AV bitstream) or at a predetermined rate (e.g., once every 200 milliseconds) less than the first rate (i.e., such that the generated AV bitstream has fewer adapter points than the input AV bitstream), then an embodiment of the method of the present invention is performed (e.g., by operating...). Figure 1 The NBMP entity 12) generates an AV bitstream in response to the input AV bitstream subject to this constraint. For example, if an adaptation point occurs once every 100 milliseconds in the input AV bitstream (i.e., the video I-frame and the corresponding audio I-frame of the input AV bitstream occur once every 100 milliseconds), then to increase the rate at which the adaptation point occurs in the generated AV bitstream, the video P-frames and audio P-frames of the input AV bitstream are remade into video I-frames and audio I-frames, and the remade I-frames are included in the generated AV bitstream in place of the P-frames, such that the adaptation point occurs more than once every 100 milliseconds in the generated AV bitstream (i.e., at least one such adaptation point occurs at the time corresponding to the remade video I-frame and the corresponding remade audio I-frame).

[0163] Figure 3 Units 3 and 7 Figure 4 Unit 3 and Figure 1 Each of the NBMP entity 12 and the playback device 16 can be implemented as a configured hardware system or as a processor programmed (or otherwise configured) with software or firmware to perform embodiments of the method of the present invention.

[0164] generally, Figure 3 Unit 3 contains at least one buffer 3A. Figure 3 Unit 3' contains at least one buffer 3A'. Figure 3 The splicer 7 includes at least one buffer (7A and / or 7D), and Figure 1 Each of the NBMP entity 12 and playback device 16 includes at least one buffer. Typically, each of these buffers (e.g., buffers 3A, 3A', 7A, and 7D, and the buffers of entity 12 and device 16) is a buffer memory coupled to receive a sequence of data packets of an AV bitstream generated (or provided to) the device containing the buffer memory, and in operation, the buffer memory stores (e.g., in a non-transitory manner) at least one segment of the AV bitstream. In typical operation of unit 3 (or 3'), the sequence of segments of the AV bitstream is asserted from buffer 3A to delivery subsystem 5 (or from buffer 3A' to transmission subsystem 5'). In typical operation of splicer 7, the sequence of segments of the AV bitstream to be spliced ​​is asserted from buffer 7A of splicer 7 to parsing subsystem 7B, and from buffer 7D of splicer 7 to parsing subsystem 7E.

[0165] Figure 3 ( Figure 4 Unit 3 and / or Figure 3 splicer 7 and / or Figure 1 The NBMP entity 12 and / or device 16 (or any component or element thereof) may be implemented in hardware, software or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits).

[0166] Some embodiments of the present invention relate to a processing unit (AVPU) configured to perform any embodiment of the methods of the present invention for generating or adapting (e.g., splicing or switching) AV bitstreams. For example, the AVPU may be an NBMP entity (e.g., Figure 1 NBMP entity 12) or production unit or audio encoder (e.g., Figure 3 or Figure 4 Unit 3). For another instance, the AVPU (e.g., an NBMP entity) can be an adapter (e.g., a splicer) configured to perform any embodiment of the AV bitstream adaptation method of the present invention (e.g., ...). Figure 3 An embodiment of a properly configured splicer 7). In another type of embodiment of the invention, the AVPU (e.g., Figure 3 Unit 3 or splicer 7) contains at least one buffer memory (e.g., Figure 3 Buffer 3A in unit 3, or Figure 3The buffer 7A or 7D in the splicer 7, or Figure 1 The source 10, or the NBMP entity 12, or the buffer memory in the device 16, wherein the buffer memory stores (e.g., in a non-transitory manner) at least one segment of an AV bitstream that has been generated by any embodiment of the method of the present invention. Instances of an AVPU include, but are not limited to, encoders (e.g., code converters), NBMP entities (e.g., NBMP entities configured to generate AV bitstreams and / or perform adaptations thereto), decoders (e.g., decoders configured to decode the contents of an AV bitstream and / or perform adaptations (e.g., splicing) of an AV bitstream to generate an adapted (e.g., spliced) AV bitstream and decode the contents of the adapted AV bitstream), codecs, AV bitstream adapters (e.g., splicers), preprocessing systems (preprocessors), postprocessing systems (postprocessors), AV bitstream processing systems, and combinations of such elements.

[0167] Embodiments of the present invention relate to a system or apparatus configured (e.g., programmed) to perform any embodiment of the methods of the present invention, and a computer-readable medium (e.g., a disk) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the methods of the present invention or steps thereof. For example, the system of the present invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data (including embodiments of the methods of the present invention or steps thereof). Such a general-purpose processor may be or include a computer system comprising input devices, memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the methods of the present invention (or steps thereof) in response to data asserted thereto.

[0168] In some embodiments, the device of the present invention is an audio encoder (e.g., an AC-4 encoder) configured to perform an embodiment of the method of the present invention. In some embodiments, the device of the present invention is an NBMP entity (e.g., an MPEG NBMP entity, which may be a CDN server or may be included in a CDN server) programmed or otherwise configured to perform an embodiment of the method of the present invention (e.g., optionally inserting or removing I-frames from the transport stream in a manner dependent on network and / or other constraints).

[0169] Embodiments of the present invention can be implemented in hardware, firmware, or software, or a combination thereof (e.g., as a programmable logic array). For example, Figure 3 Unit 3 and / or splicer 7 or Figure 1The bitstream source 10 and / or NBMP entity 12 can be implemented in appropriately programmed (or otherwise configured) hardware or firmware, for example, as a programmed general-purpose processor, digital signal processor, or microprocessor. Unless otherwise stated, the algorithms or processes included as part of embodiments of the invention are not inherently associated with any particular computer or other device. Specifically, various general-purpose machines can be used with programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized devices (e.g., integrated circuits) for performing the desired method steps. Therefore, embodiments of the invention can be implemented in one or more computer programs that execute on one or more programmable computer systems (e.g., implementing...). Figure 3 Unit 3 and / or splicer 7 or Figure 1 The programmable computer system comprises, respectively, all or some of the elements of source 10 and / or NBMP entity 12, at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0170] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logic-oriented, or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled or interpreted language.

[0171] For example, when implemented by a sequence of computer software instructions, the various functions and steps of embodiments of the present invention can be implemented by a multi-threaded sequence of software instructions running in suitable digital signal processing hardware, in which case the various means, steps and functions of the embodiments can correspond to portions of the software instructions.

[0172] Each such computer program is preferably stored or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., solid-state memory or media, or magnetic or optical media) to configure and operate the computer when the storage medium or device is read by a computer system to perform the processes described herein. The system of the present invention can also be implemented as a computer-readable storage medium configured (i.e., storing) a computer program, wherein such a storage medium causes the computer system to operate in a specific and predefined manner to perform the functions described herein.

[0173] Exemplary embodiments of the present invention include the following:

[0174] E1. A method for generating an AV bitstream, the method comprising the following steps:

[0175] Provides frames indicating content, said frames comprising frames of a first decoding type, wherein each frame indicates audio content or video content; and

[0176] Generating an AV bitstream includes remaking at least one frame of the first decoding type into a remade frame of a second decoding type different from the first decoding type, such that the AV bitstream includes segments containing the content of the remade frames, and the segments of content begin with an I-frame and include at least one P-frame following the I-frame.

[0177] If one of the frames in the first decoding type is a P-frame, then the remade frame is a remade I-frame; or if one of the frames in the first decoding type is an I-frame, then the remade frame is a remade P-frame.

[0178] E2. According to the method in E1, the step of providing the frame includes the following steps:

[0179] In the first system, an input AV bitstream containing the frame is generated; and

[0180] The input AV bitstream is delivered to the second system, and

[0181] The step of generating the AV bitstream is performed in the second system.

[0182] E3. The method according to E2, wherein the second system is a network-based media processing (NBMP) entity.

[0183] E4. The method according to any one of E1 to E3, wherein the remade frame is an audio I-frame, the frame of the first decoding type is an audio P-frame containing metadata, and the remaking step includes replacing at least some of the metadata in the metadata of the audio P-frame with different metadata copied from the previous audio I-frame, such that the remade frame contains the different metadata.

[0184] E5. The method according to any one of E1 to E3, wherein the reconstructed frame is an audio I-frame, the frame of the first decoding type is an audio P-frame containing metadata, and the reconstructing step comprises the following steps:

[0185] Modified metadata is generated by modifying metadata from previous audio I-frames; and

[0186] Replace at least some of the metadata in the metadata of the audio P-frame with the modified metadata, such that the reconstructed frame contains the modified metadata.

[0187] E6. The method according to any one of E1 to E3, wherein the remade frame is an audio I-frame, the frame of the first decoding type is an audio P-frame, and the remaking step comprises copying at least one previous P-frame into the audio P-frame.

[0188] E7. The method according to any one of E1 to E6, wherein the step of providing a frame includes the step of generating an audio I-frame and an audio P-frame indicating encoded audio content, the re-creation step includes the step of re-creating at least one of the audio P-frames into a re-created audio I-frame, and the segment of the content of the AV bitstream is a segment of encoded audio content that begins with the re-created audio I-frame.

[0189] E8. According to the method described in E7, wherein the steps of generating the audio I-frame and the audio P-frame and the step of generating the AV bitstream are performed in an audio encoder.

[0190] E9. The method according to E7 or E8, wherein the step of providing the frame includes the following steps:

[0191] In the first system, an input AV bitstream containing the audio I-frames and the audio P-frames is generated; and

[0192] The input AV bitstream is delivered to the second system, and

[0193] The step of generating the AV bitstream is performed in the second system.

[0194] E10. The method according to E9, wherein the second system is a network-based media processing (NBMP) entity.

[0195] E11. The method according to any one of E1 to E10, wherein the step of providing a frame includes the step of generating an audio I-frame indicating encoded audio content, the re-creation step includes the step of re-creation at least one of the audio I-frames into a re-creation audio P-frame, and a segment of the content of the AV bitstream is a segment of encoded audio content that begins with one of the audio I-frames and includes the re-creation audio P-frame.

[0196] E12. According to the method of E11, wherein the steps of generating the audio I-frame and generating the AV bitstream are performed in an audio encoder.

[0197] E13. The method according to E11 or E12, wherein the step of providing the frame includes the following steps:

[0198] In the first system, an input AV bitstream containing the audio I-frame is generated; and

[0199] The input AV bitstream is delivered to the second system, and

[0200] The step of generating the AV bitstream is performed in the second system.

[0201] E14. The method according to E13, wherein the second system is a network-based media processing (NBMP) entity.

[0202] E15. The method according to any one of E1 to E14, wherein the step of generating the AV bitstream is performed such that the AV bitstream satisfies at least one network constraint.

[0203] E16. The method according to E15, wherein the network constraint is the available bitrate of the AV bitstream, or the maximum time to tune to the program, or the maximum allowed segment duration of the AV bitstream.

[0204] E17. The method according to any one of E1 to E16, wherein the step of generating the AV bitstream is performed such that the AV bitstream satisfies at least one constraint, wherein...

[0205] The constraint is that the AV bitstream includes adaptation points that occur at a predetermined rate, and each adaptation point is the occurrence time of a video I-frame of the AV bitstream and at least one corresponding audio I-frame of the AV bitstream.

[0206] E18. The method according to any one of E1 to E17, wherein the AV bitstream is an MPEG-2 transport stream or based on the ISO basic media format.

[0207] E19. The method of E18, wherein each frame in the frame indicating audio content contains encoded audio data in AC-4 format. E20. A method for adapting (e.g., splicing or switching) an AV bitstream to generate an adapted (e.g., spliced) AV bitstream, wherein the AV bitstream has been generated by the method of E1.

[0208] E21. The method according to E20, wherein the AV bitstream has an adaptation point at which an I-frame of the content segment is aligned with an I-frame of a corresponding content segment of the AV bitstream, and the AV bitstream is adapted (e.g., spliced) at the adaptation point.

[0209] E22. A system for generating an AV bitstream, the system comprising:

[0210] At least one input, the at least one input being configured to receive frames indicating content, the frames comprising frames of a first decoding type, wherein each frame indicates audio content or video content; and

[0211] A subsystem, coupled and configured to generate the AV bitstream, includes remaking at least one frame of the first decoding type into a remade frame of a second decoding type different from the first decoding type, such that the AV bitstream includes segments containing the content of the remade frames, and the segments of content begin with an I-frame and include at least one P-frame following the I-frame.

[0212] If one of the frames in the first decoding type is a P-frame, then the remade frame is a remade I-frame; or if one of the frames in the first decoding type is an I-frame, then the remade frame is a remade P-frame.

[0213] E23. The system according to E22, wherein the frame indicating the content is contained in an input AV bitstream that has been delivered to the system.

[0214] E24. The system according to E22 or E23, wherein the system is a network-based media processing (NBMP) entity.

[0215] E25. The system according to any one of E22 to E24, wherein the remade frame is an audio I-frame, the frame of the first decoding type is an audio P-frame containing metadata, and the subsystem is configured to perform the remaking by replacing at least some of the metadata of the audio P-frame with different metadata copied from the previous audio I-frame, such that the remade frame contains the different metadata.

[0216] E26. The system according to any one of E22 to E24, wherein the reconstructed frame is an audio I-frame, the frame of the first decoding type is an audio P-frame containing metadata, and the subsystem is configured to perform the reconstruction, comprising:

[0217] Modified metadata is generated by modifying metadata from previous audio I-frames; and

[0218] Replace at least some of the metadata in the metadata of the audio P-frame with the modified metadata, such that the reconstructed frame contains the modified metadata.

[0219] E27. The system according to any one of E22 to E24, wherein the remade frame is an audio I-frame, the frame of the first decoding type is an audio P-frame, and the subsystem is configured to perform the remaking by copying at least one previous P-frame into the audio P-frame.

[0220] E28. The system according to any one of E22 to E27, wherein the subsystem is configured to generate the AV bitstream such that the AV bitstream satisfies at least one network constraint.

[0221] E29. The system according to E28, wherein the network constraint is the available bitrate of the AV bitstream, or the maximum time to tune to a program, or the maximum allowed segment duration of the AV bitstream.

[0222] E30. The system according to any one of E22 to E27, wherein the subsystem is configured to generate the AV bitstream such that the AV bitstream satisfies at least one constraint, wherein...

[0223] The constraint is that the AV bitstream includes adaptation points that occur at a predetermined rate, and each adaptation point is the occurrence time of a video I-frame of the AV bitstream and at least one corresponding audio I-frame of the AV bitstream.

[0224] E31. A method for generating an AV bitstream, the method comprising the following steps:

[0225] (a) Providing frames, wherein at least one of the frames is a hybrid frame comprising P-frames and I-frames, wherein the I-frames indicate an encoded version of the content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and

[0226] (b) Generate an AV bitstream comprising the P-frames by selecting at least one of the mixed frames, and including each selected P-frame in the AV bitstream such that the AV bitstream comprises a segment that begins with an I-frame and includes at least the selected P-frames following the I-frame.

[0227] E32. The method according to E31, wherein step (a) comprises generating a first I-frame sequence of an encoded version of the indicated content and a second P-frame sequence of different encoded versions of the indicated content, and wherein at least one of the mixed frames comprises one of the I-frames of the first sequence and one of the P-frames of the second sequence.

[0228] E33. A system for AV bitstream adaptation (e.g., splicing or switching), comprising:

[0229] At least one input, said at least one input being coupled to receive an AV bit stream, said AV bit stream having been generated by the method according to E1 to E21, E31 or E32; and

[0230] A subsystem, which is coupled and configured to adapt (e.g., splice or switch) the AV bitstream, thereby generating an adapted AV bitstream.

[0231] E34. An audio / video processing unit comprising:

[0232] Buffer memory; and

[0233] At least one audio / video processing subsystem coupled to the buffer memory, wherein the buffer memory stores at least one segment of an AV bitstream in a non-transitory manner, wherein the AV bitstream has been generated by the method according to any one of E1 to E21, E31 or E32.

[0234] E35. The unit according to E34, wherein the audio / video processing subsystem is configured to generate the AV bitstream.

[0235] E36. A computer program product having instructions that, when executed by a processing device or system, cause the processing device or system to perform the method according to any one of E1 to E21, E31 or E32.

[0236] Many embodiments of the invention have been described. It should be understood that various modifications may be considered. In view of the above teachings, many modifications and variations of the present embodiments of the invention are possible. It should be understood that any embodiment of the invention may be practiced in other ways within the scope of the appended claims.

Claims

1. A method for generating an output audio / video bitstream, the method comprising the following steps: (a) Providing frames, wherein at least one of the frames is a hybrid frame comprising P-frames and I-frames, wherein the I-frames indicate an encoded version of the content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and (b) Generating the output audio / video bitstream, comprising selecting at least one of the P-frames of the mixed frame, and including each selected P-frame in the output audio / video bitstream such that the output audio / video bitstream comprises a segment that begins with an I-frame and includes at least the selected P-frames following the I-frame.

2. A method for generating an output audio / video bitstream, the method comprising the following steps: (a) Providing frames, wherein at least one of the frames is a hybrid frame containing at least one data block that can be used to determine P-frames and I-frames, wherein the I-frames indicate an encoded version of content, and the P-frames indicate different encoded versions of the content, and wherein each frame in the frames indicates audio content or video content; and (b) Generating the output audio / video bitstream includes synthesizing at least one I-frame or P-frame by using at least one of the data blocks of at least one of the mixed frames, thereby generating at least one synthesized frame, and including each of the synthesized frames in the output audio / video bitstream such that the output audio / video bitstream includes a segment that begins with an I-frame and includes at least one synthesized P-frame after the I-frame, or begins with a synthesized I-frame and includes at least one P-frame after the synthesized I-frame.