Method and apparatus for personalized audio content
By receiving list files and selecting preselected elements in the HTML5 API environment to personalize the presentation of audio content, the problem of personalized audio content for audio/video experience in the prior art is solved, and efficient and reliable personalized presentation of audio content is achieved.
Patent Information
- Application Number
- CN202080080936.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-24
- Filing Date
- 2020-11-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-18
AI Technical Summary
The prior art is difficult to achieve scalable personalization of audio content for audio/video experience through the HTML5 API in an efficient and reliable manner.
By receiving a manifest file, which contains an adaptive set of reference audio bitstreams and multiple preselected elements, selecting the appropriate preselected elements and causing an audio signal rendering depending on the selected preselected elements.
The personalized presentation of audio content is realized. Users can flexibly mix multiple audio objects according to the selection of different pre-selected elements to form personalized audio signals, which improves the scalability and personalization of the audio experience.
Smart Images

Figure CN114731459B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to the following priority applications: U.S. Provisional Application 62 / 937,883 (Reference No.: D19136USP1) filed on November 20, 2019, and U.S. Provisional Application 63 / 043,179 (Reference No.: D19136USP2) filed on June 24, 2020, which are incorporated herein by reference. Technical field
[0003] This document relates to methods and devices for providing personalized audio signals to a user (especially a listener). Background art
[0004] Modern television (TV) sets enable users to load software applications onto the TV set's platform. This platform can be regarded as a browser, and the application can be a plug - in extension of the browser. For example, the software application can be provided by a content provider, and it can allow the user to select audio and / or video content from the content provider's server.
[0005] A possible context for providing personalized audio and / or video content is the HbbTV (Hybrid Broadcast Broadband TV) environment, the specification of which is ETSI TS 102 796. HbbTV utilizes the HTML5 (HyperText Markup Language) protocol, which includes an Application Programming Interface (API) to enable content providers to provide software applications for new services (e.g., in the context of Video - on - Demand VOD). The HTML5 API specifies a communication interface that allows an application (e.g., an application on a TV set) to communicate with the TV set's browser (also referred to herein as a terminal).
[0006] This document solves the technical problem of achieving scalable personalization of audio content for an audio / video experience in an efficient and reliable manner, especially via the HTML5 API. The independent claims solve this technical problem. Preferred examples are described in the dependent claims. Summary of the invention
[0007] According to one aspect, a device and / or apparatus for personalizing audio content, particularly audio content from an audio / video experience, in particular an application unit or application, is described. The device is configured to receive a manifest file for the audio content (particularly the audio / video experience). The manifest file includes at least one adaptive set of reference audio bitstreams, wherein the audio bitstreams include a plurality of different audio objects. Additionally, the manifest file includes a plurality of different preselected elements for the adaptive set, wherein the different preselected elements specify different combinations of the plurality of audio objects. The device is further configured to select a preselected element from the plurality of different preselected elements. Additionally, the device is configured to cause the rendering of an audio signal depending on the selected preselected element. In particular, metadata contained within the selected preselected element can be used to mix the plurality of audio objects to form the audio signal to be rendered.
[0008] According to another aspect, a device and / or apparatus for personalizing audio content from an audio bitstream, in particular an application unit or application, is described. The device is configured to receive an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes a pointer part having pointers to different bitstream elements of the audio bitstream segment. The device is further configured to identify a gain pointer from the pointer part, the gain pointer pointing to a first bitstream element for the gain and / or position of an audio object of the audio bitstream. The first bitstream element can be any one of the bitstream elements within the element part of the audio bitstream segment. Additionally, the device is configured to modify the value of the first bitstream element (thereby modifying the gain and / or position of the audio object for rendering) before rendering the audio object.
[0009] According to another aspect, a device and / or apparatus for implementing personalization of audio content from an audio bitstream, in particular an encoding device, is described. The device is configured to receive an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes a pointer part having pointers to the element part of the audio bitstream segment. The element part includes a first bitstream element for the gain and / or position of an audio object of the audio bitstream. The first bitstream element can be any one of the bitstream elements within the element part of the audio bitstream segment. Additionally, the device is configured to set an integrity pointer within the pointer part to a predetermined special value, wherein the predetermined special value prevents a decoder from verifying the integrity of the element part of the audio bitstream segment.
[0010] According to another aspect, a device and / or apparatus for personalizing audio content from an audio bitstream is described. The device is configured to receive an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes an element portion having different bitstream elements. The element portion includes a first bitstream element for the gain and / or position of an audio object of the audio bitstream segment. The first bitstream element can be any one of the bitstream elements within the element portion of the audio bitstream segment. The device is further configured to insert a pointer pointing to the first bitstream element into a pointer portion of the audio bitstream segment (thereby enabling modification of the value of the first bitstream element for personalization).
[0011] According to one aspect, a method for personalizing audio content (e.g., audio content from or for an audio / video experience) is described. The method includes receiving a manifest file for audio content to be presented. The manifest file includes at least one adaptive set of reference audio bitstreams, the audio bitstreams including a plurality of audio objects. Additionally, the manifest file includes a plurality of different preselected elements for the adaptive set, wherein the different preselected elements specify different combinations of the plurality of audio objects. Further, the method includes selecting a preselected element from the plurality of different preselected elements, and causing presentation of an audio signal depending on the selected preselected element.
[0012] According to another aspect, a method for personalizing audio content from an audio bitstream is described. The method includes receiving an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes a pointer portion having pointers to different bitstream elements of the audio bitstream segment. Additionally, the method includes identifying a gain pointer from the pointer portion, the gain pointer pointing to a first bitstream element for the gain and / or position of an audio object of the audio bitstream. The method further includes modifying the value of the first bitstream element before presenting the audio object.
[0013] According to another aspect, a method for personalizing audio content from an audio bitstream is described. The method includes receiving an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes a pointer portion having pointers to different bitstream elements of an element portion of the audio bitstream segment, and wherein the element portion includes a first bitstream element for the gain and / or position of an audio object of the audio bitstream segment. The method further includes setting an integrity pointer within the pointer portion to a predetermined special value, wherein the predetermined special value prevents a decoder from verifying the integrity of the element portion of the audio bitstream segment.
[0014] According to another aspect, a method for personalizing audio content from an audio bitstream is described. The method includes receiving an audio bitstream segment of the audio bitstream, wherein the audio bitstream segment includes a pointer part having pointers to different bitstream elements of an element part of the audio bitstream segment. The element part includes a first bitstream element for the gain and / or position of an audio object of the audio bitstream. The method further includes inserting a pointer to the first bitstream element into the pointer part.
[0015] It should be noted that the methods described herein may each be implemented, in whole or in part, in the form of software and / or computer-readable code on one or more processors.
[0016] According to another aspect, a software program is described. The software program may be adapted to be executed on a processor and, when executed on the processor, is used to perform the method steps outlined in this document.
[0017] According to another aspect, a storage medium is described. The storage medium may include a software program that may be adapted to be executed on a processor and, when executed on the processor, is used to perform the method steps outlined in this document.
[0018] According to another aspect, a computer program product is described. The computer program may include executable instructions that, when executed on a computer, are used to perform the method steps outlined in this document.
[0019] It should be noted that the methods and systems outlined in this patent application, including their preferred embodiments, may be used independently or in combination with other methods and systems disclosed in this document. In addition, all aspects of the methods and systems outlined in this patent application may be combined arbitrarily. In particular, the features of the claims may be combined with each other in any way. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be explained in an exemplary manner with reference to the accompanying drawings, in which:
[0021] Figure 1a An example content distribution network is shown;
[0022] Figure 1b An example content of a (HTTP-based Dynamic Adaptive Streaming over HTTP, DASH) manifest (i.e., media presentation description) file is shown;
[0023] Figure 1c An example presentation of audio components or audio objects within an audio bitstream is shown;
[0024] Figure 1d An example initialization segment of an audio bitstream or media element is shown;
[0025] Figure 1e Illustrates an example adaptive set and example preselection for implementing personalized presentation of audio content;
[0026] Figure 2a Illustrates the combined use of different adaptive sets and different preselection elements for implementing scalable personalization of audio content;
[0027] Figure 2b Illustrates example content of a manifest file;
[0028] Figure 3a Illustrates an example portion of an audio bitstream segment or frame;
[0029] Figure 3b Illustrates possible values of bitstream elements;
[0030] Figure 4a Illustrates a flowchart of an example method for providing personalized audio content (e.g., performed by a software application);
[0031] Figure 4b Illustrates a flowchart of an example method for providing personalized content (e.g., performed by a software application);
[0032] Figure 4c Illustrates a flowchart of an example method for implementing personalization of audio content (e.g., performed by a web server); and
[0033] Figure 4d Illustrates a flowchart of an example method for implementing personalization of audio content (e.g., performed by a web server). DETAILED DESCRIPTION
[0034] As described above, this document relates to providing scalable personalized audio content to listeners, particularly using HTML5 and HTML5 APIs. In this context, Figure 1a Illustrates an example content distribution (particularly broadcasting) network 100 having a web server 101 configured to provide audio and / or video content, particularly an audio bitstream 121, to a content receiver 110. The web server 101 may be operated by a content provider.
[0035] The content receiver 110 includes a terminal 111 configured to provide video and / or audio content to a decoder 113 and then to a rendering unit 114 (such as a speaker). Additionally, the content receiver 110 includes an application 112, typically provided by a content provider. The application 112 may be executed on a hardware platform (which may be integrated within a television set). The terminal 111 and the application 112 may communicate with each other via an application programming interface 122 (such as, for example, an HTLM5 API).
[0036] The content receiver 110 may be implemented using a single computing entity (such as a television set), or the content receiver 110 may be implemented within multiple computing entities (such as, for example, the entity of the terminal or browser 111 and a separate entity for the application 112).
[0037] The audio content may be provided from the server 101 to the receiver 110 using the HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) (in particular, MPEG-DASH) protocol. The DASH protocol is an adaptive bitrate streaming scheme that enables the streaming of media (in particular, video and / or audio) content from an HTTP web server 101 over the Internet. The DASH protocol is detailed in ISO / IEC 23009-1:2019 Information technology — Dynamic Adaptive Streaming over HTTP (DASH) — Part 1: Media presentation description and segment formats (see https: / / www.iso.org / standard / 79329.html ), which is incorporated herein by reference.
[0038] The DASH protocol enables the transmission of an audio bitstream 121 (for a media element) from the server 101 to the receiver 110, where the audio bitstream 121 may include multiple different audio components or audio objects (such as, for example, for different languages, for narrative content, for background music content, for audio effect content, etc.). Additionally, the DASH protocol enables the definition of different presentations that specify different combinations of one or more different audio components or audio objects. A presentation may specify
[0039] · one or more audio objects to be jointly presented from multiple different audio objects; and / or
[0040] · how to mix one or more audio objects together for presentation.
[0041] Possible ways to define a presentation or an audio experience are the so-called adaptive sets and / or the so-called preselected elements (such as Figure 1eAs shown). The DASH protocol allows different audio objects (e.g., for different languages) to be assigned to different adaptation sets 180. An adaptation set 180 may include one or more audio components or audio objects 181 (e.g., incorporated within an audio (bit) stream). For example, different adaptation sets 180 may be used to define different sets of audio objects 181, 182 for different listener groups (e.g., for different languages). To reduce the bandwidth required for the audio bit stream 121, the bit stream 121 may include only a subset of the total number of adaptation sets 180 that are available for a particular video and / or audio content (or media element).
[0042] Another way to define a presentation or audio experience is through preselection elements 190. The preselection specifies one or more audio objects 181, 182 (from the adaptation set 180) and a metadata set 191 that specifies how the one or more audio objects 181 are to be mixed together. In particular, the preselection may specify how the one or more audio objects 181, 182 of a single adaptation set 180 are to be mixed together. By providing different metadata sets 191 for different preselection elements 190, different presentations (e.g., different degrees of emphasis on narrative content or music and / or effects content) can be specified in a bitrate-efficient manner.
[0043] The DASH protocol specifies a so-called manifest file, which is an XML file that indicates and describes the different components contained within the audio bit stream 121 or media element. Figure 1b An example manifest file 140 is shown that indicates descriptions 141 for multiple different presentations, where the descriptions 141 for different presentations may be listed within the manifest file 140 according to a particular manifest file order 142. The manifest file 140 may provide a description 141 for each different presentation available within the audio bit stream 121 or media element. The descriptions 141 may be understandable by the user, thus enabling the user to select a particular presentation from the audio bit stream 121 for rendering. For example, the manifest file 140 (specifically, the descriptions 141) may indicate which languages are available and / or which types of mixes of different audio objects 181, 182 are available.
[0044] Figure 1c An example set 150 of audio objects 181 provided within the audio bit stream 121 or provided for a media element (e.g., within the adaptation set 180 of the audio bit stream 121) is illustrated. A presentation 152 can be specified in an efficient manner by separately providing indicators 153 that enable (dashed boxes) or disable (transparent boxes) different audio objects 181.
[0045] Figure 1dShows an example structure of an audio bitstream 121. The audio bitstream 121 may include an initialization segment 160 that specifies different presentations 152 available within the audio bitstream 121. In particular, the initialization segment 160 may include a plurality of presentation parts 161 that indicate the different presentations 152 available. The presentation parts 161 may be provided within the initialization segment 160 according to a specific segment order 162.
[0046] The initialization segment 160 (in particular the different presentation parts 161) may indicate so-called audio track objects, where each audio track object corresponds to a specific presentation 152. Based on the initialization segment 160 and / or based on one or more adaptation sets and / or preselected elements in the manifest file 140, a list of audio track objects for the corresponding presentation 152 list may be generated (by parsing the initialization segment 160). The list of audio track objects may be sorted according to the segment order 162 (which may be different from the manifest file order 142).
[0047] In addition, the audio bitstream 121 generally includes media or bitstream (in particular audio) segments 170 that include one or more audio objects 181. The audio bitstream segments 170 related to a specific presentation 152 (which may also be referred to as media segments) may be indicated by the presentation part 161 for the presentation 152. The audio bitstream segments 170 may correspond to a certain time segment of the audio content (e.g., corresponding to 20 ms of audio content).
[0048] As described above, this document aims to provide a mechanism for a personalized interface for providing audio tracks (i.e., audio objects), particularly in the context of a hybrid broadcast broadband television (HbbTV) environment.
[0049] The term "audio track" (or audio object 181) may refer to an interface representing a single audio track from one of the HTML media elements <audio> or <video>. A possible use of accessing the audio track 181 is to toggle its "enabled" attribute 153 to mute and unmute the track or object 181. For a detailed description, see https: / / html.spec.whatwg.org / multipage / media.html#audiotrack or https: / / developer.mozilla.org / en-US / docs / Web / API / AudioTrack which is incorporated herein by reference. An "audio track object" may be defined as a category defined by the W3C to identify entities that can be selected and / or played on their own.
[0050] A "file audio track" can be a track as defined in Section 3.1.19 of ISO / IEC 14496-12, the content of which is incorporated herein. A "file audio track" contains a sequence of access units consisting of elementary streams, as defined in Section 8.3 of that document. An "initialization segment" 160 can be defined as a sequence of bytes that contains all the initialization information required to decode a bitstream or a sequence of media segments 170, e.g., as detailed in https: / / www.w3.org / TR / 2016 / REC- media-source-20161117 / #init-segment , the content of which is incorporated herein.
[0051] Audio track elements or audio track objects can be used for personalization. Different personalized experiences can be variants derived from a common set 150 of audio objects 181, some of which are turned on or off. For example, if an English version of a documentary can be a music and effects track mixed with English dialogue, a German version can be derived by mixing the same music and effects track with German dialogue.
[0052] Traditionally, the mixing of different personalized experiences would likely occur on a mixing console located in a production studio. Due to advancements in compression technology, next-generation audio codecs are able to directly provide all different audio objects 181 in a single audio bitstream 121 to a receiver 110, which enables users to select and personalize experiences in a more flexible manner to a greater extent.
[0053] The standard for the receiver 110 defines the functionality for distributing and signaling such a multi-component stream 121 to the receiver 110. The receiver 110 can be implemented in a software environment similar to a standardized web browser. This document aims to select a function of an experience (also referred to as a presentation 152 in this document) from several different possible presentations 152, and / or aims to define the possibility of a personal experience.
[0054] As an example, for playback using the HTML5 media element in an HbbTV browser, the W3C specification of HTML5 and the HbbTV specification TS 102 796 V1.4.1 or later (which is incorporated herein by reference) jointly specify an interface capable of discovering and selecting individual presentations 152.
[0055] This document is concerned with implementing user control over narrative importance and improved accessibility, including an extensible way of outlining experiences that can be provided by the functions currently available on a television set to more advanced experiences on future television sets.
[0056] In particular, this document is concerned with combining an adaptive set 180 and preselected elements 190 in a single manifest file 140. In Figure 2aIn one example shown, an adaptive set 281 for standard audio mixing and a different adaptive set 282 for narrative audio mixing can be provided together with a plurality of different preselected elements 291, 292, 293 (and one preselection 291 for standard audio mixing) that provide different narrative mixes.
[0057] The adaptive sets 281, 282 can be defined as separate audio (bit) streams that include one or more audio objects 181, 182. Different audio streams from different adaptive sets 281, 282 include or provide different audio experiences. For example, the adaptive sets 281, 282 can include audio (bit) streams that have a plurality of audio objects 181 or individual objects 181 for multiple different languages, where the individual objects include a "standard" mix and a mix that includes only the basic elements of the narrative. Each adaptive set 281, 282 can include the same language and / or mix in multiple forms. The individual audio objects 181 within the adaptive sets 281, 282 can implement adaptive streaming using different codec settings (e.g., bitrate). For example, the adaptive sets 281, 282 for Spanish can include, for example, two audio objects 181 at 768 kbit / s and 48 kbit / s respectively, and when connected to Wi-fi, the client 110 can decode the audio object 181 at 768 kbit / s, while when connected to a mobile network, the client can decode the audio object 181 at 48 kbit / s.
[0058] In current televisions, application-level selection of audio tracks can be supported as a baseline for presenting experiences 152 from a selection of different experience options 152. The list of available audio tracks is generated by the client 110 (specifically, the application 112) from information 141 contained within the manifest file 140. The list of audio tracks (or audio track objects) can be generated from the manifest information 141 based on the adaptive sets 281, 282 and / or based on the preselected elements 291, 292, 293.
[0059] When application 112 requests media playback (i.e., a media element, in particular the audio bitstream 121 of a media element) from web server 101, application 112 passes the URL of the manifest file 140 of the media element (in particular the audio bitstream 121) to web server 101 for media playback. The manifest file 140 can be downloaded from the location specified by the URL, and the information 141 contained within the manifest file 140 can be parsed. Terminal 111 can be configured to create audio tracks (i.e., audio signals) from the adaptive sets 281, 282 included in the manifest file 140 or from the adaptive sets 281, 282 and the preselected elements 291, 292, 293. Terminal 111 can use interface 122 to provide the created audio tracks to application 112 for selection.
[0060] If only the adaptive sets 281, 282 are used, switching between different predefined experiences 152 from different adaptive sets 281, 282 can be provided. As Figure 2a shown, the different experiences that can be provided separately from different adaptive sets 281, 282 can vary greatly from each other (as shown by the triangle 200 in Figure 2a ). Providing different adaptive sets 281, 282 for different audio experiences 152 generally requires providing a dedicated audio (bit) stream for each adaptive set 281, 282. Because separate audio streams are used within different adaptive sets 281, 282, switching between different adaptive sets 281, 282 may trigger re-buffering and may not allow a seamless experience (when switching between different audio experiences 152).
[0061] In HbbTV 2.0.2, section A2.12.1 (the content of which is incorporated herein by reference), creating different audio tracks (i.e., audio signals) from the preselected elements 291, 292, 293 is described in detail. Given the fact that different preselected elements 291, 292, 293 can refer to a single adaptive set 281, there is no longer a need for separate audio streams to provide different experiences 152. In addition, since the same audio stream is used for different experiences 152, there is no longer a need for re-buffering, and thus experience selection can be provided in a seamless manner.
[0062] A manifest file 140 can be provided that includes a combination of one or more adaptive sets 281, 282 and one or more preselected elements 291, 292, 293. Different adaptive sets 281, 282 can be used to provide a basic on / off experience, and one or more preselected elements 291, 292, 293 (as Figure 2aas shown in the lower portion of
[0063] Providing a manifest file 140 that utilizes a combination of the adaptive sets 281, 282 and the preselected elements 291, 292, 293 can implement an on / off experience to select the narrative focus of the current television set using the adaptive sets 281, 282 and to progressively make advanced selections on the next-generation television using the preselected elements 291, 292, 293.
[0064] As described above, selecting different adaptive sets 281, 282 triggers a switch to a different audio stream (having multiple audio objects 181), which is then forwarded to the decoder 113 for decoding. When different preselected elements 291, 292, 293 are selected for a given adaptive set 281, 282, the decoder 112 can be reconfigured to decode the selected preselected elements 291, 292, 293 (using the same audio objects 181). Alternatively or additionally, when switching to different preselected elements 291, 292, 293, a switch of the audio objects 181 or the audio stream can be triggered.
[0065] Figure 2b An example manifest file 140 is illustrated, which includes data elements 211 for different adaptive sets 281, 282 and data elements 212 for different pre-selections. The data elements 211 for the adaptive sets 281, 282 can include pointers 213 that point to one or more data elements 212 for the corresponding adaptive sets 281, 282 or for the corresponding one or more pre-selections.
[0066] The user can change the audio gain 181 of one or more audio objects of different categories (including dialogue, important music and effects (e.g., gunshots / closing doors), less important music and effects 1 (e.g., background music, traffic), less important music and effects 2, etc.).
[0067] To be able to control the gain of audio objects 181 of different categories, the application 112 can be configured to modify the gain of one or more audio objects 181 in the audio bitstream 121. To be able to modify the gain, a pointer 303 pointing to the bitstream element 304 for the gain of each audio object 181 can be inserted into the bitstream 121, particularly into different bitstreams or media segments 170 (as Figure 3aas shown). Then, application 112 can use pointer 303 pointing to bitstream element 304 for the gain of audio object 181 to identify and modify the gain of that audio object 181. The corresponding encoder of bitstream 121 can be configured to add pointer 303 pointing to bitstream element 304 within frame or audio bitstream segment 170. Application 112 can read the pointer information inserted by the encoder and can modify each frame or audio bitstream segment 170 (for rendering) before passing frame or audio bitstream segment 170 to decoder 113. By doing so, the user can precisely define a personalized experience.
[0068] For example, application 112 can receive user input (e.g., this can be input using a user interface slider) regarding how much a dialogue object 181 should be enhanced or attenuated. Application 112 uses pointer 303 to identify bitstream element 304 for the gain of dialogue object 181. Additionally, the gain can be modified, and the modified frame or audio bitstream segment 170 can be passed to decoder 113. Decoder 113 can decode frame or audio bitstream segment 170 with the modified gain and can use the modified gain to render dialogue object 181.
[0069] Application-level bitstream-level modification can utilize the access of application 112 to bitstream 121. This may be the case when using the Media Source Extension (MSE) API on interface 122. The starting point of an AC-4 frame or audio bitstream segment 170 within the buffer can be identified by parsing the ISOBMFF (ISO / IEC Base Media File Format) file structure in the buffer.
[0070] Data pointer 303 pointing to each bitstream element 304 can be stored on a per-frame basis and can be transmitted in the ISOBMFF samples after the actual AC-4 frame. As Figure 3a shown, pointer 303 can be provided within pointer part 302, and this pointer part 302 is located behind element part 301 including bitstream element 304.
[0071] The following syntax can be used to define pointer 303 pointing to bitstream element 304 for the gain and / or position of audio object 181:
[0072] for(i = 1; i <= entry_count; i++){
[0073] unsigned int(8) kev;
[0074] unsigned int(16) frame_offset;
[0075] unsigned int(8) bitfield_length;
[0076] }
[0077] unsigned int(8) entry_count;
[0078] unsigned int(32) keyfield_id;
[0079] The above syntax can be used to provide an optional unique identifier "keyfield_id" that the application 112 can use to determine whether there is a pointer 303 to the bitstream element 304.
[0080] If the field "keyfield_id" does not exist, the presence of pointer data can be indicated by the manifest file 140 of the bitstream 121.
[0081] In the above syntax, the field "entry_count" can be used to calculate information on how much data has been added in the ISOBMFF sample. The field "entry_count" can be read before identifying the starting point in the buffer for the added pointer data (i.e., the position of the first "key" or pointer 303).
[0082] The key or pointer 303 can be assigned to a specific AC-4 bitstream element 304 and can be used to identify which AC-4 bitstream element 304 the frame position information contained within the key or pointer 303 applies to. The position of the bitstream element 304 represented by the "key" value can be given by the "frame_offset" and "bitfield_length" parameters that describe the exact position of a specific data element within the frame.
[0083] The assignment of keys or pointers is preferably known to the application 112, and there may be some special keys or pointers that can be considered. Example keys are
[0084] · Key: 0x00 - EMDF hash. Each frame is hashed, and when a bit element of the frame changes, the hash value may need to be recalculated; and / or
[0085] · key_id (Key ID): 0x01 - EMDF key ID. This pointer 305 points to the EMDF key ID. Changing this key value to the special key_id value 0x06 will cause the EMDF hash to be ignored, and the EMDF hash does not need to be recalculated by the application 112. Therefore, changing the value of key 305 can be used as an alternative to recalculating the hash value on the frame. It should be noted that the key_id 0x06 can also be changed in the encoder 101.
[0086] The application 112 may need to know the keys used for each mixed data, such as
[0087] · Key: 0x10 - Dialogue mixing level - object_gain_value (object gain value) of the dialogue;
[0088] · Key: 0x11 - Ambience 1 mixing level - object_gain_value of the ambience 1 mixing;
[0089] · Key: 0x12 - Ambience 2 mixing level - object_gain_value of the ambience 2 mixing;
[0090] · Key: 0x13 - Ambience 3 mixing level - object_gain_value of the ambience 3 mixing;
[0091] · Key: 0x14 - Object position 1;
[0092] · And so on.
[0093] Editing the key or pointer enables the application 112 to modify the mixing gain and / or object position by applying the mapping from the user interface (UI) control element to the values of the mixing gain and / or object position.
[0094] The way of encoding the object_gain_value of the bitstream element 304 usually may cause problems, and the problem can be solved as described below. The gain values of 0dB, -inf, and "repeat other object gain" are usually encoded in a special way and do not occupy the same space in the bitstream 121 as the general object gain. This indicates that these special values cannot be replaced by the general object gain, and since the general object gain cannot take the value 0 for example, the general gain cannot take any arbitrary number. If it is necessary to modify such special values, it is impossible to replace the data element alone and the bitstream 121 needs to be rewritten.
[0095] Therefore, it is recommended to encode all objects with positive or negative gains. Then the gain can be selected such that for any position of the user interface slider (i.e., any gain value), the resulting value is prevented from being 0dB.
[0096] As an example: If the maximum dialogue boost is 3 dB, encoding the dialogue with -4 dB will always ensure a non-zero value, resulting in a constant number of bits spent on this field. The level of the signal may change, which will be compensated by correctly setting the dialnorm of the dialogue to the new loudness (i.e., for the above dialogue example, if the content is typical (EBU R-128) content, effectively setting it to -27 dB).
[0097] In another example, a specific rule set (function) can be added that generates the value 314 of the specific bitstream element 304 based on the input value. This enables the use of arbitrary curves and mappings between UI inputs.
[0098] Example definitions include:
[0099] · OAMD—Object Audio Metadata (such as audio object location and gain); and / or
[0100] · MDAT—A container within an ISO base media file (MP4 file) where media data (such as audio frames) is stored. These frames are extracted based on the offset and length information stored in the MOOV container of the ISO base media file.
[0101] The positioning of AD objects can be implemented in the same way as the basic and advanced schemes described above. However, by using the method described in the basic section, each possible position will multiply the number of options, making the advanced selection preferable as the number of options increases.
[0102] Therefore, an application control method for audio processing is described. The method includes at least receiving adaptive sets 281, 282. In addition, the method includes at least receiving preselected elements 291, 292, 293. Additionally, the method includes generating a list of available audio tracks (i.e., audio signals) based on the adaptive sets 281, 282 and the preselected elements 291, 292, 293.
[0103] The application 112 can be configured to provide a basic on / off experience based on one or more adaptive sets 281, 282 on a rendering device 110 that only supports the adaptive sets 281, 282. In addition, the application 112 can be configured to provide fine-grained narrative importance control based on the preselected elements 291, 292, 293 on a rendering device 110 that supports the preselected elements 291, 292, 293 (in addition to the adaptive sets 281, 282).
[0104] The preselected elements 291, 292, 293 can be at least one of the following: dialogue, important music and effects (e.g., gunshots / closing door sounds), less important music and effects 1 (e.g., background music, traffic), less important music and effects 2, etc.
[0105] In addition, a method and apparatus for reading a bitstream 121 and for creating a list of bitstream element pointers 303 for each frame or audio bitstream segment 170 of the bitstream 121 are described. If a bitstream element 304 is included within the bitstream 121, the method can include determining the location and length of the bitstream element 304. In addition, the method includes creating a user-defined key or pointer 303 for the bitstream element 304.
[0106] Furthermore, a method for processing a bitstream 121 is described. The method includes receiving the bitstream 121. Additionally, the method includes reading one or more pointers 303 from the bitstream 121. Moreover, the method can include modifying the bitstream 121 (in particular, one or more bitstream elements 304 of the bitstream 121) based on the one or more pointers 303, and decoding the modified bitstream. The bitstream element 304 can be an object gain and / or an object position. Alternatively or additionally, the bitstream element 304 can be a recalculated hash to maintain the integrity of the frame or content segment 170 of the bitstream 121. Alternatively or additionally, the bitstream element 304 can include information related to the modified hash such that the decoder 113 ignores the integrity of the frame.
[0107] Figure 4a A flowchart of an example method 400 for personalizing audio content is shown. The audio content can be part of an audio / video experience. In particular, the audio content can be part of an HTML5 media element. The method 400 can be executed by an application 112 within a Hybrid Broadcast Broadband Television (HbbTV) system 100.
[0108] The method 400 can include receiving 401 a manifest file 140 for the audio content. The manifest file 140 can be an HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) manifest file. The audio content can be included within an audio bitstream 121. The audio bitstream 121 can include or can be an AC-4 audio bitstream (as detailed in ETSI TS 103 190, the content of which is incorporated herein by reference).
[0109] The manifest file 140 may include at least one adaptive set 281, 282 of reference audio bitstreams 121, which audio bitstreams include a plurality of audio objects 181. Additionally, the manifest file 140 may include a plurality of different preselected elements 291, 292, 293 for the adaptive sets 281, 282. The different preselected elements 291, 292, 293 may specify different combinations of the plurality of audio objects 181. In particular, the different preselected elements 291, 292, 293 may specify different schemes for mixing the plurality of audio objects 181 to form an audio signal to be presented. To this end, the different preselected elements 291, 292, 293 may include different sets of metadata 191.
[0110] The plurality of audio objects 181 may include audio objects 181 for dialogue and / or narrative content, audio objects 181 for music content, and / or audio objects 181 for audio effect content. The different audio objects 181 may be mixed according to the selected preselected elements 291, 292, 293 to provide a personalized audio experience to the user.
[0111] The method 400 may further include selecting 402 a preselected element 291 from the plurality of different preselected elements 291, 292, 293. The selection may be performed according to user input. The manifest file 140 may include a description 141 of different audio (and possibly video) experiences associated with each of the plurality of different preselected elements 291, 292, 293. The description 141 may enable the user to select an appropriate preselected element 291, 292, 293 for presentation.
[0112] Additionally, the method 400 may include determining an audio signal for presentation based on the plurality of audio objects 181 and based on the selected preselected element 291. The method 400 may further include obtaining 170 at least one audio bitstream segment (or frame) for the determined audio signal, and / or providing 170 at least one audio bitstream segment of the determined audio signal for presentation to a decoder 113. In particular, the method 400 may include causing 403 the presentation of an audio signal depending on the selected preselected element 291 (and generally according to one or more of the plurality of audio objects 181).
[0113] By providing different preselected elements 291, 292, 293 for the adaptive sets 281, 282, fine-grained personalization of audio content can be provided in an efficient manner.
[0114] Manifest file 140 may include a first adaptive set 281 that references a first audio bitstream 121 and a second adaptive set 292 that references a second audio bitstream 121. The first audio bitstream includes a first audio object set 181, and the second audio bitstream includes a second audio object set 181. The first adaptive set 281 may provide a first audio experience, and the second adaptive set 282 may provide a second audio experience. Additionally, a plurality of different preselected elements 291, 292, 293 may provide one or more intermediate audio experiences between the first audio experience and the second audio experience. In particular, compared to the second audio experience, the first audio experience may exhibit a lower degree of emphasis on dialogue and / or narrative content. The one or more intermediate audio experiences may exhibit a degree of emphasis on dialogue and / or narrative content, and the degree of emphasis is between the degree of emphasis of the first audio experience and the degree of emphasis of the second audio experience. By providing different adaptive sets for coarse-grained personalization and different preselected elements for fine-grained personalization, personalization of audio content can be provided in a particularly efficient and flexible manner.
[0115] The first audio object set 181 and the second audio object set 181 may include a common audio object set 181. On the other hand, the first adaptive set 281 and the second adaptive set 292 differ in how they combine audio objects 181 from the common audio object set 181 to form an audio signal for presentation. By doing so, different audio experiences can be provided in an efficient manner.
[0116] Figure 4b A flowchart of another example method 410 for personalizing audio content from an audio bitstream 121 (of a media element) is shown. The audio bitstream 121 may be part of an audio / video experience. Method 410 may be executed by an application 112 within the HbbTV system 100.
[0117] Method 410 may include receiving 411 an audio bitstream segment 170 of the audio bitstream 121. The audio bitstream segment 170 may include a pointer portion 302 that has pointers 303 to different bitstream elements 304 of the audio bitstream segment 170. The different bitstream elements 304 may be located in an element portion 301 of the audio bitstream segment 170. The pointer portion 302 may be located within the audio bitstream 121, downstream of the element portion 301.
[0118] Method 410 further includes identifying 412 a gain pointer 303 from the pointer portion 302, the gain pointer pointing to a first bitstream element 304 for the gain and / or position of the audio object 181 of the audio bitstream 121. The first bitstream element 304 can be any one of the bitstream elements 304 within the element portion 301 of the audio bitstream segment 170. The first bitstream element 304 can indicate a gain value (for amplification and / or attenuation) and / or a position value (for positioning in space), which can be used to mix the audio object 181 into the audio signal, particularly for attenuating or amplifying and / or positioning the audio object 181 in a personalized manner.
[0119] Additionally, method 410 can include modifying 413 the value of the first bitstream element 304 before presenting the audio object 181. Then, the modified value of the first bitstream element 304 can be used to present the audio object 181. By modifying the value, the emphasis and / or position of the audio object 181 can be modified with high precision in a personalized manner.
[0120] Method 410 can include determining a user input at a user interface (e.g., a slider), wherein the user input indicates a user value of the first bitstream element 304 that has been set by the user. Method 410 can further include setting the value of the first bitstream element 304 to the user value. By enabling the user to directly set the value of the first bitstream element 304, personalization of the audio content can be achieved in a comfortable and precise manner.
[0121] Method 410 can include determining a modified hash value of the element portion 301 of the audio bitstream segment 170, the element portion including the first bitstream element 304 having the modified value. Additionally, method 410 can include replacing the hash value of the element portion 301, which is intended to indicate the integrity of the element portion 301, particularly the hash value of the hash bitstream element of the element portion 301, with the modified hash value. By modifying the hash value, it can be ensured in an efficient and reliable manner that the audio bitstream segment 170 is decoded by the decoder 113 for presentation.
[0122] Method 410 can include verifying whether the integrity pointer 305 within the pointer portion 302 points to a hash bitstream element or exhibits a predetermined special value. The predetermined special value prevents the decoder 113 from verifying the integrity of the element portion 301 of the audio bitstream segment 170 (before decoding the audio bitstream segment 170).
[0123] Method 410 may include: if it is determined that the integrity pointer 305 points to a hash bitstream element, replacing the hash value of the hash bitstream element with a modified hash value that depends on the modified value of the first bitstream element 304. Alternatively or additionally, method 410 may include: if it is determined that the integrity pointer 305 exhibits a predetermined special value, keeping the hash value of the hash bitstream element unchanged. By doing so, it can be ensured in an efficient and reliable manner that the audio bitstream segment 170 is decoded by the decoder 113 for presentation.
[0124] Method 410 may include: if it is determined that the integrity pointer 305 exhibits a predetermined special value, ignoring the hash bitstream element when decoding and / or presenting the audio bitstream segment 170. By doing so, it can be ensured in an efficient and reliable manner that the audio bitstream segment 170 is decoded by the decoder 113 for presentation.
[0125] Method 410 may include, before identifying the gain pointer 303, determining whether there is a pointer 303 pointing to the first bitstream element 304, particularly based on the pointer part 302 and / or based on the manifest file 140 of the audio bitstream 121. If it is determined that there is a pointer 303, then only the first bitstream element 304 may be identified and the value of the first bitstream element 304 may be modified. By doing so, the efficiency of personalization can be improved.
[0126] Preferably, the value of the first bitstream element 304 is modified such that the first bitstream element 304 is replaced with a valid gain value and / or position value.
[0127] The gain pointer 303 may be identified from the pointer part 302 of the data set corresponding to the active preselected elements 291, 292, 293 of the audio bitstream 121.
[0128] Figure 4c A flowchart of a method 420 for implementing personalization of audio content from an audio bitstream 121 is shown. The audio bitstream 121 may be part of an audio / video experience and / or media element. The method 420 may be performed by an encoder or transcoder (e.g., within the network server 101).
[0129] Method 420 may include receiving 421 an audio bitstream segment 170 of the audio bitstream 121, wherein the audio bitstream segment 170 includes a pointer part 302 having pointers 303 to different bitstream elements 304 of an element part 301 of the audio bitstream segment 170. The element part 301 includes a first bitstream element 304 for the gain and / or position of the audio object 181 of the bitstream element 170. The first bitstream element 304 may be any one of the bitstream elements 304 within the element part 301 of the audio bitstream segment 170.
[0130] Additionally, method 420 includes setting 422 integrity pointer 305 within pointer portion 302 to a predetermined special value, where the predetermined special value prevents decoder 113 from verifying the integrity of element portion 301 of audio bitstream segment 170. By doing so, personalization of audio content can be achieved in a reliable and efficient manner at application 112.
[0131] Figure 4d A flowchart of an example method 430 for implementing personalization of audio content from audio bitstream 121 is shown. Audio bitstream 121 can be part of an audio / video experience and / or media element. Method 430 can be performed by an encoder or transcoder (e.g., within web server 101).
[0132] Method 430 includes receiving 431 audio bitstream segment 170 of audio bitstream 121. Audio bitstream segment 170 can include an element portion 301 having different bitstream elements 304. In particular, audio bitstream segment 170 can include a pointer portion 302 having pointers 303 to different bitstream elements 304 of element portion 301 of audio bitstream segment 170. Element portion 301 can include a first bitstream element 304 for gain and / or position of audio (bit) stream 181 of bitstream element 170. The first bitstream element 304 can be any one of the bitstream elements 304 within element portion 301 of audio bitstream segment 170.
[0133] Furthermore, method 430 includes inserting 432 pointer 303 pointing to first bitstream element 304 into pointer portion 302, enabling application 112 to identify first bitstream element 304 and modify gain and / or position values for personalization of audio content.
[0134] Additionally, devices and / or apparatuses 112, 101 configured to perform methods 400, 410, 420, 430 respectively are described.
[0135] Those of ordinary skill in the art will readily appreciate various modifications to the described embodiments in the present disclosure. Without departing from the spirit or scope of the present disclosure, the general principles defined herein can be applied to other embodiments. Therefore, the claims are not intended to be limited to the embodiments shown herein, but rather to the broadest scope consistent with the present disclosure, the principles disclosed herein, and novel features.
[0136] The methods, apparatuses, devices, and / or systems described in this document can be implemented as software, firmware, and / or hardware. Certain components can be implemented, for example, as software running on a digital signal processor or a microprocessor. Other components can be implemented, for example, as hardware and / or an application specific integrated circuit. The signals encountered in the described methods and systems can be stored on a medium such as a random access memory or an optical storage medium. These signals can be transmitted via a network such as a radio network, a satellite network, a wireless network, or a wired network (e.g., the Internet). A typical device that utilizes the methods and systems described in this document is a portable electronic device or other consumer device for storing and / or presenting audio signals.
Claims
1. A method for providing personalized audio signals, characterized in that, Wherein, The method includes - Receiving a manifest file for audio content; wherein, the manifest file includes - A first adaptive set referring to a first audio bitstream and a second adaptive set referring to a second audio bitstream, the first audio bitstream including a first audio object set, and the second audio bitstream including a second audio object set; and - A plurality of different preselected elements for the first adaptive set; wherein, the different preselected elements specify different combinations of the first audio object set; - Selecting a preselected element from the plurality of different preselected elements; and - Causing the presentation of an audio signal depending on the selected preselected element; - Wherein, the first adaptive set provides a first audio experience, and the second adaptive set provides a second audio experience; and - The plurality of different preselected elements provide one or more intermediate audio experiences between the first audio experience and the second audio experience, wherein the first audio experience shows a lower degree of emphasis on dialogue and / or narrative content compared to the second audio experience, and wherein the one or more intermediate audio experiences show a degree of emphasis on the dialogue and / or narrative content, and the degree of emphasis is between the degree of emphasis of the first audio experience and the degree of emphasis of the second audio experience.
2. The method according to claim 1, wherein Wherein, - The first audio object set and the second audio object set include a common audio object set; And - The first adaptive set and the second adaptive set differ in how to combine audio objects from the common audio object set to form the audio signal for presentation.
3. The method according to claim 1, wherein Wherein, The first audio object set includes - Audio objects for dialogue and / or narrative content; - Audio objects for music content; and / or - Audio objects for audio effect content.
4. The method according to claim 1, wherein Wherein, The manifest file includes a description of the audio experience associated with each of the plurality of preselected elements.
5. The method according to claim 1, characterized in that, Further includes - Obtaining at least one audio bitstream segment for the audio signal; and / or - Providing at least one audio bitstream segment of the audio signal for presentation to a decoder.
6. The method according to claim 1, wherein Wherein, The manifest file is an HTTP-based Dynamic Adaptive Streaming over HTTP (DASH) manifest file.
7. The method according to claim 1, wherein Wherein, The method is executed by an application within a Hybrid Broadcast Broadband Television (HbbTV) system.
8. A device for providing personalized audio signals, in particular an application unit, characterized in that, Wherein, The device is configured to - receive a manifest file for audio content; wherein, the manifest file includes - A first adaptive set referring to a first audio bitstream and a second adaptive set referring to a second audio bitstream, the first audio bitstream including a first audio object set, and the second audio bitstream including a second audio object set; and - A plurality of different preselected elements for the first adaptive set; wherein, the different preselected elements specify different combinations of the first audio object set; - Selecting a preselected element from the plurality of different preselected elements; and - Causing the presentation of an audio signal depending on the selected preselected element; - wherein the first adaptive set provides a first audio experience and the second adaptive set provides a second audio experience; and - the plurality of different preselected elements provide one or more intermediate audio experiences between the first audio experience and the second audio experience, wherein the first audio experience shows a lower degree of emphasis on dialogue and / or narrative content compared to the second audio experience, and wherein the one or more intermediate audio experiences show a degree of emphasis on the dialogue and / or narrative content that is between the degree of emphasis of the first audio experience and the degree of emphasis of the second audio experience.
9. A storage medium, characterized in that, The storage medium stores a software program, the software program including instructions for controlling one or more devices to perform the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a program, the program including computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Selection of coded next generation audio data for transport
US20170156015A1