Computer-implemented method, apparatus and computer program product for setting a playback speed of media content including audio
Patent Information
- Application Number
- CN202180008486.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-21
- Filing Date
- 2021-01-06
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2041-01-06
AI Technical Summary
然而,必须注意以使回放速度的改变不会在很大程度上使用户体验变差
[0011]有利地,本方面提供了根据音频类型来自动设置回放速度的效果。因此,本方面减少了用户在音频类型改变时必须改变回放速度的需要,从而带来了更好的用户体验。
Smart Images

Figure CN114930865B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to International Patent Application PCT / CN 2020 / 070728, filed January 7, 2020; U.S. Provisional Patent Application 62 / 978,477, filed February 19, 2020; and European Patent Application 20158755.7, filed February 21, 2020, each of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to the playback of media content, and more particularly, to a computer-implemented method, apparatus, and computer program product for setting the playback speed of media content including audio. Background Technology
[0004] Today, altering the playback speed of media content is a common feature of modern media player applications (such as streaming apps) on smartphones, smart TVs, computers, and other devices. For example, content viewers often use fast playback (i.e., playback speed greater than twice the original speed) to quickly browse media content such as videos, podcasts, and audiobooks. Similarly, sometimes content viewers slow down playback for various reasons. However, it is important to ensure that changes to playback speed do not significantly degrade the user experience.
[0005] Therefore, improvements are needed in this area. Summary of the Invention
[0006] According to a first aspect of this disclosure, a computer-implemented method is provided for setting the playback speed of media content including audio, the media content having a defined normal playback speed, the method comprising:
[0007] Receive instructions to play media content at a speed different from the media content's normal playback speed.
[0008] Analyze the audio to determine the audio type; and
[0009] Determine a playback speed that differs from the normal playback speed based on the identified audio type, and set the playback speed of the media content to the determined playback speed.
[0010] In the context of this specification, the term "normal playback speed" should be understood as the expected, unadjusted playback speed of the media content, i.e., 1.0x. In other words, normal playback speed corresponds to the playback of the media content without adjusting its playback speed. For audio, normal playback speed corresponds to the sampling rate of the recorded audio. For video, normal playback speed corresponds to the frames per second (FPS) of the recorded video.
[0011] Advantageously, this aspect provides the ability to automatically set the playback speed based on the audio type. Therefore, this aspect reduces the need for users to adjust the playback speed when the audio type changes, resulting in a better user experience.
[0012] According to some embodiments, determining a playback speed that differs from the normal playback speed includes selecting one or more predefined playback speeds based on the determined audio type.
[0013] For example, one or more predefined playback speeds can be objectively set, such as based on research on how the human auditory system typically works and at what playback speed a user can hear which types of audio while still having a good user experience. In other embodiments, one or more predefined playback speeds can be based on user input (e.g., settings set by the user to be used for all audio of a specific type in the media content).
[0014] According to some embodiments, analyzing audio to determine the audio type includes: analyzing the audio to determine whether the audio includes dialogue; and / or analyzing the audio to determine whether the audio includes music.
[0015] Advantageously, this embodiment allows for improved user experience during, for example, fast playback mode of media content. Typically, when using high playback speeds, users sometimes find it difficult to understand dialogue within the media content. Furthermore, when the media content (for the user) is a foreign-language film, reducing the playback speed to allow the user to understand the dialogue can be advantageous. For example, if the media content is a film (where dialogue is often crucial for understanding the plot), the determined playback speed for media content including this audio type can generally be lower than for other audio types. If the media content involves a musical performance such as the Eurovision Song Contest, the dialogue can be considered less important, and the playback speed for media content including this audio type can generally be set to a relatively high value. Moreover, for many types of media content such as movies or talk shows, music is considered less important by users, and therefore, the determined playback speed for media content including this audio type can be increased.
[0016] According to some embodiments, the method further includes setting the playback speed to a default playback speed if the audio type cannot be determined. Advantageously, the default playback speed may be a predetermined playback speed for a fast playback mode used for the media content.
[0017] According to some embodiments, the following steps are repeatedly performed while playing media content: analyzing audio to determine the audio type, determining a playback speed different from the normal playback speed based on the determined audio type, and setting the playback speed of the media content to the determined playback speed.
[0018] Advantageously, this allows for automatic adjustment of playback speed at runtime based on the audio type of the current section of the media content.
[0019] According to some embodiments, the method further includes the following steps:
[0020] Select one of one or more predefined audio time stretching algorithms based on the determined audio type, and
[0021] Set the audio time stretching algorithm for the media content to the selected audio time stretching algorithm.
[0022] In the context of this specification, the term "audio time stretching algorithm" should be understood as an algorithm used to change the speed / duration of audio. Audio time stretching algorithms may or may not affect pitch. Audio time stretching algorithms can be applied in the time domain or the frequency domain of audio.
[0023] In some cases, a particular audio time-stretching algorithm may not be suitable or optimal for certain audio types. For example, an audio time-stretching algorithm that allows pitch scaling of audio may make dialogue difficult to understand (the chipmunk effect). In these embodiments, or in other embodiments where a better audio time-stretching algorithm exists for certain audio types, it may be advantageous to automatically change the audio time-stretching algorithm used for media content based on the determined audio type.
[0024] According to some embodiments, the method further includes: if the audio type cannot be determined, setting the audio stretching algorithm of the media content as the default audio stretching algorithm. Advantageously, a default audio stretching algorithm with average performance for most audio types can be selected.
[0025] According to some embodiments, the following steps are repeatedly performed while playing media content: analyzing audio to determine the audio type, selecting one or more predefined audio time stretching algorithms based on the determined audio type, and setting the audio time stretching algorithm of the media content to the selected audio time stretching algorithm.
[0026] Advantageously, this allows the audio time-stretching algorithm to be automatically adjusted at runtime based on what audio type the current part of the media content is.
[0027] According to some embodiments, the step of analyzing audio to determine the audio type includes: for at least one audio type, determining a confidence score related to the audio including the audio type; and determining whether the confidence score exceeds a threshold confidence score.
[0028] Advantageously, this can increase the flexibility of the methods described herein.
[0029] For example, the audio analysis method or software used can, for instance, output a confidence value for each (or every other, every two, every four, etc.) audio frame of the media content, indicating that the audio includes a certain audio type. If any analyzed type receives a confidence score above a threshold, the audio can be considered to include that audio type. If none of the output confidence scores are above the threshold, the audio can be considered not to include the analyzed audio type. Accordingly, based on the output confidence scores, it can be determined which playback speed and, optionally, an audio time-stretching algorithm should be applied. If more than one audio type receives a confidence score above the threshold, a decision can be made based on which audio type receives the highest confidence score or which group of audio types is considered part of the current audio frame.
[0030] In a second aspect, this disclosure provides a computer program product including instructions adapted to perform the method of the first aspect when executed by a device with processing capabilities.
[0031] In a third aspect, this disclosure provides an apparatus configured to determine the playback speed of media content including audio, the apparatus including circuitry configured to perform the method of the first aspect.
[0032] The second and third aspects can usually have the same characteristics and advantages as the first aspect. Attached Figure Description
[0033] The above and other objects, features and advantages of this disclosure will be better understood through the following illustrative and non-limiting detailed description of various embodiments of the present disclosure with reference to the accompanying drawings, wherein the same reference numerals will be used for similar elements, in which:
[0034] Figure 1 A method for setting the playback speed of media content including audio, according to an embodiment, is shown.
[0035] Figure 2A method for setting the playback speed of media content including audio, according to an embodiment, is shown.
[0036] Figure 3 An example of a device for setting the playback speed of media content including audio, according to an embodiment, is illustrated.
[0037] Figure 4 It shows Figure 3 Examples of the device. Detailed Implementation
[0038] This disclosure will now be described more fully below with reference to the accompanying drawings, in which embodiments of the disclosure are illustrated. The systems and apparatus disclosed herein will be described during operation.
[0039] This disclosure stems from the understanding that many users speed up media playback to save time, such as when trying to keep up with a TV series or podcast, while in certain scenarios, users want to slow it down to avoid missing anything important. Typically, users must manually select the playback speed. A typical use case is when a user sets the speed to the fastest (e.g., 2.0x) to save time, but in conversational scenarios, the user might find the audio pitch too high and difficult to understand, or the conversation too fast. The user then needs to manually adjust the speed back to normal or some other, less fast rate (e.g., 1.5x, or even lower, such as 0.75x). At the end of the conversation, the user must again manually adjust the speed to the highest to save time. Many users eventually give up on manually adjusting the speed, resulting in a poor user experience. Similar manual selection of playback speed can be performed for other types of media content, such as when playing music or when the audio of the media content is noisy.
[0040] The inventors have recognized that users will appreciate methods and devices that can automatically adjust playback speed based on the current audio type of the media content. Intelligent and automatic adjustment of playback speed can be achieved based on an audio classification mechanism (audio analysis method / software). This will be described below.
[0041] Figure 1 A flowchart illustrating a computer-implemented method for setting the playback speed of media content, including audio, is provided as an example. Optionally, the method can also be used to determine an audio time-stretching algorithm to use when the media content is not played at normal speed (1.0x). Figure 1 In the diagram, the dotted part is optional.
[0042] The method begins by receiving an instruction from S02 to play the media content at a speed different from the normal playback speed of the media content. This can be based on user input, such as an instruction to play the media content at a speed exceeding the normal playback speed of the media content (i.e., playback speed > 1.0 times) or at a speed lower than the normal playback speed (i.e., playback speed < 1.0 times).
[0043] The next step involves analyzing the S03 audio to determine the audio type. Advantageously, this step is performed for each audio frame; however, in some embodiments, analysis is performed using audio from multiple audio frames, or using a single audio frame and performing analysis every other audio frame, every two audio frames, every four audio frames, etc.
[0044] Analysis S03 may involve determining at least one of the following: the pitch of the audio, the harmonic structure of the audio, the zero-crossing rate of the audio, the periodicity of the audio, the chroma of the audio, the spectral width of the audio, and the spectral envelope of the audio.
[0045] Analysis S03 can be performed on one or more defined audio types. In this paper's examples, dialogue and music are used as examples. For instance, audio characteristics such as a certain pitch, a certain harmonic structure, a certain linear predictive coding, and a certain zero-crossing rate can indicate dialogue. Audio characteristics such as a certain duration, a certain periodicity, a certain chroma, and a certain spectral width can indicate background music. For example, the most significant difference between music and dialogue lies in the spectral width of the audio, where the spectral range of music can be much wider than that of human speech.
[0046] According to some embodiments, analyzing S03 audio to determine the audio type includes: analyzing the audio to determine whether the audio includes dialogue; and / or analyzing the audio to determine whether the audio includes music. However, it should be understood that the methods described herein can be used for other audio types, such as applause, forest sounds, war scenes, noise, VoIP, etc.
[0047] The step of analyzing the S03 audio to determine the audio type may optionally include, for at least one audio type, determining a confidence score related to the audio of the media content including the audio type, and determining that the audio includes the audio type when the confidence score exceeds a threshold confidence score. For example, the confidence score may correspond to a certainty that the audio includes the audio type at 50%, 66%, 75%, etc. If more than one audio type obtains a confidence score higher than the threshold confidence score, the audio type with the highest confidence score can be used to adjust the playback speed and / or audio time stretching algorithm. In other embodiments, the playback speed and / or audio time stretching algorithm may be adjusted to a specific combination of values / algorithms for the determined audio types with confidence scores higher than the threshold.
[0048] After analyzing the audio, a playback speed that differs from the normal playback speed is determined based on the identified audio type. For example, determining a playback speed that differs from the normal playback speed may include selecting one or more predefined playback speeds based on the identified audio type. The predefined playback speed may be set based on user input, or set, for example, by a content provider (e.g., received in the metadata of the media content) or by the provider of the software implementing the methods discussed herein, in a hard-coded setting. The predefined playback speed may also be based on an AI or machine learning algorithm that receives input from multiple users and determines a predefined playback speed for a specific audio type for certain types of media content (e.g., based on metadata such as content type, media content length, etc.).
[0049] The one or more predefined playback speeds may include a playback speed for each possible result from audio analysis step S03. The one or more predefined playback speeds may include a playback speed associated with multiple results from audio analysis step S03, and / or a playback speed for an audio type that is not mapped to a playback speed in the one or more predefined playback speeds.
[0050] After the playback speed S04 has been determined, the method includes setting the playback speed of the media content to the determined playback speed at S06. Optionally, if the audio type cannot be determined, the playback speed can be set to a default playback speed at S06. The default playback speed can be, for example, 1.5x, 2.0x, or 2.5x. The default playback speed can also be based on the metadata of the media content (indicated in the metadata), where, for example, the default playback speed is different for different types of media content. The default playback speed can also be based on user input (as described below). Figure 3 (To be further described). The default playback speed may be based on an AI or machine learning algorithm that receives input from multiple users and determines a preferred default playback speed for certain types of media content (e.g., based on metadata such as content type, media content length, etc. (as indicated in the metadata)).
[0051] Examples of different playback speeds will now be discussed by way of example.
[0052] For example, if it is determined that the audio includes dialogue (or may include dialogue), the playback speed can be changed compared to when the audio does not include dialogue or compared to the default playback speed. The predefined playback speed for dialogue can be based on user input (e.g., a setting set by the user to be used for all dialogue in the media content), or it can be automatically set, for example, based on the media content's metadata. For example, if the media content is a movie (where dialogue may be crucial for understanding the plot), the predefined playback speed for dialogue is typically lower than the default playback speed. If the media content involves musical performances, dialogue may be considered unimportant, and the predefined playback speed for dialogue can be higher than the default playback speed. In some cases, the predefined playback speed for dialogue can be lower than the normal playback speed, such as when the media content (for the user) is a foreign language film.
[0053] Similar to the discussion about dialogue, in some situations, defining a specific playback speed for media content that includes music in its audio can be beneficial. For example, in a film, scenes where music tracks are included as audio may be less important to the plot and can be played at a higher speed. In other situations, the portion of media content that includes music as audio may be considered as important as, or more important than, all other portions of the media content that do not include dialogue. The predefined playback speed for music can be based on user input (e.g., settings set by the user for all portions of media content where the audio includes music), or it can be automatically set, for example, according to the media content's metadata (as indicated in the metadata), or it may be a hard-coded value.
[0054] Optionally, the audio stretching algorithm S06 can be selected based on the determined audio type, and this audio stretching algorithm can be set as the audio time stretching algorithm for the media content in S07. As mentioned above, a certain audio time stretching algorithm may not be suitable or optimal for some audio types. Advantageously, setting the audio time stretching algorithm of S07 based on the determined audio type allows for improved user experience.
[0055] An example of an audio time stretching algorithm will now be discussed.
[0056] According to some embodiments, if the audio type cannot be determined, a default audio time-stretching algorithm can be used. In some embodiments, the default audio time-stretching algorithm can be an algorithm that affects the pitch of the audio, such as the WSOLA algorithm, which has average performance for most audio types. In other embodiments, another algorithm can be used, such as a suitable frame-based method. The default audio time-stretching algorithm can be user-defined or automatically set according to metadata (indicated in the metadata), as discussed above in conjunction with predefined playback speeds. In some embodiments, the default audio time-stretching algorithm can be hard-coded and cannot be changed by the user.
[0057] In some cases, the default audio time-stretching algorithm is not suitable or optimal for dialogue. For example, the default audio time-stretching algorithm may allow pitch scaling of the audio, making the dialogue difficult to understand (chipmunk effect). In these embodiments, or in other embodiments where a better audio time-stretching algorithm exists for dialogue compared to the one used as the default, it may be advantageous to automatically change the audio time-stretching algorithm used for the audio when it is determined that the audio of the media content includes dialogue.
[0058] According to some embodiments, when it is determined in analysis step S03 that the audio includes dialogue, the audio time-stretching algorithm for the media content is set to either the temporal pitch-synchronous overlay (TD-PSOLA) algorithm or the pointer interval control overlay (PICOLA) algorithm. These examples of audio time-stretching algorithms have good performance in changing the playback speed of speech (i.e., dialogue) and, by being specifically designed to preserve the timbre of the dialogue, still maintain a reasonable understanding of the resulting speed-adjusted audio.
[0059] According to some embodiments, when it is determined in analysis step S03 that the audio includes music, the audio time-stretching algorithm for the media content is set in S07 to a predefined audio time-stretching algorithm that allows changes in the pitch of the audio. This allows the rendered audio to sound as if the pitch of the music has been raised (or lowered) by, for example, an octave, but the music remains clear and smooth. Any suitable audio time-stretching algorithm with adjusted playback speed for music can be used. For example, a predefined audio time-stretching algorithm for music could be the WSOLA algorithm, which is based on waveform similarity. The WSOLA algorithm performs well in changing the playback speed of music without causing the sound to become blurred or inaudible.
[0060] According to some embodiments, the predefined audio time-stretching algorithm for music differs from the predefined audio time-stretching algorithm for dialogue. This allows for greater flexibility when changing the playback speed of media content.
[0061] As discussed above, analysis step S03 can output audio corresponding to multiple audio types. In these cases, the playback speed and optional audio time stretching algorithm can be set accordingly. This will now be illustrated with an example of the case where the audio includes both dialogue and music. However, other mixtures of audio types are also possible.
[0062] Generally, dialogue can be considered the most important audio type for viewers, and therefore, when the audio includes both dialogue and music, playback settings for dialogue can be used. As mentioned above, this can be implemented in different ways for certain types of media content (e.g., based on user preferences or based on media content metadata).
[0063] Therefore, in some embodiments, when it is determined in analysis step S03 that the audio corresponds to both dialogue and music, the audio time-stretching algorithm for the media content is set in S07 to either the temporal pitch-synchronous overlay (TD-PSOLA) algorithm or the pointer interval control overlay (PICOLA) algorithm. Similarly, when it is determined that the audio includes both dialogue and music, a predefined playback speed for dialogue can also be used.
[0064] In other embodiments, music is considered the most important audio type, and accordingly, the WSOLA algorithm can be selected as the audio time stretching algorithm, and a predefined playback speed for music can be used.
[0065] According to some embodiments, the process is repeated while playing media content. Figure 1 The steps in the method are steps S03, S04, and S06, and optionally steps S05 and S07. When media content is received as a streaming media transmission, the audio analysis step S03 can advantageously be performed in real time. However, as mentioned above, it is unnecessary to perform the analysis step S03 for each audio frame to reduce the computational power required to execute the method.
[0066] Streaming media refers to video or audio content that is sent over the Internet in compressed form and played immediately, rather than being saved to a hard drive. Using the methods disclosed herein, it is possible to adjust playback speed and / or set audio stretching algorithms in real time during streaming, which can improve the user experience.
[0067] In some embodiments, when media content is received as a streaming media transmission, the method may further include adjusting the streaming transmission rate based on a determined playback speed of the media content. In this way, the bit rate required for the streaming media transmission can be optimized based on the playback speed of the media content.
[0068] Figure 2Another flowchart of a method for setting the playback speed of media content including audio, according to an embodiment, is shown. Optionally, the method can also be used to determine an audio time-stretching algorithm to use when the media content is not played at normal speed (1.0x). Figure 2 In the diagram, the dotted part is optional.
[0069] The method begins by receiving an instruction in S102 to play the media content at a speed different from the normal playback speed of the media content. This can be based on user input, such as an instruction that the media content should be played at a speed exceeding the normal playback speed of the media content (i.e., playback speed > 1.0 times).
[0070] The next step involves analyzing one or more audio frames of the S103 media content to determine the audio type. Advantageously, this step is performed one audio frame at a time; however, in some embodiments, audio from multiple audio frames is analyzed together, or a single audio frame is analyzed and then every other audio frame, every two audio frames, every four audio frames, etc.
[0071] Analysis S103 may involve determining at least one of the following: the pitch of the audio, the harmonic structure of the audio, the zero-crossing rate of the audio, the periodicity of the audio, the chroma of the audio, the spectral width of the audio, and the spectral envelope of the audio.
[0072] Analysis S103 can be performed on one or more defined audio types. In this paper's examples, dialogue and music are used as examples. For instance, audio characteristics such as a certain pitch, a certain harmonic structure, a certain linear predictive coding, and a certain zero-crossing rate can indicate dialogue. Audio characteristics such as a certain duration, a certain periodicity, a certain chroma, and a certain spectral width can indicate music.
[0073] It should be understood that the method described in this article can be used for other audio types, such as applause, forest sounds, etc., in which case the method can be adjusted accordingly.
[0074] The step of analyzing the audio in S103 to determine the audio type may optionally include, for at least one audio type, determining a confidence score related to the audio in the media content including the audio type, and determining that the audio includes the audio type when the confidence score exceeds a threshold confidence score. For example, the confidence score may correspond to a certainty that the audio includes the audio type at 50%, 66%, 75%, etc. If more than one audio type obtains a confidence score higher than the threshold confidence score, the audio type with the highest confidence score can be used to adjust the playback speed and / or audio time stretching algorithm. In other embodiments, the playback speed and / or audio time stretching algorithm may be adjusted to a specific combination of values / algorithms for the determined audio types with confidence scores higher than the threshold.
[0075] If it is determined that the audio in S104 does not include any of the audio types defined in one or more audio types, then the default playback speed and default audio time stretching algorithm in S105 are used. The default playback speed can be, for example, 1.5x, 2.0x, or 2.5x.
[0076] In some embodiments, the default audio time stretching algorithm can be an algorithm that affects the pitch of the audio, such as the WSOLA algorithm, which has average performance for most audio types.
[0077] If at some point during the playback of media content it is determined that the audio to be played by S104 includes any of the defined audio types, the playback speed used to play the media content and, optionally, the audio time stretching algorithm used for the played audio can be adjusted accordingly.
[0078] exist Figure 2 In the example, the defined audio types include at least dialogue and optionally music. However, as clearly stated above, this is merely an example and other / additional audio types can be defined.
[0079] In one embodiment, upon determining that the audio in S106 includes dialogue, the method includes setting the playback speed of the media content in S109 to a predefined playback speed (hereinafter referred to as the first predefined playback speed), which may differ from the default playback speed. The first predefined playback speed may be lower than the default playback speed. Advantageously, the dialogue may then be easier for the listener / user to understand. In other embodiments, depending on the circumstances, the first predefined playback speed may be higher than the default playback speed, for example, when the media content corresponds to a nature documentary or where the dialogue is considered unimportant other content (e.g., media content including moving pictures, where the graphic content is the primary stimulus).
[0080] Optionally, the method includes: upon determining that the audio includes dialogue, setting the audio time-stretching algorithm used for playing back the media content to a predefined audio time-stretching algorithm (hereinafter referred to as the first predefined audio time-stretching algorithm) in S110, which may differ from the default audio time-stretching algorithm. Advantageously, the first predefined audio time-stretching algorithm can be tailored to dialogue to improve the user experience. In some embodiments, the first predefined audio time-stretching algorithm is a temporal pitch-synchronous overlay (TD-PSOLA) algorithm, such as a pointer-interval-controlled overlay (PICOLA) algorithm.
[0081] Optionally, the specified audio type includes music, wherein when it is determined that one or more audio frames analyzed in S102 include music, at least one of the audio time-stretching algorithm and playback speed is changed. When accelerating media content where the audio includes music, pitch shifting can advantageously be allowed to improve the user experience. Pitch shifting raises / lowers the pitch of the music (e.g., by an octave), but the auditory experience is that the music track remains clear and smooth even with the adjusted playback speed. Furthermore, for some media content, the music in the audio (such as background music) can indicate important or less important content. Therefore, the playback speed can also be adjusted.
[0082] Therefore, the method may include, upon determining that the audio in S107 includes music, setting the playback speed of the media content in S111 to a predefined playback speed (hereinafter referred to as the second predefined playback speed), which is different from the first predefined playback speed and / or the default playback speed. Alternatively or additionally, the method may include, upon determining that the audio in S107 includes music, setting the audio time-stretching algorithm used for playing back the media content in S112 to a predefined audio time-stretching algorithm (hereinafter referred to as the second predefined audio time-stretching algorithm), which may be different from the first predefined audio time-stretching algorithm. The second predefined audio time-stretching algorithm may also be different from the default audio time-stretching algorithm. In other embodiments, the second predefined audio time-stretching algorithm is the same as the default audio time-stretching algorithm. In some embodiments, the second predefined audio time-stretching algorithm may be a waveform similarity-based superposition WSOLA algorithm.
[0083] In some embodiments, the method includes, upon determining that the audio in S108 corresponds to both dialogue and music, setting the audio time-stretching algorithm for playing back the media content to a first predefined audio time-stretching algorithm in S111, and setting the playback speed of the media content to a first predefined playback speed in S112. In this embodiment, dialogue is considered more important than music, and the playback speed and audio time-stretching algorithm are selected accordingly. In other embodiments (not included) Figure 2In the above-discussed implementation, the audio type that yields the highest confidence score is used to set the playback speed and / or audio time-stretching algorithm. In other embodiments (not included)... Figure 2 In the Chinese version, music is considered more important than dialogue, and playback speed and / or audio time stretching algorithms are set accordingly.
[0084] As discussed above, this disclosure can be used in real-time scenarios where playback speed and / or audio time-stretching algorithms are continuously updated based on the characteristics of the audio in the media content. Therefore, Figure 2 The method includes the step of determining whether there are more audio frames in the media content at S113 (i.e., the media content being played has not yet ended). In this case, the method is iterated again by re-analyzing the audio at S103 (i.e., one or more new audio frames in the media content) to determine the audio type. Otherwise, the method ends at S114.
[0085] Figure 3 The implementation is illustrated schematically. Figure 1 and / or Figure 2 Device 200 for performing methods. For example, device 200 includes devices configured to perform methods. Figure 1 and / or Figure 2 The circuitry for the method. This circuitry may include one or more processors. Typically, Figure 1 and / or Figure 2 The method can be implemented in device 200 as software, firmware, hardware, or a combination thereof. Figure 3 The device includes a playback speed controller 202. The playback speed controller 202 receives media content 201, including audio. The media content 201 can be received during streaming media transmission, i.e., as streaming media content. The playback speed controller 202 determines the playback speed. The device further includes a decoder 204 that decodes the media content according to the determined playback speed. The decoded media content is then sent to an audio analysis unit 206, which analyzes the audio of the media content and informs the playback speed controller what audio types are currently included in the media content. Subsequently, the playback speed controller 202 uses the information received from the audio analysis unit 206 to determine the playback speed for the next one or more frames of the media content 201.
[0086] In some embodiments, the streaming rate of the media transmission can be adjusted based on the playback speed of the media content. This can be done by the device 200 notifying the streaming provider of the current application's playback speed to optimize the bitrate used to transmit the media content 201 to the device based on the playback speed.
[0087] The audio analysis unit 206 can operate in real time, which can be advantageous from the user's perspective in the case of streaming media. Therefore, playback speed and, optionally, audio stretching algorithms can be adjusted whenever the audio type in the media content 201 indicates that this should be done.
[0088] Optionally, the audio analysis unit also informs the audio renderer 208 what audio types are currently included in the media content. The audio renderer 208 then selects an appropriate audio time-stretching algorithm for rendering and outputs the rendered audio 210 for playback. By placing the audio analysis unit 206 before the audio renderer 208, any interference introduced during pitch changes can be avoided. Therefore, the audio analysis unit 206 will only process the raw audio data to obtain a score, without introducing any potential interference as applied audio effects during audio time-stretching.
[0089] Figure 4 This is illustrated by example. Figure 3 The device 200 further includes means 302 for playing back media content to a user. Means 302 may include a display and / or speakers. Figure 4 Device 200 further includes a user interface 304. Here, the user can select whether to use a manual speed setting or an automatic speed setting. Each setting (manual or automatic) includes editing symbols 306, 308, where the user can define, for example, the playback speed to be used. In the manual setting, the playback speed selected by the user is applied directly to the media content. In the automatic setting, the algorithm described herein is used to determine which playback speed and, optionally, which audio time-stretching algorithm should be applied based on the audio characteristics of the media content. In this case, user interface 304 can be configured (e.g., via editing symbol 308) to set at least one of the following: a first predefined playback speed, a second predefined playback speed, and a default predefined playback speed. In some embodiments, the first / second / default audio time-stretching algorithm can also be set by the user. Furthermore, device 200 can be implemented to present recommended playback speeds and / or audio time-stretching algorithms for different types of content. Such recommended playback speed / audio time-stretching algorithms can be based on AI or machine learning algorithms that receive input from multiple users and determine preferred playback speed / audio time-stretching algorithms for certain types of media content (e.g., based on metadata such as content type, media content length, etc.).
[0090] By way of example, the following scenarios are applicable when using methods / devices / computer program products as described in this article:
[0091] • If the user sets a very fast playback speed (e.g., 2.0x, above the threshold speed), the playback speed will be automatically slowed down (e.g., 1.5x) when a conversation is detected.
[0092] • When PICOLA is used as the second predetermined audio time stretching algorithm, automatically switch to WSOLA when the scene is not a dialogue.
[0093] • When a user attempts to learn a foreign language by watching foreign TV shows / movies, slow down the dialogue playback speed to below normal (e.g., 0.75 times).
[0094] • When users are watching a music performance, speed up the dialogue scene (e.g., 2.0x) while the judges are commenting, and reset it to normal speed (1.0x) or between 1.0x and 2.0x during the music performance itself.
[0095] • When listening to content without video (such as podcasts and FM radio), the playback speed can be adjusted based on whether a conversation is in progress.
[0096] After studying the foregoing description, further embodiments of this disclosure will become apparent to those skilled in the art. Although embodiments and examples have been disclosed in this specification and accompanying drawings, this disclosure is not limited to these specific examples. Many modifications and variations can be made without departing from the scope of this disclosure as defined by the appended claims. Any reference numerals appearing in the claims should not be construed as limiting their scope.
[0097] Furthermore, based on a study of the accompanying drawings, this disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing this disclosure. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The simple fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be advantageously utilized.
[0098] The systems and methods disclosed above can be implemented as software, firmware, hardware, or a combination thereof. In hardware implementations, the division of tasks among the functional units mentioned in the above description does not necessarily correspond to the division of physical units; rather, a physical component may have multiple functions, and a task may be performed collaboratively by multiple physical components. Some or all of the components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or application-specific integrated circuits. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage medium includes any volatile and non-volatile, removable and non-removable medium implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital universal disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transmission medium.
[0099] Various aspects of this disclosure can be understood from the following enumerated example embodiments (EEE):
[0100] EEE 1. A computer-implemented method for setting the playback speed of media content including audio, the media content having a defined normal playback speed, the method comprising:
[0101] Receive an instruction to play the media content at a speed different from the normal playback speed of the media content.
[0102] Analyze the audio to determine the audio type; and
[0103] A playback speed different from the normal playback speed is determined based on the determined audio type, and the playback speed of the media content is set to the determined playback speed.
[0104] EEE 2. The method according to EEE 1, wherein determining a playback speed different from the normal playback speed includes: selecting one or more predefined playback speeds based on the determined audio type.
[0105] EEE 3. The method according to EEE 2, wherein the one or more predefined playback speeds are received in the metadata of the media content.
[0106] EEE 4. The method according to any one of EEE 1 to 3, wherein analyzing the audio to determine the audio type comprises:
[0107] Analyze the audio to determine whether it includes dialogue; and / or
[0108] The audio is analyzed to determine whether it includes music.
[0109] EEE 5. The method according to any one of EEE 1 to 4, further comprising:
[0110] If the audio type cannot be determined, the playback speed is set to the default playback speed.
[0111] EEE 6. The method according to any one of EEE 1 to 5, wherein the following steps are repeatedly performed while playing the media content: analyzing the audio to determine an audio type, determining a playback speed different from the normal playback speed based on the determined audio type, and setting the playback speed of the media content to the determined playback speed.
[0112] EEE 7. The method according to any one of EEE 1 to 6 further includes the following steps:
[0113] Select one or more predefined audio time stretching algorithms based on the determined audio type, and set the audio time stretching algorithm of the media content as the selected audio time stretching algorithm.
[0114] EEE 8. The method according to EEE 7, wherein an instruction for the one or more predefined audio time stretching algorithms is received in the metadata of the media content.
[0115] EEE 9. The method according to any one of EEE 7 to 8 further comprises:
[0116] If the audio type cannot be determined, the audio stretching algorithm of the media content is set to the default audio stretching algorithm.
[0117] EEE 10. The method according to EEE 9, wherein the default audio time stretching algorithm affects the pitch of the audio content.
[0118] EEE 11. The method according to any one of EEE 7 to 10, wherein, when it is determined that the audio includes dialogue, the audio time stretching algorithm of the media content is set to either the temporal pitch-synchronous overlay TD-PSOLA algorithm or the pointer interval control overlay PICOLA algorithm.
[0119] EEE 12. The method according to any one of EEE 7 to 11, wherein, when it is determined that the audio includes music, the audio time stretching algorithm of the media content is set to a waveform similarity superposition WSOLA algorithm.
[0120] EEE 13. The method according to any one of EEE 7 to 12, wherein, when it is determined that the audio corresponds to both dialogue and music, the audio time stretching algorithm of the media content is set to either the temporal pitch synchronization superposition TD-PSOLA algorithm or the pointer interval control superposition PICOLA algorithm.
[0121] EEE 14. The method according to any one of EEE 7 to 13, wherein the following steps are repeatedly performed while playing the media content: analyzing the audio to determine an audio type, selecting one or more predefined audio time stretching algorithms based on the determined audio type, and setting the audio time stretching algorithm of the media content to the selected audio time stretching algorithm.
[0122] EEE 15. The method according to any one of EEE 1 to 14, wherein the step of analyzing the audio to determine the audio type comprises: for at least one audio type, determining a confidence score related to the audio including the audio type and determining whether the confidence score exceeds a threshold confidence score.
[0123] EEE 16. The method according to any one of EEE 1 to 15 further includes the step of: receiving the media content as a streaming media transmission, wherein the step of analyzing the audio is performed in real time.
[0124] EEE 17. The method according to any one of EEE 1 to 16 further comprises the following steps:
[0125] Receive the media content as a streaming media transmission;
[0126] The streaming speed of the media transmission is adjusted based on the determined playback speed of the media content.
[0127] EEE 18. A computer program product comprising instructions adapted to perform the method according to any one of EEE 1 to 17 when executed by a processing-capable device.
[0128] EEE 19. A computer-readable storage medium storing a computer program product according to EEE 18.
[0129] EEE 20. An apparatus configured to set a playback speed of media content including audio, the apparatus comprising circuitry configured to perform a method for setting the playback speed according to any one of EEE 1 to 17.
Claims
1. A computer-implemented method for setting the playback speed of media content including audio, the media content having a defined normal playback speed, the method comprising: Receive an instruction to play the media content at a speed different from the normal playback speed of the media content; Analyze the audio to determine the audio type; as well as Based on the determined audio type, a playback speed different from the normal playback speed is determined, and the playback speed of the media content is set to the determined playback speed. The method further includes: Select one or more predefined audio time stretching algorithms based on the determined audio type, and Set the audio time stretching algorithm of the media content to the selected audio time stretching algorithm. In this context, the metadata of the media content receives instructions for one or more predefined audio time-stretching algorithms.
2. The method according to claim 1, wherein, Determining a playback speed that differs from the normal playback speed includes selecting one or more predefined playback speeds based on the determined audio type.
3. The method according to claim 2, wherein, Receive one or more predefined playback speeds in the metadata of the media content.
4. The method according to any one of claims 1 to 3, wherein, Analyzing the audio to determine the audio type includes: Analyze the audio to determine whether it includes dialogue; and / or The audio is analyzed to determine whether it includes music.
5. The method according to any one of claims 1 to 3, further comprising: If the audio type cannot be determined, the playback speed is set to the default playback speed.
6. The method according to any one of claims 1 to 3, wherein, While playing the media content, the following steps are repeated: analyzing the audio to determine the audio type, determining a playback speed different from the normal playback speed based on the determined audio type, and setting the playback speed of the media content to the determined playback speed.
7. The method of claim 1, further comprising: If the audio type cannot be determined, the audio stretching algorithm of the media content is set to the default audio stretching algorithm.
8. The method according to claim 7, wherein, The default audio time stretching algorithm affects the pitch of the audio content.
9. The method according to claim 1, wherein, When it is determined that the audio includes dialogue, the audio time stretching algorithm of the media content is set to either the time-domain pitch-synchronous overlay TD-PSOLA algorithm or the pointer interval control overlay PICOLA algorithm.
10. The method according to claim 1, wherein, When it is determined that the audio includes music, the audio time stretching algorithm of the media content is set to the WSOLA algorithm based on waveform similarity.
11. The method according to claim 1, wherein, When it is determined that the audio corresponds to both dialogue and music, the audio time stretching algorithm of the media content is set to either the time-domain pitch-synchronous superposition TD-PSOLA algorithm or the pointer interval control superposition PICOLA algorithm.
12. The method according to claim 1, wherein, While the media content is being played, the following steps are repeated: analyzing the audio to determine the audio type, selecting one or more predefined audio time stretching algorithms based on the determined audio type, and setting the audio time stretching algorithm of the media content to the selected audio time stretching algorithm.
13. The method according to any one of claims 1 to 3, wherein, The step of analyzing the audio to determine the audio type includes: for at least one audio type, determining a confidence score related to the audio including the audio type, and determining whether the confidence score exceeds a threshold confidence score.
14. The method according to any one of claims 1 to 3, further comprising the following steps: The media content is received as a streaming media transmission, wherein the step of analyzing the audio is performed in real time.
15. The method according to any one of claims 1 to 3, further comprising the following steps: Receive the media content as a streaming media transmission; The streaming speed of the media transmission is adjusted based on the determined playback speed of the media content.
16. A computer program product comprising instructions adapted to perform the method according to any one of claims 1 to 15 when executed by a device having processing capabilities.
17. A computer-readable storage medium storing a computer program, the computer program including instructions adapted to perform the method according to any one of claims 1 to 15.
18. An apparatus configured to set a playback speed of media content including audio, the apparatus comprising circuitry configured to perform a method for setting the playback speed according to any one of claims 1 to 15.
Citation Information
Patent Citations
Variable speed playback
US20130031266A1
Content-based audio playback speed controller
US20170004858A1