Vocal attenuation mechanism in on-device apps
The system addresses the challenge of vocal track separation in media playback by using machine learning to predict quality scores and adjust track volumes, improving the singing experience through high-quality instrumental playback.
Patent Information
- Application Number
- JP2025530749
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-04
- Filing Date
- 2023-12-04
- Publication Date
- 2025-11-14
AI Technical Summary
Existing media playback systems lack the ability to effectively separate and attenuate vocal tracks from musical accompaniment, limiting interactive experiences such as singing along with music.
A system and method for providing music playback with vocal attenuation, using machine learning models to predict quality scores and enable users to adjust vocal and accompaniment track volumes, and a graphical user interface to manage vocal attenuation features.
Enables users to interactively adjust vocal and accompaniment volumes, enhancing the singing experience by ensuring high-quality playback of instrumental tracks.
Smart Images

Figure 2025537401000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to media playback, and more particularly to providing a system and method for music playback with vocal track reduction. [Background technology]
[0002] Media consumption is the frequent use of electronic devices. Users can use many types of electronic devices to access, collect, stream, and download their favorite music for playback. Now, it may be desirable to participate in an interactive experience with a media item. As an example, a user may want to sing along with their favorite music. [Brief explanation of the drawings]
[0003] [Figure 1] 1 illustrates a simplified network diagram in block diagram form according to one or more embodiments.
[0004] [Figure 2] 1 illustrates, in flowchart form, an exemplary method for media track suppression, according to one or more embodiments.
[0005] [Figure 3] 1 illustrates, in flowchart form, an exemplary method for media track quality analysis in accordance with one or more embodiments.
[0006] [Figure 4] 1 illustrates, in flowchart form, an exemplary method for media track quality analysis of a collection of media items, according to one or more embodiments.
[0007] [Figure 5] 1 illustrates an exemplary graphical user interface in accordance with one or more embodiments.
[0008] [Figure 6] FIG. 1 illustrates an exemplary system diagram of an electronic device in accordance with one or more embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0009] The present disclosure is directed to a system, method, and computer-readable medium for providing music playback with vocal reduction. Generally, techniques are disclosed for providing audio signal processing for separating tracks of a media file having an audio component. Additionally, techniques are disclosed for providing vocal attenuation functionality for media items having an audio component on a user device. In some embodiments, the user device may be a mobile computing device, a tablet computing device, or the like.
[0010] According to one or more embodiments, the disclosed technology addresses the need in the art for providing media playback with vocal attenuation. In one embodiment, a user device provides audio playback of a media file, such as a song, while allowing a user to adjust the volume levels of the media file's separate tracks. For example, while a song is playing, a user can adjust the volume of the vocal track, the accompaniment track, or both.
[0011] In accordance with one or more embodiments, the disclosed technology addresses the need in the art for providing attenuation of a media item or collection of media items on a client media device. A media item may include one or more audio components, such as vocals and musical accompaniment (e.g., instruments).
[0012] According to one or more embodiments, the disclosed techniques may include a model that flags media content based on one or more quality control metrics. The model may include, for example, a vocal suppression model that extracts instrument-only or accompaniment components of a media file. The extracted instrument-only components may be analyzed or processed by the model to predict a quality score for the media file. The predicted score may be subjective, objective, or a combination thereof. In some embodiments, a media item may be provided with a quality label. If the predicted quality score for the media file is below a threshold, vocal attenuation for the media file is disabled. If the predicted quality score is above the threshold, vocal attenuation for the media file is enabled. The threshold may be a predefined threshold in some embodiments. The predefined threshold for quality of the media file may be defined as a perceptual quality metric. In some embodiments, the perceptual quality metric threshold may be determined through human experimentation and / or feedback.
[0013] According to one or more embodiments, the decay quality model may consider additional factors that affect or may affect the decay quality of a media item, including, but not limited to, the genre of the media item (e.g., classical, jazz, rock, pop), the length of the media item being less than a certain threshold (e.g., 10 seconds), or non-musical content (e.g., spoken word, nature sounds, sound effects).
[0014] According to one or more embodiments, media items may be displayed on a graphical user interface. The graphical user interface may include a graphical representation of the media item, such as a song, music track, or the like. Additionally, a collection of media items may be displayed on the graphical user interface. The collection of media items may include, for example, one or more songs from an album, a playlist, or the like. Additionally, according to one or more embodiments, a media item or a graphical representation of a collection of media items may be represented by an icon. The icon may provide a visual indication to a user via the graphical user interface. In some embodiments, the icon may indicate that a vocal attenuation feature is available for the media item. Additionally or alternatively, the icon may indicate that a vocal attenuation feature is available for a collection of media items. In another embodiment, the icon may indicate that a vocal attenuation feature is available for a subset of media items in the collection of media items. In some embodiments, the icon may be a selectable icon. When selected, the icon may initiate a vocal attenuation feature for one or more associated media items. The vocal attenuation feature may be performed by an application running on the user device or on another device (e.g., a server, cloud-based). Alternatively, the vocal attenuation feature may be a service provided by an operating system, such as a mobile operating system.
[0015] In the following description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the disclosed concepts. As part of this description, some of the drawings of the present disclosure depict structures and devices in block diagram form to avoid obscuring the novel aspects of the disclosed embodiments. In this regard, references to elements in the drawings labeled with a number (e.g., 100) without an associated identifier should be understood to refer to all instances of the element in the drawing with the identifier (e.g., 100a and 100b). Additionally, as part of this description, some of the figures of the present disclosure may be provided in the form of flow diagrams. The boxes in any particular flow diagram may be presented in a particular order. However, it should be understood that the particular flow of any flow diagram is used only to illustrate one embodiment. In other embodiments, any of the various components shown in the flow diagrams may be omitted, or these components may be performed in a different order or simultaneously. Additionally, other embodiments may include additional steps not shown as part of the flow diagrams. The language used herein has been chosen solely for purposes of readability and explanation, and not to limit or restrict the disclosed subject matter. References to "one embodiment" or "one embodiment" in this disclosure mean that a particular feature, structure, or characteristic described with respect to an embodiment is included in at least one embodiment, and multiple references to "one embodiment" or "one embodiment" should not necessarily be understood as all referring to the same embodiment or to different embodiments.
[0016] It will be appreciated that in developing any actual implementation (as in any development project), numerous decisions must be made to achieve the developer's specific goals (e.g., conformance with system and business-related constraints), and that these goals may vary from implementation to implementation. It will also be appreciated that such a development effort may be complex and time-consuming, but would nevertheless be a routine undertaking for one skilled in the art of image capture having the benefit of this disclosure.
[0017] For purposes of this disclosure, media items are referred to as "songs." However, in one or more embodiments, media items referred to as "songs" may be any type of media item, such as audio media items, video media items, visual media items, text media items, podcasts, interviews, radio stations, etc.
[0018] 1, there is shown a simplified block diagram of a media service 100 connected to a client device 140, for example, via a network 150. Client device 140 may be a multifunction device such as a mobile phone, a tablet computer, a personal digital assistant, a portable music / video player, a wearable device, or any other electronic device that includes a media playback system.
[0019] Media service 100 may include one or more servers or other computing or storage devices in which various modules and storage may be included. While media service 100 is shown in an exemplary manner as including various components, in one or more embodiments, the various components and functionality may be distributed across multiple network devices, such as servers, network storage, etc. Furthermore, additional components may be used, and some combination of the functionality of any of the components may be combined. Generally, media service 100 may include one or more memory devices 112, one or more storage devices 114, and one or more processors 116, such as central processing units (CPUs) or graphical processing units (GPUs). Furthermore, processor 116 may include multiple processors of the same or different types. Memory 112 may each include one or more different types of memory that may be used in conjunction with processor 116 to perform device functions. For example, memory 112 may include cache, ROM, and / or RAM. Memory 112 may store various programming modules, including media management module 102, decay quality module 103, and decay module 104A, during execution.
[0020] The media service 100 can store media files, media file data, music catalog data, such as information about songs, albums, artists and creators, publishers, etc. Additional data can include, but is not limited to, media file attenuation quality data (e.g., quality metrics, thresholds), model training data (e.g., attenuation model training data, attenuation quality model training data), and quality flag data. The media service 100 can store this data in the media store 105 in storage 114. The storage 114 can include one or more physical storage devices. The physical storage devices can be located in a single location or distributed across multiple locations, such as multiple servers. Media files can include label data to indicate the availability of vocal attenuation and can be stored in the media store 105. In one or more embodiments, the label data can include information about songs or other media items, such as music videos, indicating the availability or non-availability of vocal attenuation. Additionally or alternatively, the label data can include information about complete albums or song collections, indicating the availability or non-availability of vocal attenuation.
[0021] In another embodiment, the media store 105 may include model training data for creating datasets for training models such as the attenuation quality module 103. The model training data may include labeled training data that a machine learning model uses to learn and then predict quality scores for media items. In some embodiments, the training data may consist of data pairs of input data and output data. For example, the input data may include media items to be processed along with corresponding quality labels. The quality labels may be objective or subjective. Furthermore, the quality labels may be obtained through detailed experiments with human listeners. Objective labels may indicate quality measures obtained, for example, by calculating classical source separation metrics. Subjective labels may indicate perceptual quality scores assigned by human annotators, for example.
[0022] In another embodiment, the media store 105 may include model training data for creating a dataset for training a model, such as the attenuation module 104A of the media service 100 or the attenuation module 104B of the client device 140. The model training data may include labeled training data that the machine learning model uses to learn and is then enabled to extract the instrument-only components of a song. In some embodiments, the training data may consist of data pairs of input data and output data. For example, the input data may include a song (e.g., a mix) with instrumental accompaniment and vocals, and the output data may include a song (e.g., an instrument) without vocals. Candidate pairs for a music catalog may include a pair of songs (song-1, song-2) if one or more of the following parameters are met: a) song-1 and song-1 are by the same artist or creator; b) song-1 and song-2 are metadata equivalents (e.g., they appear on the same album); c) song-1 and song-2 are approximately the same length in seconds (+ / - 1 second); d) one of song-1 and song-2 is tagged as instrumental; and e) neither song-1 nor song-2 is within a pre-defined excluded genre (e.g., comedy, sound effects, karaoke).
[0023] Returning to media service 100, memory 112 includes modules containing computer-readable code executable by processor 116 to cause media service 100 to perform various tasks. As shown, memory 112 may include media management module 102, decay quality module 103, and decay module 104A. According to one or more embodiments, media management module 102 manages a library or catalog of media content. The library may be user-specific, such as specific to a user of client device 140. Media management module 102 may provide media content in response to requests from client devices, such as client device 140.
[0024] The memory 112 also includes a damped quality module 103. In one or more embodiments, the attention quality module 103 may include a machine learning model, including a quality flagger tool, used to estimate the quality of an audio-suppressed song. The quality flagger tool may then assign a flag to the media item based on this estimation, where the flag is positive if the resulting audio is of good quality when the audio or vocals of the media item are suppressed. Alternatively, a negative flag may be assigned to the media item if the resulting audio is of poor quality when the audio or vocals of the media item are suppressed. The damped quality module's model may be trained using a training dataset, as described herein. The quality flagger tool may use the model to predict a score for a media item based on the quality of the media item when it is played with the vocals suppressed. The quality of an audio-suppressed media item may be based on multiple factors, including, but not limited to, the presence of audible artifacts or vocals that are still partially audible. The quality threshold may be predefined by a user of the system.
[0025] In some embodiments, the decay quality module may predict a quality metric for a collection of media items, such as an album or a playlist. In one example, all songs on an album may be given a positive flag, with vocal decay enabled except for one. In this example, the presence of one song with a negative flag may be considered undesirable, particularly if the song is assigned a quality score slightly below the threshold applied to the album's songs. This may be corrected by determining the ratio between the album's approved songs and the number of songs on the album. If this ratio is above a coverage threshold (e.g., 80%), the quality score of the negatively flagged track is compared to a quality threshold to obtain a difference. If this difference is below a tolerance threshold (i.e., the reduced quality is slightly low) or compared to a second threshold (i.e., a threshold lower than the first threshold), the flag may be changed from negative to positive for the media item, such that the song is considered available for vocal decay. In some embodiments, this adjusted flag using the second threshold or tolerance threshold may be utilized in an album experience context. That is, if the song in question is experienced in isolation from the album, the song may remain unflagged for vocal attenuation, but when experienced in the context of the album (i.e., when the user is listening to other songs from the same album consecutively within the same listening session), the song may be flagged for vocal attenuation.
[0026] Additional metrics may be considered by the quality flagger tool when evaluating media items, including, but not limited to, media items having a duration of less than 10 seconds, media items belonging to a predefined set of excluded genres (e.g., classical, instrumental), or the content of the media item not including music (e.g., spoken word, nature sounds, sound effects).
[0027] The memory 112 also includes an attenuation module 104A. In one or more embodiments, the attention module 104A may include a machine learning model for attenuating specific audio components of a media file, such as isolating the vocal track of a song that includes vocals and instrumental accompaniment. In some embodiments, the attenuation module may be configured to receive an audio item and output a modified version of the audio item having separately modifiable audio tracks. That is, according to some embodiments, the audio tracks do not need to be pre-separated.
[0028] A model for attenuating vocals may be created using a training dataset, as described herein. In some embodiments, an attenuation module may be provided on a user's mobile device, such as attenuation module 104B of client device 140. Additionally, attenuation module 104 may be used to adjust the audio characteristics of media items received from a network device, obtained locally on the client device, or stored, etc. The attenuation module may be used to provide playback of media items, such as songs with reduced or attenuated vocals, via a mobile device with high-quality instrumental accompaniment. In some embodiments, the attenuation module provides functionality that allows a user to modify the sound characteristics of the attenuated track separately from the rest of the audio item to which the isolated track belongs. As an example, a user can reduce the volume of a vocal track of a song without reducing the volume of the rest of the song's audio components. Similarly, in some embodiments, the attenuation module may allow a user to modify the sound characteristics of the rest of a media item while maintaining the audio characteristics of the isolated portion.
[0029] 2 illustrates in flowchart form an exemplary method for attenuating media items. The method may be implemented by attenuation module 104B on a client device, such as client device 140 of FIG. 1. For purposes of explanation, the following steps are described in the context of FIG. 1. However, it should be understood that various actions may be taken by alternative components. Furthermore, various actions may be performed in a different order. Furthermore, some actions may be performed simultaneously, some may not be required, or other actions may be added.
[0030] The flowchart begins at 205, where the media player 126 retrieves a media item on the client device 140. According to one or more embodiments, the media item may be retrieved from a media service, as shown in Figure 1. The media item may include at least one audio component consisting of vocals (e.g., singing) with instrumental accompaniment.
[0031] The flowchart continues at 210, where the media item is applied to a source separation method to generate separated audio tracks (e.g., vocals, instruments). As an example, the media item may be applied to an artificial neural network to perform the source separation method and separate the background (i.e., accompaniment) from the vocals. The neural network may generate an ideal mask for separating a target source, such as the background or vocals. In some embodiments, the technique includes providing a modified media item such that at least a portion of the media item (i.e., a particular audio track, such as a vocal track) is modifiable separately from the remainder of the media item.
[0032] The flowchart continues at 215, where the media player 126 obtains the modified media item from the source separation process. The modified media item may include one or more separated audio tracks, such as a background instrumental accompaniment track and a vocal track. The modified media item may be stored in the media store 128 of the storage 124, for example, by the client device 140.
[0033] The flowchart continues at 220, where the media player 126 provides the ability to modify the characteristics of the separated audio track during playback of the media item. In some embodiments, the listener may be provided with a user interface component, such as a volume slider, to adjust the volume level of the media item during playback. The listener can then effectively reduce the volume and hear only the instrumental accompaniment portion of the song (e.g., a sing-along feature). Alternatively, the listener may adjust the instrumental portion to a lower setting and increase the vocal track during playback. In some embodiments, printed lyrics of the vocal track may be shown to the listener. For example, the lyrics may be displayed on a portion of the listener's device in time with the media item during playback.
[0034] In some embodiments, an initial determination may be made as to whether to provide attenuation for a particular media item. FIG. 3 illustrates in flowchart form an exemplary method for generating a quality flag for a media item that includes at least one audio component. The quality flag may indicate a media item for which vocal attenuation is to be provided according to a prediction that the media item is suitable for attenuation. This method may be implemented by an attenuation quality module 103 on a media serving device, such as the media serving device 100 of FIG. 1. For purposes of explanation, the following steps are described in the context of FIG. 1. However, it should be understood that various actions may be taken by alternative components. Furthermore, various actions may be performed in different orders. Furthermore, some actions may be performed simultaneously, some may not be required, or other actions may be added.
[0035] The flowchart begins with module 103 obtaining a media item. In one example, the media item may be obtained from media store 105 in storage 114. The media item may include a purely audio media item or a media item that includes an audio component, such as a video item. The media item may include an audio component consisting of vocals and instrumental accompaniment.
[0036] The flowchart continues at 310, where module 103 obtains a quality metric for the media item. The quality metric provides an indication as to whether the media item should be considered suitable for, and therefore made available for, the vocal attenuation functionality described herein. In some embodiments, the quality metric may be provided in the form of a positive flag or some value that indicates the suitability of the media item for attenuation. The quality metric may be provided in several ways, such as user-generated, based on a set of quality parameters applied to the media item, etc.
[0037] In some embodiments, as shown at 315, module 103 applies the media item to a quality network having a machine learning model (e.g., an artificial neural network) to predict a quality metric. The quality metric may be a score assigned to the media file that predicts how a human listener will perceive the media item when heard with vocal attenuation enabled (e.g., whether the listener will perceive the media item positively or negatively). In some embodiments, the quality network may provide a value representing the predicted percentage of users who will find the media item suitable for attenuation.
[0038] The flowchart then determines whether the quality metric meets a quality threshold. The quality threshold may indicate an acceptable level of the media file when listening to sans vocals. If the quality metric meets the threshold, the flowchart proceeds to block 320, where access to attenuation of the media item is provided. For example, the media item may be available for attenuation on a playback device, and / or playback of the media item may be associated with a user interface component that allows a user to modify the audio characteristics of a particular portion of the media item separately and independently from the rest of the media item according to the attenuation. If the quality metric does not meet the threshold, the media item is left as is 325, and no attenuation functionality is provided for the media item.
[0039] According to some embodiments, decay of a media item may be provided based on a quality metric. However, in some embodiments, a tolerance threshold or secondary threshold may also be utilized in certain contexts, such as for a particular song in the context of an album. FIG. 4 illustrates in flowchart form an exemplary method for generating a quality flag for a collection of media items that include at least one audio component. The method may be implemented by a decay quality module 103 on a media service device, such as the media service device 100 of FIG. 1. For purposes of explanation, the following steps are described in the context of FIG. 1. However, it should be understood that various actions may be taken by alternative components. Furthermore, various actions may be performed in different orders. Furthermore, some actions may be performed simultaneously, some may not be required, or other actions may be added.
[0040] The flowchart begins at block 405, where media items are obtained for a media collection. A media collection may include a set of media items provided as part of a collection, such as a record, a series, etc. In some embodiments, the media items may be obtained in the context of a device requesting playback of one or more media items from the media collection.
[0041] Flowchart 400 continues at block 410, where a quality metric is obtained for each of the media items. The quality metric may be provided in the form of a value that represents the quality level of the media item with respect to decay (i.e., how well the media item performs with respect to decay). In some embodiments, the quality metric may be obtained from a user-provided value, by applying rule-based parameters to determine the value, or by subjecting the media items to a network trained to predict quality metrics.
[0042] At block 415, a determination is made as to whether a threshold portion of the media collection meets a quality threshold. As described above, whether attenuation is provided for individual media items may be based on a quality threshold. In some embodiments, at block 415, a determination is made as to which portion of the media items in the collection meet the quality threshold. The threshold portion may indicate a percentage or occupancy of the collection at which a reduced or acceptable threshold should be applied to the remainder of the items (i.e., items that do not meet the quality threshold) to provide access to attenuation for more items in the collection, for example, to improve the interactive experience of the collection. Thus, if the threshold portion does not meet the quality threshold, the flowchart ends at block 430, and attenuation is provided only for media items in the collection for which the media items meet the quality threshold.
[0043] Returning to block 415, if a determination is made that the threshold portion of the media item has been met, the flowchart proceeds to block 420, where a reduced quality threshold is applied to the remainder of the media item. For example, if media items must individually have a quality threshold of 0.6 to provide attenuation, and 8 of the 10 songs in an album have a quality metric of at least 0.6 (e.g., compared to the 70% threshold portion), the remaining 2 songs can be compared to a quality threshold of 0.5.
[0044] The flowchart ends at block 425, where access to decay is provided to media items that meet either the original quality metric or the reduced quality metric. That is, eight songs having at least six quality metrics can be provided with the decay feature. Additionally, if either of the remaining two songs in the album meet the reduced threshold of 0.5, they will also be provided with the decay feature.
[0045] FIG. 5 illustrates an exemplary graphical user interface according to one or more embodiments. Specifically, FIG. 5 illustrates a multifunction device 500 with a touchscreen that displays media content, such as in media player 126 of FIG. 1. Media management module 102 can generate a graphical user interface that includes a graphical representation 502 of one or more songs or other media items. The graphical representation may include album art or a representation of the artist. The graphical representation of the one or more songs may be modified or enhanced with a graphical indication of whether attenuation is available for the one or more songs. In another embodiment, a graphical indication may be provided to indicate whether attenuation is available for a collection of media items, such as a playlist, album, or the like. In at least one embodiment, the graphical indication may be a selectable icon. Upon selection of a vocal attenuation indicator, vocal attenuation for the corresponding song or track may be initiated as described above. During playback of the media item, lyrics may be displayed in portion 504 along with the vocal attenuation. In some embodiments, the lyrics may be displayed based on timestamp data for the media item. Playback controls 506 may be provided to the user, such as pausing, rewinding, or fast-forwarding the media item being played. Characteristics of the media item being played may also be controlled, as shown in portions 508 and 510. During playback of a media item with vocal attenuation, the user may manipulate a slider to control the volume level of the song or instrumental portions of the media item. Additionally or alternatively, the user may manipulate a slider to control the volume level of the vocal portion of the media item. Manipulation may be received as input from the user via the touchscreen, such as by tapping "+" to increase the volume or by tapping "-" to decrease the volume of either the song or vocals. Alternatively, the user may manipulate the volume level by, for example, sliding a finger left or right on the slider to decrease / increase the volume level accordingly.
[0046] 6, a simplified functional block diagram of an exemplary multifunction device 600 is shown, according to one embodiment. Multifunction device 600 may represent representative components for devices such as media service 100, media service 100, and client device 140 of FIG. 1. Multifunction electronic device 600 may include a processor 605, a display 610, a user interface 615, graphics hardware 620, device sensors 625 (e.g., proximity sensor / ambient light sensor, accelerometer, and / or gyroscope), a microphone 630, audio codec(s) 635, speaker(s) 640, communications circuitry 645, digital image capture circuitry 650 (e.g., including a camera system), video codec(s) 655 (e.g., supporting a digital image capture unit), memory 660, a storage device 665, and a communications bus 670. Multifunction electronic device 600 may be, for example, a digital camera or a personal electronic device such as a personal digital assistant (PDA), a personal music player, a mobile phone, or a tablet computer.
[0047] Processor 605 can execute instructions necessary to perform or control the operation of numerous functions performed by device 600 (e.g., image generation and / or processing as disclosed herein). Processor 605 can, for example, drive display 610 and receive user input from user interface 615. User interface 615 can allow a user to interact with device 600. For example, user interface 615 can take various forms, such as buttons, a keypad, a dial, a click wheel, a keyboard, a display screen, and / or a touch screen. Processor 605 can also be a system-on-chip, such as found in mobile devices, and can include a dedicated graphics processing unit (GPU). Processor 605 can be based on a reduced instruction-set computer (RISC) or complex instruction-set computer (CISC) architecture, or any other suitable architecture, and can include one or more processing cores. Graphics hardware 620 can be dedicated computing hardware for processing graphics and / or assisting processor 605 in processing graphics information. In one embodiment, graphics hardware 620 may include a programmable GPU.
[0048] Image capture circuit 650 may include two (or more) lens assemblies 680A and 680B, each having a distinct focal length. For example, lens assembly 680A may have a shorter focal length relative to the focal length of lens assembly 680B. Each lens assembly may have a distinct associated sensor element 690. Alternatively, two or more lens assemblies may share a common sensor element. Image capture circuit 650 may capture still images and / or video images. Output from image capture circuit 650 may be processed, at least in part, by video codec(s) 655 and / or processor 605 and / or graphics hardware 620, and / or a dedicated image processing unit or pipeline incorporated within circuit 665. Images so captured may be stored in memory 660 and / or storage 665.
[0049] Sensor and camera circuitry 650 can capture still and video images, which may be processed, at least in part, by video codec(s) 655 and / or processor 605 and / or graphics hardware 620, and / or a dedicated image processing unit incorporated within circuitry 650, in accordance with this disclosure. Images so captured may be stored in memory 660 and / or storage 665. Memory 660 may include one or more different types of media used by processor 605 and graphics hardware 620 to perform device functions. For example, memory 660 may include memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage 665 may store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. Storage 665 may include one or more non-transitory computer-readable storage devices including, for example, magnetic disks and tapes (fixed, floppy, and removable), optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM). Memory 660 and storage 665 may be used to tangibly store computer program instructions or code organized into one or more modules and written in any desired computer programming language. For example, when executed by processor 605, such computer program code may perform one or more of the methods described herein.
[0050] The scope of the disclosed subject matter should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the terms "comprising" and "wherein," respectively.
Claims
1. Retrieving a media item from a media service; applying the media items to a separator network; obtaining a modified media item having a separated audio track from the separator network; receiving a selection of the media item for playback; A method comprising:
2. providing the ability to modify one or more characteristics of the separated audio track during playback of the media item; The method of claim 1 further comprising:
3. The method of claim 2 , wherein the one or more characteristics include a volume output of the separated audio track during playback of the media item.
4. The method of claim 1 , wherein the separated audio track is a vocal track.
5. The method of claim 1 , wherein the media items include instruments and vocals.
6. The method of claim 1 , wherein the graphical representation of the media item includes at least one graphical indicator.
7. The method of claim 6 , wherein the at least one graphical indicator indicates that vocal attenuation is available for the media item.
8. The method of claim 1 , wherein the media item is a song.
9. The method of claim 1 , wherein the media item is an album.
10. Acquiring a media item; obtaining at least one quality metric for the media item; applying the media item to a quality network to obtain the at least one quality metric; determining whether the at least one quality metric meets a quality threshold; A method comprising:
11. enabling access to a vocal attenuation feature when the at least one quality metric meets the quality threshold; The method of claim 10 further comprising:
12. disabling access to a vocal attenuation feature when the at least one quality metric is below the quality threshold; 12. The method of claim 10 or 11, further comprising:
13. 13. The method of any one of claims 10 to 12, wherein the media items include at least one audio component.
14. The method of any one of claims 10 to 13, wherein the quality threshold is a predefined quality threshold.
15. obtaining a plurality of media items for a media collection; obtaining a quality metric for each of the plurality of media items; determining whether a threshold portion of the media collection meets a quality threshold; A method comprising:
16. responsive to determining that the threshold portion is below a quality threshold, enabling access to a vocal attenuation function only for media items of the plurality of media items having a quality metric above the quality threshold; 16. The method of claim 15, further comprising:
17. applying a reduced quality threshold to media items of the plurality of media items having a quality metric below the quality threshold; enabling access to a vocal attenuation feature for media items of the plurality of media items having quality metrics above the quality threshold and the reduced quality threshold; 17. The method of claim 16, further comprising:
18. The method of claim 17 , wherein the quality threshold and the reduced quality threshold are predefined thresholds.
19. 19. The method of any one of claims 15 to 18, wherein the media collection is an album.
20. 19. The method of any one of claims 15 to 18, wherein the media collection is a playlist.
Citation Information
Patent Citations
Singing voice separation with deep u-net convolutional networks
US20200043517A1