System and method for audio source segmentation to cancel human voices
By extracting segments from the audio source and using a voice removal algorithm to remove human voices in real time, the problem of poor karaoke experience in existing technologies is solved, achieving high-quality audio source playback while maintaining the integrity of non-human voice components.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARMAN BECKER AUTOMOTIVE SYST GMBH
- Filing Date
- 2025-10-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to effectively remove human voice components from audio sources in real time, resulting in a poor karaoke experience. Furthermore, conventional methods may damage non-human voice components, reducing audio quality.
By receiving an audio source, extracting segments, and using a voice removal algorithm to remove human voice components in real time, a modified segment is generated. During playback, crossfade-in/fade-out or sequential playback is performed to reduce playback gaps and maintain the integrity of non-human voice components.
It achieves near real-time removal of human voices from the audio source, improving the quality of the karaoke experience while keeping other components of the audio source unchanged, thus providing an improved karaoke experience.
Smart Images

Figure CN121963676A_ABST
Abstract
Description
Systems and methods for audio source segmentation to remove human voices Technical Field
[0001] The various implementation schemes generally involve audio processing, and more specifically, audio source segmentation to eliminate human voices. Background Technology
[0002] Modern vehicles include in-vehicle infotainment (IVI) systems that receive audio and video input from various sources. IVI systems include various output devices, such as displays and speakers located throughout the vehicle. IVI systems take input (such as audio input) selected by the user from local or remote audio sources and play the audio input using the output devices in the vehicle.
[0003] A karaoke experience can be provided by an IVI system and involves singing along to a pre-recorded audio accompaniment played by the IVI system's audio output device. The user sings along to the pre-recorded accompaniment, and in some instances, the user's voice is captured using a microphone and reproduced using the same audio output device that played the pre-recorded accompaniment. In some cases, users prefer to utilize an audio source with lead vocals and / or background vocals removed. Some pre-recorded audio accompaniments are created specifically for karaoke experiences by pre-processing songs to remove vocal components. Pre-processing is typically performed by personnel such as audio engineers or producers, or by automatic vocal removal algorithms, and the pre-processed song is provided as an audio source to the audio playback system. In other examples, pre-recorded audio accompaniments for karaoke experiences are created by recording an instrumental version of a song without lead vocals and / or backing vocals. In either scenario, creating a version of a song for a karaoke experience requires pre-processing or pre-recording of the song for the karaoke experience. Another technique for providing a karaoke experience is to play a song and allow the user to sing an unmodified version of it. However, the karaoke experience provided by audio sources that include human voices can lead to a poor karaoke experience for many users.
[0004] Some karaoke experiences offer mechanisms for real-time suppression of vocals during the karaoke experience. One technique for real-time vocal suppression is to perform mid-frequency ducking on the audio source, which reduces the volume of the mid-frequency component of the audio signal, which typically contains vocals. However, when using mid-frequency ducking, other components in the audio besides vocals (such as instrumental components) are eliminated, thus reducing the quality of the karaoke experience. Center channel ducking or suppression is a technique used for 5.1, 7.1, or other multi-channel audio sources with a discrete center channel. However, many audio sources, including music, are typically two-channel audio sources lacking a discrete center channel.
[0005] One drawback of using conventional techniques to remove vocals from audio sources for a karaoke experience is the inability to utilize many vocal removal algorithms in real time. Vocal removal algorithms typically require significant processing time, preventing their application in real-time to streaming audio sources. Furthermore, using pre-recorded karaoke versions of songs does not allow users to experience karaoke from all audio sources played by the audio playback system. Using unmodified audio sources containing vocals for a karaoke experience results in a poor user experience. Performing mid-frequency ducking or center channel ducking also reduces the quality of the karaoke experience by eliminating components other than vocals from the audio source.
[0006] As previously stated, there is a need in the art for more effective techniques to process audio sources that provide users with an acceptable karaoke experience. Summary of the Invention
[0007] In various embodiments, a computer-implemented method includes: receiving an audio source for playback; extracting a first segment of the audio source, the first segment including a first portion of the audio source; removing a first vocal component from the first segment to create a first modified segment; and causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
[0008] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the vocal components of an audio source (such as a song containing the vocal components that the user desires for a karaoke experience) are essentially removed in real time. By removing the vocal components from the song in real time, a karaoke experience is provided for any number of audio sources available for streaming playback. Furthermore, by utilizing a vocal removal algorithm instead of techniques such as mid-frequency or center channel ducking to remove the vocal components of the song, the quality of the karaoke version of the audio source is improved because other non-vocal components of the audio source remain unchanged when playing the karaoke version. Therefore, playing an audio source without vocal components along with sound input captured by one or more microphones provides an improved karaoke experience. These technical advantages provide one or more technical advancements superior to existing methods. Attached Figure Description
[0009] By referring to various embodiments, one can understand in detail the aforementioned features of the various embodiments in the more specific description of the inventive concept briefly described above, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concept and should therefore not be construed as limiting the scope in any way, and other equally effective embodiments exist.
[0010] Figure 1 shows a block diagram of a computing device configured to implement one or more aspects of the present disclosure.
[0011] Figure 2 shows a block diagram of an IVI system configured to implement one or more aspects of this disclosure.
[0012] Figure 3 shows an example of an audio source processed according to one or more aspects of this disclosure.
[0013] Figure 4 shows an example of an audio source processed according to one or more aspects of this disclosure.
[0014] Figure 5 shows another example of an audio source processed according to one or more aspects of this disclosure.
[0015] Figure 6 is a flowchart of method steps for processing an audio source according to one or more aspects of this disclosure. Detailed Implementation
[0016] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details. For illustrative purposes, multiple instances of the same object are indicated where necessary by reference numerals identifying the object and parenthetical numbers identifying the instances.
[0017] Figure 1 shows a block diagram of an audio playback system configured to implement one or more aspects of the present disclosure. As shown, the audio playback system 100 includes, but is not limited to, a computing device 110, an audio source 120, an input module 130, and an output module 140. The computing device 110 includes, but is not limited to, a processing unit 112 and a memory 114. The memory 114 includes, but is not limited to, an audio playback application 116 and a data storage area 118. The data storage area 118 includes, but is not limited to, a voice cancellation algorithm 122.
[0018] In operation, computing device 110 executes audio playback application 116 to control audio playback. In one example, audio is played from one or more vehicle components or sources, either inside or outside the vehicle. Specifically, processing unit 112 executes audio playback application 116 and causes audio to be played on one or more output devices associated with audio playback system 100. Audio playback application 116 receives audio sources 120, such as terrestrial or satellite radio signals, music or other content obtained from streaming audio services, audio files stored on storage devices associated with the vehicle, or audio content streamed from another device, such as a Bluetooth device to which computing device 110 is connected.
[0019] The audio playback application 116 also combines the audio source 120 played by the audio playback system 100 to provide a karaoke experience for the user. For example, the audio playback application 116 receives audio input from the input module 130, such as human voice input detected by a microphone associated with the audio playback system 100. The audio playback application 116 plays the audio input along with the audio source 120 on an audio output device (such as one or more speakers). In some cases, in addition to playing the audio source 120 and the audio input, the audio playback application 116 also plays video content on a display in the vehicle or switches interior or exterior lighting to enhance the karaoke experience.
[0020] The computing device 110 includes a processing unit 112 and a memory 114. In various embodiments, the computing device 110 is a device including one or more processing units 112, such as a system-on-a-chip (SoC). In various embodiments, the computing device 110 is a mobile computing device wirelessly connected to other devices in the vehicle, such as a tablet computer, mobile phone, media player, etc. In some embodiments, the computing device 110 is a host unit included in a vehicle system. Alternatively or additionally, the computing device 110 may be a detachable device installed as part of a separate console within a part of the vehicle. Typically, the computing device 110 is configured to coordinate the overall operation of the audio playback system 100. The embodiments disclosed herein are contemplated as any technically feasible system configured to implement the functionality of the audio playback system 100 via the computing device 110. The functionality and technology of the audio playback system 100 are also applicable to other types of transportation, including consumer vehicles, commercial trucks, airplanes, helicopters, spacecraft, ships, submarines, etc.
[0021] Processing unit 112 may include one or more central processing units (CPUs), digital signal processing units (DSPs), microprocessors, application-specific integrated circuits (ASICs), neural processing units (NPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and so on. Processing unit 112 typically includes a programmable processor that executes program instructions to manipulate input data and generate outputs. In some embodiments, processing unit 112 may include any number of processing cores and other modules to facilitate program execution.
[0022] Memory 114 includes memory modules or a collection of memory modules. Memory 114 typically includes memory chips, such as random access memory (RAM) chips, which store applications and data for processing by processing unit 112. In various embodiments, memory 114 includes non-volatile memory, such as optical drives, magnetic drives, flash memory drives, or other storage devices. Audio playback application 116 within memory 114 is executed by processing unit 112 to realize the overall functionality of computing device 110 and thus coordinate the operation of the entire audio playback system 100.
[0023] Audio playback application 116 processes audio source 120 and / or audio input received from input module 130 to reproduce audio signals. In various embodiments, audio playback application 116 plays the audio source along with voice input from one or more occupants or users of the vehicle via output module 140. Voice input is acquired via input module 130 to provide a karaoke experience. Additionally, audio playback application 116 processes audio source 120 to remove vocal components from audio source 120, thereby providing an improved karaoke experience. Audio playback application 116 removes vocal components from audio source 120 by splitting audio source 120 into one or more segments. These segments are provided to a vocal removal algorithm, which removes vocals from the segments to generate modified segments. Audio playback application 116 then plays the modified segments via output module 140. The modified segments are played sequentially in the order they originally existed in audio source 120.
[0024] Audio playback application 116 buffers audio source 120 by extracting an initial segment from audio source 120 and processing that initial segment using voice cancellation algorithm 122. The initial segment provides a buffer that allows audio playback application 116 to process one or more additional or subsequent segments of audio source 120 using voice cancellation algorithm 122 during playback of the initial segment. In some implementations, the length of the initial segment of audio source 120 is chosen to provide sufficient processing time for audio playback application 116 to process subsequent segments of audio source 120, such that processing of subsequent segments is completed before playback of the initial segment is finished. Subsequent segments may be the same size as the initial segment or constitute the entire remainder of audio source 120.
[0025] The audio playback application 116 also utilizes techniques to reduce the user's perception of any potential gaps between modified segments played by the audio playback system 100. In one example, segments extracted from audio source 120 overlap in time, such that the end of the first segment overlaps in time with the beginning of the second segment. In this scenario, the second segment follows the first segment from audio source 120. The audio playback application 116 processes these segments to remove vocal components from the respective segments to produce the modified segments. The audio playback application 116 then causes the audio playback system 100 to crossfade in and out of the modified segments to create a smooth transition between the two segments from audio source 120. In another example, the audio playback application 116 plays the first modified segment sequentially, then the second modified segment, without crossfading in or out. The size of the segments extracted from audio source 120 and modified by the audio playback application 116 can be varied using different techniques. For example, the first segment may be relatively small compared to subsequent segments, allowing the audio playback application 116 to process the first segment to remove vocal components and begin playing the modified first segment to reduce or eliminate any perceived delay by the user when the audio playback system 100 plays the audio source 120. The second segment may be larger than the first segment, but only large enough that the audio playback application 116 can complete processing the second segment before the modified first segment finishes playing. In some instances, if the computing device 110 has sufficient processing resources to complete processing the larger second segment before the modified first segment finishes playing, the second segment may include the entire remainder of the audio source 120. Therefore, the audio playback application 116 processes the second segment before the modified first segment finishes playing to eliminate any gaps in the playback of the audio source 120 between the modified first segment and the modified subsequent segments. In another embodiment, the extracted segments are of equal size, except perhaps the last segment of the audio source 120. In another scenario, the size of the audio segment extracted from the audio source 120 is selected based on the minimum or maximum input size supported by the voice removal algorithm 122 used by the audio playback application 116 to remove human voice components from the corresponding segment.
[0026] To remove vocal components from audio source 120, audio playback application 116 provides the extracted audio segments to vocal removal algorithm 122. Computing device 110 executes vocal removal algorithm 122 to remove vocal components from the audio source. Audio playback application 116 may utilize more than one vocal removal algorithm 122 selected based on attributes of audio source 120 and / or user preferences. For example, some vocal removal algorithms 122 are configured to remove vocal components from certain types of content or music genres better than other algorithms. Therefore, audio playback application 116 analyzes metadata associated with audio source 120 to determine the content type or genre of audio source 120 and selects a vocal removal algorithm 122 based on the content type or genre. In another example, the user selects different karaoke modes or configuration parameters associated with karaoke modes provided by audio playback application 116. A first mode removes all vocals from audio source 120 according to user preferences. A second mode removes only the main vocals, but secondary or backup vocals remain in audio source 120. In one example, the audio playback application 116 selects the karaoke mode and voice removal algorithm 122 based on the detected presence of other occupants in the vehicle. For instance, if more than one or two occupants are detected in the vehicle, the audio playback application 116 can select the voice removal algorithm 122 to remove all sounds from the audio source 120 when all occupants in the vehicle want to participate in the karaoke experience. If only one occupant is detected in the vehicle, the audio playback application 116 can select the voice removal algorithm 122 to remove only the occupant's voice from the audio source 120. Therefore, the audio playback application 116 selects the appropriate voice removal algorithm 122 based on the selected mode or user preference regarding which voices should be removed from the audio source 120. In another example, different voice removal algorithms 122 provide different performance or output results. Therefore, the user can select different voice removal algorithms 122 provided by the audio playback application 116 to drive the karaoke mode based on the performance characteristics of the selected voice removal algorithm 122.
[0027] Data storage area 118 is part of memory 114 that locally stores various types of data, including voice removal algorithm 122 and other data (not shown), such as content items, data tables, etc. For example (a table that maps audio tones to events) and / or application data associated with the audio playback application 116 ( For example(This includes security application data, metadata, etc.). In various embodiments, data storage area 118 may include volatile memory and may correspond to a portion of non-volatile memory. In some embodiments, computing device 110 may synchronize data between volatile and non-volatile memory, such that copies of the data are stored in both volatile and non-volatile memory. In some embodiments, data storage area 118 stores downloaded audio files obtained from a network source or other remote source. The audio files can be played via output module 140 through audio playback application 116.
[0028] As described above, the voice removal algorithm 122 in the data storage area 118 includes one or more algorithms utilized by the audio playback application 116 to remove the main voice and / or secondary voices from the audio source 120 played by the output module 140. The audio playback application 116 may utilize multiple voice removal algorithms 122 depending on user preferences or the detected voice characteristics of the audio source 120. Furthermore, some voice removal algorithms 122 operate to remove only the main voice in the input, while others operate to remove all voices from the audio source. Therefore, the audio playback application 116 selects a specific voice removal algorithm 122 for removing voice components from the audio source 120 based on a selected karaoke mode or the user's selection of the voice removal algorithm 122.
[0029] Audio source 120 includes one or more data sources that provide audio signals for reproduction. Audio source 120 includes pre-recorded audio accompaniment, such as songs. In various embodiments, audio source 120 is included in devices within a vehicle, such as an entertainment subsystem included in the vehicle's head unit, a rear-seat entertainment console, a device installed in the vehicle, etc. In some embodiments, audio source 120 is included in a mobile device, wearable device, and / or other portable device connected to audio playback application 116. Additionally, audio source 120 can be located remotely from the vehicle. In such instances, a remote data source streams audio source 120 to computing device 110, and then audio playback application 116 transmits audio source 120 to an output device associated with output module 140 for reproduction.
[0030] Input module 130 includes one or more means for performing measurements and / or acquiring data relating to certain subjects in the environment. In various embodiments, input module 130 generates sensor data relating to a user and / or non-user objects in the environment. In some embodiments, input module 130 is coupled to and / or included within computing device 110 and transmits sensor data to processing unit 112.
[0031] In various embodiments, input module 130 includes audio sensors, such as built-in microphones and / or microphone arrays, for recording sounds within the vehicle's cabin. Vehicle occupant sensors include, for example, optical sensors, such as RGB cameras, infrared cameras, depth cameras, and / or camera arrays, including two or more of such cameras oriented towards the vehicle's seating area. Cabin sensors include, for example, pressure sensors integrated into the seat positions within the vehicle, which detect when an occupant is seated in a specific seating position within the vehicle. In some embodiments, input module 130 includes touch sensors, position sensors, etc. For example Accelerometers and / or inertial measurement units (IMUs) or other types of sensors that record the presence, body position, and / or movement of users within the vehicle.
[0032] In some implementations, the input module 130 includes physiological sensors, such as a heart rate monitor, an electroencephalogram (EEG) system, a radio sensor, a thermal sensor, and a skin conductance sensor. For example The input module 130 includes devices capable of receiving input, such as a keyboard, mouse, touchscreen, and other input devices for providing input data to the computing device 110. In various embodiments, the input module 130 is associated with a specific console, such as a personalized screen mounted to a portion of the seat or a console-specific input component.
[0033] Output module 140 includes one or more devices capable of providing output, such as a display screen or a speaker. In various embodiments, one or more of input module 130 or output module 140 are incorporated into or located outside computing device 110. In some embodiments, computing device 110, input module 130, or output module 140 may be components of an IVI system or entertainment subsystem included in a vehicle.
[0034] Vehicle system
[0035] Figure 2 illustrates an example IVI system 200 including the audio playback system 100 of Figure 1 according to various embodiments. As shown, the IVI system 200 includes, but is not limited to, an input module 130, a computing device 110, and an output module 140. The input module 130 includes, but is not limited to, one or more microphones 222, occupant-facing sensors 226, and cabin sensors 228. The computing device 110 includes, but is not limited to, an audio playback application 116. The output module 140 includes, but is not limited to, a speaker 230, a display 232, and a human-machine interface (HMI) 234. The audio playback application 116 includes, but is not limited to, an input processing module 236 and an output generation module 238.
[0036] In some embodiments, the computing device 110 may be integrated into the vehicle's host unit. The host unit is a component of the vehicle, installed in any location within the vehicle's passenger compartment in any technically feasible manner. In some embodiments, the host unit includes any number and type of instruments and applications and provides any number of input and output mechanisms. For example, the host unit enables users ( For example The driver and / or passengers can control the IVI system. The main unit supports any number of input and output data types and formats known in the art. For example, the main unit may include built-in Bluetooth (for hands-free calling and / or audio streaming), USB connectivity, voice recognition, camera input via input module 130, video output via output module 140 for any number and type of displays 232, and any number of audio outputs. Generally, any number of sensors, displays, receivers, transmitters, etc., can be integrated into the main unit or implemented externally. Additionally, the computing device 110 may be located in other locations within the vehicle, such as behind an interior panel hidden in a storage compartment not visible to passengers.
[0037] In operation, audio playback application 116 receives audio source 120 and causes speaker 230 associated with output module 140 to play a modified version of audio source 120 processed by audio playback application 116. Audio source 120 includes songs, radio stations, or other audio sources that can be played or streamed by computing device 110. In one scenario, a user of IVI system 200 activates karaoke mode of audio playback application 116 via HMI 234 and selects audio source 120. The modified version of audio source 120 is a version of audio source 120 from which the audio playback application 116 has removed the lead vocals or all vocal components. To remove vocal components from audio source 120, audio playback application 116 extracts multiple segments from audio source 120. These multiple segments are provided as input to vocal removal algorithm 122, which removes the lead vocals and / or secondary vocal components from the input and returns the modified segments. Audio playback application 116 sequentially plays modified segments to provide a karaoke experience for occupants of the vehicle implementing IVI system 200. In one example, audio playback application 116 plays a first modified segment via output module 140, while subsequent audio segments are processed by voice removal algorithm 122 to generate the next consecutive modified audio segment. Audio playback application 116 completes processing of the next consecutive modified audio segment before the first modified segment finishes playing. In this way, the next consecutive modified audio segment is ready to play before the first modified segment finishes playing, allowing audio playback application 116 to remove human voice components from audio source 120 in near real-time. In some examples, the only delay experienced by the user is the processing time of audio playback application 116 processing the first segment extracted from audio source 120.
[0038] The audio playback application 116 also detects audio input from one or more microphones 222 of the input module 130. Audio input refers to voice input obtained from one or more microphones 222 within the vehicle, such as from a vehicle occupant participating in a karaoke experience. The audio playback application 116 causes the speaker 230 of the output module 140 to play the audio input in addition to playing the audio source 120. In some cases, the audio playback application 116 modifies the audio input by applying compression, reverb, auto-tuning, or other effects. The audio playback application 116 plays the audio input along with the audio source 120 on an audio output device, such as one or more speakers. In some cases, in addition to playing the audio source 120 and the audio input, the audio playback application 116 also plays video content on a display within the vehicle or switches interior or exterior lighting to enhance the karaoke experience.
[0039] The audio playback application 116 also detects the number and / or location of occupants within the vehicle based on input received from the input module 130. For example, the audio playback application 116 detects seat positions within the vehicle based on sensor data from one or more microphones 222, occupant-facing sensors 226, or cabin sensors 228. For instance, the audio playback application 116 determines that more than one occupant is present in the vehicle and selects a voice cancellation algorithm 122 that removes both the master and secondary voices from the audio source 120. As another example, the audio playback application 116 determines that only one occupant is present in the vehicle and selects a voice cancellation algorithm 122 that removes only the master voice from the audio source 120. Additionally, the audio playback application 116 can apply lighting effects using interior or exterior vehicle lighting customized based on the number of detected occupants or the detected seat positions of the vehicle occupants. These lighting effects or other customizations can be defined by a user profile stored in the data storage area 118.
[0040] Input module 130 includes various types of sensors, one or more microphones 222, occupant-facing sensors 226, and cabin sensors 228. In some cases, input module 130 also includes, but is not limited to, vehicle sensors such as external cameras, external microphones, accelerometers, etc. Occupant-facing sensors 226 include cameras or motion sensors oriented to detect the presence of occupants within the vehicle. In some cases, occupant-facing sensors 226 may also detect users based on facial recognition, allowing audio playback application 116 to identify user profiles specifying karaoke experience preferences, such as the selection of a particular voice cancellation algorithm 122. Cabin sensors 228 include other types of sensors, such as pressure sensors, temperature sensors, or other types of sensors that also detect the presence of occupants within the vehicle. In various embodiments, input module 130 provides audio playback application 116 with a combination of sensor data, which can use input acquired by one or more microphones 222 and sensor data from occupant-facing sensors 226 and cabin sensors 228 to determine the number of occupants within the vehicle or the seating positions of the occupants. Additionally, when a user selects karaoke mode in the vehicle, the input module 130 provides audio input from one or more microphones 222, which can be played using the vehicle's speakers 230.
[0041] Output module 140 includes various types of output devices, including but not limited to speaker 230, display 232, and HMI 234. Output module 140 performs one or more actions in response to output signals from computing device 110 or other subsystems within the vehicle. For example, output module 140 receives audio output from computing device 110, which may include multiple audio outputs mixed together by computing device 110. Output module 140 uses speaker 230 within the vehicle to play the audio output. For example, audio playback application 116 mixes audio source 120 with audio input detected by one or more microphones 222 and transmits the audio output, including both audio source 120 and audio input, to output module 140, which uses speaker 230 to play the audio. As another example, output module 140 receives additional information from computing device 110, causing display 232 or HMI 234 to display notifications, messages, alarms, or other information.
[0042] Figure 3 illustrates an example of an audio source 120 processed according to one or more aspects of this disclosure. Figure 3 shows how an audio playback application 116 extracts segments from the audio source 120 and processes the segments to remove vocal components to generate modified segments, which are then played to provide a karaoke experience.
[0043] Figure 3 depicts an audio source 120 provided to an audio playback application 116. The audio source 120 represents a song retrieved from a data storage area 118 or a song streamed from a streaming audio source or a terrestrial or satellite radio station. Therefore, when the audio playback application 116 receives the audio source 120, it extracts one or more audio segments 302 from the audio source 120. In the example of Figure 3, the user activates a karaoke mode provided by the audio playback application 116, which causes the audio playback application 116 to extract audio segments 302 from the audio source 120 and remove vocal components from the corresponding audio segments 302. As shown in Figure 3, the audio playback application 116 first extracts an audio segment 302a, which has a length of t milliseconds, where t represents the time slice of the audio segment 302a. The audio playback application 116 provides the audio segment 302a as input to the vocal removal algorithm 122 and receives the modified audio segment in which vocal components have been removed.
[0044] The audio playback application 116 also extracts audio segment 302b following audio segment 302a from the audio source 120. As shown in Figure 3, audio segments 302a and 302b overlap in time, so that once audio segments 302a and 302b are modified by the voice cancellation algorithm 122, the audio playback application 116 can play the modified audio segment by minimizing or eliminating the perceptual gaps between one or more audio segments, including audio segments 302a and 302b. In the example in Figure 3, the audio playback application 116 extracts audio segment 302 from the audio source 120 every t / 3 milliseconds, and the size of audio segment 302 is t milliseconds, which causes audio segments 302a and 302b to overlap in time. The audio playback application 116 continues to extract additional audio segments 302, such as audio segment 302c extracted t / 3 milliseconds after the start of audio segment 302b, which overlaps temporally with one or more of audio segments 302a or 302b, and so on. Examples of this disclosure may utilize other levels of temporal overlap between adjacent segments. Furthermore, according to examples of this disclosure, the temporal overlap between audio segments 302 extracted from audio source 120 does not require the removal of vocal components from audio source 120.
[0045] In one example, audio segments 302a and 302b are the same size. The size of audio segment 302 is chosen such that audio segment 302 can be processed by audio playback application 116 to eliminate vocal components within a time amount equal to or less than the time required to play the previous audio segment 302. In other words, the size of audio segment 302 is chosen such that audio playback application 116 processes the subsequent segment before the previous segment has finished playing. By completing the processing of a segment before the previous segment has finished playing, playback gaps are eliminated, and vocal components are eliminated from the audio source essentially in real time.
[0046] Figure 4 illustrates an example of an audio source 120 processed according to one or more aspects of this disclosure. Figure 4 shows additional details about how an audio playback application 116 extracts segments from the audio source 120, processes the segments to remove vocal components to generate modified segments, and generates modified audio sources 410 that are played to provide a karaoke experience.
[0047] Figure 4 illustrates audio segment 402, which includes audio segments 402a, 402b, and 402c. Audio segments 402a, 402b, and 402c represent audio segments 302a, 302b, and 302c that have had their human voice components removed by the human voice removal algorithm 122. Therefore, the human voice removal algorithm 122 outputs audio segments 402a, 402b, and 402c with their human voice components removed. As discussed in Figure 3 above, audio segments 302 overlap with each other in time. Therefore, audio segments 402 also overlap with each other in time. The temporal overlap of audio segments 402 allows the audio playback application 116 to play audio segment 302 in a manner that minimizes the playback gaps perceived by the user.
[0048] In the example of Figure 4, audio playback application 116 extracts sub-segments from audio segment 402. For example, it extracts audio sub-segment 408a from audio segment 402a, audio sub-segment 408b from audio segment 402b, and audio sub-segment 408c from audio segment 402c. In one example, the sizes of audio sub-segments 408a, 408b, and 408c are chosen such that they match the amount of temporal overlap between audio segments 302. Once the voice removal algorithm 122 has finished processing audio sub-segment 402a to create a modified segment without voice components, playback of audio sub-segment 408a begins. In the case where the first segment is playing after Karaoke mode is activated, audio playback application 116 may cause the entire audio segment 402a to be played from the beginning of audio segment 402a because no previous audio segment 402 is being played. In some examples, audio playback application 116 may cause the unmodified portion of audio source 120 to be played before audio segment 402.
[0049] Once the audio playback application 116 processes and outputs audio segment 402b from the voice removal algorithm 122, it extracts audio sub-segment 408b from the audio segment 402b. Then, once the audio playback application 116 detects that the playback of audio sub-segment 408a has been completed, it initiates playback of audio sub-segment 408b. Once the audio playback application 116 processes and outputs audio segment 402c from the voice removal algorithm 122, it extracts audio sub-segment 408c from the audio segment 402c. Then, once the audio playback application 116 detects that the playback of audio sub-segment 408b has been completed, it initiates playback of audio sub-segment 408c. This process can continue until the playback of audio source 120 and subsequent audio sources 120 is completed, or until the user disables the karaoke mode provided by the audio playback application 116. By playing audio segments 408a, 408b, and 408c sequentially, the audio playback application 116 reduces or eliminates perceived discontinuities or playback gaps between audio segments 408a, 408b, and 408c during playback.
[0050] Therefore, the modified audio source 410 is generated by sequentially playing audio sub-segments 408a, 408b, 408c, and subsequent sub-segments generated by the audio playback application 116. The modified audio source 410 represents the audio source 120 from which the audio playback application 116 has removed the human voice component. In some implementations, the introduced playback delay is equivalent to the processing time of the first audio segment processed by the audio playback application 116. It should be understood that the audio playback application 116 can process a given audio source 120 without extracting audio sub-segments from audio segment 302, but can directly provide audio segment 302 to the human voice removal algorithm 122 to remove the human voice content, and then sequentially play the modified audio segments output by the human voice removal algorithm 122.
[0051] Figure 5 illustrates another example of an audio source 120 processed according to one or more aspects of this disclosure. Figure 5 shows additional details about how the audio playback application 116 extracts segments from the audio source 120, processes the segments to remove vocal components to generate modified segments, and generates modified audio sources 510 that are played to provide a karaoke experience.
[0052] Similar to the examples in Figures 3 and 4, Figure 5 shows audio segment 402, which includes audio segments 402a, 402b, and 402c. Audio segments 402a, 402b, and 402c represent audio segments 302a, 302b, and 302c that have had their human voice components removed by the human voice removal algorithm 122. Therefore, the human voice removal algorithm 122 outputs audio segments 402a, 402b, and 402c with their human voice components removed. As discussed in Figures 3 and 4 above, audio segments 302 overlap with each other in time. Therefore, audio segments 402 also overlap with each other in time. The temporal overlap of audio segments 402 allows the audio playback application 116 to play audio segments 302 in a way that minimizes the playback gaps perceived by the user.
[0053] In the example of Figure 5, audio playback application 116 extracts sub-segment 508 from audio segment 402. In the example of Figure 5, the size of sub-segment 508 extracted from audio segment 402 is larger than in the example of Figure 4 to illustrate the concept of various sizes of audio segments and sub-segments that can be utilized. Additionally, various additional playback techniques can be utilized, such as crossfading the modified segment output by voice removal algorithm 122. Audio sub-segment 508a is extracted from audio segment 402a, audio sub-segment 508b is extracted from audio segment 402b, and audio sub-segment 508c is extracted from audio segment 402c. In this example, the sizes of audio sub-segments 508a, 508b, and 508c are chosen such that they are greater than the amount of time overlap between audio segments 302, so that the modified audio source 510 is generated by crossfading audio sub-segments 508a, 508b, and 508c with each other. Once the voice removal algorithm 122 has finished processing the audio segment 402a to create a modified segment without any vocal components, playback of the audio segment 508a begins. In the case where the first segment is playing after Karaoke mode is activated, the audio playback application 116 may cause the entire audio segment 402a to play from the beginning, since no previous audio segment 402 is being played. In some examples, the audio playback application 116 may cause the unmodified portion of the audio source 120 to play before the audio segment 402.
[0054] Once the audio playback application 116 processes and outputs the audio segment 402b from the voice cancellation algorithm 122, it extracts the audio sub-segment 508b from the audio segment 402b. Then, once the audio playback application 116 detects that the content generated by the playback of the audio sub-segment 508a is also included in the audio sub-segment 508b, it initiates the playback of the audio sub-segment 508b, but performs crossfade-in / fade-out playback with the audio sub-segment 508a, as indicated by the crossfade-in / fade-out area 509a, to further reduce or eliminate the perceptibility of playback gaps. The audio playback application 116 determines that the content included in the playback of the audio sub-segment 508a is also included in the audio sub-segment 508b based on the corresponding start and end timestamps associated with the audio sub-segment 508a and the audio sub-segment 508b. Therefore, once the playback of audio segment 508a occurs at a timestamp within audio source 120 that is also within audio segment 508b, the audio playback application 116 begins crossfading in and out of audio segments 508a and 508b. The audio playback application 116 crossfades in and out of audio segments 508a and 508b by gradually decreasing the volume of audio segment 508a while gradually increasing the playback volume of audio segment 508b.
[0055] Once the audio playback application 116 processes and outputs the audio segment 402c from the voice cancellation algorithm 122, the audio playback application 116 extracts the audio sub-segment 508c from the audio segment 402c. Then, once the audio playback application 116 detects that the playback of the audio sub-segment 508b is nearing completion or that the content being played is also contained within the audio sub-segment 508c, the audio playback application 116 initiates the playback of the audio sub-segment 508c, but performs crossfade-in and crossfade-out playback with the audio sub-segment 508b, as indicated by the crossfade-in and crossfade-out area 509b, thereby reducing or eliminating the perceptible playback gaps.
[0056] Therefore, a modified audio source 510 is generated by playing audio segments 508a, 508b, and 508c, as well as subsequent segments processed by the voice removal algorithm 122 utilized by the audio playback application 116. The modified audio source 510 results in an audio source 120 from which the audio playback application 116 has removed the human voice component. In some implementations, the introduced playback delay is equivalent to the processing time of the first audio segment processed by the audio playback application 116.
[0057] Figure 6 is a flowchart of method steps for processing audio source 120 according to one or more aspects of this disclosure. Although the method steps are described with respect to the system of Figures 1 to 5, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of various embodiments.
[0058] As shown, method 600 begins at step 602, where audio playback application 116 receives audio source 120 for playback. Audio source 120 is selected by the user or automatically or randomly by audio playback application 116. In some implementations, the user selects a karaoke mode provided by audio playback application 116 of IVI system 200 and selects a song via a user interface provided by IVI system 200.
[0059] At step 604, the audio playback application 116 buffers the playback of the audio source 120 by extracting the initial audio segment 302 or the first segment from the audio source 120 and providing the first segment to the voice removal algorithm 122. The voice removal algorithm 122 removes the main voice and / or secondary voice from the initial segment to generate an initial audio segment 402 without voice components. The voice removal algorithm 122 is selected based on user selection, user preference, or the number of occupants in the vehicle. For example, if the audio playback application 116 detects only one occupant in the vehicle, the audio playback application 116 selects the voice removal algorithm 122 that removes only the main voice but allows the secondary voice to be retained. If the audio playback application 116 detects more than one occupant in the vehicle, the audio playback application 116 selects the voice removal algorithm 122 that removes both the main voice and the secondary voice. Playback does not begin until the voice removal algorithm 122 has processed the initial segment. At step 606, the audio playback application causes the initial audio segment 402 to play.
[0060] At step 608, the audio playback application 116 extracts a subsequent audio segment 302 from the audio source 120. The subsequent audio segment 302 overlaps temporally with the initial audio segment 302. In other words, the subsequent audio segment 302 includes a portion of the end of the initial audio segment 302 of the audio source 120, as well as a portion of the audio source 120 that follows but is not part of the initial audio segment 302. It should be noted that in all implementations, the subsequent audio segment 302 does not need to overlap temporally with the initial audio segment 302.
[0061] At step 610, the audio playback application 116 processes the subsequent audio segment 302 using the voice removal algorithm 122 to remove the human voice component and generates a subsequent audio segment 402 with the human voice component removed. As described above, various techniques can be used to select the size of the subsequent audio segment 302. For example, the size of the subsequent audio segment 302 is the same as the initial audio segment 302. As another example, the audio playback application 116 calculates the time required to process the audio segment 302 using the voice removal algorithm 122 and selects the size of the subsequent audio segment 302 such that the generation of the subsequent audio segment 402 is completed simultaneously with or before the generation of the initial audio segment 402 without the human voice component.
[0062] In optional step 612, the audio playback application 116 crossfades in and out of the subsequent audio segment and the initial audio segment. The audio playback application 116 crossfades in and out the portions of the subsequent audio segment 402 that overlap with the initial audio segment 402 in time, to reduce or eliminate the user's perception of gaps between segments. In some embodiments, the audio playback application 116 does not crossfade in and out of the subsequent audio segment 402 with the initial audio segment 402, but instead plays the subsequent audio segment 402 when the initial audio segment 402 has finished playing.
[0063] At step 614, the audio playback application 116 causes at least a portion of the subsequent audio segment 402 to be played. In some examples, the audio playback application 116 initiates playback of the subsequent audio segment 402 after the initial audio segment 402 has finished playing. Then, method 600 returns to step 608, where the audio playback application 116 extracts the next subsequent audio segment 302 after the subsequent audio segment 302 from the audio source 120. Thus, the subsequent audio segment 302 then becomes the initial audio segment 302, as described in steps 608, 610, and 612 of method 600, and the next subsequent audio segment 302 becomes the subsequent audio segment 302, as described in steps 608, 610, 612, and 614 of method 600. Method 600 continues until playback of the audio source 120 has finished or has been interrupted by the user or another event.
[0064] In summary, the audio playback system plays an audio source (such as a song from a local or remote source) along with audio input (such as vocal input from a user). The audio playback system segments the audio source. It provides segments of the audio source to a vocal removal algorithm to remove vocal components from the segments. The initial segment of the audio source is processed by the vocal removal algorithm and played using an audio output device associated with the IVI system. The initial segment provides a buffer of audio from which vocal components have been removed. Subsequent segments are provided to the vocal removal algorithm. Subsequent segments are processed and played by the vocal removal algorithm until all segments have been processed and played. In some implementations, segments corresponding to the audio source overlap in time. The overlapping segments are processed by the vocal removal algorithm to produce segments with vocal components removed. In some cases, sub-segments of the overlapping segments are played by the audio playback system. In other scenarios, the playback of overlapping sub-segments uses crossfading to further reduce or eliminate playback discontinuities between sub-segments. For example, an audio playback system fades in and out overlapping segments, making the listener perceive that the audio source is being played as a continuous audio stream.
[0065] At least one technical advantage of the disclosed technology over the prior art is that, using the disclosed technology, the vocal removal algorithm essentially removes the vocal components of the songs for which the user desires a karaoke experience in real time. By removing the vocal components from the songs in real time, a karaoke experience is provided for any number of audio sources available for streaming playback. Furthermore, capturing vocal input within the vehicle using microphones allows the vocal input to be played along with the song. Therefore, playing audio sources without vocal components along with sound input captured by one or more microphones provides an improved karaoke experience. These technical advantages provide one or more technical advancements superior to existing methods.
[0066] 1. In some embodiments, a computer-implemented method includes: receiving an audio source for playback; extracting a first segment of the audio source, the first segment including a first portion of the audio source; removing a first vocal component from the first segment to create a first modified segment; and causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
[0067] 2. The computer-implemented method as described in Clause 1, wherein removing the first human voice component from the first segment comprises performing the human voice removal algorithm on the first segment to produce a first modified segment.
[0068] 3. The computer-implemented method as described in Clause 1 or 2, further comprising detecting the number of users, wherein the voice cancellation algorithm is selected based on the number of users.
[0069] 4. The computer-implemented method as described in any one of Clauses 1 to 3, wherein the number of users includes the number of occupants of the vehicle.
[0070] 5. A computer-implemented method as described in any one of Clauses 1 to 4, further comprising: extracting a second segment of the audio source, the second segment comprising a second portion of the audio source following the first portion; removing a second vocal component from the second segment to create a second modified segment; and causing at least one sub-segment of the second modified segment to be played after the first modified segment using one or more audio output devices.
[0071] 6. A computer-implemented method as described in any one of Clauses 1 to 5, wherein the first segment and the second segment overlap in time.
[0072] 7. A computer-implemented method as described in any one of Clauses 1 to 6, wherein causing the playback of the sub-segment of the second modified segment after the sub-segment of the first modified segment comprises crossfading the playback of the sub-segment of the first segment with the playback of the sub-segment of the second segment after the first segment.
[0073] 8. A computer-implemented method as described in any one of Clauses 1 to 7, wherein the sub-segment causing the playback of the second modified segment after the first modified segment includes at least one sub-segment causing the playback of the second modified segment after the playback of at least one of the sub-segments of the first modified segment is completed.
[0074] 9. A computer-implemented method as described in any one of Clauses 1 to 8, further comprising selecting the size of the first segment based on the processing time required to remove the first human voice component from the first segment.
[0075] 10. A computer-implemented method as described in any one of clauses 1 to 9, wherein the size of the second segment is different from the size of the first segment.
[0076] 11. The computer-implemented method of any one of clauses 1 to 10, wherein the processing time required to remove the second vocal component from the second segment following the first segment is less than the playback time of the first segment.
[0077] 12. In some embodiments, one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: receiving an audio source for playback; extracting a first segment of the audio source, the first segment comprising a first portion of the audio source; removing a first vocal component from the first segment to create a first modified segment; and causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
[0078] 13. One or more non-transitory computer-readable media as described in Clause 12, wherein removing the first human voice component from the first segment comprises performing the human voice removal algorithm on the first segment to produce a first modified segment.
[0079] 14. One or more non-transitory computer-readable media as described in Clause 12 or 13, wherein the step further comprises: extracting a second segment of the audio source, the second segment comprising a second portion of the audio source following the first portion; removing a second vocal component from the second segment to create a second modified segment; and causing the second modified segment to be played after the first modified segment using one or more audio output devices.
[0080] 15. One or more non-transitory computer-readable media as described in any one of clauses 12 to 14, wherein the first segment and the second segment overlap in time.
[0081] 16. One or more non-transitory computer-readable media as described in any one of clauses 12 to 15, wherein causing the playback of the sub-segment of the second modified segment after the sub-segment of the first modified segment includes crossfading the playback of the sub-segment of the first segment with the playback of the sub-segment of the second segment after the first segment.
[0082] 17. One or more non-transitory computer-readable media as described in any one of clauses 12 to 16, wherein the sub-segment causing the playback of the second modified segment after the first modified segment includes at least one sub-segment causing the playback of the second modified segment after the playback of at least one of the sub-segments of the first modified segment is completed.
[0083] 18. One or more non-transitory computer-readable media as described in any one of clauses 12 to 17, further comprising selecting the size of the first segment based on the processing time required to remove the first human voice component from the first segment.
[0084] 19. One or more non-transitory computer-readable media as described in any one of clauses 12 to 18, wherein the playback of the first segment is delayed based on the processing time.
[0085] 20. In some embodiments, a system includes: one or more audio output devices; a memory storing an audio playback application; and a processor coupled to the memory, the processor executing the audio playback application by performing the following steps: receiving an audio source for playback; extracting a first segment of the audio source, the first segment including a first portion of the audio source; removing a first vocal component from the first segment to create a first modified segment; and causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
[0086] Any and all combinations of any claim element referenced in any claim and / or any element described in this application in any manner fall within the scope of the invention and protection.
[0087] Various embodiments have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
[0088] Aspects of this disclosure may be embodied as a system, method, or computer program product. Therefore, aspects of this disclosure may take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, all of which are generally referred to herein as a “module,” “system,” or “computer.” Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or group of circuits. Additionally, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0089] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media may include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that may contain or store programs for use with or in connection with an instruction execution system, device, or apparatus.
[0090] Various aspects of this disclosure are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed via a processor of a computer or other programmable data processing apparatus, these instructions enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this respect, each box in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative implementations, the functions indicated in the boxes may occur in a different order than indicated in the drawings. For example, two boxes shown consecutively may actually be executed substantially simultaneously, or depending on the functionality involved, the boxes may sometimes be executed in reverse order. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified functions or actions.
[0092] Although the foregoing describes an embodiment of this disclosure, other and more embodiments of this disclosure may be devised without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.
Claims
1. A computer-implemented method, comprising: Receive audio sources for playback; Extract a first segment from the audio source, the first segment comprising a first part of the audio source; Remove the first human voice component from the first segment to create the first modified segment; And causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
2. The computer-implemented method of claim 1, wherein removing the first human voice component from the first segment comprises performing the human voice removal algorithm on the first segment to produce a first modified segment.
3. The computer-implemented method of claim 2, further comprising detecting the number of users, wherein the voice cancellation algorithm is selected based on the number of users.
4. The computer-implemented method of claim 3, wherein the number of users includes the number of occupants of the vehicle.
5. The computer-implemented method as described in claim 1, further comprising: Extract a second segment from the audio source, the second segment comprising a second portion of the audio source following the first portion; Remove the second vocal component from the second segment to create a second modified segment; And causing at least one sub-segment of the second modified segment to be played after the first modified segment using one or more audio output devices.
6. The computer-implemented method of claim 5, wherein the first segment and the second segment overlap in time.
7. The computer-implemented method of claim 5, wherein causing the playback of the sub-segment of the second modified segment after the sub-segment of the first modified segment includes cross-fading the playback of the sub-segment of the first segment with the playback of the sub-segment of the second segment after the first segment.
8. The computer-implemented method of claim 5, wherein the sub-segment causing the playback of the second modified segment after the first modified segment includes at least one sub-segment causing the playback of the second modified segment after the playback of at least one of the sub-segments of the first modified segment is completed.
9. The computer-implemented method of claim 5, further comprising selecting the size of the first segment based on the processing time required to remove the first human voice component from the first segment.
10. The computer-implemented method of claim 6, wherein the size of the second segment is different from the size of the first segment.
11. The computer-implemented method of claim 10, wherein the processing time required to remove the second voice component from the second segment following the first segment is less than the playback time of the first segment.
12. A non-transitory computer-readable medium storing one or more instructions, which, when executed by one or more processors, cause the one or more processors to perform the following steps: receiving an audio source for playback; extracting a first segment of the audio source, the first segment comprising a first portion of the audio source; removing a first human voice component from the first segment to create a first modified segment; and causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.
13. One or more non-transitory computer-readable media as claimed in claim 12, wherein removing the first human voice component from the first segment comprises performing the human voice removal algorithm on the first segment to produce a first modified segment.
14. The one or more non-transitory computer-readable media of claim 12, wherein the step further comprises: Extract a second segment from the audio source, the second segment comprising a second portion of the audio source following the first portion; Remove the second vocal component from the second segment to create a second modified segment; And cause the second modified segment to be played after the first modified segment using one or more audio output devices.
15. One or more non-transitory computer-readable media as claimed in claim 14, wherein the first segment and the second segment overlap in time.
16. One or more non-transitory computer-readable media as claimed in claim 14, wherein causing the playback of the sub-segment of the second modified segment after the sub-segment of the first modified segment includes crossfading the playback of the sub-segment of the first segment with the playback of the sub-segment of the second segment after the first segment.
17. One or more non-transitory computer-readable media as claimed in claim 14, wherein the sub-segment causing the playback of the second modified segment after the first modified segment includes at least one sub-segment causing the playback of the second modified segment after the playback of at least one of the sub-segments of the first modified segment is completed.
18. One or more non-transitory computer-readable media as claimed in claim 14, further comprising selecting the size of the first segment based on the processing time required to remove the first human voice component from the first segment.
19. One or more non-transitory computer-readable media as claimed in claim 18, wherein the playback of the first segment is delayed based on the processing time.
20. A system comprising: One or more audio output devices; The memory stores the audio playback application; and a processor coupled to the memory, the processor performing the audio playback application by executing the following steps: receiving an audio source for playback; extracting a first segment of the audio source, the first segment including a first portion of the audio source; removing a first vocal component from the first segment to create a first modified segment; And causing at least one sub-segment of the first modified segment to be played using one or more audio output devices.