Real-time human voice removal from audio source

By identifying and modifying the left, right, and center channels of an audio source, and combining mid-frequency band processing and microphone capture, the problem of real-time voice removal in existing technologies has been solved, improving the quality of the karaoke experience and the efficiency of computing resource utilization.

CN121963675APending Publication Date: 2026-05-01HARMAN BECKER AUTOMOTIVE SYST GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARMAN BECKER AUTOMOTIVE SYST GMBH
Filing Date
2025-10-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies cannot effectively remove human voice components from audio sources in real time, resulting in a decline in the quality of the karaoke experience. In particular, human voice components cannot be preserved in multi-channel audio sources. Furthermore, existing algorithms consume high computational resources and cannot be applied to streaming audio sources in real time.

Method used

By identifying the left and right channels of the audio source, modified left and right channels are generated respectively, and a center channel is generated. The human voice component is removed using mid-band ducking, attenuation or compression technology, and combined with the microphone to capture human voice input in the car, an improved karaoke experience is provided in real time.

Benefits of technology

It enables real-time attenuation or removal of the main vocal component of the audio source while preserving instrument content. It is suitable for stereo and multi-channel audio sources, improves the quality of the karaoke experience, and supports real-time application of streaming audio sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963675A_ABST
    Figure CN121963675A_ABST
Patent Text Reader

Abstract

Various embodiments disclose a computer-implemented method comprising: receiving an audio source for playback by an audio playback system; identifying a left channel and a right channel associated with the audio source; generating a modified left channel, including subtracting the right channel from the left channel; generating a modified right channel, including subtracting the left channel from the right channel; the modified left channel is caused to play on a left channel speaker of the audio playback system, and the modified right channel is caused to play on a right channel speaker of the audio playback system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments typically involve audio processing, and more specifically, real-time removal of human voices from an audio source. Background Technology

[0002] Modern vehicles include in-vehicle infotainment (IVI) systems that receive audio and video input from various sources. IVI systems include various output devices, such as displays and speakers throughout the vehicle. The IVI system receives user-selected input (such as audio input) from local or remote audio sources and uses the vehicle's output devices to play that audio input.

[0003] A karaoke experience can be provided by an IVI system and involves one or more users singing along to a pre-recorded audio performance, which is played back by the IVI system's audio output device. Users sing along to the pre-recorded audio performance, and in some instances, a microphone is used to capture the user's voice, which is then reproduced using the same audio output device that played the pre-recorded audio performance. In some cases, users prefer to use an audio source from which lead vocals and / or background vocals have been removed. Pre-recorded audio performances are created specifically for use with the karaoke experience by preprocessing the audio source to remove vocal components. Preprocessing is typically performed by personnel such as audio engineers or producers or by automatic vocal removal algorithms, and the pre-processed audio source is provided as an audio source to the audio playback system. In other examples, pre-recorded audio performances for use with the karaoke experience are created by recording an instrumental version of the audio source without lead and / or secondary vocals. In either scenario, creating a version of the audio source for use in the karaoke experience requires preprocessing or pre-recording the audio source used for that karaoke experience. Another technology used to provide a karaoke experience involves playing an audio source and allowing users to sing along to an unmodified version of that audio source. However, using an audio source that includes vocals results in a poor karaoke experience for many users.

[0004] Some karaoke experiences offer mechanisms for real-time suppression of the vocal component of the audio source played during the karaoke experience. One technique for real-time vocal suppression is mid-band ducking, which reduces the volume of the mid-frequency components of the audio signal, typically contained within the vocal component. However, by ducking, other components of the audio besides the vocal component (such as instrumental components) are removed, degrading the quality of the karaoke experience. Additionally, in the case of 5.1, 7.1, or other multi-channel audio sources, the vocal component is typically contained in the center channel of the multi-channel audio source. Therefore, the center channel component can be removed or ducked, reducing the volume of the channel that typically contains the vocal component. However, this is generally not available for 5.1, 7.1, or other multi-channel audio sources.

[0005] One drawback of conventional techniques for removing vocal components from audio sources to provide a karaoke experience is the inability to utilize many vocal removal algorithms in real time. Vocal removal algorithms typically require significant processing time, making them unsuitable for real-time use on audio sources being streamed for playback. Furthermore, using pre-recorded karaoke versions of audio sources does not provide users with a karaoke experience for all audio sources played by the audio playback system. A disadvantage of performing mid-band ducking on the left and right channels of an audio source is that components other than vocals are removed by these techniques, degrading the quality of the karaoke experience. A disadvantage of performing center-channel ducking on audio sources containing a discrete center channel is that the discrete center channel is generally unusable for music.

[0006] As previously stated, there is a need in the art for more efficient technologies for processing audio sources that can provide users with an acceptable karaoke experience. Summary of the Invention

[0007] In various embodiments, a computer-implemented method includes: receiving an audio source for playback by an audio output device; identifying a left channel and a right channel associated with the audio source; playing a modified left channel, comprising subtracting the right channel from the left channel, on the left channel of the audio output device; and playing a modified right channel, comprising subtracting the left channel from the right channel, on the right channel of the audio output device.

[0008] At least one technical advantage of the disclosed technology over the prior art is that it attenuates the lead vocal component of the audio source that the user expects for their karaoke experience in real time and with less computational resources than vocal removal algorithms. This provides a karaoke experience for virtually any audio source being streamed for playback by attenuating or removing the lead vocal component in real time. Additionally, it preserves the instrumental content of the audio source by avoiding the use of mid-band ducking in the left and right channels separately. The disclosed technology can also remove the vocal component of stereo content when 5.1, 7.1, or other multichannel audio formats with a discrete center channel are unavailable. Furthermore, using a microphone to capture in-vehicle vocal input allows that vocal input to be played along with the audio source. Therefore, playing an audio source without lead vocals along with vocal input captured by one or more microphones provides an improved karaoke experience. These technical advantages provide one or more technological advancements superior to existing methods. Attached Figure Description

[0009] To gain a detailed understanding of the features described above in the various embodiments, a more specific description of the inventive concept briefly outlined above can be obtained by referring to the various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings illustrate only typical embodiments of the inventive concept and should therefore not be construed as limiting the scope in any way, and other equivalent embodiments exist.

[0010] Figure 1 A block diagram of a computing device configured to implement one or more aspects of this disclosure is shown.

[0011] Figure 2 A block diagram of an IVI system configured to implement one or more aspects of this disclosure is shown.

[0012] Figure 3 An example of an audio source processed in accordance with one or more aspects of this disclosure is shown.

[0013] Figure 4 This is a flowchart of method steps for processing an audio source according to one or more aspects of this disclosure. Detailed Implementation

[0014] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to those skilled in the art that the inventive concepts may be practiced without one or more of these specific details. For purposes of explanation, multiple instances of similar objects are symbolized using reference numerals to identify the object and bracketed numerals to identify the instance when necessary.

[0015] Figure 1A block diagram of an audio playback system configured to implement one or more aspects of the present disclosure is shown. As shown, the audio playback system 100 includes, but is not limited to, a computing device 110, an audio source 120, an input module 130, and an output module 140. The computing device 110 includes, but is not limited to, a processing unit 112 and a memory 114. The memory 114 includes, but is not limited to, an audio playback application 116.

[0016] In operation, computing device 110 executes audio playback application 116 to control audio playback. In one example, audio is played from one or more vehicle components or sources, either inside or outside the vehicle. Specifically, processing unit 112 executes audio playback application 116 and causes audio to be played on one or more output devices associated with audio playback system 100. Audio playback application 116 receives audio sources 120, such as terrestrial or satellite radio signals, music or other content obtained from streaming audio services, audio files stored on storage devices associated with the vehicle, or audio content streamed from another device, such as a Bluetooth device to which computing device 110 is connected.

[0017] The audio playback application 116 also combines the audio source 120 played by the audio playback system 100 to provide a karaoke experience for the user. For example, the audio playback application 116 receives audio input from the input module 130, such as human voice input detected by a microphone associated with the audio playback system 100. The audio playback application 116 plays the audio input along with the audio source 120 on an audio output device such as one or more speakers. In some cases, in addition to playing the audio source 120 and the audio input, the audio playback application 116 also plays video content on a display in the vehicle or switches interior or exterior lighting to enhance the karaoke experience.

[0018] The computing device 110 includes a processing unit 112 and a memory 114. In various embodiments, the computing device 110 is a device (such as a system-on-a-chip (SoC)) including one or more processing units 112. In various embodiments, the computing device 110 is a mobile computing device wirelessly connected to other devices in the vehicle, such as a tablet computer, mobile phone, media player, etc. In some embodiments, the computing device 110 is a host unit included in a vehicle system. Alternatively or additionally, the computing device 110 may be a detachable device installed in part of the vehicle as part of a separate console. Typically, the computing device 110 is configured to coordinate the overall operation of the audio playback system 100. The embodiments disclosed herein contemplate any technically feasible system configured to implement the functionality of the audio playback system 100 via the computing device 110. The functionality and technology of the audio playback system 100 are also applicable to other types of vehicles, including consumer vehicles, commercial trucks, airplanes, helicopters, spacecraft, small boats, submarines, etc.

[0019] Processing unit 112 may include one or more central processing units (CPUs), digital signal processing units (DSPs), microprocessors, application-specific integrated circuits (ASICs), neural processing units (NPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), etc. Processing unit 112 typically includes a programmable processor that executes program instructions to manipulate input data and generate outputs. In some embodiments, processing unit 112 may include any number of processing cores and other modules for facilitating program execution.

[0020] Memory 114 includes memory modules or a collection of memory modules. Memory 114 typically includes memory chips such as random access memory (RAM) chips that store application programs and data for processing by processing unit 112. In various embodiments, memory 114 includes non-volatile memory such as optical drives, magnetic drives, flash drives, or other storage devices. Audio playback application 116 within memory 114 is executed by processing unit 112 to enable the overall functionality of computing device 110 and thus coordinate the operation of audio playback system 100 as a whole.

[0021] Audio playback application 116 processes audio source 120 and / or audio input received from input module 130 to reproduce audio signals. In various embodiments, audio playback application 116 plays audio source 120 along with audio input from one or more occupants or users of the vehicle via output module 140. Audio input is obtained via input module 130 to provide a karaoke experience. Additionally, audio playback application 116 processes audio source 120 to remove vocal components from audio source 120, providing an improved karaoke experience. Audio source 120 includes a stereo input signal comprising a left channel and a right channel. Audio playback application 116 removes vocal components from audio source 120 in real time by performing processing operations on the left and right channels to generate modified left and right channels, respectively. The modified left channel is generated based on the left and right channels of the stereo input. The modified right channel is also generated based on the left and right channels of the stereo input. Additionally, a center channel is generated, which combines the left and right channels of the stereo input. The modified left channel is played on the left channel of the output module 140, for example, via one or more left channel speakers. The modified right channel is played on the right channel of the output module 140, for example, via one or more right channel speakers.

[0022] A modified left channel is generated by identifying the left and right channels of the stereo input corresponding to audio source 120. The right channel is then subtracted from the left channel to create the modified left channel. Subtracting the right channel from the left channel has the effect of removing any content present in either channel (typically including vocal components), but allows other content (typically including instrumental content) to remain in the modified left channel. A modified right channel is generated by identifying the left and right channels of the stereo input corresponding to audio source 120. The left channel is then subtracted from the right channel to create the modified right channel. Subtracting the left channel from the right channel has the effect of removing any content present in either channel (typically including vocal components), but allows other content (typically including instrumental content) to remain in the modified right channel.

[0023] Audio playback application 116 generates a center channel based on the left and right channels. The left and right channels are added together to create the center channel. In one example, the center channel is played by both the left and right channel speakers of output module 140. In another example, the center channel is played by the center channel speaker of output module 140. When the user enables the karaoke mode provided by audio playback application 116, or when audio playback application 116 detects audio input via input module 130 during a karaoke experience, audio playback application 116 generates a modified center channel based on the center channel created from the left and right channels of audio source 120. The modified center channel is output to output module 140 for playback. One or more processing techniques, such as mid-band ducking, mid-band attenuation, compression, or other real-time audio processing techniques that remove or attenuate vocal components from the center channel, are used to generate the modified center channel in which vocal components are removed or attenuated from the center channel. The audio playback application 116 causes the output module 140 to play the modified center channel, which may involve using the left and right channels of the output module 140 to play the modified center channel.

[0024] In some implementations, the audio playback application 116 plays the modified center channel only when voice input is detected from the input module 130. In this scenario, when voice input is not received by the input module 130, the unmodified center channel is played. In some examples, the audio playback application 116 plays the modified center channel when a karaoke mode is selected by the user in the audio playback application 116 via a user interface provided by the audio playback application 116. In other examples, the audio playback application 116 plays the modified center channel when more than one occupant of the vehicle is detected and whenever karaoke mode is enabled in the audio playback application 116. In another scenario, the user can select when to provide voice input, such as via a button on the microphone 222 or another user input device. In this case, the audio playback application 116 plays the modified center channel when the user indicates that voice input is being provided. In some implementations, when no more voice input is detected after a threshold time period or when the detection of the termination of voice input is detected, the audio playback application 116 resumes outputting the unmodified center channel for playback by the output module 140.

[0025] Audio source 120 includes one or more data sources that provide audio signals for reproduction. Audio source 120 includes pre-recorded audio performances, such as songs. In various embodiments, audio source 120 is included in a device within the vehicle (such as an entertainment subsystem included in the vehicle's head unit, a rear-seat entertainment console, a device installed in the vehicle, etc.). In some embodiments, audio source 120 is included in a mobile device, wearable device, and / or other portable device connected to audio playback application 116. Additionally, audio source 120 may be located remotely from the vehicle. In such instances, a remote data source streams audio source 120 to computing device 110, and then audio playback application 116 transmits audio source 120 to an output device associated with output module 140 for reproduction.

[0026] Input module 130 includes one or more means for performing measurements and / or acquiring data relating to certain objects in the environment. In various embodiments, input module 130 generates sensor data relating to a user and / or objects in the environment that are not the user. In some embodiments, input module 130 is coupled to and / or included within computing device 110 and sends the sensor data to processing unit 112.

[0027] In various embodiments, the input module 130 includes audio sensors, such as built-in microphones and / or microphone arrays for recording sound within the vehicle's cabin. Vehicle occupant sensors include optical sensors, such as RGB cameras, infrared cameras, depth cameras, and / or camera arrays, comprising two or more cameras oriented towards the seating area of ​​the vehicle. Cabin sensors include, for example, pressure sensors integrated into the seating position within the vehicle, which detect when an occupant is seated in a specific seating position within the vehicle. In some embodiments, the input module 130 includes touch sensors, position sensors (…), and other sensors that register the presence, body position, and / or movement of a user within the vehicle. For example accelerometers and / or inertial measurement units (IMUs) or other types of sensors.

[0028] In some embodiments, the input module 130 includes physiological sensors, such as a heart rate monitor, an electroencephalogram (EEG) system, a radio sensor, a thermal sensor, and a skin conductance sensor. For example The input module 130 also includes devices capable of receiving input, such as a keyboard, mouse, touchscreen, and other input devices for providing input to the computing device 110. In various embodiments, the input module 130 is associated with a specific console, such as a personalized screen mounted to part of the seat or a console-specific input component.

[0029] Output module 140 includes one or more devices capable of providing output, such as a display screen or a speaker. In various embodiments, one or more of input module 130 or output module 140 are incorporated into computing device 110 or external to computing device 110. In some embodiments, computing device 110, input module 130, or output module 140 may be components of an IVI system or entertainment subsystem included in a vehicle.

[0030] Vehicle system

[0031] Figure 2 The illustrations show various embodiments including Figure 1An example IVI system 200 is provided for the audio playback system 100. As shown, the IVI system 200 includes, but is not limited to, an input module 130, a computing device 110, and an output module 140. The input module 130 includes, but is not limited to, one or more microphones 222, occupant-facing sensors 226, and cabin sensors 228. The computing device 110 includes, but is not limited to, an audio playback application 116. The output module 140 includes, but is not limited to, a speaker 230, a display 232, and a human-machine interface (HMI) 234. The audio playback application 116 includes, but is not limited to, an input processing module 234 and an output generation module 238.

[0032] In some embodiments, the computing device 110 may be integrated into the vehicle's host unit. The host unit is a component of the vehicle, installed in any location within the vehicle's passenger compartment in any technically feasible manner. In some embodiments, the host unit includes any number and type of instruments and applications, and provides any number of input and output mechanisms. For example, the host unit enables the user ( For example The driver and / or passengers can control the IVI system. The main unit supports any number of input and output data types and formats, as known in the art. For example, the main unit may include built-in Bluetooth, USB connectivity, voice recognition, camera input via input module 130, video output via output module 140 for any number and type of displays 232, and any number of audio outputs. Typically, any number of sensors, displays, receivers, transmitters, etc., can be integrated into the main unit or implemented externally to the main unit. Additionally, the computing device 110 may be located elsewhere in the vehicle, such as behind an interior panel hidden in a manager not visible to passengers.

[0033] In operation, the audio playback application 116 receives the audio source 120 and causes the speaker 230 associated with the output module 140 to play a modified version of the audio source 120 that has been processed by the audio playback application 116. The audio source 120 includes songs, radio stations, or other audio sources that can be played or streamed by the computing device 110. In one scenario, a user of the IVI system 200 activates the karaoke mode of the audio playback application 116 via the HMI 236 and selects the audio source 120. The modified version of the audio source 120 is a version of the audio source 120 from which the main or all vocal components have been removed by the audio playback application 116. To remove the vocal components from the audio source 120, the audio playback application 116 identifies the left and right channels in the stereo audio signal corresponding to the audio source 120. The audio playback application 116 then generates the modified left channel, modified right channel, and center channel based on the left and right channels. If the center channel signal is played using the output module 140, which does not include a center channel speaker, the center channel is also referred to as the phantom center channel. The audio playback application 116 outputs modified left and right channels to the output module 140 for playback. The audio playback application 116 also outputs the center channel or a modified center channel to the output module 140 for playback, depending on whether human voice input is detected via the input module 130.

[0034] Audio playback application 116 generates a modified left channel by identifying the left channel signal of audio source 120 and subtracting the right channel signal of audio source 120 from the left channel. Audio playback application 116 generates a modified right channel signal of audio source 120 and subtracts the left channel signal of audio source 120 from the right channel. Because vocal components are typically present in both the left and right channel signals of audio source 120, subtracting opposite signals from the left and right channels has the effect of removing vocal components. Therefore, the modified left channel and modified right channel represent signals from which vocal components have been removed or attenuated. Audio playback application 116 generates a center channel by adding the left and right channels of audio source 120. In many stereo signals corresponding to audio source 120, primary and secondary vocals exist in both the left and right channels. Therefore, adding the left and right channels produces a center channel in which vocals are present. The audio playback application 116 generates a modified center channel by performing one or more processing operations on the center to remove the vocal component.

[0035] For example, audio playback application 116 performs mid-band ducking to reduce the level of the mid-band in the center channel to produce a modified center channel. The mid-band can represent a frequency range such as 250 Hz to 4 kHz. In some examples, the mid-band represents a narrower frequency range, such as 500 Hz to 2 kHz. As another example, audio playback application 116 performs mid-band attenuation to reduce the level of the mid-band to produce a modified center channel. As another example, audio playback application 116 performs mute on certain frequencies in the mid-band to reduce or remove vocal components in the center channel to produce a modified center channel. As another example, audio playback application 116 performs compression to reduce the dynamic range of the mid-band to produce a modified center channel. As yet another example, audio playback application 116 completely mutes the center channel, so that only the modified left and modified right channels are output for playback by output module 140.

[0036] When the karaoke mode of the audio playback application 116 is activated, the audio playback application 116 outputs modified left and right channels to the output module 140 for playback. The output module 140 plays the modified left channel on one or more left channel speakers. The output module 140 plays the modified right channel on one or more right channel speakers. When the karaoke mode is activated, the audio playback application 116 outputs the center channel to the output module 140 for playback. The output module 140 plays the center channel on the left and right channel speakers. In some implementations, when the karaoke mode of the audio playback application 116 is activated, the audio playback application 116 outputs a modified center channel to the output module 140 for playback, and the output module 140 plays the modified center channel on the left and right channel speakers.

[0037] In one scenario, when no human voice input is detected by one or more microphones 222 of input module 130, audio playback application 116 outputs an unmodified center channel to output module 140 for playback. When audio playback application 116 detects human voice input provided by input module 130 via one or more microphones 222, audio playback application 116 outputs a modified center channel to output module 140 for playback. Then, audio playback application 116 outputs the modified center channel to output module 140 until no human voice input is detected by one or more microphones 222 within a threshold time period.

[0038] Input obtained by input module 130 includes voice input from one or more microphones 222 within the vehicle, such as voice input from occupants participating in the karaoke experience. Audio playback application 116 causes the speaker 230 of output module 140 to play the voice input in addition to playing the audio source 120. In some cases, audio playback application 116 modifies the voice input by applying compression, reverb, auto-tuning, or other effects to the audio input. Audio playback application 116 plays the voice input along with the audio source 120 on an audio output device such as one or more speakers. In some cases, in addition to playing the audio source 120 and the voice input, audio playback application 116 also plays video content on a display within the vehicle or switches interior or exterior lighting to enhance the karaoke experience.

[0039] The audio playback application 116 also detects the number and / or location of occupants within the vehicle based on input received from the input module 130. For example, the audio playback application 116 detects seat positions within the vehicle based on sensor data from one or more microphones 222, occupant-facing sensors 226, or cabin sensors 228. For instance, when more than one occupant is detected in the vehicle by the input module 130, the audio playback application 116 determines that more than one occupant is present and outputs a modified center channel to the output module 140 for playback. As another example, the audio playback application 116 determines that only one occupant is present in the vehicle and outputs an unmodified center channel until audio input is detected via one or more microphones 222. Additionally, the audio playback application 116 can apply lighting effects using interior or exterior vehicle lighting customized based on the number of occupants detected or the detected seat positions of the occupants. These lighting effects or other customizations can be defined by user profiles stored in the data storage.

[0040] Input module 130 includes various types of sensors, including one or more microphones 222, occupant-facing sensors 226, and cabin sensors 228. In some cases, input module 130 also includes, but is not limited to, vehicle sensors such as outward-facing cameras, external microphones, accelerometers, etc. Occupant-facing sensors 226 include cameras or motion sensors oriented to detect the presence of occupants within the vehicle. In some cases, occupant-facing sensors 226 may also detect users based on facial recognition, enabling audio playback application 116 to identify user profiles specifying karaoke experience preferences, such as the selection of specific voice removal algorithms. Cabin sensors 228 include other types of sensors, such as pressure sensors, temperature sensors, or other types of sensors that also detect the presence of occupants within the vehicle. In various embodiments, input module 130 provides audio playback application 116 with a combination of sensor data, which can use input obtained from one or more microphones 222 and sensor data from occupant-facing sensors 226 and cabin sensors 228 to determine the number of occupants within the vehicle or the seating positions of the occupants. Additionally, when the karaoke mode is selected by a user in the vehicle, the input module 130 provides audio input from one or more microphones 222, which can be played using the vehicle's speakers 230.

[0041] Output module 140 includes various types of output devices, including but not limited to speaker 230, display 232, and HMI 234. Speaker 230 includes one or more left channel speakers and one or more right channel speakers. In some examples, speaker 230 also includes a center channel speaker. Output module 140 performs one or more actions in response to output signals from computing device 110 or other subsystems within the vehicle. For example, output module 140 receives audio output from computing device 110, which may include multiple audio outputs mixed together by computing device 110. Output module 140 uses speaker 230 within the vehicle to play the audio output. For example, audio playback application 116 mixes audio source 120 with audio input detected by one or more microphones 222 and transmits an audio output including both audio source 120 and audio input to output module 140, which uses speaker 230 to play the audio. As another example, output module 140 receives additional information from computing device 110 and causes display 232 or HMI 234 to display notifications, messages, alarms, or other information.

[0042] Figure 3 An example of an audio source 120 processed according to one or more aspects of this disclosure is shown. Figure 3This demonstrates how an audio source 120, which includes a stereo signal, is isolated into a left and right channel and processed by an audio playback application 116 to remove vocal components to enhance the karaoke experience.

[0043] like Figure 3 As shown, audio source 120 represents a stereo signal including left and right channels. Therefore, the left channel L and right channel R of audio source 120 are isolated by audio playback application 116. Audio playback application 116 adds L and R to generate center channel C. Audio playback application 116 also generates a modified left channel L' by subtracting R from L. Audio playback application 116 generates a modified right channel R' by subtracting L from R. L', R', and C are provided to output module 140 for playback by, for example, speakers 230 in a vehicle. Output module 140 can play L' in one or more left channel speakers. Output module 140 can play R' in one or more right channel speakers. Output module 140 can play C in the center channel speaker. In the absence of a center channel speaker in output module 140, C can be played in both the left and right channel speakers to create a phantom center channel speaker. Additionally, when human voice input is detected or when the karaoke mode is selected by the user, the audio playback application 116 generates a modified C and outputs it to the output module 140, which attenuates or removes the human voice component of the C.

[0044] Figure 4 This is a flowchart of method steps for processing audio source 120 according to one or more aspects of this disclosure. Although referenced... Figures 1 to 3 The system describes the method steps, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the various embodiments.

[0045] As shown in the figure, method 400 begins at step 402, where the audio playback application 116 receives the audio source 120 for playback. The audio source 120 is selected by the user or automatically or randomly by the audio playback application 116. In some implementations, the user selects a karaoke mode provided by the audio playback application 116 of the IVI system 200 and selects a song via a user interface provided by the IVI system 200.

[0046] At step 404, the audio playback application 116 isolates the left channel L and the right channel R from the audio source 120. At step 406, the audio playback application 116 generates a modified left channel L' from the left channel L of the audio source 120. The modified left channel L' is created by subtracting the right channel R from the left channel L.

[0047] At step 408, the audio playback application 116 causes the output module 140 to play the modified left channel L'. The output module 140 plays the modified left channel L' on one or more left channel speakers.

[0048] At step 410, the audio playback application 116 generates a modified right channel R' from the right channel R of the audio source 120. The modified right channel R' is created by subtracting the left channel L from the right channel R.

[0049] At step 412, the audio playback application 116 causes the modified right channel R' to be played. The modified right channel R' is created by subtracting the left channel L from the right channel R. The output module 140 plays the modified right channel R' on one or more right channel speakers.

[0050] At step 414, the audio playback application 116 causes the center channel C corresponding to the audio source 120 to be played. The audio playback application 116 generates the center channel C by adding the contents of the left channel L and the right channel R isolated from the audio source 120. The audio playback application 116 outputs the center channel C to the output module 140, which creates a phantom center channel by playing the center channel C via the center channel speaker or by playing the center channel C through the left channel speaker and the right channel speaker.

[0051] At step 416, the audio playback application 116 determines whether human voice input has been detected via the input module 130. Human voice input may be provided by one or more occupants of the vehicle via one or more microphones 222 of the input module 130. If no human voice input is detected, method 400 returns to or remains at step 414, where the audio playback application 116 plays the center channel via the output module 140. If human voice input is detected at step 416, method 400 proceeds to step 418. In some examples, when a user enables karaoke mode via the audio playback application 116, the audio playback application 116 proceeds to step 418 instead of waiting for human voice input to be detected, or the audio playback application 116 proceeds to step 418 in addition to waiting for human voice input to be detected.

[0052] At step 418, the audio playback application 116 generates a modified center channel from the audio source 120. The audio playback application 116 generates the modified center channel by applying one or more processing techniques to the center channel C to attenuate, mute, or otherwise remove vocal components from the center channel C. For example, the audio playback application 116 generates a modified center channel in which vocal components are removed or attenuated from the center channel using one or more processing techniques (such as mid-band ducking, mid-band attenuation, compression, or other real-time audio processing techniques that remove or attenuate vocal components from the center channel).

[0053] At step 420, the audio playback application 116 causes the output module 140 to play the modified center channel. If the output module 140 does not include a center channel speaker, the modified center channel is played by the left and right channel speakers of the output module 140. If the output module 140 includes a center channel speaker, the modified center channel is played by the center channel speaker.

[0054] Then, method 400 returns to step 416, where the audio playback application 116 determines whether voice input has been detected via input module 130 or whether the user has enabled karaoke mode via audio playback application 116. In some implementations, audio playback application 116 determines that no more voice input has been detected when no more voice input is detected after a threshold time period. In this scenario, method 400 returns to step 414, where audio playback application 116 outputs an unmodified center channel C for playback by output module 140. If voice input is detected within the threshold time period, method 400 continues to steps 418 and 418420, where audio playback application 116 generates and outputs a modified center channel.

[0055] It should be understood that, Figure 4 In method 400, steps 406 and 410 can be executed simultaneously or in different orders. Similarly, steps 408 and 412 can be executed simultaneously or in different orders. Additionally, steps 408, 412, and 414 can be executed simultaneously or in different orders. Furthermore, steps 408, 412, and 420 can be executed simultaneously or in different orders.

[0056] In summary, the audio playback system enables the playback of audio sources, such as songs or instrumental tracks from local or remote sources, along with audio inputs, such as vocal input from a user. The left and right channels associated with the audio sources are identified and isolated separately. A modified left channel is generated, consisting of the left channel minus the right channel. A modified right channel is generated, consisting of the right channel minus the left channel. A center channel is generated, which involves adding the right channel to the left channel. If vocal input is detected or the user selects a karaoke mode, a modified center channel is generated, from which the vocal input is removed or attenuated. The modified left channel is output to one or more left channel output devices, such as speakers, for playback. The modified right channel is output to one or more right channel output devices, such as speakers, for playback. The center channel, or a modified center channel, is output to one or more output devices (such as speakers) corresponding to the center channel for playback.

[0057] At least one technical advantage of the disclosed technology over the prior art is that it attenuates the main vocal component of the audio source that the user expects for their karaoke experience in real time. However, some secondary or background vocals are retained in the audio source processed according to the disclosed technology. A karaoke experience can be provided with virtually any audio source that is streamed for playback by attenuating or removing the main vocal component of the audio source in real time. Furthermore, using a microphone to capture in-vehicle vocal input allows that vocal input to be played along with the audio source. Therefore, playing an audio source without a main vocal component along with vocal input captured by one or more microphones provides an improved karaoke experience. These technical advantages provide one or more technological advancements superior to existing methods.

[0058] 1. In some embodiments, a computer-implemented method includes: receiving an audio source for playback by an audio playback system; identifying a left channel and a right channel associated with the audio source; generating a modified left channel including the right channel subtracted from the left channel; generating a modified right channel including the left channel subtracted from the right channel; such that the modified left channel is played on a left channel speaker of the audio playback system; and such that the modified right channel is played on a right channel speaker of the audio playback system.

[0059] 2. The computer-implemented method as described in Clause 1, further comprising: generating a center channel including the left channel added to the right channel; and causing the center channel to be played on at least one speaker of the audio playback system.

[0060] 3. The computer-implemented method as described in Clause 1 or 2, further comprising: generating a modified center channel by removing a human voice component from the center channel, and causing the modified center channel to be played on the at least one speaker of the audio playback system.

[0061] 4. The computer-implemented method of any one of Clauses 1 to 3, wherein generating the modified center channel includes mute, attenuate, or ducking the mid-frequency band components of the center channel.

[0062] 5. A computer-implemented method as described in any one of Clauses 1 to 4, wherein generating the modified center channel comprises compressing the center channel by reducing the dynamic range of the center channel to generate the modified center channel.

[0063] 6. The computer-implemented method of any one of Clauses 1 to 5, further comprising detecting human voice input from a microphone coupled to the audio playback system, wherein in response to detecting the human voice input, the modified center channel is played.

[0064] 7. The computer-implemented method of any one of Clauses 1 to 6, further comprising: detecting the termination of the human voice input; and, in response to detecting the termination of the human voice input, causing the center channel to be played on the at least one speaker of the audio playback system.

[0065] 8. A computer-implemented method as described in any one of Clauses 1 to 7, wherein detecting the voice input includes detecting user input via a microphone or user input device.

[0066] 9. A computer-implemented method as described in any one of Clauses 1 to 8, wherein the at least one speaker of the audio playback system comprises a center channel speaker.

[0067] 10. A computer-implemented method as described in any one of clauses 1 to 9, wherein the at least one speaker of the audio playback system comprises the left channel speaker and the right channel speaker.

[0068] 11. The computer-implemented method of any one of Clauses 1 to 10, further comprising causing human voice input received from at least one microphone coupled to the audio playback system to be played on at least one speaker of the audio playback system.

[0069] 12. In some embodiments, one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: receiving an audio source for playback by an audio playback system; identifying a left channel and a right channel associated with the audio source; generating a modified left channel including the right channel subtracted from the left channel; generating a modified right channel including the left channel subtracted from the right channel; causing the modified left channel to be played on a left channel speaker of the audio playback system; and causing the modified right channel to be played on a right channel speaker of the audio playback system.

[0070] 13. One or more non-transitory computer-readable media as described in Clause 12, wherein the steps further comprise: generating a center channel by adding the left channel to the right channel, generating a modified center channel by removing a vocal component from the center channel, and causing the modified center channel to be played on at least one speaker of the audio playback system.

[0071] 14. One or more non-transitory computer-readable media as described in Clause 12 or 13, wherein generating the modified center channel includes muting, attenuating, or ducking the mid-frequency band components of the center channel.

[0072] 15. One or more non-transitory computer-readable media as described in any one of Clauses 12 to 14, wherein the step further comprises detecting human voice input from a microphone coupled to the audio playback system, wherein in response to detecting the human voice input, the modified center channel is played.

[0073] 16. One or more non-transitory computer-readable media as described in any one of Clauses 12 to 15, wherein the step further comprises: detecting termination of the human voice input; and, in response to detecting the termination of the human voice input, causing the center channel to be played on the at least one speaker of the audio playback system.

[0074] 17. One or more non-transitory computer-readable media as described in any one of Clauses 12 to 16, wherein generating the modified center channel comprises compressing the center channel by reducing the dynamic range of the center channel to generate the modified center channel.

[0075] 18. One or more non-transitory computer-readable media as described in any one of Clauses 12 to 17, wherein the generation of the modified center channel is performed in response to a user selection of a karaoke mode.

[0076] 19. In some embodiments, a system includes one or more audio output devices, a memory storing an audio playback application, and a processor coupled to the memory, the processor performing the audio playback application by performing the following steps: receiving an audio source for playback by an audio playback system, identifying a left channel and a right channel associated with the audio source, generating a modified left channel including the right channel subtracted from the left channel, generating a modified right channel including the left channel subtracted from the right channel, such that the modified left channel is played on a left channel speaker of the audio playback system, and such that the modified right channel is played on a right channel speaker of the audio playback system.

[0077] 20. The system as described in Clause 19, wherein the one or more audio output devices, the memory, and the processor are integrated into the vehicle.

[0078] Any element of any claim recited in any claim and / or any element described in this application, in any form and in any combination thereof, falls within the contemplated scope and contemplated protection of this invention.

[0079] The description of the various embodiments has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0080] Various aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may be collectively referred to herein as a “module,” a “system,” or a “computer.” Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or a collection of circuits. Additionally, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media on which computer-readable program code is embodied.

[0081] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, device, or apparatus.

[0082] The foregoing description of aspects of this disclosure refers to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed via a processor of a computer or other programmable data processing apparatus, the instructions enable the implementation of functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions mentioned in the blocks may not appear in the order shown in the drawings. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or sometimes in reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action.

[0084] While the foregoing refers to embodiments of this disclosure, other and further embodiments of this disclosure may be conceived without departing from the basic scope of this disclosure, the scope of which is defined by the appended claims.

Claims

1. A computer-implemented method, comprising: Receive audio sources for playback by an audio playback system; Identify the left and right channels associated with the audio source; Generate a modified left channel, which includes subtracting the right channel from the left channel; Generate a modified right channel, which includes subtracting the left channel from the right channel; This causes the modified left channel to be played on the left channel speaker of the audio playback system; as well as This causes the modified right channel to be played on the right channel speaker of the audio playback system.

2. The computer-implemented method as described in claim 1, further comprising: Generate a center channel, which includes the addition of the right channel and the left channel; as well as This causes the center channel to be played on at least one speaker of the audio playback system.

3. The computer-implemented method as described in claim 2, further comprising: The modified center channel is generated by removing the human voice component from the center channel; as well as This causes the modified center channel to be played on at least one speaker of the audio playback system.

4. The computer-implemented method of claim 3, wherein generating the modified center channel includes mute, attenuating, or ducking the mid-frequency components of the center channel.

5. The computer-implemented method of claim 3, wherein generating the modified center channel comprises compressing the center channel by reducing the dynamic range of the center channel to generate the modified center channel.

6. The computer-implemented method of claim 3, further comprising detecting human voice input from a microphone coupled to the audio playback system, wherein in response to detecting the human voice input, the modified center channel is played.

7. The computer-implemented method of claim 6, further comprising: Detect the termination of the human voice input; as well as In response to the detection of the termination of the human voice input, the center channel is played on at least one speaker of the audio playback system.

8. The computer-implemented method of claim 6, wherein detecting the voice input includes detecting user input via a microphone or user input device.

9. The computer-implemented method of claim 2, wherein the at least one speaker of the audio playback system comprises a center channel speaker.

10. The computer-implemented method of claim 2, wherein the at least one speaker of the audio playback system comprises the left channel speaker and the right channel speaker.

11. The computer-implemented method of claim 1, further comprising causing human voice input received from at least one microphone coupled to the audio playback system to be played on at least one speaker of the audio playback system.

12. One or more non-transitory computer-readable media, the non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: Receive audio sources for playback by an audio playback system; Identify the left and right channels associated with the audio source; Generate a modified left channel, which includes subtracting the right channel from the left channel; Generate a modified right channel, which includes subtracting the left channel from the right channel; This causes the modified left channel to be played on the left channel speaker of the audio playback system; as well as This causes the modified right channel to be played on the right channel speaker of the audio playback system.

13. One or more non-transitory computer-readable media as claimed in claim 12, wherein the step further comprises: The center channel is generated by adding the left channel to the right channel; as well as The modified center channel is generated by removing the human voice component from the center channel; as well as This causes the modified center channel to be played on at least one speaker of the audio playback system.

14. One or more non-transitory computer-readable media as claimed in claim 13, wherein generating the modified center channel includes muting, attenuating, or ducking the mid-frequency band components of the center channel.

15. One or more non-transitory computer-readable media as claimed in claim 13, wherein the step further comprises detecting human voice input from a microphone coupled to the audio playback system, wherein in response to detecting the human voice input, the modified center channel is played.

16. One or more non-transitory computer-readable media as claimed in claim 15, wherein the step further comprises: Detect the termination of the human voice input; as well as In response to the detection of the termination of the human voice input, the center channel is played on at least one speaker of the audio playback system.

17. One or more non-transitory computer-readable media as claimed in claim 13, wherein generating the modified center channel comprises compressing the center channel by reducing the dynamic range of the center channel to generate the modified center channel.

18. One or more non-transitory computer-readable media as claimed in claim 13, wherein the generation of the modified center channel is performed in response to a user selection of a karaoke mode.

19. A system comprising: One or more audio output devices; Storage device; stores audio playback applications. as well as A processor, coupled to the memory, executes the audio playback application by performing the following steps: Receive audio sources for playback by an audio playback system; Identify the left and right channels associated with the audio source; Generate a modified left channel, which includes subtracting the right channel from the left channel; Generate a modified right channel, which includes subtracting the left channel from the right channel; This causes the modified left channel to be played on the left channel speaker of the audio playback system; as well as This causes the modified right channel to be played on the right channel speaker of the audio playback system.

20. The system of claim 19, wherein the one or more audio output devices, the memory, and the processor are integrated into the vehicle.