Dynamic effect karaoke
By analyzing the song audio signals and detecting the singer information in the car, and automatically adjusting the audio effects of the karaoke system, the problem of manual adjustment of the existing system is solved, and the sound quality and user experience are improved.
Patent Information
- Application Number
- CN202380079430.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-11-02
- Publication Date
- 2025-06-27
AI Technical Summary
Existing karaoke systems require manual adjustment of audio effect parameters, especially when noise is disturbed in the on-board environment, it is difficult to achieve automated adjustments, affecting the sound quality and user experience.
By analyzing the audio signal of the song, the relevant settings are automatically adjusted, multiple microphones are used to detect the number of singers participating in karaoke and the seat position, and the audio effects are dynamically allocated, such as automatic gain control, to ensure that the audio effects of each contributor are consistent.
It realizes automatic adjustment of audio effects under different songs and environmental conditions, improves sound quality and user experience, and reduces the need for manual adjustment.
Smart Images

Figure CN120226072A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 425,428, filed on November 15, 2022, the content of which is incorporated herein by reference. Background of the Invention
[0003] The present invention relates to vocal accompaniment for dynamically processing recorded audio content, particularly dynamic effects in a karaoke system.
[0004] Traditional "karaoke" is an interactive entertainment typically provided in clubs and bars where people sing along with recorded music using a microphone. This music is usually an instrumental version of a well - known pop song. The lyrics are typically displayed on a video screen along with moving symbols, color - changing, or music video images to guide the singer. Hardware and some software (such as smartphone apps) systems include functions such as digital signal processing of the user's voice, for example, adding reverb and tuning their voice to a specified pitch.
[0005] Karaoke systems are usually equipped with an electro - acoustic system with a microphone, amplifier, and speakers to enhance the singer's voice. The karaoke device plays a special music track with the lead vocal track missing. The music itself incorporates many audio effects applied during the music production process. Therefore, for the voice of a karaoke singer, the electro - acoustic system should also apply some audio effects so that the contributions of the music and the karaoke are very well - matched in style.
[0006] Therefore, karaoke devices usually provide various audio effects that can be selected to enrich the singer's voice. For example, an equalizer for emphasizing relevant frequencies, a compressor for reducing volume variations, a reverb effect with an adjustable reverb time, a delay that adds a decaying echo of the voice to the output signal, a chorus effect that creates the illusion of multiple singers singing simultaneously, an exciter that adds a brightening effect to the voice, a pitch transformation or harmonizer effect, and so on.
[0007] In - vehicle karaoke, sometimes also known as "Carpool Karaoke" (referring to the television show starring James Corden), is a type of karaoke that can be performed by passengers or the vehicle driver. Commercial products can support multiple functions, such as capturing the voices of multiple passengers in the vehicle and reducing feedback sounds from the vehicle's speakers. Summary of the Invention
[0008] Traditional karaoke systems require adjustment of individual effects for different types of songs. For example, the reverb time for a slow song may be longer compared to the optimal reverb time for a faster-paced song. Similarly, it is desirable to match the delay time between two successive echoes to the rhythm of the song. Depending on the instruments of the song, the equalizer may need to be adjusted to blend the sounds into the mix. Ballads or emotional songs may require volume variations, while in energetic songs, the dynamics of the sound should be compressed. Additionally, these potential adjustments can also vary within a song related to the current section of the song, with different effects perhaps being chosen for the chorus compared to the verse. Currently available karaoke systems require manual adjustment of these effects. There is a desire for a more automated system to assist in automatically adjusting the relevant parameters without necessarily requiring any manual adjustment.
[0009] The vehicle environment poses many challenges to karaoke systems, including a relatively "quiet" acoustic environment, as well as the presence of road noise and other environmental noises that are loud and vary, for example, with speed and road type, and may also include external noises such as sirens and construction pile drivers. Very generally, one or more systems and methods described in this document dynamically change the processing characteristics of a user's (i.e., the singer's) input based on the characteristics of the audio being sung and / or the acoustic environment in which the user is singing. Advantages of system processing can include an improved user experience (e.g., more fun or engaging to use the system) and / or higher-quality audio output (e.g., the combined presented audio and captured audio results in more desirable and / or pleasing characteristics).
[0010] In this document, the term "karaoke" should be interpreted broadly to include any situation in which a system is configured to present an acoustic signal to one or more users and capture the audio produced by one or more users during the presentation of the acoustic signal. For purposes of discussion, the presented acoustic signal may be referred to below as a "song", without meaning that the audio signal includes singing or spoken words, nor limiting it to backing music; and the captured audio (or its processed version) may be referred to as the user's "voice", without meaning that the captured audio necessarily includes lyrics or other spoken or sung words. Finally, it is not required to present the text of a song or the like to the user of a karaoke system, nor is it required that the captured audio be presented to the user along with the song.
[0011] In one aspect, a computer-implemented karaoke system adjusts relevant settings based on the attributes of a song, such as automatically determined by analyzing the audio signal of the song.
[0012] In some embodiments, a Karaoke system is deployed for use in a vehicle, e.g., for use by a driver and / or one or more passengers of the vehicle. For example, multiple microphones can be used as dedicated microphones for distributed speakers, or multiple microphones can be used in a close proximity in an array configuration to process a beamformer to focus on a specific direction where a passenger is speaking. Using multiple microphones, the system can also detect the number of singers participating in the Karaoke and their seats in the vehicle. The system may then assign different audio effects (e.g., automatic gain control (AGC)) to individual contributors to ensure a consistent level for each singing contributor. For example, a singer in the back seat may be assigned a typical effect for a background singer. Some pitch transformation may also be applied, e.g., transposition of an octave.
[0013] In some embodiments, a selected song is analyzed according to basic attributes such as speed, volume dynamics, music style, genre, song structure, etc. Based on these attributes, a set of effects is selected and configured. In some embodiments, some of this information may already be available from a database (e.g., pitch frequency, rhythm, and polyphony can be obtained from a MIDI file, and information about the genre can be obtained from a database), so no automatic extraction is required. In some embodiments, a predefined set of audio effects for a song can be prepared by manual tuning, e.g., a user manually tunes their favorite song to ensure an optimal (i.e., most desirable) set of audio effects.
[0014] In some embodiments, the audio effects can be changed within a song. For example, different effect settings may be used for the chorus and the verse. For example, based on the repetition of the chorus, the determination of the chorus and the verse can be automatically determined. As another example, different effects can be applied at the end of a song compared to during the song, e.g., a delay effect is introduced at the end when the vocals break off, and the delay effect does not persist as this may be disturbing to the singer.
[0015] In some embodiments, the audio effects can be adjusted according to the background noise. For example, in high-noise situations, a higher playback gain and less reverb can be applied. The background noise can be estimated using the same microphones as for Karaoke. Such audio effects are particularly applicable to in-vehicle applications where the background noise may be high and vary over time. In some embodiments, the microphones and speakers are not necessarily dedicated to Karaoke, e.g., integrated into an audio entertainment system, a hands-free phone system, and / or a voice assistant system.
[0016] In some embodiments, a Karaoke system is configured to interact with other Karaoke systems at other locations to form a distributed Karaoke system, enabling users to participate from multiple locations. For an in-vehicle Karaoke system, the vehicle is connected via a mobile communication system, and drivers and / or passengers in multiple vehicles can provide vocals for a song. For example, the playback of the song and the vocals can be synchronized or otherwise coordinated by the system. The voice of the singer is not only played in the local vehicle but also transmitted to other vehicles and added to the Karaoke track together with the voices from afar. The audio playback of two vehicles can be synchronized as much as possible, and the remaining mismatch in synchronization (which may be inevitable) is taken into account in the reverberation effect. The sound from the remote vehicle can be processed with different audio effects and presented as a chorus on the surround speakers. For example, the music in two vehicles A and B starts simultaneously and is synchronized. Then, the voice of the singer from vehicle A is fed into the effect section of the remote vehicle B, where, for example, a surround effect, reverberation, etc. are generated. The voice of the singer in vehicle B will also be transmitted to vehicle A.
[0017] In one aspect, generally, a method for dynamically modifying audio in response to user input for presentation together with the playback of a source song includes processing a microphone signal to generate an audio voice signal representing the user input.
[0018] Based on the characteristics of the source song, parameter values for one or more audio modification methods are determined, and the audio voice signal is processed using the audio modification method configured according to the determined parameter values to generate an enhanced vocal signal. The advantage of this modification is that the user does not have to manually readjust the parameters when changing songs, which can be particularly beneficial when the user is busy with other tasks, such as driving a vehicle.
[0019] The enhanced vocal signal and the source song are combined to generate an audio drive signal, and the audio drive signal is provided to the user for acoustic presentation.
[0020] An acoustic signal is obtained at the microphone to generate the microphone signal. The acoustic signal includes at least the user's voice and the acoustic presentation of the audio drive signal. Processing the microphone signal can then include removing the acoustic presentation in the audio drive signal based on an adaptation using at least one of a reference of the audio drive signal and the source song. The acoustic signal can include ambient noise, and the processing of the microphone signal includes noise reduction.
[0021] The audio modification methods include one or more of reverberation, echo, excitation, and pitch modification processing.
[0022] The characteristics of the source song include one or more of genre, rhythm, pitch, and time signature.
[0023] Determining parameter values includes determining time-varying parameter values that vary during a song. Such time-variation can be beneficial because different parts of a song (e.g., chorus and verse) may require different processing.
[0024] The original vocal signal and the vocals-removed song signal can be determined to correspond to the original source song, and combining the enhanced vocal signal and the source song includes combining the enhanced vocal signal and the vocals-removed song signal.
[0025] Processing an audio sound signal to determine a sound level of a user input. The sound level can indicate the presence or absence of sound, or can indicate the sound volume or energy.
[0026] Determining that a user input is present in the sound signal during a first period and absent during a second period.
[0027] Forming an audio drive signal includes combining the audio sound signal and the vocals-removed song signal during the first period to produce the audio drive signal.
[0028] Forming an audio drive signal includes combining the original vocal signal and the vocals-removed song signal during the second period to produce the audio drive signal. This presentation of the original vocals can be beneficial when the user forgets the lyrics and starts singing at a lower level.
[0029] Forming the audio drive signal further includes combining the original vocal signal during a first interval at an attenuation level based on the determined sound level (e.g., based on a history of the sound level or time filtering). This attenuated presentation of the original vocals can be beneficial when the user may be unsure of the lyrics and starts singing at a lower level.
[0030] Determining the original vocal signal and the vocals-removed song signal includes receiving the vocal signal and the vocals-removed signal before playing the source song.
[0031] Determining the original vocal signal and the vocals-removed song signal includes processing the original source song to mix the vocal component and the vocals-removed component of the original source song.
[0032] Detecting sound in a microphone signal and, during a period when no sound is detected in the microphone signal, providing a signal corresponding to the original source song, including providing at least some of the original vocal signal for acoustic presentation to a user.
[0033] Collecting a microphone signal inside the cabin of a first vehicle and presenting the audio drive signal inside the cabin of the first vehicle.
[0034] Receiving a remote vocal signal from a second vehicle and combining the enhanced vocal signal, the remote vocal signal, and the source song to produce an audio drive signal.
[0035] An enhanced vocal signal is provided for presentation in a second vehicle.
[0036] The presentation of the song in the first vehicle and the second vehicle is synchronized.
[0037] During the presentation of the song to the user, parameter values of one or more audio modification methods are determined at the first vehicle based on the characteristics of the source song.
[0038] Before presenting the song to the user, parameter values of one or more audio modification methods are determined based on the characteristics of the source song.
[0039] In another aspect, generally, instructions are stored on a non - transitory machine - readable medium. When a processor executes these instructions, the processor is caused to perform all the steps of any of the above - mentioned methods.
[0040] In another aspect, generally, an audio processing system includes a processor configured to perform all the steps of any of the above - mentioned methods. The audio processing system may include an in - vehicle audio processing system. The audio processing system may be integrated into at least one of an audio entertainment system, a hands - free telephone system, or a voice assistant system.
[0041] Other features and advantages of the present invention will be apparent from the following description and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic diagram of a vehicle - based karaoke system.
[0043] Figure 2 is a schematic block diagram of an audio processing system.
[0044] Figure 3 is a schematic block diagram of a sound processing module.
[0045] Figure 4 is a schematic block diagram of an audio processing system with dynamic remixing. DETAILED DESCRIPTION
[0046] Referring to Figure 1 , the audio processing system 100 provides a way for the user 110 to "sing along" with recorded audio presented in an acoustic environment, generally providing an enjoyable experience of a combination of user vocal input and recorded audio for the user or other listeners. This process is commonly referred to as "karaoke", although this term is used to describe the system, the characteristics of the system should not be inferred from the usage of this term. Specific examples of the karaoke system are described below in the context where the acoustic environment is a vehicle 120 (or multiple vehicles) and the user and / or listeners are drivers and / or passengers in the vehicle.
[0047] In Figure 1 the vehicle example shown, user 110 is the driver of the vehicle, where audio is played through one or more speakers 103 in the vehicle cabin and the audio is captured at one or more microphones 102 within the cabin. Although a single microphone is shown, multiple microphones may also be used, such as to form a directional microphone array. Further, although the figure and discussion focus on a single user, examples of the system may be multiple users in the vehicle providing vocal input together. In some examples, there are other modes of monitoring the user. One such example is via camera 104, which can be used to detect and analyze the user's vocal output (e.g., based on lip position, etc.), especially in cases of high noise or high music volume. Some of the examples described below relate to users in multiple vehicles, such as users in a first vehicle 120 and a second vehicle 120A, and the processing in multiple vehicles can be coordinated, e.g., to synchronize input and / or output processing across multiple vehicles. Finally, the song played as part of this process may be stored locally in the vehicle (e.g., in the entertainment system), or may be streamed from a server 160 communicating with the vehicle, or may be streamed from the user's smartphone in the vehicle (e.g., via a Bluetooth wireless link), e.g., stored on the smartphone or retrieved by an application executing on the smartphone from a remote server (e.g., via a cellular data link). In some examples, metadata for the song may be available, e.g., including genre, pre-computed acoustic or musical parameters that may vary for the song and the lyrics text.
[0048] In at least some examples, the audio processing described below is performed in an entertainment system 101, which generally includes a general-purpose computer processor and / or a digital signal processing (DSP) processor under software control (e.g., software instructions stored in the entertainment system in a non-transitory machine-readable medium). User input to the entertainment system (e.g., manual or via voice command), which is not shown, can be used to initiate audio processing, select a desired song, adjust the volume, etc. Further, the entertainment system 101 can receive input from a vehicle system 121, e.g., providing vehicle states (e.g., vehicle speed, navigation status, etc.) that may affect system operation. In some examples, the audio processing is performed on the same platform as an in-vehicle communication (ICC) system, e.g., utilizing the same microphones and speakers for enhancing communication between users (e.g., drivers and passengers) in the vehicle.
[0049] Referring to Figure 2, Audio processing includes audio input processing 210, followed by sound effect processing 220, and then audio output processing 240. The audio input processing receives an audio input from vehicle microphone 102, which captures acoustic signals within the vehicle cabin. Typically, microphone 102 captures the user's voice 257, vehicle noise 259 (e.g., road noise or interfering speech from non-singers or voice assistants), and acoustic feedback 255 from speaker 103 to microphone 102. The audio input processing can perform one or both of acoustic "echo" removal and noise reduction. Echo removal removes as much of the signal emitted by the speaker as possible. This removal can be based on signal 245 representing the acoustic signal emitted by speaker 103, or it can be based on the audio signal 205 of the song being played, and can use an adaptive removal method, e.g., adjusting the removal filter (e.g., using the least mean square (LMS) method). For example, various noise reduction methods may be based on spectral subtraction or adaptive Wiener filtering. The audio input processing can also include acoustic beamforming or microphone selection to improve the capture of the user's voice and / or exclude interfering noise and / or speaker feedback. The audio input processing can also include automatic gain control and / or equalization, e.g., to match the level of song 205, or perhaps match a target level representing the original vocals excluded from the song provided with the song.
[0050] The output of the audio input processing 210 is an audio vocal signal 215, which ideally represents only the user's voice and may actually include residual noise and song signals that have been significantly attenuated. Then, the vocal signal 215 is processed by the sound effect processing module 220. As discussed further below, this module introduces audio effects that enhance the user's vocal input. The module receives the song signal 205, allowing for dynamic adjustment of the effects to match the entire song and / or vary within the song, e.g., having different effects in the verses and choruses. In some examples described further below, the song signal 205 includes data as well as the audio signal, e.g., providing metadata for adjusting the sound effects. In Figure 2 some examples not shown, acoustic conditions such as noise level or other noise characteristics (e.g., spectral shape) are provided to the sound effect processing module 220, e.g., to modify the parameters of certain sound effects, e.g., reducing the amount of reverb in high-noise situations. The output of this module is an enhanced vocal signal 225. In some examples, the sound effect processing is also directly controlled by the user, e.g., by setting parameters using manual or voice command inputs.
[0051] Then, the enhanced vocal signal 225 is combined with the song signal 205, as shown in this example, as signal summation 230, to produce a combined signal 235, which is passed to the audio output processing module 240. For example, the output processing module can amplify the combined signal to a user-controlled level and adjust typical parameters, such as balance and front-to-back attenuation, in the case of multiple speakers. The output of the audio output processing module 245 is an audio drive signal 245, which drives the speakers and can be fed back to the audio input processing module 210 for "echo" removal. In some examples, the audio output processing adapts to the vehicle's environment, for example, increasing the overall compression level in a high-noise environment.
[0052] Reference Figure 3 , in one example of the sound processing module 220, the song signal 205 is first processed by the song analysis module 310, and the analysis information is passed to the sound modification module 320, which processes the sound signal 215 to produce an enhanced sound signal 225. In one example, the song analysis module can provide an on / off indicator to indicate whether a song is playing or whether the song being played is suitable for karaoke due to the lack of vocals. The song analysis module 310 can make this determination, for example, based on spectral signal energy determination, which can distinguish non-song (e.g., just voice) audio from song audio. In another example, the song analysis module can include a broad classification 313 of the song as a whole, for example, by classifying the song type as "rock", "folk", "a cappella chorus", etc. The song classification can also track different parameters, such as tempo (e.g., beats per minute). Many other parameters can be determined from the acoustic signal and / or provided from the metadata provided for the song (such as level (power, amplitude), pitch, time signature, chorus-versus-verse indicator) and data regarding excluded lyrics (e.g., representation of words excluded from the song audio signal). The characterization of the song can also be done before the song is played. The feature data can be stored, for example, as metadata along with the song signal audio.
[0053] Sound modification can implement one or more controllable modifications of the sound signal. Such modifications can include adding reverb 322 or echo 323. Such modifications can be parameterized, for example, including the characteristics of the echo (e.g., time delay of the echo) or reverb characteristics (e.g., impulse response over time). Similarly, an exciter 324 (i.e., for adding harmonic distortion) can be parameterized by its gain, and pitch modification can be parameterized by the target key of the song (e.g., automatically detected in the song audio or provided in the metadata of the song).
[0054] Figure 3Not shown is an alternative that includes modifying the song audio based on the captured voice signal 215. For example, such song modification can include time modification to match the user's singing rate, or selection of portions of the song based on recognition or matching of the sung lyrics to the original lyrics. Such tracking may be beneficial in a vehicle environment where the driver may not be able to safely read the lyrics while driving!
[0055] Reference Figure 4 , in some examples, the system is able to adapt to whether the user is singing, so if they are singing, a karaoke experience is provided, and if they are not singing (e.g., because they have forgotten the lyrics), the original audio with vocals is provided. In this example, the functionality of the audio input processing 210 is the same as Figure 2 the example of. The voice signal 215 (which may or may not actually include the user's voice) is passed to a voice detector 450, which determines the voice level, which can be a binary indicator of whether the user is singing, or in some examples, the voice level can represent the acoustic volume of the singing (e.g., amplitude, energy). In some examples, the voice is detected only during periods when there is an original voice, for example, preventing the introduction of the user's voice at other times. Based on the voice level output of the voice detector 450, the dynamic mixer 440 passes, attenuates, or completely blocks the original vocals 435 while passing the vocals-removed song audio 432. In some examples, during the period when the user is singing, the output signal 445 of the mixer 440 corresponds to Figure 3 the song audio 205 (in some examples, including a attenuated version of the original vocals 435 based on the user's singing level), while during the period when the user is not singing, the signal 445 corresponds to the original song including its original vocals. In some examples, the attenuation of the original vocals is based on the voice level using the history (e.g., time filtering) of the voice level (whether binary or representing acoustic volume). For example, in use, when the user forgets the lyrics of the song and stops singing, the original vocals provide an audio cue for the lyrics, and if the user starts singing only at a lower level because they are unsure of the lyrics, the original vocals are included at an attenuated level. In Figure 4 the example shown, the optional sound effect processing 420 can operate in the Figure 2 and Figure 3 way of the sound effect processing 220, and if the voice detector 450 indicates that the user is not singing, the processing is optionally prohibited. In Figure 4In the example shown, the song source 401 provides a song with original vocals in audio form, and the automatic demixer 430 divides the audio into vocals 435 and the remaining audio as signal 432 before or during the playback of the song, such that the combination of signals 432 and 435 (e.g., sum) produces the original song audio. Various demixers can be used, e.g., as described in "Audio Demixing Challenge 2021" by Mitsufuji et al., arXiv:2108.13559v3 (2022). Note that various forms of demixing can be used, e.g., for multi-part vocals (e.g., Capella), specific vocal parts can be removed while other vocal parts are retained. In some examples, demixing is performed before the song is played, e.g., to reduce the computation required during the playback of the song. For example, the demixed song can be provided remotely to the vehicle and / or stored in the song library in the vehicle. In some examples, the presence of vocals (e.g., signal energy) in the original vocals 435 is used to determine the time periods during which the user is expected to sing, allowing the voice detector 450 to detect the user's voice only during these time periods and as a timeout basis after the absence of the user's voice, at which point the original vocals should be re-introduced. Note that in this document, when referring to combining a first signal with a second signal to produce a third signal, such combination is not restricted as a fourth signal can also be combined with the first and second signals to produce the third signal. Additionally, the specific order or method of forming the third signal should not be construed as limited to a particular sequence of operations as long as the third signal includes at least some of the other signals combined to form it.
[0056] As described above, in some examples, users in multiple vehicles coupled via a mobile communication system (e.g., cellular data) participate in providing vocals for a signal song. In some examples, each user in a vehicle has a cellular smartphone coupled to the entertainment system of each vehicle, and vehicle-to-vehicle communication is through the smartphones, while in some examples, vehicle-to-vehicle mobile communication is integrated into the vehicle. There are many examples where the entertainment systems in multiple vehicles coordinate to provide a multi-car experience. In one example, the same song is stored in each vehicle and played on two vehicles in a time-synchronized manner, with each vehicle picking up the vocals. The local vocals are played as described above and transmitted to the remote vehicle, where they are combined with the local vocals. The remote vocals are also received in the local system and added to the signal output in the local vehicle. While precise synchronization is desired, in cases where perfect synchronization cannot be achieved, this mismatch can be considered a reverberation effect such that the remote vocals are only heard as a delayed reverberation without a direct (i.e., non-delayed) component. In some such examples, in addition to this reverberation effect, the sound processing module can be modified to use the remote vocals. In some examples, the remote vocals can be provided as background singers, e.g., whose contribution is attenuated or spatially placed differently from the local vocals in the acoustic environment.
[0057] In some multi-vehicle examples, one vehicle may contribute a subset of the vocal parts rather than the synchronized audio, and these parts are then presented in another vehicle, where the user there presents additional vocal parts. This approach can be achieved by removing different vocal parts (for replacement by users in these vehicles) through different vehicles.
[0058] Referring again to Figure 3 , the mapping from the output of the song analysis to the parameters of the vocal modification can be based on various techniques. One technique utilizes predetermined data adjusted for different situations. For example, the settings of the reverb time and the exciter gain can be parameterized according to the genre of the song, as shown in the following table:
[0059]
[0060] In some examples, the mapping from the features of the song to the sound modification parameters is a learned mapping. For example, it is implemented as a decision tree, a neural network, and / or a regression model. This learned mapping can be based on, for example, training data that includes songs and manual user settings of the songs according to user preferences, so that the automated system basically imitates the human selection of settings. In some examples, this learned mapping can adapt to the preferences of a specific user, for example, based on monitoring the user's modification of the automatically set parameters.
[0061] As described above, the vehicle environment is an example where the system can be deployed. The dynamic adjustment of the vocal effects can also be used in more traditional environments such as nightclubs. In addition, the multi-location synchronization (or time shift) implementation may include all or part of the locations as non-vehicle locations.
[0062] The implementation of the above technologies can utilize software, hardware, or a combination of hardware and software. The implementation using software can utilize processor instructions stored on a non-transitory machine-readable medium, and the instructions are executed by a processor or a processing system to execute the method. The instructions can include high-level languages, intermediate representations (such as "bytecode"), or machine-level instructions. The processor executing the instructions can be a physical processor or can also involve a virtual processor that accepts the instructions and causes the physical processor hosting the virtual processor to execute the method steps. The processing system can include components executed in one or more vehicles (or other user devices), and can also include a processor based on a remote server (such as "the cloud"). The processing system can include a dedicated processor (such as a digital signal processor (DSP) for performing audio processing functions) and a high-performance digital processor (such as a graphics processing unit (GPU) or a tensor processor), and such a processing system can be used for functions implemented by artificial neural networks, such as sound removal. The hardware implementation can include dedicated circuits, for example, for implementing audio processing such as echo removal or vocal removal.
[0063] Multiple embodiments of the present invention have been described. However, it should be understood that the above description is intended to illustrate rather than limit the scope of the present invention, which is defined by the scope of the appended claims. Therefore, other embodiments are also within the scope of the appended claims. For example, various modifications can be made without departing from the scope of the present invention. In addition, some of the above steps may be independent of the order, so they can be executed in an order different from the described order.
Claims
1. A method for dynamically modifying user input (257) for presentation along with the playback of a source song (205), the method comprising: Processing a microphone signal to generate an audio sound signal (215) representing the user input; Determining parameter values of one or more audio modification methods (322 - 325) based on characteristics (312 - 313) of the source song; Processing the audio sound signal (215) with an audio modification method (322 - 325) configured according to the determined parameter values to generate an enhanced vocal signal (225); Combining the enhanced vocal signal (225) and the source song (205) to generate an audio drive signal (245); And Providing the user with the audio drive signal (245) for acoustic presentation.
2. The method according to claim 1 further comprises obtaining an acoustic signal at a microphone to generate the microphone signal, wherein, The acoustic signal at least includes the user's voice and the acoustic presentation of the audio drive signal.
3. The method according to claim 2, wherein Processing the microphone signal includes removing the acoustic presentation of the audio drive signal based on an adaptation using at least one of a reference of the audio drive signal and the source song.
4. The method according to any one of claims 1 to 3, wherein The acoustic signal includes ambient noise, and processing the microphone signal includes noise reduction.
5. The method according to any one of claims 1 to 4, wherein The audio modification method includes one or more of reverb, echo, excitation, and pitch modification processing.
6. The method according to any one of claims 1 to 5, wherein The characteristics of the source song include one or more of genre, rhythm, pitch, and time signature.
7. The method according to any one of claims 1 to 6, wherein Determining the parameter values includes determining time-varying parameter values that vary during the song.
8. The method according to any one of claims 1 to 7, further comprising: Determining an original vocal signal (435) corresponding to the original source song (401) and a vocal-removed song signal (432); And Wherein, combining the enhanced vocal signal (225) and the source song (205) includes combining the enhanced vocal signal and the vocal-removed song signal.
9. A method for dynamically modifying user input (257) for presentation along with the playback of a source song (401), the method comprising: Determining an original vocal signal (435) corresponding to the original source song (401) and a vocal-removed song signal (432); Processing a microphone signal to generate an audio sound signal (215); Processing the audio sound signal to determine the sound level of the user input, including determining the presence of the user input during a first period and the absence of the user input during a second period; Forming an audio drive signal (245), including combining the audio sound signal (215) and the vocal-removed song signal (432) during the first period to generate the audio drive signal (245); and combining the original vocal signal (435) and the vocal-removed song signal (432) during the second period to generate the audio drive signal (245); And Providing the user with the audio drive signal (245) for acoustic presentation.
10. The method according to claim 9, wherein, Forming the audio drive signal further includes combining the original vocal signal (435) at an attenuation level based on the determined sound level during a first interval.
11. The method according to any one of claims 8 to 10, wherein, Determining that the original vocal signal (435) and the vocals-removed song signal (432) includes receiving the vocal signal and the vocals-removed signal before playing the source song.
12. The method according to any one of claims 8 to 10, wherein Determining that the original vocal signal and the vocals-removed song signal includes processing the original source song to mix the vocal component and the vocals-removed component of the original source song.
13. The method according to any one of claims 8 to 12, further comprising detecting sound in the microphone signal and providing a signal corresponding to the original source song (401) during a period when no sound is detected in the microphone signal, including providing at least some of the original vocal signal (435) to the user for acoustic rendering.
14. The method according to any one of claims 1 to 13, wherein, The microphone signal is collected in the cabin of the first vehicle, and the audio drive signal is presented in the cabin of the first vehicle.
15. The method according to claim 14, further comprising receiving a remote vocal signal from a second vehicle and combining the enhanced vocal signal (225), the remote vocal signal, and the source song (205) to generate the audio drive signal (245).
16. The method according to claim 15, further comprising providing the enhanced vocal signal for presentation in the second vehicle.
17. The method according to any one of claims 15 and 16, further comprising synchronizing the presentation of the song in the first vehicle and the second vehicle.
18. The method according to claim 14, wherein, During the presentation of the song to the user, determining parameter values of one or more audio modification methods (322-325) based on characteristics (312-313) of the source song at the first vehicle.
19. The method according to claim 14, wherein, Before presenting the song to the user, determining parameter values of one or more audio modification methods (322-325) based on characteristics (312-313) of the source song.
20. A non-transitory machine-readable medium, including instructions stored thereon, which when executed by a processor cause the processor to perform all the steps of any one of claims 1 to 19.
21. An audio processing system, including a processor configured to perform all the steps of any one of claims 1 to 19.
22. The audio processing system according to claim 21, including an in-vehicle audio processing system.
23. The audio processing system according to claim 21 or 22, wherein the audio processing system is integrated into at least one of an audio entertainment system, a hands-free phone system, and a voice assistant system.